Table document identification method and device, electronic equipment and storage medium

By performing table detection and line detection on wired table documents, cell intersection points and structure information are determined, and combined with text detection and recognition, the problem of low recognition accuracy of wired table documents in the prior art is solved, achieving higher recognition accuracy and calculation efficiency.

CN120088804APending Publication Date: 2025-06-03IFLYTEK CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202411704009.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-11-26
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

In the prior art, the recognition accuracy of wired table documents is low, especially in scenarios where complex tables or texts are large, and the problem of misidentification or incomplete recognition of the table structure is prone to occur.

Method used

By performing table detection of the image to be recognized, obtaining the table image, and performing table line detection, and determining the table cell intersection information and structure information. Then, match the text detection box with the table cell structure information to perform text recognition to obtain the text recognition results for each cell.

Benefits of technology

Improve the recognition accuracy of wired table documents, ensure that each cell can correctly correspond to its text line area, reduce computing resource consumption, and improve computing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088804A_ABST
    Figure CN120088804A_ABST
Patent Text Reader

Abstract

The invention provides a table document recognition method and device, electronic equipment and a storage medium, and the method comprises the steps: carrying out the table detection of a to-be-recognized document image, and obtaining a table image; performing table line detection on the table image to obtain a transverse line segment and a vertical line segment; determining table cell intersection point information based on the transverse line segments and the vertical line segments, and determining table cell structure information according to the intersection point information; performing text detection on the to-be-recognized document image to obtain a detection result; and matching each text detection box in the detection result with each cell in the table cell structure information to obtain a text line region corresponding to each cell, and performing text recognition on the text line region to obtain a text recognition result of each cell. According to the table structure identification method and device, table line detection is carried out on the table image, the transverse line segments and the vertical line segments are obtained according to detection, the intersection point information and the structure information of the table cells can be accurately determined, and the accuracy of table structure identification is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular, to a method, apparatus, electronic device, and storage medium for recognizing tabular documents. Background Art

[0002] With the continuous development of deep learning technology, more and more visual algorithm models are applied to specific tasks such as image classification, detection, and recognition. Especially in the field of text recognition, the detection and recognition technology for the content of documents containing wired tables has shown extremely high demand in actual enterprise production and operation. For example, in the daily operation of an enterprise, it is often necessary to recognize various images such as office documents and bills. In these document images, in addition to containing regular text paragraphs, they often contain a large number of wired tables. Accurately detecting and recognizing such wired table document images can not only improve the automation processing level of the enterprise, but also help to greatly reduce the human input and lower the enterprise operation cost.

[0003] Currently, for the recognition of wired table documents, an end-to-end recognition scheme is usually adopted. First, the cell structure information is output from the table image, and then the recognition result of the table text is obtained by combining optical character recognition (OCR). However, this scheme often needs to model the table structure into a specific format (such as HTML or LaTeX sequence) for recognition, which not only places higher requirements on the training data, but also the model usually adopts a heavier Transformer structure, resulting in a large model volume and being difficult to be efficiently deployed on the local side. In addition, the model is limited by the maximum length of sequence recognition, and in the case of complex tables or scenes with a large amount of text, it is easy to have problems such as misrecognition or incomplete recognition of the table structure, resulting in a low recognition accuracy. Summary of the Invention

[0004] The present invention provides a method, apparatus, electronic device, and storage medium for recognizing tabular documents to solve the defect of low recognition accuracy of wired table documents in the related art.

[0005] The present invention provides a method for recognizing tabular documents, including: Performing table detection on a document image to be recognized to obtain a table image; Performing table line detection on the table image to obtain a plurality of horizontal line segments and a plurality of vertical line segments; Based on the plurality of horizontal line segments and the plurality of vertical line segments, determining table cell intersection information, and determining table cell structure information according to the table cell intersection information; Performing text detection on the document image to be recognized to obtain a detection result; Match each text detection box in the detection results with each cell in the table cell structure information to obtain the text line area corresponding to each cell, and perform text recognition on the text line area to obtain the text recognition results of each cell.

[0006] According to a table document recognition method provided by the present invention, the table cell intersection information includes an intersection coordinate array and an intersection sorting array. Determining the table cell intersection information based on the multiple horizontal line segments and the multiple vertical line segments includes: Based on the straight lines where each horizontal line segment is located and the straight lines where each vertical line segment is located, obtain all intersections of the table cells, and use the information of each intersection to construct an intersection coordinate array. The information of each intersection includes the index of the horizontal line segment where each intersection is located, the index of the vertical line segment where each intersection is located, and the coordinates of each intersection. Sort each intersection in the intersection coordinate array based on a preset order to obtain an intersection sorting array.

[0007] According to a table document recognition method provided by the present invention, the table cell structure information includes a cell corner point information array and a cell row-column index array. Determining the table cell structure information based on the table cell intersection information includes: Based on the table cell intersection information, determine the four corner points of each cell, and use the corner point information of each cell to construct a cell corner point information array. Based on the cell corner point information array, determine the row-column indexes of each cell, and use the row-column indexes of each cell to construct a cell row-column index array.

[0008] According to a table document recognition method provided by the present invention, determining the four corner points of each cell based on the table cell intersection information includes: Traverse each intersection in the intersection sorting array of the table cell intersection information; When the currently traversed intersection meets the first preset condition, determine the currently traversed intersection as the upper left corner point of the cell, and based on the second preset condition, search to the right along the horizontal line segment where the upper left corner point of the cell is located to find the upper right corner point of the cell; After determining the upper right corner point of the cell, based on the third preset condition, search downward along the vertical line segment where the upper left corner point of the cell is located to find the lower left corner point of the cell; Based on the index of the horizontal line segment where the upper left corner point of the cell is located and the index of the vertical line segment where the upper right corner point of the cell is located, match and obtain the lower right corner point of the cell from the intersection coordinate array of the table cell intersection information; The upper left corner point, the upper right corner point, the lower left corner point, and the lower right corner point of the cell are the four corner points of the cell.

[0009] A table document recognition method provided by the present invention, determining the row and column indexes of each cell based on the cell corner point information array, includes: Based on the cell corner point information array, determining the minimum height and minimum width of all cells; Based on the minimum height and the minimum width, respectively grouping all horizontal line segments and all vertical line segments to obtain a horizontal line segment grouping dictionary and a vertical line segment grouping dictionary, where the key of the grouping dictionary is the line segment index and the value is the group number of the line segment; Based on the cell corner point information array, determining the line segment index of each cell, and applying the line segment index to match the row and column group numbers of each cell from the horizontal line segment grouping dictionary and the vertical line segment grouping dictionary, and taking the row and column group numbers of each cell as the row and column indexes of each cell.

[0010] A table document recognition method provided by the present invention, matching each text detection frame in the detection result with each cell in the table cell structure information to obtain the text line area corresponding to each cell, includes: Traversing each text detection frame and each cell, and calculating the text box overlap degree, cell overlap degree, and minimum bounding rectangle of the current text detection frame and the current cell; When the text box overlap degree is greater than or equal to a preset threshold, determining that the current text detection frame is the text line area corresponding to the current cell; When the cell overlap degree is greater than or equal to a preset threshold, determining that the minimum bounding rectangle is the text line area corresponding to the current cell; When both the text box overlap degree and the cell overlap degree are less than the preset threshold, determining the text line area corresponding to the current cell based on the current text detection frame and the minimum bounding rectangle.

[0011] A table document recognition method provided by the present invention, performing table detection on the document image to be recognized to obtain a table image, includes: Based on a table detection model, performing table detection on the document image to be recognized to obtain the coordinates of each corner point of the table and a table detection frame; Performing geometric transformation based on the coordinates of each corner point of the table and the table detection frame to obtain the table image.

[0012] A table document recognition method provided by the present invention, performing text recognition on the text line area to obtain the text recognition result of each cell, includes: Based on the text recognition model, perform text recognition on the text line regions to obtain the text recognition results of the respective cells; The text recognition model includes a first text recognition model and a second text recognition model. The first text recognition model is trained based on text images containing horizontal text, and the second text recognition model is trained based on mixed text images containing horizontal and vertical text.

[0013] According to a table document recognition method provided by the present invention, it further includes: Perform text recognition on each text detection box located outside the table in the detection results to obtain text recognition results outside the table; Integrate the text recognition results outside the table with the text recognition results of the respective cells to obtain a document recognition result.

[0014] The present invention also provides a table document recognition device, including: A table detection unit for performing table detection on a document image to be recognized to obtain a table image; A table line detection unit for performing table line detection on the table image to obtain a plurality of horizontal line segments and a plurality of vertical line segments; A table structure analysis unit for determining table cell intersection information based on the plurality of horizontal line segments and the plurality of vertical line segments, and determining table cell structure information according to the table cell intersection information; A text detection unit for performing text detection on the document image to be recognized to obtain a detection result; A text recognition unit for matching each text detection box in the detection result with each cell in the table cell structure information to obtain a text line region corresponding to each cell, and performing text recognition on the text line region to obtain the text recognition results of the respective cells.

[0015] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the table document recognition method as described in any one of the above is implemented.

[0016] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the table document recognition method as described in any one of the above is implemented.

[0017] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the table document recognition method as described in any one of the above is implemented.

[0018] The table document recognition method, device, electronic device, and storage medium provided by the present invention first perform table detection to quickly locate and extract the table image from the document image to be recognized, which can reduce the complexity of subsequent processing. Then, table line detection is performed on the table image to obtain the horizontal and vertical line segments of the table. Based on these horizontal and vertical line segments, the intersection information and structural information of the table cells can be accurately determined, ensuring the accuracy of table structure recognition, thereby providing a reliable basis for subsequent text matching and recognition. By matching the text detection frame with the table cells, it can be ensured that each cell can be correctly corresponding to its text line area, thereby improving the accuracy of text recognition and further improving the recognition accuracy of wired table documents. In addition, the present invention does not adopt a complex deep learning model, but determines the intersection information and structural information of the table cells according to line segment rules, thereby reducing the consumption of computing resources and improving the computing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the present invention or related technologies, the following will briefly introduce the drawings required for use in the embodiments or related technology descriptions. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0020] Figure 1 is a schematic flowchart of the table document recognition method provided by the present invention; Figure 2 is a schematic diagram of the output result of the table detection model provided by the present invention; Figure 3 is a schematic diagram of the visualization result of the table line detection model provided by the present invention; Figure 4 is a schematic diagram of the visualization result of step 133 provided by the present invention; Figure 5 is a schematic diagram of the visualization result of step 134 provided by the present invention; Figure 6 is a schematic diagram of the visualization result of the OCR detection and recognition model provided by the present invention; Figure 7 is a schematic diagram of the visualization result of the excel table provided by the present invention; Figure 8 is a schematic flowchart of the general detection and recognition method applied to wired table documents provided by the present invention; Figure 9 is a schematic diagram of the structure of the table document recognition device provided by the present invention; Figure 10 is a schematic diagram of the structure of the electronic device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0021] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention in conjunction with the accompanying drawings in the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts fall within the protection scope of the present invention.

[0022] Currently, table recognition is mainly divided into wired table recognition and wireless table recognition. Among them, wireless table recognition refers to the extraction and recognition of the structure and text content of a table lacking obvious border lines in an image; wired table recognition refers to the extraction and recognition of the structure and text content of a table with clear border lines in an image. The present invention mainly proposes a set of practical and feasible general technical solutions for wired table documents.

[0023] Current wired table recognition usually includes two-stage and end-to-end recognition methods: The two-stage scheme uses a method of first detecting and then recognizing, that is, first detecting the table in the image, then recognizing the table structure, and finally combining OCR to obtain the recognition result of the table text; while the end-to-end scheme directly outputs the cell structure information of the table through the table image, and finally combines OCR to obtain the recognition result of the table text.

[0024] However, both of these schemes have certain defects: The existing end-to-end scheme needs to model the table structure into an HTML or LaTeX sequence recognition task, which not only places higher requirements on the training data, but also the model usually adopts a heavier Transformer structure, resulting in a large model volume and being difficult to be efficiently deployed on the local side. In addition, the model is limited by the maximum length of sequence recognition, and in the case of complex tables or scenarios with a large amount of text, it is prone to problems such as incorrect structure recognition or incomplete recognition, resulting in a low recognition accuracy. And the existing two-stage scheme has an overly complex design for the cell structure recognition model, with high requirements for data annotation, which is not conducive to rapid deployment and application.

[0025] In addition, although traditional table line detection algorithms (such as the Hough transform algorithm and morphological operations) can achieve the recognition of the table structure to a certain extent, the recognition accuracy is not high, and the generality and generalization ability are limited. For unseen wired table data, these traditional methods often have incorrect structure recognition due to improper threshold setting, and it is difficult to meet the requirements of actual applications.

[0026] In response to this, the present invention adopts a two-stage scheme to provide a complete method for recognizing tabular document. By using deep learning algorithms to train a table detection model and a table line detection model respectively, and combining with a table cell intersection detection algorithm, the structural information of the wired table cells can be easily obtained; at the same time, combining with a text detection model and a text recognition model to detect and recognize the text content of the tabular document; finally, using a text line-cell matching algorithm to output the text recognition results of the entire table cells. This method has strong completeness and good versatility, is suitable for most wired table scenarios, and at the same time, the model can be easily made lightweight, can be easily deployed to the local and cloud, and has low requirements for the training data set, thus overcoming the above-mentioned defects.

[0027] Figure 1 is a schematic flowchart of the method for recognizing tabular document provided by the present invention, as Figure 1 shown, the method includes: Step 110, perform table detection on the document image to be recognized to obtain a table image.

[0028] It should be noted that the document image to be recognized refers to the original image that needs to be subjected to table recognition and text extraction. This image can be a scanned paper document, a screenshot of an electronic document, or any image containing table and text information. Table detection refers to the process of recognizing and extracting the table area from the document image to be recognized.

[0029] Specifically, the table detection of the document image to be recognized can be achieved through the following steps: First, preprocess the image to be recognized, such as denoising, binarization, enhancing contrast, etc., to improve the accuracy of table detection. Then, use image processing techniques to extract features in the image, such as edges, corners, textures, etc. These features can help identify the structure of the table. Next, use machine learning or deep learning algorithms to analyze the extracted features and identify the table area in the image. Finally, the recognized table area can be further processed, such as thinning the edges, correcting the shape, etc., to obtain a more accurate table image. Here, the table image refers to the image that only contains the table area extracted from the document image to be recognized after table detection.

[0030] Step 120, perform table line detection on the table image to obtain a plurality of horizontal line segments and a plurality of vertical line segments.

[0031] Specifically, after detecting and locating the table image, the table line detection can be further performed on the table image. Here, the table line detection refers to accurately extracting the line information that constitutes the table from the recognized table image. These lines, namely the table lines, are the key elements that define the table structure and cell layout. Through table line detection, the algorithm can identify the horizontal and vertical line segments in the table, providing a basis for subsequent determination of cell intersection information and structural information.

[0032] Specifically, a horizontal and vertical line segmentation binary image of the table can be obtained through a table line detection model. Then, the minimum bounding rectangles are fitted to the horizontal and vertical lines in the binary image, and the starting and ending coordinates of the table line segments (including horizontal and vertical line segments) can be approximately obtained from the four-point coordinates of the average rectangle. Here, the table line detection model can be implemented using a deep learning semantic segmentation model, such as the HRNet or Unet segmentation model. Through these segmentation models, the horizontal and vertical line segments of the table can be recognized. It should be understood that in a table document, the horizontal and vertical line segments are the basic elements that make up the table structure. Horizontal line segments are usually located at the top, bottom, and between each cell of the table, and are used to separate different rows. Visually, they appear as horizontal lines. Vertical line segments are located on the left, right, and between each cell of the table, and are used to separate different columns. Visually, they appear as vertical lines.

[0033] Step 130: Based on the multiple horizontal line segments and the multiple vertical line segments, determine the intersection information of table cells, and based on the intersection information of table cells, determine the structural information of table cells.

[0034] Specifically, horizontal and vertical line segments are the basic elements that make up the table structure. These line segments visually divide the table into different rows and columns, and their intersections define the corner points of the cells. Therefore, by detecting and identifying these lines, all the intersection information of the table cells can be determined.

[0035] Specifically, by traversing all the detected horizontal and vertical line segments and finding their intersections, all the intersections of the table cells can be obtained, which together define the boundaries of the table cells. The intersection information of table cells refers to the position and coordinate information of the intersections of all the horizontal and vertical line segments in the table document, and these intersections can also be referred to as the corner points of all the cells in the table.

[0036] After obtaining the intersection information of table cells, the arrangement and combination of these intersections can be further analyzed to determine the structural information of table cells. Specifically, traverse all the intersections in a certain order. According to the horizontal and vertical line segments where each intersection is located, sequentially search and determine the four corner point information of each cell. Subsequently, group all the horizontal and vertical line segments. According to the group numbers of the horizontal and vertical line segments where the four corner points of each cell are located, determine the starting row and column numbers and the ending row and column numbers of each cell. Based on the four corner point information and the row and column numbers of each cell, the structural information of table cells can be obtained. Here, the structural information of table cells refers to the arrangement and layout information of all the cells in the table document, and these information define the number of rows and columns of the table, as well as the position and size of each cell.

[0037] Step 140: Perform text detection on the document image to be recognized to obtain a detection result.

[0038] It should be noted that text detection is a technology in the field of computer vision, aiming to automatically identify and locate text regions in an image. Through text detection, the algorithm can identify which regions in the image contain text and determine the positions and ranges of these text regions.

[0039] Specifically, while performing table detection and cell structure parsing on the document image to be recognized, text detection can be synchronized. Specifically, it can be achieved through the following steps: First, preprocess the image, including operations such as denoising, enhancing contrast, and binarization, to improve the accuracy of text detection; then, use image processing techniques to extract features in the image, which are related to the texture, shape, color, etc. of the text region; next, use machine learning or deep learning algorithms to analyze the extracted features to identify the text regions in the image. In addition, the recognized text regions can be further processed, such as refining the edges and correcting the shapes, to obtain more accurate text detection frames.

[0040] It can be understood that during the text detection process, the detection result refers to the position and range information of all text regions in the image recognized by the algorithm. These detection results are usually represented in the form of text detection frames, and each detection frame contains the coordinate information of a text region.

[0041] Step 150: Match each text detection frame in the detection result with each cell in the table cell structure information to obtain the text line region corresponding to each cell, and perform text recognition on the text line region to obtain the text recognition result of each cell.

[0042] Specifically, according to the table cell structure information determined in Step 130, the exact position and range of each cell can be obtained. Then, the algorithm traverses all text detection frames and all cells, and calculates the overlap degree between the text detection frame and the cell. Finally, according to the calculation result of the overlap degree, the text line region corresponding to each cell is determined. Here, the text line region refers to the region where the text content located within the matching cell is, which can be the region identified by the text detection frame matching the cell, or the region after correcting the text detection frame matching the cell. The embodiments of the present invention do not make specific limitations on this.

[0043] After matching the text line area of ​​each cell, text recognition can be performed on these text line areas to obtain the text recognition results corresponding to each cell. Here, text recognition (also known as optical character recognition, OCR) is a computer vision technology used to extract and recognize text information from an image. The text recognition result of each cell refers to the text content extracted and recognized from the text line area of ​​the cell after the text recognition process.

[0044] The method provided by the embodiment of the present invention firstly detects the table to quickly locate and extract the table image from the document image to be identified, thereby reducing the complexity of subsequent processing, and then performs table line detection on the table image to obtain the horizontal and vertical line segments of the table. Based on these horizontal and vertical line segments, the intersection information and structural information of the table cells can be accurately determined to ensure the accuracy of the table structure recognition, thereby providing a reliable basis for subsequent text matching and recognition. By matching the text detection box with the table cell, it can be ensured that each cell can correctly correspond to its text row area, thereby improving the accuracy of text recognition, and then improving the recognition accuracy of wired table documents. In addition, the present invention does not adopt a complex deep learning model, but determines the intersection information and structural information of the table cells according to the line segment rules, thereby reducing the consumption of computing resources and improving computing efficiency.

[0045] Based on any of the above embodiments, step 110 specifically includes: Step 111, based on the table detection model, performing table detection on the document image to be recognized to obtain the coordinates of each corner point of the table and the table detection frame; Step 112, performing geometric transformation based on the coordinates of each corner point of the table and the table detection frame to obtain the table image.

[0046] Specifically, the table detection model can adopt the CenterNet structure, which is a target detection network architecture. The backbone network of the model can flexibly select network structures such as resnet18, mobilenet or shufflenet_v2 according to the usage scenario to adapt to cloud or local scenarios. In order to facilitate and better correct the distorted table, the model models the minimum horizontal rectangle and the coordinates of the four corner points of the supervised table; the training data set of the model includes collecting several real wired table images, marking the area where the table is located, and marking the coordinates of the four corner points of the table. At the same time, in order to better detect the positions of the four corner points of the distorted table, these real tables and some collected background images are cut out to artificially synthesize a batch of data. The synthesis process controls the rotation and scaling of the table to increase the diversity of the synthesized image.

[0047] Figure 2 is a schematic diagram of the output result of the table detection model provided by the present invention, such asFigure 2 As shown, for an image containing a table, after being detected by a trained table detection model, the model can directly output the positions of the four corner points of the table (i.e., Figure 2 the four corner points of the green wireframe in Figure 2 ) and the table detection box of the minimum horizontal rectangle (i.e.,

[0048] the blue rectangular box in ). Considering that the table may be distorted during the image acquisition process, therefore, according to the four corner point coordinates output by the model, geometric transformation of the image can be performed to correct the table within the green wireframe according to the shape of the blue rectangular box, and the corrected table image can be obtained. Performing OCR recognition on the corrected table image helps to improve the text detection and recognition rate.

[0049] Specifically, the table line detection model can be implemented by using a deep learning semantic segmentation model, such as the HRNet or Unet segmentation model. By detecting the table image through the table line detection model, the horizontal and vertical line segments of the table can be recognized. During the training process of the table line detection model, the ground truth of the background pixels can be set to 0, the horizontal line pixels can be set to 1, and the vertical line pixels can be set to 2 for the model to distinguish and learn. The loss function for model training is CrossEntorpy Loss combined with Dice Loss (a loss function for semantic segmentation tasks). The training dataset of the model can use several real table data annotated by labelme, and at the same time, several synthetic data images are artificially synthesized according to these real data and background images.

[0050] Figure 3 is a schematic diagram of the visualization result of the table line detection model provided by the present invention. As shown in Figure 3, for a table image displayed on the left, the horizontal and vertical line segmentation binary image of the table (i.e., the black and white image shown in the middle of Figure 3) can be obtained through the table line detection model. By fitting the minimum circumscribed rectangle to the line segments in the binary image, the starting coordinates and ending coordinates of the table line segments can be approximately obtained from the four point coordinates of the average rectangle (the green dots in Figure 3 represent the starting points of the line segments, and the red dots represent the ending points of the line segments). It should be understood that both the horizontal and vertical line segments are obtained by connecting pixels. By fitting the minimum circumscribed rectangle to these pixels, the corresponding line segments can be obtained.

[0051] Based on any of the above embodiments, the intersection information of the table cells includes an intersection coordinate array and an intersection sorting array; correspondingly, in step 130, determining the intersection information of the table cells based on the multiple horizontal line segments and the multiple vertical line segments includes: Step 131, based on the straight lines where each horizontal line segment is located and the straight lines where each vertical line segment is located, obtain all the intersections of the table cells, and apply the information of each intersection to construct an intersection coordinate array, where the information of each intersection includes the index of the horizontal line segment where each intersection is located, the index of the vertical line segment where it is located, and the coordinates of each intersection; Step 132, sort each intersection in the intersection coordinate array according to a preset order to obtain an intersection sorting array.

[0052] Specifically, according to the intersection information of the horizontal and vertical line segments of the table, the embodiments of the present invention propose a set of table cell parsing solutions based on line segment rules. First, the intersection information of all the horizontal and vertical line segments in the table can be determined. Through the straight lines where all the horizontal line segments and vertical line segments of the table are located, all the intersection coordinate arrays all_intersection_points of the table cells can be obtained. The key of each element in this array is composed of the index of the horizontal line segment where the intersection is located and the index of the vertical line segment where it is located, and the value is the coordinate of the intersection. For example, all_intersection_points[i][j] represents the intersection coordinate of the straight line where the i-th horizontal line segment is located and the straight line where the j-th vertical line segment is located.

[0053] After obtaining the intersection coordinate array, sort all the intersections in the preset order from left to right and then from top to bottom to obtain the first intersection sorting array intersection_table_hv_points_info. The elements of this array store the intersection coordinates and the indexes of the vertical and horizontal lines where the intersection is located. At the same time, sort all the intersections in the order from top to bottom and then from left to right to obtain the second intersection sorting array intersection_table_vh_points_info. The elements of this array store the intersection coordinates and the indexes of the horizontal and vertical lines where the intersection is located. It should be understood that there are two intersection sorting arrays, namely the first intersection sorting array and the second intersection sorting array, and these two arrays sort the intersections in different orders. According to the first intersection sorting array, all the intersections on the same horizontal line segment can be determined, and according to the second intersection sorting array, all the intersections on the same vertical line segment can be determined.

[0054] Based on any of the above embodiments, the table cell structure information includes a cell corner point information array and a cell row and column index array; correspondingly, in step 130, determining the table cell structure information according to the table cell intersection information includes: Step 133: Based on the intersection information of the table cells, determine the four corner points of each cell, and construct an array of cell corner point information by applying the corner point information of each cell; Specifically, according to the determined intersection information of the table cells (including the intersection coordinate array and the intersection sorting array), the four corner point coordinates of each cell can be further determined.

[0055] Furthermore, in Step 133, the determining of the four corner points of each cell based on the intersection information of the table cells includes: Step 1331: Traverse each intersection in the intersection sorting array of the intersection information of the table cells; Step 1332: When the currently traversed intersection meets the first preset condition, determine the currently traversed intersection as the upper left corner point of the cell, and based on the second preset condition, search to the right along the horizontal line segment where the upper left corner point of the cell is located to find the upper right corner point of the cell; where the first preset condition is that it is not the last intersection of the vertical line segment of the table where it is located and not the last intersection of the horizontal line segment of the table where it is located, and the second preset condition is that it is not the last intersection of the vertical line segment of the table where it is located; Step 1333: After determining the upper right corner point of the cell, based on the third preset condition, search downward along the vertical line segment where the upper left corner point of the cell is located to find the lower left corner point of the cell; the third preset condition is that it is not the last intersection of the horizontal line segment of the table where it is located; Step 1334: Based on the index of the horizontal line segment where the upper left corner point of the cell is located and the index of the vertical line segment where the upper right corner point of the cell is located, match and obtain the lower right corner point of the cell from the intersection coordinate array of the intersection information of the table cells; The upper left corner point, the upper right corner point, the lower left corner point, and the lower right corner point of the cell are the four corner points of the cell.

[0056] Specifically, according to the first intersection point sorting array and the second intersection point sorting array, the intersection points can be traversed from the first row to the second-to-last row and from the first column to the second-to-last column of the table to determine whether each traversed intersection point is the top-left point tl_point of a cell. The determination principle of tl_point (i.e., the first preset condition) is that this point cannot be the last one among all the intersection points of the vertical line segment where it is located (obtained using the second intersection point sorting array intersection_table_vh_points_info), nor can it be the last one among the intersection points of the horizontal line segment where it is located (obtained using the first intersection point sorting array intersection_table_hv_points_info). If this principle is not satisfied, continue to traverse and judge the next intersection point; if this point is the top-left point of the cell, then search to the right from this point for the top-right point tr_point of the cell. The determination principle of tr_point (i.e., the second preset condition) is that this point cannot be the last one among the intersection points of the vertical line segment where it is located. If this principle is not satisfied, continue to search to the right until it is found. If it cannot be found, then tr_point must be on the right boundary line segment of the table.

[0057] At this point, the coordinates of the two points tl_point and tr_point of the table cell have been determined. Among them, the indexes of the horizontal line segment and the vertical line segment where tl_point is located can be respectively expressed as start_h_segment_index and start_v_segment_index, and the indexes of the horizontal line segment and the vertical line segment where tr_point is located can be respectively expressed as start_h_segment_index and end_v_segment_index. Let: tl_point_info = [tl_point, start_h_segment_index, start_v_segment_index] tr_point_info = [tr_point, start_h_segment_index, end_v_segment_index] Among them, tl_point and tr_point respectively represent the coordinates of the upper left corner point and the upper right corner point. Continue to search down for the lower left corner point bl_point of the table cell. The specific method is to first obtain v_points_info = intersection_table_vh_points_info[start_v_segment_index], that is, all the intersection information of the vertical line segment to which the upper left corner point tl_point of the cell belongs. Let tl_point_index be the index of tl_point in v_points_info, and then search down to confirm bl_point. The determination principle of bl_point (i.e., the third preset condition) is: this point cannot be the last one of all the intersection points of the horizontal line segment where it is located (obtained using the first intersection sorting array intersection_table_hv_points_info). If this principle is not satisfied, continue to search down until it is found. If it cannot be found, then bl_point must be on the lower boundary line segment of the table.

[0058] At this point, the coordinates of the lower left corner point bl_point of the table cell can be determined. The indexes of the horizontal and vertical line segments intersecting with bl_point are respectively represented as end_h_segment_index and start_v_segment_index. Let: bl_point_info = [bl_point, end_h_segment_index, start_v_segment_index] Through end_h_segment_index and end_v_segment_index, the lower right corner point of the table cell can be directly obtained from the corner point coordinate array, that is: br_point = all_intersection_points[end_h_segment_index][end_v_segment_index] Let br_point_info = [br_point, end_h_segment_index, end_v_segment_index]. At this point, the coordinate information of the four corner points of this single cell, as well as the index information of the corresponding horizontal and vertical line segments, are obtained and shown as follows: cell_points_info = [tl_point_info, tr_point_info, bl_point_info, br_point_info] In this way, by traversing in a loop, an array cells_points_info of the four-point information of all cells in the table is obtained. The key of each element in this array is the index of the cell, and the value is the information of the four corner points of the cell. Figure 4 It is a schematic diagram of the visualization result of step 133 provided by the present invention, as Figure 4 shown. The numbers 0, 1, 2... in the figure are the indexes of each cell.

[0059] Step 134: Based on the array of cell corner point information, determine the row and column indexes of each cell, and apply the row and column indexes of each cell to construct an array of cell row and column indexes.

[0060] Specifically, in step 134, the determining of the row and column indexes of each cell based on the array of cell corner point information includes: Step 1341: Based on the array of cell corner point information, determine the minimum height and minimum width of all cells; Step 1342: Based on the minimum height and the minimum width, group all horizontal line segments and all vertical line segments respectively to obtain a horizontal line segment grouping dictionary and a vertical line segment grouping dictionary. The key of the grouping dictionary is the line segment index, and the value is the group number of the line segment; Step 1343: Based on the array of cell corner point information, determine the line segment indexes of each cell, and apply the line segment indexes to match the row and column group numbers of each cell from the horizontal line segment grouping dictionary and the vertical line segment grouping dictionary, and use the row and column group numbers of each cell as the row and column indexes of each cell.

[0061] Specifically, first, according to the array of cell corner point information obtained in step 133, the width and height of each cell can be calculated, so that the minimum width min_cell_width and minimum height min_cell_height of the table cells can be obtained.

[0062] Then, group all the horizontal line segments of the table. First, create a horizontal line segment grouping dictionary valid_h_groups, where the key is the index of each horizontal line segment and the value corresponds to the group number. Initialize the group number group_h_index to start from 0. Traverse all the horizontal line segments. The adjacent two horizontal line segments are current_h_segment and next_h_segment. Calculate the vertical distance v_dis1 from the starting point of the current_h_segment line segment to the next_h_segment, and calculate the vertical distance v_dis2 from the end point of the current_h_segment line segment to the next_h_segment. If and , it is considered that current_h_segment and next_h_segment are in the same group, and the corresponding group numbers are equal; otherwise, the group number group_h_index is incremented by 1, and the group number of next_h_segment is the new group_h_index.

[0063] Next, group all the vertical segments of the table. First, create a vertical segment grouping dictionary valid_v_groups, where the key is the index of each vertical segment and the value corresponds to the group number. Initialize the group number group_v_index to start from 0. Traverse all the vertical segments, with two adjacent vertical segments being current_v_segment and next_v_segment. Calculate the horizontal distance h_dis1 from the starting point of the current_v_segment to the horizontal direction of next_v_segment, and calculate the horizontal distance h_dis2 from the ending point of the current_v_segment to the horizontal direction of next_v_segment. If and , it is considered that current_v_segment and next_v_segment are in the same group, and the corresponding group numbers are equal; otherwise, the group number group_v_index is incremented by 1, and the group number of next_v_segment is the new group_v_index. It should be understood that the purpose of this step is to divide the vertical segments that can be regarded as on the same vertical line but are actually disconnected in the table into the same group.

[0064] Finally, determine the starting row number and column number, as well as the ending row number and column number of each cell. First, create an array cells_row_col_indices for cell row and column indices. Then, traverse the array cells_points_info of cell corner point information. According to the four-point information cell_points_info = [tl_point_info, tr_point_info, br_point_info, bl_point_info] of each cell in it, the starting horizontal segment index start_h_segment_index, the ending horizontal segment index end_h_segment_index, the starting vertical segment index start_v_segment_index, and the ending vertical segment index end_v_segment_index of the cell can be obtained. Based on these segment indices, the starting row and column numbers and the ending row and column numbers of the cell can be matched from the horizontal segment grouping dictionary and the vertical segment grouping dictionary, that is: the starting row number start_row of the cell = valid_h_groups[start_h_segment_index], the ending row number of the cell is end_row = valid_h_groups[end_h_segment_index] - 1, the starting column number start_col of the cell = valid_v_groups[start_v_segment_index], and the ending column number is end_col = valid_v_groups[end_v_segment_index] - 1. These row and column numbers constitute the row and column indices of the cell. Add [start_row, end_row, start_col, end_col] to the array cells_row_col_indices to construct the array of cell row and column indices.

[0065] Figure 5 is a schematic diagram of the visualization result of step 134 provided by the present invention, as Figure 5 shown. In the figure, the row and column indices of each cell are represented in red font at the lower left corner of each cell. By representing the row and column indices of the cell in the Figure 5 form shown, it is convenient to output the text content in the table to an excel table after parsing and recognizing the table. Therefore, the first two indices represent the row where the cell is located, and the last two indices represent the column where the cell is located. For example, the row and column index "0,0,1,6" indicates that the cell is located in row 0 and columns 1 to 6. Specifically, reference can be made to the Figure 7 shown excel table.

[0066] Based on any of the above embodiments, step 140 specifically includes: Based on the text detection model, perform text detection on the text image to be recognized to obtain the detection result.

[0067] Specifically, the text detection model can adopt the pan++ model. The backbone network of the model can be flexibly selected from network structures such as resnet18, mobilenet, or shufflenet_v2 according to the usage scenario to adapt to cloud or local scenarios; the training dataset of the model can adopt several real text detection data labeled by labelImg and open-source datasets.

[0068] Furthermore, in step 150, the text recognition of the text line area to obtain the text recognition results of each cell includes: Based on the text recognition model, perform text recognition on the text line area to obtain the text recognition results of each cell; The text recognition model includes a first text recognition model and a second text recognition model. The first text recognition model is trained based on text images containing horizontal text, and the second text recognition model is trained based on mixed text images containing horizontal and vertical text.

[0069] Specifically, the text recognition model can adopt the CRNN model. The backbone network of the model can be flexibly selected from network structures such as resnet18, mobilenet, or shufflenet_v2 according to the usage scenario to adapt to cloud or local scenarios; the dataset of the model adopts real collected text line annotation data and open-source datasets, and an artificial synthetic text line dataset is generated by collecting thousands of fonts combined with background images and corpora. In order to be able to recognize both horizontal and vertical text, two CRNN models are trained. One CRNN model has a training dataset of text lines that are all horizontal text images, and the other CRNN model has a training dataset of mixed horizontal and vertical text line images.

[0070] Furthermore, the method further includes: Perform text recognition on each text detection box outside the table in the detection result to obtain the text recognition result outside the table; Integrate the text recognition result outside the table with the text recognition results of each cell to obtain the document recognition result.

[0071] It should be noted that in a table document image, in addition to the text inside the table that needs to be recognized, there may also be text content outside the table. Therefore, when performing text detection and recognition on the document image to be recognized, the text inside and outside the table can be detected and recognized. Finally, by integrating the text recognition result outside the table and the text recognition results of each cell, the recognition result of the entire document can be obtained.

[0072] Figure 6 It is a schematic diagram of the visualization result of the OCR detection and recognition model provided by the present invention. As Figure 6 shown, for a table image, the trained OCR detection and recognition model (including a text detection model and a text recognition model) is used to perform text detection and recognition on it. The text detection box results of the entire image are stored in the array text_boxes. The elements in this array are the four-point coordinates of the text detection box. The text line recognition results are stored in the array text_recog_results, and the elements are the recognized strings of the text lines. Figure 6 The green rectangular boxes shown in [Figure] are the output text detection boxes. Each detection box has a corresponding index number, and the recognized text string is marked on each detection box. It should be understood that when performing text detection and recognition on the entire table image, the text detection boxes outside the table can be distinguished first, and text recognition can be performed on these text detection boxes to obtain the text recognition results outside the table. For the text recognition of each cell, text recognition can be performed after the text detection box of each cell is matched.

[0073] Based on any of the above embodiments, in step 150, the matching of each text detection box in the detection result with each cell in the table cell structure information to obtain the text line area corresponding to each cell includes: Step 151, traverse each text detection box and each cell, and calculate the text box overlap degree, cell overlap degree, and minimum bounding rectangle of the current text detection box and the current cell; Step 152, when the text box overlap degree is greater than or equal to a preset threshold, determine that the current text detection box is the text line area corresponding to the current cell; Step 153, when the cell overlap degree is greater than or equal to a preset threshold, determine that the minimum bounding rectangle is the text line area corresponding to the current cell; Step 154, when both the text box overlap degree and the cell overlap degree are less than the preset threshold, determine the text line area corresponding to the current cell based on the current text detection box and the minimum bounding rectangle.

[0074] Specifically, in order to obtain the text recognition results of each cell more accurately, the embodiment of the present invention combines the cell corner point information array cells_points_info and the text detection box array text_boxes, designs a text box-cell matching algorithm, and combines secondary correction detection and recognition to obtain the recognition result cells_contents of the final table. First, create a cell-text detection box matching array cells_text_boxes. The detailed process is as follows: Initialize the cell index cell_id = 0 and the text detection box index text_id = 0. Loop through all cells and text detection boxes, and calculate the text box overlap iou_text and the cell overlap iou_ceil between the current text detection box text_box and the current cell detection box ceil_box, as well as the minimum bounding rectangle intersection_box of their intersection. Among them, iou_text and iou_ceil are the ious (overlap degrees) relative to the text box and the cell, and the calculation formulas are as follows: Among them, represents the area of the cell region, which can be calculated based on the corner coordinates of the cell; represents the area of the text detection box region, which can be calculated based on the corner coordinates of the text detection box.

[0075] When iou_text >= 0.999, it indicates that the cell basically encloses the detection box. At this time, cells_text_boxes[cell_id] can directly contain text_boxes[text_id], that is, the text line region corresponding to the current cell is the current text detection box.

[0076] When iou_cell >= 0.999, it indicates that the detection box basically encloses the cell. In this case, cells_text_boxes[cell_id] can be made to contain the corrected intersection detection box, that is, the current text detection box is corrected according to the shape of the minimum bounding rectangle, and the corrected intersection detection box is the text line region corresponding to the current cell.

[0077] When iou_text < 0.999 and iou_cell < 0.999, let the width and height of intersection_box be intersection_box_w and intersection_box_h respectively, and the width and height of the text detection box be text_box_w and text_box_h respectively. Only when intersection_box_h / text_box_h > 0.6 and min(intersection_box_w, intersection_box_h) >= 5, that is, when the ratio of the height of the bounding rectangle to the height of the text detection box is large enough, make cells_text_boxes[cell_id] contain the corrected intersection detection box.

[0078] Detect all text detection boxes cells_text_boxes in the table cells, extract all text line region images through perspective transformation, use two models of text line recognition CRNN to recognize the text line results contained in each table cell, and merge the recognition results of multiple text lines in the same cell, adding line breaks between different text lines.

[0079] Finally, organize the cells_contents array, where the elements are the information of all cells, including the starting and ending row and column numbers of the cells, the four-point coordinates of the table cells, and the merged text strings recognized inside the cells. Combining the text recognition results outside the table with cells_contents, the detection and recognition results of the entire table document can be output.

[0080] Figure 7 It is a schematic diagram of the visualization result of the excel table provided by the present invention, as Figure 7 shown. For the cells_contens array, third-party libraries such as the xlwt module in python or OpenXLSX in c++ can be used to generate excel table files accordingly, and the restored visualization result of the excel table is as Figure 7 shown.

[0081] Based on any of the above embodiments, Figure 8 It is a schematic diagram of the process of the general detection and recognition method for wired table documents provided by the present invention, as Figure 8 shown. The method includes: Step S1, perform table detection on the wired table document image to be recognized to obtain the table main body image.

[0082] Specifically, based on the table detection model, perform table detection on the wired table document image to obtain the four corner point coordinates of the table and the table detection box in the form of a horizontal rectangle; perform perspective transformation correction according to the four corner point coordinates of the table and the table detection box to obtain the table main body image.

[0083] Step S2, based on the table line detection model, perform table line detection on the table main body image to obtain the horizontal line segments and vertical line segments of the table.

[0084] S3, perform table structure parsing based on line segment rules to obtain the table cell structure.

[0085] S31, determine the intersection information Specifically, according to the straight lines where all the horizontal line segments in the table are located and the straight lines where all the vertical line segments are located, obtain the array of all intersection point coordinates of the table cells. Sort all the intersection point coordinates in the intersection point coordinate array in the order from left to right and then from top to bottom to obtain the first intersection point sorting array. At the same time, obtain the second intersection point sorting array in the order from top to bottom and then from left to right.

[0086] S32. Determine the cell coordinates Specifically, traverse the intersection points in the first intersection point sorting array and the second intersection point sorting array in the order from the first row to the penultimate row and from the first column to the penultimate column. Determine whether each traversed intersection point is the upper left corner point of the cell. If the currently traversed intersection point does not meet the first preset condition, it is determined that the currently traversed intersection point is not the upper left corner point of the cell, and continue to traverse the next intersection point. If the currently traversed intersection point meets the first preset condition, it is determined that the currently traversed intersection point is the upper left corner point of the cell, and based on the second preset condition, search for the upper right corner point of the cell to the right of the currently traversed intersection point.

[0087] After finding the upper right corner point of the cell, based on the third preset condition, search downward along the vertical line segment where the upper left corner point of the cell is located to find the lower left corner point of the cell. Based on the index of the horizontal line segment where the lower left corner point is located and the index of the vertical line segment where the upper right corner point is located, combined with the intersection point coordinate array, determine the lower right corner point of the cell. According to the determined upper left corner point, upper right corner point, lower left corner point, and lower right corner point of the cell, determine the four-point information array of the table cell.

[0088] S33. Determine the row and column indices of the cell Specifically, according to the four-point information array of the table cells, calculate the width and height of each cell to obtain the minimum width and height among all cells. Group all the horizontal line segments and the vertical line segments of the table respectively, determine the group numbers of each horizontal line segment and each vertical line segment to obtain the horizontal line segment grouping dictionary and the vertical line segment grouping dictionary. The key of the dictionary is the index of each line segment, and the value corresponds to the group number.

[0089] According to the four-point information array of the table cells, determine the starting horizontal line segment index, ending horizontal line segment index, starting vertical line segment index, and ending vertical line segment index of the cell. According to these indices, combined with the horizontal line segment grouping dictionary and the vertical line segment grouping dictionary, determine the starting row number and column number, ending row number and column number of each cell. These row and column numbers are the row and column indices of the cell.

[0090] S4. Text detection and recognition Specifically, based on the text detection model, text detection is performed on the wired table document image to be recognized, and the detection result is obtained. The detection result includes the four-point coordinates of each text detection box, which are stored in the text detection box array. The text outside the table in the detection result is recognized through the text recognition model to obtain the text recognition result outside the table.

[0091] S5, integrate the table cell recognition results, use the text line-cell matching algorithm, and output the text recognition result of the entire table cell Specifically, according to the four-point information array of the table cells and the text detection box array, loop through all cells and text detection boxes, calculate the text overlap degree, cell overlap degree, and minimum bounding rectangle of the current text detection box and the current cell. Based on the text overlap degree, cell overlap degree, minimum bounding rectangle, and the current text detection box, determine the detection box corresponding to each cell. Use perspective transformation to extract all text line region images from all detection boxes in each cell, and use the text recognition model to recognize the text line results contained in each cell. For multiple text line recognition results of the same cell, merge them and add line breaks between different text lines. Here, the text line recognition model can recognize horizontal and vertical texts simultaneously. Integrate the text recognition results of all cells and the text recognition results outside the table, and output the recognition result of the entire document.

[0092] The embodiment of the present invention proposes a complete and general solution for the detection and recognition of wired table documents, with strong generality, capable of processing most wired tables, capable of recognizing horizontal and vertical texts, low data acquisition and training costs for each model, and can be deployed on the cloud and local. Moreover, the table structure algorithm uses a line segment rule-based solution, which is accurate and effective, and does not adopt a complex deep learning model, with high computational efficiency.

[0093] Based on any of the above embodiments, Figure 9 is the structural schematic diagram of the table document recognition device provided by the present invention, as Figure 9 shown, the device includes: A table detection unit 910, configured to perform table detection on the document image to be recognized to obtain a table image; A table line detection unit 920, configured to perform table line detection on the table image to obtain a plurality of horizontal line segments and a plurality of vertical line segments; A table structure analysis unit 930, configured to determine table cell intersection information based on the plurality of horizontal line segments and the plurality of vertical line segments, and determine table cell structure information according to the table cell intersection information; A text detection unit 940, configured to perform text detection on the document image to be recognized to obtain a detection result; The text recognition unit 950 is configured to match each text detection box in the detection result with each cell in the table cell structure information, obtain the text line area corresponding to each cell, and perform text recognition on the text line area to obtain the text recognition result of each cell.

[0094] The device provided by the embodiment of the present invention first performs table detection to quickly locate and extract the table image from the document image to be recognized, which can reduce the complexity of subsequent processing. Then, it performs table line detection on the table image to obtain the horizontal and vertical line segments of the table. Based on these horizontal and vertical line segments, the intersection information and structure information of the table cells can be accurately determined, ensuring the accuracy of table structure recognition, thereby providing a reliable basis for subsequent text matching and recognition. By matching the text detection boxes with the table cells, it can ensure that each cell can be correctly corresponding to its text line area, thereby improving the accuracy of text recognition and further improving the recognition accuracy of wired table documents. In addition, the present invention does not adopt a complex deep learning model, but determines the intersection information and structure information of the table cells according to line segment rules, thereby reducing the consumption of computing resources and improving the computing efficiency.

[0095] Based on any of the above embodiments, the intersection information of the table cells includes an intersection coordinate array and an intersection sorting array. The table structure analysis unit 930 includes an intersection information determination unit, and the intersection information determination unit is configured to: Based on the lines where each horizontal line segment is located and the lines where each vertical line segment is located, obtain all intersections of the table cells, and apply the information of each intersection to construct an intersection coordinate array. The information of each intersection includes the index of the horizontal line segment where each intersection is located, the index of the vertical line segment where each intersection is located, and the coordinates of each intersection; Sort each intersection in the intersection coordinate array based on a preset order to obtain an intersection sorting array.

[0096] Based on any of the above embodiments, the table cell structure information includes a cell corner point information array and a cell row-column index array. The table structure analysis unit 930 includes a structure information determination unit, and the structure information determination unit includes: A corner point determination subunit, configured to determine the four corner points of each cell based on the intersection information of the table cells, and apply the corner point information of each cell to construct a cell corner point information array; An index determination subunit, configured to determine the row-column index of each cell based on the cell corner point information array, and apply the row-column index of each cell to construct a cell row-column index array.

[0097] Based on any of the above embodiments, the corner point determination subunit is specifically configured to: Traverse each intersection in the intersection sorting array of the intersection information of the table cells; When the currently traversed intersection point satisfies the first preset condition, determine the currently traversed intersection point as the upper left corner point of the cell, and based on the second preset condition, search to the right along the horizontal line segment where the upper left corner point of the cell is located to find the upper right corner point of the cell; After determining the upper right corner point of the cell, based on the third preset condition, search downward along the vertical line segment where the upper left corner point of the cell is located to find the lower left corner point of the cell; Based on the index of the horizontal line segment where the upper left corner point of the cell is located and the index of the vertical line segment where the upper right corner point of the cell is located, match and obtain the lower right corner point of the cell from the intersection coordinate array of the table cell intersection information; The upper left corner point, the upper right corner point, the lower left corner point, and the lower right corner point of the cell are the four corner points of the cell.

[0098] Based on any of the above embodiments, the index determination subunit is specifically configured to: Based on the cell corner point information array, determine the minimum height and minimum width of all cells; Based on the minimum height and the minimum width, group all horizontal line segments and all vertical line segments respectively to obtain a horizontal line segment grouping dictionary and a vertical line segment grouping dictionary, where the key of the grouping dictionary is the line segment index and the value is the group number of the line segment; Based on the cell corner point information array, determine the line segment index of each cell, and apply the line segment index to match and obtain the row and column group numbers of each cell from the horizontal line segment grouping dictionary and the vertical line segment grouping dictionary, and use the row and column group numbers of each cell as the row and column indexes of each cell.

[0099] Based on any of the above embodiments, the text recognition unit 950 includes a text area determination subunit, and the text area determination subunit is used for: Traverse each text detection box and each cell, and calculate the text box overlap degree, cell overlap degree, and minimum bounding rectangle of the current text detection box and the current cell; When the text box overlap degree is greater than or equal to a preset threshold, determine the current text detection box as the text line area corresponding to the current cell; When the cell overlap degree is greater than or equal to a preset threshold, determine the minimum bounding rectangle as the text line area corresponding to the current cell; When both the text box overlap degree and the cell overlap degree are less than the preset threshold, determine the text line area corresponding to the current cell based on the current text detection box and the minimum bounding rectangle.

[0100] Based on any of the above embodiments, the table detection unit 910 is specifically configured to: Based on the table detection model, perform table detection on the document image to be recognized, and obtain the coordinates of each corner point of the table and the table detection frame; Perform geometric transformation based on the coordinates of each corner point of the table and the table detection frame to obtain the table image.

[0101] Based on any of the above embodiments, the text recognition unit 950 includes a cell text recognition subunit, and the cell text recognition subunit is used for: Based on the text recognition model, perform text recognition on the text line area to obtain the text recognition results of each cell; The text recognition model includes a first text recognition model and a second text recognition model. The first text recognition model is trained based on text images containing horizontal text, and the second text recognition model is trained based on mixed text images containing horizontal text and vertical text.

[0102] Based on any of the above embodiments, the device further includes a text integration unit, and the text integration unit is used for: Perform text recognition on each text detection frame outside the table in the detection result to obtain the text recognition result outside the table; Integrate the text recognition result outside the table with the text recognition results of each cell to obtain the document recognition result.

[0103] Figure 10 Illustrates a schematic physical structure diagram of an electronic device, as Figure 10 shown. The electronic device may include: a processor 1010, a communication interface 1020, a memory 1030, and a communication bus 1040. Among them, the processor 1010, the communication interface 1020, and the memory 1030 complete mutual communication through the communication bus 1040. The processor 1010 can call the logical instructions in the memory 1030 to execute the table document recognition method, and the method includes: performing table detection on the document image to be recognized to obtain a table image; performing table line detection on the table image to obtain a plurality of horizontal line segments and a plurality of vertical line segments; based on the plurality of horizontal line segments and the plurality of vertical line segments, determining the intersection information of table cells, and according to the intersection information of table cells, determining the cell structure information of the table; performing text detection on the document image to be recognized to obtain a detection result; matching each text detection frame in the detection result with each cell in the cell structure information of the table to obtain the text line area corresponding to each cell, and performing text recognition on the text line area to obtain the text recognition results of each cell.

[0104] In addition, when the logical instructions in the above-mentioned memory 1030 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the related technology, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0105] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the table document recognition method provided by the above-mentioned various methods. The method includes: performing table detection on the document image to be recognized to obtain a table image; performing table line detection on the table image to obtain a plurality of horizontal line segments and a plurality of vertical line segments; based on the plurality of horizontal line segments and the plurality of vertical line segments, determining table cell intersection information, and according to the table cell intersection information, determining table cell structure information; performing text detection on the document image to be recognized to obtain a detection result; matching each text detection box in the detection result with each cell in the table cell structure information to obtain a text line region corresponding to each cell, and performing text recognition on the text line region to obtain a text recognition result of each cell.

[0106] In yet another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the table document recognition method provided by the above-mentioned various methods. The method includes: performing table detection on the document image to be recognized to obtain a table image; performing table line detection on the table image to obtain a plurality of horizontal line segments and a plurality of vertical line segments; based on the plurality of horizontal line segments and the plurality of vertical line segments, determining table cell intersection information, and according to the table cell intersection information, determining table cell structure information; performing text detection on the document image to be recognized to obtain a detection result; matching each text detection box in the detection result with each cell in the table cell structure information to obtain a text line region corresponding to each cell, and performing text recognition on the text line region to obtain a text recognition result of each cell.

[0107] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0108] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the related technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0109] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or equivalently replace some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present invention.

Claims

1. A table document recognition method, characterized in that: include: Perform table detection on the document image to be recognized to obtain a table image; Performing table line detection on the table image to obtain a plurality of horizontal line segments and a plurality of vertical line segments; Determine table cell intersection information based on the multiple horizontal line segments and the multiple vertical line segments, and determine table cell structure information according to the table cell intersection information; Performing text detection on the document image to be recognized to obtain a detection result; Each text detection box in the detection result is matched with each cell in the table cell structure information to obtain a text row area corresponding to each cell, and text recognition is performed on the text row area to obtain a text recognition result for each cell.

2. The table document recognition method according to claim 1, characterized in that: The table cell intersection information includes an intersection coordinate array and an intersection sort array, and determining the table cell intersection information based on the multiple horizontal line segments and the multiple vertical line segments includes: Based on the straight lines where each horizontal line segment is located and the straight lines where each vertical line segment is located, all intersections of the table cells are obtained, and the information of each intersection is applied to construct an intersection coordinate array, wherein the information of each intersection includes the index of the horizontal line segment where each intersection is located and the index of the vertical line segment where each intersection is located and the coordinates of each intersection; Based on a preset order, the intersection points in the intersection point coordinate array are sorted to obtain an intersection point sorting array.

3. The table document recognition method according to claim 1, characterized in that: The table cell structure information includes a cell corner point information array and a cell row and column index array, and determining the table cell structure information according to the table cell intersection information includes: Based on the table cell intersection information, four corner points of each cell are determined, and the corner point information of each cell is applied to construct a cell corner point information array; Based on the cell corner point information array, the row and column indexes of each cell are determined, and the row and column index array of each cell is constructed by applying the row and column indexes of each cell.

4. The table document recognition method according to claim 3, characterized in that: The determining of four corner points of each cell based on the table cell intersection information includes: Traversing each intersection point in the intersection point sorting array of the table cell intersection point information; In the case where the currently traversed intersection meets the first preset condition, determining the currently traversed intersection as the upper left corner of the cell, and based on the second preset condition, searching for the upper right corner of the cell to the right along the horizontal line segment where the upper left corner of the cell is located; After determining the upper right corner of the cell, based on a third preset condition, searching for the lower left corner of the cell downward along the vertical line segment where the upper left corner of the cell is located; Based on the index of the horizontal line segment where the upper left corner of the cell is located and the index of the vertical line segment where the upper right corner of the cell is located, the lower right corner of the cell is matched from the intersection coordinate array of the table cell intersection information; The upper left corner point of the cell, the upper right corner point of the cell, the lower left corner point of the cell and the lower right corner point of the cell are four corner points of the cell.

5. The table document recognition method according to claim 3, wherein determining the row and column indexes of each cell based on the cell corner point information array comprises: Based on the cell corner point information array, determine the minimum height and minimum width of all cells; Based on the minimum height and the minimum width, all horizontal line segments and all vertical line segments are grouped respectively to obtain a horizontal line segment grouping dictionary and a vertical line segment grouping dictionary, wherein the key of the grouping dictionary is the line segment index and the value is the line segment group number; Based on the cell corner point information array, the line segment index of each cell is determined, and the line segment index is applied to match the row and column group numbers of each cell from the horizontal line segment grouping dictionary and the vertical line segment grouping dictionary, and the row and column group numbers of each cell are used as the row and column indexes of each cell.

6. The table document recognition method according to claim 1, characterized in that: The matching of each text detection frame in the detection result with each cell in the table cell structure information to obtain a text row area corresponding to each cell includes: Traverse each text detection box and each cell, calculate the text box overlap, cell overlap and minimum bounding rectangle of the current text detection box and the current cell; When the text box overlap is greater than or equal to a preset threshold, determining that the current text detection box is a text row area corresponding to the current cell; When the cell overlap is greater than or equal to a preset threshold, determining the minimum bounding rectangle to be the text row area corresponding to the current cell; When the text box overlap and the cell overlap are both smaller than a preset threshold, a text row area corresponding to the current cell is determined based on the current text detection box and the minimum circumscribed rectangle.

7. The table document recognition method according to any one of claims 1 to 6, characterized in that: The step of performing table detection on the document image to be recognized to obtain a table image includes: Based on the table detection model, performing table detection on the document image to be recognized to obtain the coordinates of each corner point of the table and the table detection frame; The table image is obtained by performing geometric transformation based on the coordinates of each corner point of the table and the table detection frame.

8. The table document recognition method according to any one of claims 1 to 6, characterized in that: The performing text recognition on the text line area to obtain the text recognition results of each cell includes: Based on the text recognition model, performing text recognition on the text line area to obtain text recognition results of each cell; The text recognition model includes a first text recognition model and a second text recognition model. The first text recognition model is trained based on a text image containing horizontal text, and the second text recognition model is trained based on a mixed text image containing horizontal text and vertical text.

9. The table document recognition method according to any one of claims 1 to 6, characterized in that: Also includes: Performing text recognition on each text detection frame outside the table in the detection result to obtain a text recognition result outside the table; The out-of-table text recognition result is integrated with the text recognition result of each cell to obtain a document recognition result.

10. A table document recognition device, characterized in that: include: A table detection unit, used for performing table detection on the document image to be recognized to obtain a table image; A table line detection unit, used to perform table line detection on the table image to obtain a plurality of horizontal line segments and a plurality of vertical line segments; a table structure parsing unit, configured to determine table cell intersection information based on the plurality of horizontal line segments and the plurality of vertical line segments, and determine table cell structure information according to the table cell intersection information; A text detection unit, used to perform text detection on the document image to be recognized to obtain a detection result; The text recognition unit is used to match each text detection box in the detection result with each cell in the table cell structure information to obtain the text row area corresponding to each cell, and perform text recognition on the text row area to obtain the text recognition results of each cell.

11. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the table document recognition method according to any one of claims 1 to 9 is implemented.

12. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the table document recognition method according to any one of claims 1 to 9 is implemented.

Citation Information

Cited By

  • Image processing method and device

    CN121415429A

  • Vector diagram general table identification method and system based on AI model, terminal and medium

    CN121861688A