Table recognition method and device and electronic equipment
This table recognition method, which combines OCR algorithms and Hough line detection with geometric constraints, solves the problem of strong dependence on data annotation in existing technologies, and achieves efficient and accurate table recognition. It is suitable for document digitization and automatic parsing of complex tables.
Patent Information
- Application Number
- CN202511368543.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-24
- Publication Date
- 2025-10-28
AI Technical Summary
Existing table recognition methods rely on deep learning models, which require a large amount of labeled data for training, resulting in a strong dependence on data labeling.
The OCR algorithm is used to identify the text position in the table image, generate a text area mask, convert it to the HSV color space to design a multi-scale color mask, perform Canny edge detection and multi-scale Hough line detection, combine geometric constraints to reconstruct the table structure, determine the intersection of cell rows and columns, and generate the cell position matrix of the table.
It achieves efficient and accurate table recognition without the need for deep learning models, reduces development costs, can quickly adapt to new table formats, reduces dependence on data annotation, and runs on ordinary CPUs.
Smart Images

Figure CN120853205A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer vision technology, specifically to a table recognition method, apparatus, and electronic device. Background Technology
[0002] In document digitization, table recognition is a crucial step in achieving data structuring. The table recognition methods used in related technologies primarily rely on deep learning models (such as Faster R-CNN, YOLO, and Transformer), but these methods require extensive training with labeled data and are highly dependent on data annotation.
[0003] There is currently no effective technical solution to the problem that related technologies require a large amount of labeled data for training when recognizing tables, and are highly dependent on data labeling. Summary of the Invention
[0004] The main objective of this disclosure is to provide a table recognition method, apparatus, and electronic device to solve the problem in related technologies that table recognition requires a large amount of labeled data for training and is highly dependent on data labeling.
[0005] To achieve the above objectives, a first aspect of this disclosure provides a table recognition method, comprising: The OCR algorithm is used to identify the text location in the table image, generate a text region mask, filter the table image, and obtain a pure table background image without text. The plain table background image is converted to the HSV color space, and multi-scale color masks are designed to address the differences in the shades of the table lines, thereby enhancing the contrast of the candidate line segment areas and obtaining the enhanced image. Canny edge detection is performed on the enhanced image to extract edge features, and multi-scale Hough line detection is used to generate a list of candidate line segments containing both horizontal and vertical directions. By applying geometric constraints to all candidate line segments in the candidate line segment list, the table structure is reconstructed, and a minimum line segment list representing the table structure is generated. The row and column intersections of cells are determined based on a list of minimum line segments, and the intersections are sorted according to their coordinates to generate a cell position matrix of the table, thereby enabling the parsing of the table structure.
[0006] Optionally, the plain table background image is converted to the HSV color space, and a multi-scale color mask is designed to address the differences in the shades of the table lines, enhancing the contrast of candidate line segment regions to obtain an enhanced image, including: Convert the plain table background image from RGB to HSV color space; For table lines of varying shades, different HSV threshold ranges are set, and pixel values are remapped in the RGB space to obtain the remapped image. Convert the remapped image to grayscale.
[0007] Furthermore, different HSV threshold ranges are set for table lines of varying shades, and pixel value remapping is performed in the RGB space, including: For dark table lines, set the brightness threshold range to [0, 180], the saturation threshold range to [0, 45], and the hue threshold range to [200, 255]. Areas that meet the HSV threshold range for deep table lines are set to white in RGB space, while areas that do not meet the HSV threshold range for deep table lines retain their original pixel values in RGB space to highlight high contrast. For light table lines, set the brightness threshold range to [0, 180], the saturation threshold range to [0, 30], and the hue threshold range to [200, 240]. Areas that meet the HSV threshold range for shallow table lines are set to white in RGB space, while areas that do not meet the HSV threshold range for shallow table lines are set to black to enhance low contrast.
[0008] Optionally, Canny edge detection is performed on the enhanced image to extract edge features, and multi-scale Hough line detection is used to generate a list of candidate line segments containing both horizontal and vertical directions, including: Configure the detection parameters for Canny edge detection. The detection parameters for Canny edge detection include low threshold, high threshold, and kernel size. Based on the detection parameters of Canny edge detection, Canny edge detection is performed on the enhanced image to extract edge features and obtain a binary image; Based on the binary image, determine the binary image of the dark table line edge and the binary image of the light table line edge; Configure deep table line detection parameters, which include the input binary image of the deep table line edge, the first distance resolution, the first angle resolution, the first accumulator threshold, the first minimum line segment length, and the first maximum line segment gap; Configure shallow table line detection parameters, which include the input shallow table line edge binary image, second distance resolution, second angle resolution, second accumulator threshold, second minimum line segment length, and second maximum line segment gap. Based on the deep and shallow table line detection parameters, the OpenCV HoughLinesP algorithm is called to perform a two-parameter Hough line transform on the binary image, generating a candidate line segment list covering all types of lines.
[0009] Optionally, geometric constraints are used to process all candidate line segments in the candidate line segment list, including: Using geometric constraints, calculate the extreme values of coordinates for all candidate line segments, and generate the top, bottom, left, and right border lines of the table based on the extreme values of coordinates, thus filling in the missing border lines of the table. If the distance between the horizontal line segment and the vertical border line is less than a preset first distance threshold, the horizontal coordinate of the line segment endpoint is corrected; if the distance between the vertical line segment and the horizontal border line is less than a preset second distance threshold, the vertical coordinate of the line segment endpoint is corrected to align the line segment endpoints and obtain a corrected list of line segments. For non-collinear parallel line segments, if two horizontal line segments are parallel and the distance between them is less than or equal to a preset third distance threshold, the ordinates of the two horizontal line segments are uniformly revised to the larger ordinate value of the two horizontal line segments. If two vertical line segments are parallel and the distance between them is less than or equal to a preset fourth distance threshold, the abscissas of the two vertical line segments are uniformly revised to the larger abscissa value of the two vertical line segments, so as to merge adjacent parallel line segments. If the angle between the lines containing the two line segments is less than or equal to a preset angle, or the Euclidean distance between the two line segments is less than or equal to a preset fifth distance threshold, then the two line segments are merged, and the minimum, maximum, minimum, and maximum values of the horizontal coordinates of the merged line segment are taken to form a new endpoint, thus forming a continuous line segment. If there are no other line segments within the area that is a preset sixth distance threshold from the line segment, then the line segment is set as an isolated noise line segment and the isolated noise line segment is removed. For the horizontal and vertical line segments in the corrected line segment list, the Euclidean distance between the two endpoints of the horizontal line segment and the nearest vertical line segment is detected, and the x-coordinate is updated to the x-coordinate of the nearest vertical line segment. Similarly, the Euclidean distance between the two endpoints of the vertical line segment and the nearest horizontal line segment is detected, and the y-coordinate is updated to the y-coordinate of the nearest horizontal line segment to optimize the alignment of the line segment endpoints.
[0010] Optionally, the row and column intersections of the cells are determined based on the list of minimum line segments, and the row and column intersections are sorted according to their coordinates to generate a cell position matrix of the table, including: Generate a list of intersection points based on the intersection types of horizontal and vertical line segments in the minimum line segment list; Sort all intersection points in the intersection point list in ascending order of their ordinates; if the ordinates of intersection points are the same, sort them in ascending order of their abscissas. Take the first coordinate from the sorted intersection list as the top-left endpoint of the table, and determine the top-right and bottom-right endpoints of the table. Verify that the bottom-left endpoint exists in the intersection list. If the verification is successful, save the cell formed by the top-left and bottom-right endpoints, and remove the top-left, top-right, bottom-right, and bottom-left endpoints from the intersection list. Repeat the above steps to remove endpoints until the intersection list is empty, generating a matrix of cell positions for the table.
[0011] Furthermore, based on the intersection types of horizontal and vertical line segments in the minimum line segment list, a corresponding list of intersection points is generated, including: If a horizontal line segment and a vertical line segment intersect only at their endpoints, the intersection type is "L" and one intersection point is generated. If the endpoint of a horizontal line segment intersects the middle of a vertical line segment, or if the endpoint of a vertical line segment intersects the middle of a horizontal line segment, the intersection type is "T", generating two identical intersection points; If both the horizontal and vertical line segments intersect at the middle, the intersection type is "+", generating 4 identical intersection points.
[0012] A second aspect of this disclosure provides a form recognition device, comprising: The recognition unit is used to identify the text location in the table image using the OCR algorithm, generate a text region mask, filter the table image, and obtain a pure table background image without text. The conversion unit is used to convert the pure table background image to the HSV color space. It designs multi-scale color masks for the differences in the shades of the table lines to enhance the contrast of the candidate line segment areas and obtain the enhanced image. The detection unit is used to perform Canny edge detection on the enhanced image, extract edge features, and generate a list of candidate line segments containing both horizontal and vertical directions using multi-scale Hough line detection. The processing unit is used to process all candidate line segments in the candidate line segment list using geometric constraints, reconstruct the table structure, and generate a minimum line segment list that represents the table structure. The parsing unit is used to determine the row and column intersections of cells based on the minimum line segment list, sort the row and column intersections according to their coordinates, generate the cell position matrix of the table, and realize the parsing of the table structure.
[0013] A third aspect of this disclosure provides a computer-readable storage medium storing computer instructions for causing a computer to perform the table recognition method provided in any of the first aspects.
[0014] A fourth aspect of this disclosure provides an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor to cause the at least one processor to perform the table recognition method provided in any of the first aspects.
[0015] In the table recognition method provided in this disclosure, an OCR algorithm is used to identify the text position in the table image, generate a text region mask, and filter the table image to obtain a pure table background image without text; by filtering the table image through the text region mask, the interference of character edges on line detection is avoided, and the purity of subsequent Hough line detection is improved. The plain table background image is converted to the HSV color space. Multi-scale color masks are designed to address the differences in the shades of the table lines, thereby enhancing the contrast of candidate line segment regions and obtaining an enhanced image. Multi-scale color masks are used to highlight the differences in the image of lines of different colors and solid / dim types, providing clearer image features for subsequent edge detection. Canny edge detection is performed on the enhanced image to extract edge features, and multi-scale Hough line detection is used to generate a list of candidate line segments containing horizontal and vertical directions. Geometric constraints are then applied to all candidate line segments in the list to reconstruct the table structure, generating a list of minimum line segments representing the table structure. Based on this minimum line segment list, the row and column intersections of each cell are determined, and these intersections are sorted by coordinates to generate a cell position matrix, thus resolving the table structure. This disclosure does not require deep learning models or labeled data; instead, it achieves table recognition based on Hough line detection and geometric constraints, reducing development costs and allowing for rapid adaptation to new table formats. It also solves the problem in related technologies where table recognition requires extensive training with labeled data and is highly dependent on labeled data. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the specific embodiments or related technologies of this disclosure, the accompanying drawings used in the description of the specific embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is a schematic flowchart of a table recognition method provided in an embodiment of the present disclosure; Figure 2 A visual illustration of the candidate line segment list provided in an embodiment of this disclosure, wherein the right side is a magnified view of a portion of the left side; Figure 3A visual illustration of the minimum line segment list provided in the embodiments of this disclosure; Figure 4 A block diagram of a table recognition device provided in an embodiment of this disclosure; Figure 5 A block diagram of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0018] In order to enable those skilled in the art to better understand the solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the embodiments described are only part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present disclosure.
[0019] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be used interchangeably where appropriate for the embodiments of this disclosure described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0020] It should be noted that, unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other. This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.
[0021] In document digitization, table recognition is a crucial step in achieving data structuring. The table recognition methods used in related technologies primarily rely on deep learning models (such as Faster R-CNN, YOLO, and Transformer), but these methods require extensive training with labeled data and are highly dependent on data annotation.
[0022] Hough line detection, as a traditional computer vision method, does not require data training, but it has shortcomings such as sensitivity to text interference and weak ability to detect multiple types of lines (e.g., it cannot effectively distinguish between solid lines, dashed lines, and lines of different colors).
[0023] To address the aforementioned problems, this disclosure provides a table recognition method that can be applied to the automatic parsing of table parameters in scenarios such as electronic document processing, official documents, and financial reports. It solves the problem of requiring a large amount of labeled data for training and its strong dependence on data labeling in related technologies. Furthermore, it also addresses the poor adaptability to complex tables (such as mixed solid / dashed lines, borderless tables, and merged cells), achieving efficient and accurate table recognition without the need for deep learning models. Figure 1 As shown, the method includes the following steps S11 to S15: Step S11: Use the OCR algorithm to identify the text position in the table image, generate a text region mask, filter the table image, and obtain a pure table background image without text; text region mask preprocessing: use the OCR algorithm to identify the text position in the table image, generate a mask matrix to filter the text region, obtain a pure table background image without text, filter the table image with the text region mask to avoid interference from character edges on line detection, and improve the purity of subsequent Hough line detection.
[0024] Specifically, the table image is an image containing a table. This embodiment of the disclosure does not limit the size of the table. The specific type of the table image can be an image with solid table lines, an image with dashed table lines, a borderless table image, or a borderless image with a mix of solid and dashed lines. For borderless images with mixed real and virtual elements, text region masking preprocessing is performed. The text positions in a table image are identified using an OCR algorithm, generating a text region mask matrix. Pixel-by-pixel filtering of the table image yields a pure table background image without text. Specifically, the OCR algorithm detects text bounding boxes, creating a mask of the same size as the image (filling text areas with white and non-text areas with black). Bitwise operations are used to remove text pixels, avoiding interference from character edges in subsequent line detection. For example, for an image of a document containing text, after masking, only the mesh structure formed by table lines is retained, significantly improving the purity of Hough line detection.
[0025] Step S12: Convert the pure table background image to the HSV color space, design a multi-scale color mask for the differences in the shades of the table lines, enhance the contrast of the candidate line segment areas, and obtain the enhanced image; the multi-scale color mask can be a dual-scale color mask. For example, for lines of different shades, a dual-scale color mask can be designed, where dark lines highlight high contrast and light lines enhance low contrast, thereby enhancing the contrast of the candidate line segment areas and achieving differentiated highlighting of different types of lines.
[0026] In one optional embodiment of this disclosure, step S12 includes: Convert the plain table background image from RGB to HSV color space; For table lines of varying shades, different HSV threshold ranges are set, and pixel values are remapped in the RGB space to obtain the remapped image. The HSV threshold ranges specifically include luminance (V) threshold range, saturation (S) threshold range, and hue (H) threshold range. A binary mask is generated using the HSV threshold ranges (white indicates that the HSV threshold range is met, and black indicates that the HSV threshold range is not met). Pixel values are then remapped in the original RGB image based on the binary mask. Convert the remapped image to grayscale.
[0027] In one optional embodiment of this disclosure, different HSV threshold ranges are set for table lines of varying shades, and pixel value remapping is performed in the RGB space, including: For dark table lines, set the brightness threshold range to [0, 180], the saturation threshold range to [0, 45], and the hue threshold range to [200, 255]. Regions that meet the HSV threshold range for dark table lines are set to white [255, 255, 255] in RGB space, while regions that do not meet the HSV threshold range for dark table lines retain their original pixel values in RGB space to highlight high contrast; for regions that do not meet the HSV threshold range for dark table lines, the original RGB pixel values of the image are retained to maintain the visual salience of dark lines. For light table lines, set the brightness threshold range to [0, 180], the saturation threshold range to [0, 30], and the hue threshold range to [200, 240]. Light table lines can be dashed lines, light gray, or light blue line segments. The regions that meet the HSV threshold range for light table lines are set to white [255, 255, 255] in RGB space, while the regions that do not meet the HSV threshold range for light table lines are set to black to enhance low contrast. The contrast of light lines is enhanced by using a black background.
[0028] In this embodiment of the disclosure, color depth is separated by HSV threshold range and combined with RGB pixel remapping to achieve targeted enhancement of different types of lines. Dark lines are highlighted with high contrast by retaining the original value, while light lines are enhanced with low contrast by using a black and white background. For example, light gray lines are converted into white lines and black solid lines retain the original value. By highlighting the differences in different colors and solid / dark types of lines in the image, clearer image features are provided for subsequent edge detection, thereby improving the accuracy of edge detection.
[0029] Step S13: Perform Canny edge detection on the enhanced image to extract edge features, and use multi-scale Hough line detection to generate a candidate line segment list containing horizontal and vertical directions; combine Canny edge detection and two-parameter Hough line transform to generate a candidate line segment list covering all types of lines, achieving accurate capture of all types of lines.
[0030] In one optional embodiment of this disclosure, step S13 includes: Configure the detection parameters for Canny edge detection. The detection parameters for Canny edge detection include low threshold, high threshold, and kernel size. Specifically, in the detection parameters of Canny edge detection, the low threshold can be 50: this controls the detection sensitivity of weak edges, ensuring that edge points of light table lines (such as light gray dashed lines) are effectively captured, reducing the false negative rate; The high threshold can be 100: to filter strong edge points, reduce background noise interference, and at the same time retain weak edges connected to strong edges through a dual threshold connection mechanism to ensure the continuity of discontinuous lines such as dashed lines. The kernel size can be 3: This specifies the convolution kernel size of the Sobel operator, controls the neighborhood range of gradient calculation, balances edge localization accuracy and noise resistance, and has good performance for fine edges of table lines.
[0031] Based on the detection parameters of Canny edge detection, Canny edge detection is performed on the enhanced image to extract edge features and obtain a binary image; after Canny edge detection, the grayscale image is converted into a binary image (edges).
[0032] Based on the binary image, determine the binary image of the dark table line edge and the binary image of the light table line edge; Configure deep table line detection parameters, which include the input binary image of the deep table line edge, the first distance resolution, the first angle resolution, the first accumulator threshold, the first minimum line segment length, and the first maximum line segment gap; Specifically, the parameters for deep table line detection are configured as follows: edges: Binary image of the edges of the deep table lines, data after filtering of the deep table lines, and binary image after Canny edge detection; rho=1: Distance resolution, in pixels. Setting it to 1 pixel controls the distance accuracy of line segment detection, enabling precise positioning of line segments. theta=π / 180: Angular resolution in radians, set to 1 degree precision to capture horizontal / vertical line segments; threshold=30: Accumulator threshold, only retains line segments with ≥30 votes, filtering out random noise; minLineLength=35: Minimum line segment length. Only line segments with a length of ≥35 pixels are retained to eliminate interference from short lines and adapt to solid lines in dark tables. maxLineGap=15: Maximum line gap, allowing connection of discontinuous line segments (such as dark dashed lines) with a distance ≤15 pixels.
[0033] Configure shallow table line detection parameters, which include the input shallow table line edge binary image, second distance resolution, second angle resolution, second accumulator threshold, second minimum line segment length, and second maximum line segment gap. Specifically, an enhanced parameter combination is used to improve the detection capability of light-colored lines. The light-colored line detection parameter configuration is as follows: edges: Binary image of the edges of the shallow table lines, data after filtering of the shallow table lines, and binarized image data after Canny edge detection; rho=1: Maintain the same distance resolution as the dark table lines; theta=π / 180: Maintain the same angular resolution as the dark table lines; threshold=40: Accumulator threshold, only retains line segments with ≥40 votes. Compared with dark table lines, the voting threshold is increased to 40 to filter light-colored background noise. minLineLength=45: Minimum line segment length. Only line segments with a length of ≥45 pixels are retained. Compared with dark table lines, the minimum line segment length is increased to 45 pixels to eliminate interference from short lines. maxLineGap=15: Maintain the same maximum gap threshold as the deep table lines.
[0034] Based on the detection parameters for deep and shallow table lines, the OpenCV HoughLinesP algorithm is invoked to perform a two-parameter Hough line transform on the binary image, generating a candidate line segment list covering all line types. For deep table lines, long segment detection parameters are used to capture thick solid lines. Compared to deep table lines, shallow table lines have improved accumulator thresholds and minimum segment lengths to filter noise, while maintaining the detection parameters rho, theta, and maxLineGap consistent with those for deep table lines to ensure accuracy in angle and gap connections. The combined results generate a candidate line segment list covering all line types.
[0035] After merging the two types of test results, a result such as Figure 2 The list of candidate line segments shown on the left includes both horizontal and vertical lines, covering candidate line segments of different line widths and solid / dashed types. Figure 2 The right side of the image shows a magnified view of the candidate line segment list, illustrating redundant and isolated line segments.
[0036] This disclosure, through the aforementioned differentiated parameter configuration, achieves accurate detection of table lines of different widths, solid / dashed, and light / dark types, ultimately obtaining a complete list of candidate line segments, providing high-quality input for subsequent line segment correction. This disclosure supports the recognition of complex tables such as solid lines, dashed lines, and borderless tables. On datasets that simultaneously contain dashed lines and borderless cells, it can improve recognition accuracy and adaptability, achieving efficient and accurate table recognition without the need for deep learning models.
[0037] Step S14: Process all candidate line segments in the candidate line segment list using geometric constraints to reconstruct the table structure and generate a minimum line segment list representing the table structure; process all candidate line segments using geometric constraints, including calculating coordinate extreme values to fill missing borders, correcting line segment endpoint alignment based on distance thresholds, merging adjacent parallel line segments, and removing isolated noise line segments to generate a minimum line segment list containing only the core structure, thereby realizing line segment correction and table structure reconstruction.
[0038] In an optional embodiment of this disclosure, step S14, which processes all candidate line segments in the candidate line segment list using geometric constraints, includes: Using geometric constraints, the extreme values of coordinates for all candidate line segments are calculated. Based on these extreme values, the top, bottom, left, and right border lines of the table are generated to fill in the missing border lines. A coordinate system is established with the top left corner of the table as the origin, the positive x-axis pointing to the right, and the positive y-axis pointing downwards. The extreme values of the horizontal and vertical coordinates of all candidate line segments are calculated to generate virtual top, bottom, left, and right border lines to fill in the missing lines on the four sides of the table, thus solving the problem of most tables omitting borders by nature. If the distance between the horizontal line segment and the vertical border line is less than a preset first distance threshold, the x-coordinate of the line segment endpoints is corrected; if the distance between the vertical line segment and the horizontal border line is less than a preset second distance threshold, the y-coordinate of the line segment endpoints is corrected to align the line segment endpoints and obtain a corrected list of line segments. In the initial correction, the first and second distance thresholds can be the same or different. For example, both the first and second distance thresholds can be set to 10 pixels. If the distance between the horizontal line segment and the vertical border line is less than 10 pixels, the x-coordinate of the horizontal line segment is forcibly corrected to the x-coordinate of the border. Similarly, the y-coordinate of the vertical line segment is corrected. If the distance between the vertical line segment and the horizontal border line is less than 10 pixels, the y-coordinate of the vertical line segment is forcibly corrected to the y-coordinate of the border to ensure that the line segment endpoints are aligned.
[0039] For non-collinear parallel line segments, if two horizontal line segments are parallel and the distance between them is less than or equal to a preset third distance threshold, the y-coordinates of the two horizontal line segments are uniformly revised to the larger y-coordinate value of the two horizontal line segments. If two vertical line segments are parallel and the distance between them is less than or equal to a preset fourth distance threshold, the x-coordinates of the two vertical line segments are uniformly revised to the larger x-coordinate value of the two vertical line segments to merge adjacent parallel line segments. The third and fourth distance thresholds can be the same or different. For example, both the third and fourth distance thresholds can be set to 5 pixels. If two line segments are parallel and the distance between them is ≤ 5 pixels, then: for horizontal line segments, the y-coordinates of the two line segments are uniformly revised to the larger y-value; for vertical line segments, the x-coordinates of the two line segments are uniformly revised to the larger x-value. This can eliminate redundancy of closely spaced parallel line segments caused by detection errors and line segment width.
[0040] If the angle between the lines containing the two line segments is less than or equal to a preset angle, or the Euclidean distance between the two line segments is less than or equal to a preset fifth distance threshold, then the two line segments are merged. The minimum, maximum, minimum, and maximum values of the horizontal coordinate of the merged line segment are taken as the new endpoints to form a continuous line segment. For example, the preset angle can be set to 5° and the fifth distance threshold can be set to 10 pixels. If the two line segments meet the following conditions, they are merged into one line segment: The angle between the lines containing the line segment is ≤5° (considered collinear); The Euclidean distance between the endpoints of a line segment is ≤10 pixels (for end-to-end connections, overlaps, or endpoint spacing ≤10 pixels). Take the minimum and maximum x-coordinates of the merged horizontal line segment, or the minimum and maximum y-coordinates of the vertical line segment, as the new endpoints to form continuous line segments, such as merging short line segments split by dashed lines.
[0041] If no other line segment exists within the area of the line segment at a preset sixth distance threshold, the line segment is classified as an isolated noise line segment and discarded. The sixth distance threshold can be set to 10 pixels. Taking a horizontal line segment as an example, the endpoint coordinates of the horizontal line segment are (A1.x, y) and (A2.x, y), and a rectangular area with a detection interval of {x∈[A1.x-10,A2.x+10],y∈[y-10, y+10]} is constructed. If no other line segment exists within this rectangular area, the current horizontal line segment is determined to be an isolated noise line segment and discarded. Similarly, for a vertical line segment, the detection interval {x∈[x-10, x+10],y∈[B1.y-10, B2.y+10]} is constructed with the endpoint coordinates (x, B1.y) and (x, B2.y). If no other line segment exists within the detection interval, the current vertical line segment is determined to be an isolated noise line segment and discarded.
[0042] For the horizontal and vertical line segments in the corrected line segment list, the Euclidean distance between the two endpoints of the horizontal line segment and the nearest vertical line segment is detected, and the x-coordinate is updated to the x-coordinate of the nearest vertical line segment. Similarly, the Euclidean distance between the two endpoints of the vertical line segment and the nearest horizontal line segment is detected, and the y-coordinate is updated to the y-coordinate of the nearest horizontal line segment to optimize the alignment of the line segment endpoints. A second coordinate correction is then performed on the initially corrected line segment list to further optimize endpoint alignment. Specifically, for horizontal line segments, the Euclidean distance between the two endpoints and the nearest vertical line segment is detected, and the x-coordinate is updated to the x-coordinate of the nearest vertical line segment. For vertical line segments, the distance between the two endpoints and the nearest horizontal line segment is detected, and the y-coordinate is updated to the y-coordinate of the nearest horizontal line segment. By eliminating the misalignment of the endpoints of horizontal and vertical line segments, the integrity of the table's line topology is ensured, generating a minimal line segment list containing only the core structure. Figure 3 This is a visualization of the minimum list of line segments obtained after the candidate line segment list has been corrected and the table structure reconstructed.
[0043] Step S15: Determine the row and column intersections of the cells based on the minimum line segment list, and sort the intersections by coordinates to generate a cell position matrix for the table, thus parsing the table structure. This involves analyzing geometric intersections based on the minimum line segment list, constructing cells by coordinate sorting, identifying different intersection types, and generating a cell position matrix through coordinate verification and iterative processing, thereby parsing the table structure.
[0044] In one optional embodiment of this disclosure, step S15 includes: Generate a list of intersection points based on the intersection types of horizontal and vertical line segments in the minimum line segment list; Sort all intersection points in the intersection point list in ascending order of their ordinates; if the ordinates of intersection points are the same, sort them in ascending order of their abscissas. Take the first coordinate from the sorted intersection list as the top-left endpoint of the table, and determine the top-right and bottom-right endpoints of the table. Verify that the bottom-left endpoint exists in the intersection list. If the verification is successful, save the cell formed by the top-left and bottom-right endpoints, and remove the top-left, top-right, bottom-right, and bottom-left endpoints from the intersection list. Specifically, the steps to create a cell are as follows: ① Take the first coordinate from the sorted list as the top-left endpoint A1[A1.x, A1.y], and determine the top-right endpoint A2[A2.x, A2.y] that satisfies the following conditions: A2.x > A1.x and A1.y = A2.y; ②Determine the lower right endpoint A3[A3.x, A3.y] based on the upper right endpoint A2, with the following requirements: A3.y > A1.y and A3.x = A2.x; ③ Verify that the lower left endpoint A4[A1.x, A3.y] exists in the intersection list, and that line segment A3A4 is a sub-segment of the smallest horizontal line segment, which satisfies: If a horizontal line segment S exists with endpoint coordinates (x1, y) and (x2, y), and x1≤A4.x≤A3.x≤x2, then the verification passes, the resulting cell {A1, A3} is saved, and {A1, A2, A3, A4} is removed from the intersection list; if a horizontal line segment S does not exist, then step ② is repeated.
[0045] Repeat the endpoint removal steps described above until the intersection list is empty, generating the table's cell position matrix. That is, iteratively repeat the cell formation steps described above until the intersection list is empty, ultimately generating the table's cell position matrix.
[0046] In one optional embodiment of this disclosure, a corresponding list of intersection points is generated based on the intersection types of horizontal and vertical line segments in the minimum line segment list, including: If a horizontal line segment and a vertical line segment intersect only at their endpoints, the intersection type is "L" and one intersection point is generated. If the endpoint of a horizontal line segment intersects the middle of a vertical line segment, or if the endpoint of a vertical line segment intersects the middle of a horizontal line segment, the intersection type is "T", generating two identical intersection points; If both the horizontal and vertical line segments intersect at the middle, the intersection type is "+", generating 4 identical intersection points.
[0047] This disclosure is based on Hough line detection and geometric constraints, which can quickly adapt to new table formats and solve the problem that related technologies require a large amount of labeled data for training when recognizing tables, and are highly dependent on data labeling. Moreover, it does not require high-performance hardware such as GPUs or dedicated acceleration devices, and can run smoothly in a normal CPU environment. Compared with the deep learning methods used in related technologies, this disclosure can reduce hardware costs and energy consumption, and can be applied to resource-constrained scenarios such as embedded devices and mobile terminals. At the same time, it avoids graphics card driver compatibility issues and significantly lowers the deployment threshold. This disclosure can be applied to automated batch document processing, shortening the time required for table recognition in images. Furthermore, it does not require labeled data or retraining and can be directly applied to scenarios such as electronic documents, e-government reports, and financial statements.
[0048] As can be seen from the above description, this disclosure achieves the following technical effects: This disclosure does not require deep learning models and labeled data. Instead, it achieves table recognition based on Hough line detection and geometric constraints, which reduces development costs and can quickly adapt to new table formats. It solves the problem that related technologies require a large amount of labeled data for training when recognizing tables and are highly dependent on data labeling. By filtering the table image using a text region mask, interference from character edges on line detection is avoided, thus improving the purity of Hough line detection. By separating color depth using the HSV threshold range and combining it with RGB pixel remapping, targeted enhancement of different types of lines is achieved. Dark lines are highlighted with high contrast by retaining their original values, while light lines are enhanced with low contrast by using a black and white background. By highlighting the differences in different colors and types of lines in the image, clearer image features are provided for edge detection, thus improving the accuracy of edge detection.
[0049] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.
[0050] This disclosure also provides a table recognition device for implementing the above method embodiments, such as... Figure 4 As shown, the form recognition device 20 includes: The recognition unit 21 is used to recognize the text position in the table image using the OCR algorithm, generate a text region mask, filter the table image, and obtain a pure table background image without text. The conversion unit 22 is used to convert the pure table background image to the HSV color space, design multi-scale color masks for the differences in the shades of the table lines, enhance the contrast of the candidate line segment areas, and obtain the enhanced image. The detection unit 23 is used to perform Canny edge detection on the enhanced image, extract edge features, and generate a list of candidate line segments containing horizontal and vertical directions using multi-scale Hough line detection. Processing unit 24 is used to process all candidate line segments in the candidate line segment list using geometric constraints, reconstruct the table structure, and generate a minimum line segment list representing the table structure. Parsing unit 25 is used to determine the row and column intersections of cells based on the minimum line segment list, sort the row and column intersections according to coordinates, generate the cell position matrix of the table, and realize the parsing of the table structure.
[0051] The specific methods of execution of each unit in the above device embodiments have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0052] The present disclosure also provides an electronic device, such as Figure 5 As shown, the electronic device includes one or more processors 31 and a memory 32. Figure 5 Take a processor 31 as an example.
[0053] The controller may also include an input device 33 and an output device 34.
[0054] The processor 31, memory 32, input device 33, and output device 34 can be connected via a bus or other means. Figure 5 Taking the example of a connection between China and Israel via a bus.
[0055] Processor 31 can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or combinations of the above types of chips. The general-purpose processor can be a microprocessor or any conventional processor.
[0056] The memory 32, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as the program instructions / modules corresponding to the control method in this embodiment. The processor 31 executes various functional applications and data processing of the server by running the non-transitory software programs, instructions, and modules stored in the memory 32, thereby implementing the table recognition method of the above-described method embodiment.
[0057] The memory 32 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created by the use of the processing device operated by the server. Furthermore, the memory 32 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, the memory 32 may optionally include memory remotely located relative to the processor 31, and these remote memories can be connected to a network connection device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0058] Input device 33 can receive input numerical or character information, and generate key signal inputs related to user settings and function control of the server's processing device. Output device 34 may include display devices such as a display screen.
[0059] One or more modules are stored in memory 32, and when executed by one or more processors 31, they perform actions such as... Figure 1 The method shown.
[0060] Those skilled in the art will understand that all or part of the processes in the above method embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), random access memory (RAM), flash memory (FM), hard disk drive (HDD), or solid-state drive (SSD), etc.; the storage medium can also include combinations of the above types of memory.
[0061] Although embodiments of the present disclosure have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present disclosure, and such modifications and variations all fall within the scope defined by the appended claims.
Claims
1. A table recognition method, characterized in that, include: The OCR algorithm is used to identify the text location in the table image, generate a text region mask, filter the table image, and obtain a pure table background image without text. The pure table background image is converted to the HSV color space, and a multi-scale color mask is designed to address the differences in the shades of the table lines, thereby enhancing the contrast of the candidate line segment regions and obtaining an enhanced image. The enhanced image is subjected to Canny edge detection to extract edge features, and multi-scale Hough line detection is used to generate a list of candidate line segments containing both horizontal and vertical directions. Geometric constraints are used to process all candidate line segments in the candidate line segment list to reconstruct the table structure and generate a minimum line segment list that represents the table structure. Based on the list of minimum line segments, the row and column intersections of the cells are determined, and the intersections are sorted according to their coordinates to generate a cell position matrix of the table, thereby realizing the parsing of the table structure.
2. The method according to claim 1, characterized in that, The process of converting the plain table background image to the HSV color space, designing multi-scale color masks to address the differences in line depth, and enhancing the contrast of candidate line segment regions to obtain an enhanced image includes: Convert the pure table background image from RGB to HSV color space; For table lines of varying shades, different HSV threshold ranges are set, and pixel values are remapped in the RGB space to obtain the remapped image. The remapped image is converted into a grayscale image.
3. The method according to claim 2, characterized in that, The method of setting different HSV threshold ranges for table lines of varying shades and remapping pixel values in the RGB space includes: For dark table lines, set the brightness threshold range to [0, 180], the saturation threshold range to [0, 45], and the hue threshold range to [200, 255]. Areas that meet the HSV threshold range for deep table lines are set to white in RGB space, while areas that do not meet the HSV threshold range for deep table lines retain their original pixel values in RGB space to highlight high contrast. For light table lines, set the brightness threshold range to [0, 180], the saturation threshold range to [0, 30], and the hue threshold range to [200, 240]. Areas that meet the HSV threshold range for shallow table lines are set to white in RGB space, while areas that do not meet the HSV threshold range for shallow table lines are set to black to enhance low contrast.
4. The method according to claim 1, characterized in that, The enhanced image is subjected to Canny edge detection to extract edge features, and multi-scale Hough line detection is used to generate a candidate line segment list containing both horizontal and vertical directions, including: Configure the detection parameters of the Canny edge detection, which include low threshold, high threshold, and kernel size; Based on the detection parameters of the Canny edge detection, Canny edge detection is performed on the enhanced image to extract edge features and obtain a binary image; Based on the binary image, determine the binary image of the dark table line edge and the binary image of the light table line edge; Configure deep table line detection parameters, which include the input binary image of the deep table line edge, a first distance resolution, a first angle resolution, a first accumulator threshold, a first minimum line segment length, and a first maximum line segment gap; Configure shallow table line detection parameters, which include the input shallow table line edge binary image, second distance resolution, second angle resolution, second accumulator threshold, second minimum line segment length and second maximum line segment gap; Based on the deep table line detection parameters and the shallow table line detection parameters, the OpenCV HoughLinesP algorithm is called to perform a two-parameter Hough line transform on the binary image, generating a candidate line segment list covering all types of lines.
5. The method according to claim 1, characterized in that, The process of using geometric constraints to process all candidate line segments in the candidate line segment list includes: Using the geometric constraints, calculate the extreme values of the coordinates of all candidate line segments, and generate the top, bottom, left, and right border lines of the table based on the extreme values of the coordinates to fill in the missing border lines of the table. If the distance between the horizontal line segment and the vertical border line is less than a preset first distance threshold, the horizontal coordinate of the line segment endpoint is corrected; if the distance between the vertical line segment and the horizontal border line is less than a preset second distance threshold, the vertical coordinate of the line segment endpoint is corrected to align the line segment endpoints and obtain a corrected list of line segments. For non-collinear parallel line segments, if two horizontal line segments are parallel and the distance between them is less than or equal to a preset third distance threshold, the ordinates of the two horizontal line segments are uniformly revised to the larger ordinate value of the two horizontal line segments. If two vertical line segments are parallel and the distance between them is less than or equal to a preset fourth distance threshold, the abscissas of the two vertical line segments are uniformly revised to the larger abscissa value of the two vertical line segments, so as to merge adjacent parallel line segments. If the angle between the lines containing the two line segments is less than or equal to a preset angle, or the Euclidean distance between the two line segments is less than or equal to a preset fifth distance threshold, then the two line segments are merged, and the minimum, maximum, minimum, and maximum values of the horizontal coordinates of the merged line segment are taken to form a new endpoint, thus forming a continuous line segment. If there are no other line segments within the area that is a preset sixth distance threshold from the line segment, then the line segment is set as an isolated noise line segment and the isolated noise line segment is removed. For the horizontal and vertical line segments in the corrected line segment list, the Euclidean distance between the two endpoints of the horizontal line segment and the nearest vertical line segment is detected, and the horizontal coordinate is updated to the horizontal coordinate of the nearest vertical line segment. The Euclidean distance between the two endpoints of the vertical line segment and the nearest horizontal line segment is detected, and the vertical coordinate is updated to the vertical coordinate of the nearest horizontal line segment to optimize the alignment of the line segment endpoints.
6. The method according to claim 1, characterized in that, The step of determining the row and column intersections of cells based on the list of minimum line segments, and sorting the row and column intersections according to their coordinates to generate a cell position matrix of the table includes: Based on the intersection type of horizontal and vertical line segments in the minimum line segment list, generate a corresponding list of intersection points; Sort all intersection points in the intersection point list in ascending order of their ordinates; if the ordinates of intersection points are the same, sort them in ascending order of their abscissas. Take the first coordinate from the sorted intersection list as the top left corner of the table, and determine the top right corner and bottom right corner of the table. Verify that the bottom left corner exists in the intersection list. If the verification is successful, save the cell formed by the top left corner and bottom right corner, and remove the top left corner, top right corner, bottom right corner and bottom left corner from the intersection list. Repeat the above steps to remove endpoints until the intersection list is empty, generating a matrix of cell positions for the table.
7. The method according to claim 6, characterized in that, The step of generating a corresponding intersection point list based on the intersection types of horizontal and vertical line segments in the minimum line segment list includes: If a horizontal line segment and a vertical line segment intersect only at their endpoints, the intersection type is "L" and one intersection point is generated. If the endpoint of a horizontal line segment intersects the middle of a vertical line segment, or if the endpoint of a vertical line segment intersects the middle of a horizontal line segment, the intersection type is "T", generating two identical intersection points. If both the horizontal and vertical line segments intersect at the middle, the intersection type is "+", generating 4 identical intersection points.
8. A form recognition device, characterized in that, include: The recognition unit is used to recognize the text location in the table image using an OCR algorithm, generate a text region mask, filter the table image, and obtain a pure table background image without text. The conversion unit is used to convert the pure table background image to the HSV color space, design multi-scale color masks for the differences in the shades of the table lines, enhance the contrast of the candidate line segment areas, and obtain an enhanced image. The detection unit is used to perform Canny edge detection on the enhanced image, extract edge features, and generate a list of candidate line segments containing horizontal and vertical directions using multi-scale Hough line detection. The processing unit is used to process all candidate line segments in the candidate line segment list using geometric constraints, reconstruct the table structure, and generate a minimum line segment list representing the table structure. The parsing unit is used to determine the row and column intersections of the cells based on the list of minimum line segments, sort the row and column intersections according to their coordinates, generate a cell position matrix of the table, and realize the parsing of the table structure.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing a computer to perform the table recognition method according to any one of claims 1 to 7.
10. An electronic device, characterized in that, The electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor to cause the at least one processor to perform the table recognition method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Method and system for detecting form frame lines in form documents
CN110210409A
Table data extraction method, device and equipment and computer storage medium
CN112418180A
Table identification method and device, storage medium and electronic equipment
CN114724154A
Outer wall visual inspection method and system based on deep learning
CN120495942A