Detection Method of Image Tables Based on MSER under Natural Photographing Conditions
Through Retinex image enhancement and gamma transformation processing, combined with morphology and connectivity domain analysis, table detection and perspective transformation correction are used using the MSER algorithm, which solves the problem of difficult table image recognition by mobile terminals under natural conditions, and realizes high robust automatic table detection and recognition.
Patent Information
- Application Number
- CN202111408000.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-25
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2041-11-25
AI Technical Summary
The table images taken by mobile terminals under natural conditions have problems such as uneven lighting, shadows, low contrast and distortion, which makes it difficult to identify tables.
The Retinex image enhancement algorithm is used to eliminate the influence of lighting imbalance, and the contrast enhancement is performed in combination with image gamma transformation. The text area is extracted using morphology and connectivity domain analysis method. The table area is detected through the MSER algorithm, and perspective transformation correction is performed to determine the row and column coordinates of the table cells.
It improves the detection and recognition of table images taken under natural conditions, and realizes automatic detection and recognition under various conditions, which has good practicality.
Smart Images

Figure CN114120303B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing and optical character recognition (OCR), and particularly to a method for detecting image tables based on MSER under natural photographing conditions. Background Art
[0002] As a structured way of organizing information, tables are highly refined, concise, and standardized, and are often used for data record statistics, experimental result analysis, etc. At present, many table documents are provided in the form of pictures, and restoring the table information in the picture form into digital data is the basis for further processing and data analysis.
[0003] The devices for obtaining table images mainly include scanners, dedicated devices such as high-speed document cameras, and handheld mobile terminal devices. Due to the advantages of popularity and convenience of mobile terminal devices, people are more inclined to use mobile terminals to obtain table images and recognize table contents at any time. Affected by uneven illumination, photographing angles, etc., the table images taken by mobile terminal devices under natural conditions have problems such as shadows, low contrast, and distortion of table images, making the recognition of tables in images taken under natural conditions more difficult than that of tables obtained by dedicated devices.
[0004] Among them, the Retinex image enhancement algorithm is based on the "retina-cortex theory". The basic principle is to estimate the illumination component from the original image according to the illumination-reflection model of the image, and then try to eliminate (or reduce) the illumination component to obtain the reflection component of the object, so as to obtain the true appearance of the object. The illumination-reflection model represents the natural scene image f(x, y) as the product of the light source illumination field function i(x, y) and the object reflection field function r(x, y) in the scene:
[0005] f(x, y) = i(x, y) · r(x, y);
[0006] Using the center surround Retinex algorithm, the estimated value i'(x, y) of the illumination component i(x, y) can be expressed as
[0007] i′(x, y) = F(x, y) * f(x, y);
[0008] where "*" is the convolution operation, and F(x, y) is the center surround function
[0009]
[0010] where K is the normalization factor.
[0011] Substitute the estimated illumination component value \(i'(x, y)\) for \(i(x, y)\) into \(f(x, y)=i(x, y)\cdot r(x, y)\), and the Retinex enhanced output image after eliminating the influence of the illumination component can be obtained.
[0012] \(R(x,y)=\ln[f(x,y)]-\ln[F(x,y)*f(x,y)]\);
[0013] The Maximally Stable Extrema Region (MSER) algorithm was proposed by Matas et al. in 2004. The MSER algorithm first binarizes the image. The binarization threshold takes 256 different thresholds within the interval \([0, 255]\), and the binarized image undergoes a process from all white to all black. Let \(R(i)\) represent a connected region in the binarized image when the threshold is \(i\). When the threshold changes from \(i\) to \(i + \Delta\), the connected region becomes \(R(i+\Delta)\).
[0014]
[0015] Where \(|R(i)|\) represents the area of region \(R(i)\), and \(|R(i + \triangle)-R(i)|\) represents the area change of the connected region when \(i\) changes to \(i+\triangle\). When \(q(i)\) is a local minimum, it means that the area change of the connected region in the binarized image is the smallest within a certain gray threshold change range, and at this time \(R(i)\) is called the maximally stable extrema region.
[0016] The MSER algorithm can well adapt to various complex illumination environments and is a classic algorithm commonly used for extracting stable feature regions in natural images, such as character detection, license plate detection, and traffic sign detection, etc. Summary of the Invention
[0017] The technical problem to be solved by the present invention is to provide a method for detecting tables in images under natural photographing conditions based on the MSER algorithm, aiming to solve the problems of unsatisfactory preprocessing effect of images taken by mobile terminals such as mobile phones and the robustness of table detection and recognition.
[0018] To solve the above problems, the technical solution adopted by the present invention is: the method for detecting an image table based on MSER under natural photographing conditions includes the following steps:
[0019] S1: Convert the collected image into a grayscale image, and use the Retinex image enhancement algorithm to eliminate the influence of uneven illumination in the image to obtain the grayscale image \(R(x, y)\);
[0020] S2: Use image gamma transformation to perform contrast enhancement processing on the grayscale image \(R(x, y)\) obtained in step S1 to obtain the enhanced image \(H(x, y)\);
[0021] S3: Extract the text region in the grayscale image H(x, y) using image morphological processing methods and connected component analysis, and calculate the average height of the characters in the image;
[0022] S4: Detect the overall table region in the grayscale image H(x, y) using the Maximally Stable Extremal Regions (MSER) algorithm, and determine the minimum detection area of the MSER algorithm according to the average height of the characters in the image H(x, y) obtained in the step S3, so as to obtain the contour of the overall table region;
[0023] S5: Use the minimum circumscribed rectangle and the circumscribed rectangle of the contour of the overall table region in the step S4 to calculate the perspective transformation matrix, perform a correction operation on the image, and calculate the coordinate values of each corrected table region in the image;
[0024] S6: Detect each table cell after image distortion correction using the Maximally Stable Extremal Regions (MSER) algorithm, merge the rows and columns of each table cell, and determine the row coordinates and column coordinates of the table cells.
[0025] Adopt the above technical solution, use the Retinex image enhancement algorithm based on the retinocortical theory and image gamma transformation to perform image enhancement on the image taken under natural photographing conditions; use morphological and connected component analysis methods to extract the text region in the image to obtain the average height parameter of the characters in the image; use the region detection MSER method to detect the overall table region in the image, and the minimum detection area of the Maximally Stable Extremal Regions (MSER) algorithm is determined according to the average character height. Use the minimum circumscribed rectangle and the circumscribed rectangle of the contour of the overall table region to calculate the perspective transformation matrix, and perform a distortion correction operation on the image. Detect each table cell after correction, merge the rows and columns of each table cell, and determine the row coordinates and column coordinates of the table cells; this method has good robustness to images taken under various conditions, can be used for automatic detection and recognition of tables in images, and has good practicability.
[0026] As a preferred technical solution of the present invention, the specific steps of performing image enhancement processing on the grayscale image R(x, y) using image gamma transformation in the step S2 include:
[0027] S21: Perform histogram statistical analysis on the grayscale image R(x, y), and record the gray level corresponding to the histogram peak as Maxval, and record the minimum gray level of the grayscale image R(x, y) as Minval;
[0028] S22: Perform gray level transformation on the pixel r(x, y) in the grayscale image R(x, y) to h(x, y) to obtain the enhanced image H(x, y), where the transformation formula is:
[0029]
[0030] As a preferred technical solution of the present invention, in step S3, a morphological method and a connected component analysis method are used to extract the text region in the image, and the average height value of the characters in the image is calculated. The specific steps are as follows:
[0031] S31: Obtain the binarization threshold of the enhanced grayscale image H(x, y) using the Otsu algorithm. Convert the grayscale image H(x, y) into a binary image by obtaining the binarization threshold, and invert the binary image, that is, perform amplitude inversion on the binary image to obtain a binary image B(x, y) with a background of 0 and tables and characters of 1. Then perform a closing operation on the binary image B(x, y). The size of the closing operation kernel can be selected as 5×5 or 7×7;
[0032] S32: Use an image connected component extraction method to extract the connected regions in the binary image, and calculate the width w and height h of each connected region. Set thresholds according to the scale size (width W, height H) of the text in the image relative to the image. Set the thresholds Tw = W / 10 and Th = H / 20; Filter out the connected regions where w > Tw or h > Th to obtain the heights of the connected regions that meet the conditions;
[0033] S33: Calculate the average value of the heights of all connected regions that meet the conditions, and use the average height value as the average height value H_char of the characters in the image.
[0034] As a preferred technical solution of the present invention, in step S4, the MSER method for region detection is used to detect the overall table region in the image. The specific steps are as follows:
[0035] S41: Set each parameter for the initialization of the MSER object, where:
[0036] _min_area is the minimum table area detected by the MSER algorithm. This value is determined according to the average character height value H_char in the image obtained in step S3. Set _min_area = k1×H_char×H_char, indicating that the minimum area of the table to be detected should be greater than k1 times the average character area in the image; The value range of k1 is 50 to 300;
[0037] _max_area is the maximum table area detected by the MSER algorithm. This value is determined according to the maximum ratio k2 of the area of the table to be detected to the area of the image. Set it as _max_area = k2×W×H, where W and H are the width and height of the image respectively, and 0 < k2 < 1;
[0038] S42: Obtain the contour coordinate values [Cent1, Cent2, …, CentN] of each region in the image and the coordinate values [Box1, Box2, …, BoxN] of the circumscribed rectangles corresponding to each contour by using the MSER algorithm;
[0039] S43: Apply the non-maximum suppression (NMS) algorithm to all the circumscribed rectangles of the contours obtained in step S42, filter out all the inscribed rectangles within the circumscribed rectangle of the overall table contour, and obtain the overall contour Contour_Table of the table in the image.
[0040] As a preferred technical solution of the present invention, the specific steps in step S5 are as follows:
[0041] S51: Obtain the minimum circumscribed rectangle RectminBox and the circumscribed rectangle RectBox of the table contour Contour_Table, and the corresponding four corner points are [SrcPoint1, SrcPoint2, SrcPoint3, SrcPoint4] and the four corner points [DstPoint1, DstPoint2, DstPoint3, DstPoint4] of the minimum circumscribed rectangle;
[0042] S52: [SrcPoint1, SrcPoint2, SrcPoint3, SrcPoint4] and [DstPoint1, DstPoint2, DstPoint3, DstPoint4] form 4 pairs of points with a perspective mapping relationship. Substitute the coordinates of the 4 pairs of points into the perspective transformation formula to obtain the perspective transformation coefficient matrix M of the image.
[0043]
[0044] S53: Use the perspective transformation matrix M to perform distortion correction on the image H(x, y). The mapping relationship between the pixel coordinate points (X, Y) in the corrected image I(x, y) and the corresponding coordinate points (x, y) in the original image is:
[0045]
[0046]
[0047] As a preferred technical solution of the present invention, the specific steps in step S6 are as follows:
[0048] S61: Detect each table cell in the distortion-corrected image I(x, y) by using the MSER algorithm;
[0049] S611: Set each parameter of the MSER algorithm, where:
[0050] _min_area: The minimum cell area for MSER detection, set to k3×H_char×H_char, indicating that the minimum area of the table cells to be detected should be greater than k3 times the average area of the characters in the image; the value of k3 ranges from 5 to 10;
[0051] _max_area: The maximum cell area for MSER detection, set to k4×H_char×H_char, indicating that the maximum area of the table cells to be detected should be less than k4 times the average area of the characters in the image; the value of k4 ranges from 100 to 200;
[0052] S612: Use the MSER algorithm to obtain the contour coordinate values of each region in the image and the bounding rectangles corresponding to each contour;
[0053] S613: Apply the non-maximum suppression (NMS) algorithm to all the bounding rectangles of the contours, filter out all the inscribed rectangles within the bounding rectangles of the table cell contours, and obtain the bounding rectangles of the N cell contours of the table in the image. The top-left corner points of the bounding rectangles are [Top_left_x_i, Top_left_y_i], where i = 1, 2,..., N;
[0054] S62: Let the sequence [Label_x_1,..., Label_x_N] be the column labels of each table cell, and set Label_x_1 = 1; sort the ordinates Top_left_x_i (i = 1, 2,..., N) of the top-left corner points of the bounding rectangles in ascending order to obtain the sequence Top_left_sort_x_i (i = 1, 2,..., N), and calculate the difference Delt_x_i+1 between the (i + 1)-th and i-th points of the sorted sequence. The formula is:
[0055] Delt_x_i+1 = Top_left_sort_x_i+1 - Top_left_sort_x_i;
[0056] If Delt_x_i+1 > H_char, the column coordinate Label_x_i+1 of the (i + 1)-th cell is Label_x_i+1, otherwise Label_x_i+1 = Label_x_i; perform iterative calculation to obtain the column labels [Label_x_1,..., Label_x_N] of each table cell;
[0057] S63: Let the sequence [Label_y_1,..., Label_y_N] be the row labels of each cell in the table, and set Label_y_1 = 1; sort the vertical coordinates Top_left_y_i (i = 1, 2,..., N) of the upper left corner points of the circumscribed rectangles in ascending order to obtain the sequence Top_left_sort_y_i (i = 1, 2,..., N). The formula for calculating Delt_y_i+1 is:
[0058] Delt_y_i+1 = Top_left_sort_y_i+1 - Top_left_sort_y_i;
[0059] If Delt_y_i+1 > H_char, the row coordinate Label_y_i+1 of the (i + 1)-th cell is Label_y_i+1, otherwise Label_y_i+1 = Label_y_i; through iterative calculation, the row labels [Label_y_1,..., Label_y_N] of each cell in the table can be obtained.
[0060] As a preferred technical solution of the present invention, in step S1, the Retinex image enhancement algorithm is used to process the image taken under natural photographing conditions to eliminate the influence of uneven distribution of the illumination component in the image on the subsequent table structure extraction and OCR recognition performance; the specific steps are as follows: convert the table image collected by the mobile terminal from the RGB space to the grayscale space to obtain a grayscale image; then use the single-scale Retinex algorithm to enhance the grayscale image to eliminate the influence of uneven illumination on the overall performance of the algorithm.
[0061] As a preferred technical solution of the present invention, the specific steps of using the single-scale Retinex algorithm to enhance the grayscale image in step S1 are as follows:
[0062] S11: Use large-scale Gaussian blur to calculate the estimated value i'(x, y) of the illumination component at the pixel point (x, y) of the grayscale image. The size of the Gaussian kernel template is size×size, and the formula for size is:
[0063]
[0064] where H and W are the height and width of the grayscale image, and the value of k ranges from 9 to 13;
[0065] S12: Take the natural logarithm ln[f(x, y)] of the grayscale value f(x, y) at the pixel point (x, y) of the grayscale image;
[0066] S13: In the logarithmic domain, subtract the illumination component i'(x, y) from f(x, y) to obtain the high-frequency reflection component R(x, y) of the grayscale image. R(x, y) is the single-scale Retinex output image R(x, y) of the grayscale image. The formula is: R(x, y) = ln[f(x, y)] - ln[i'(x, y)]. Usually, the Gaussian scale (size) takes values such as 5, 7, 9, etc. The scale (size) of the large-scale Gaussian in this technical solution takes values from 90 to 150.
[0067] As a preferred technical solution of the present invention, the specific steps of performing grayscale transformation on the pixel r(x, y) in the grayscale image R(x, y) to h(x, y) in step S22 are as follows:
[0068] S221: Compress the grayscale dynamic range of the image R(x, y) from 0 - 255 to Minval - Maxval;
[0069]
[0070] S222: Normalize the grayscale value r'(x, y) to 0 - 1, perform gamma transformation on the normalized image with gamma coefficient γ = 3, and then map the gamma-transformed normalized image to 0 - 255 integer type to obtain h(x, y). The mapping formula is:
[0071] h(x, y) = int(255 * [(r'(x, y) - minVal) / (maxVal - minVal) γ )
[0072] Compared with the prior art, the beneficial effects of the technical solution of the present invention are as follows: The detection method of the image table based on MSER under natural photographing conditions uses the Retinex image enhancement algorithm based on the theory of the retina cerebral cortex and image gamma transformation to enhance the image taken under natural photographing conditions; uses morphological and connected component analysis methods to extract the text area in the image and obtain the average height parameter of the characters in the image; uses the MSER algorithm to detect the overall table area in the image, and the minimum detection area of the MSER algorithm is determined according to the average character height. Use the minimum circumscribed rectangle and circumscribed rectangle of the overall table area contour to calculate the perspective transformation matrix and perform distortion correction operation on the image; detect each table cell after correction, merge the rows and columns of each table cell, and determine the row coordinates and column coordinates of the table cells; The present invention has good robustness to images taken under various conditions, can be used for automatic detection and recognition of tables in images, and has good practicability. Description of the Drawings
[0073] The technical solution of the present invention will be further described below in conjunction with the drawings:
[0074] Figure 1 Flow chart of the detection method of the image table based on MSER under natural photographing conditions of the present invention;
[0075] Figure 2 Original image of the table to be recognized taken by a handheld mobile terminal under natural lighting conditions, which is the detection method of the image table based on MSER under natural photographing conditions of the present invention;
[0076] Figure 3 Enhanced image after Retinex processing and gamma transformation in the detection method of the image table based on MSER under natural photographing conditions of the present invention;
[0077] Figure 4 Minimum bounding rectangle and bounding rectangle of the table in the image detected by using the MSER algorithm in the detection method of the image table based on MSER under natural photographing conditions of the present invention;
[0078] Figure 5 Table image after perspective transformation in the detection method of the image table based on MSER under natural photographing conditions of the present invention;
[0079] Figure 6 Each table cell of the table to be recognized in the image detected by using the MSER algorithm in the detection method of the image table based on MSER under natural photographing conditions of the present invention. Detailed implementation manners
[0080] The technical solutions of the present invention will be described clearly and completely below in conjunction with the embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0081] Embodiment: As Figure 1 shown, the detection method of the image table based on MSER under natural photographing conditions includes the following steps:
[0082] S1: Convert the collected image into a grayscale image, and use the Retinex image enhancement algorithm to eliminate the influence of uneven illumination in the image, obtaining the grayscale image R(x, y); the specific steps are as follows: convert the table image collected by the mobile terminal from the RGB space to the grayscale space to obtain a grayscale image, as Figure 2 shown; then use the single-scale Retinex algorithm to perform enhancement processing on the grayscale image to eliminate the influence of uneven illumination on the overall performance of the algorithm;
[0083] The specific steps for enhancing a grayscale image using the single-scale Retinex algorithm are as follows:
[0084] S11: Calculate the estimated illumination component value i'(x, y) at the pixel point (x, y) of the grayscale image using large-scale Gaussian blur. The size of the Gaussian kernel template is size×size, and the formula for size is:
[0085]
[0086] where H and W are the height and width of the grayscale image, and the value of k ranges from 9 to 13;
[0087] S12: Take the natural logarithm ln[f(x, y)] of the grayscale value f(x, y) at the pixel point (x, y) of the grayscale image;
[0088] S13: In the logarithmic domain, subtract the illumination component i'(x, y) from f(x, y) to obtain the high-frequency reflection component R(x, y) of the grayscale image. R(x, y) is the single-scale Retinex output image R(x, y) of the grayscale image; the formula is:
[0089] R(x, y) = ln[f(x, y)] - ln[i′(x, y)];
[0090] S2: Use image gamma transformation to perform contrast enhancement on the grayscale image R(x, y) output by the Retinex algorithm to obtain the enhanced grayscale image H(x, y);
[0091] The specific steps for performing image enhancement on the grayscale image R(x, y) using image gamma transformation in step S2 include:
[0092] S21: Perform histogram statistical analysis on the grayscale image R(x, y) to obtain the grayscale level corresponding to the histogram peak, denoted as Maxval, and the minimum grayscale level of the grayscale image R(x, y), denoted as Minval;
[0093] S22: Perform grayscale transformation on the pixel r(x, y) in the grayscale image R(x, y) to h(x, y) to obtain the enhanced image H(x, y), where the transformation formula is:
[0094]
[0095] The specific steps for performing grayscale transformation on the pixel r(x, y) in the grayscale image R(x, y) to h(x, y) in step S22 are as follows:
[0096] S221: Compress the grayscale dynamic range of the image R(x, y) from 0 to 255 to Minval to Maxval;
[0097]
[0098] S222: Normalize the grayscale value r'(x, y) to 0 - 1, perform gamma transformation on the normalized image with gamma coefficient γ = 3, and then map the gamma-transformed normalized image to an integer type of 0 - 255 to obtain h(x, y); the mapping formula is:
[0099] h(x, y) = int(255 * [(r′(x, y) - minVal) / (maxVal - minVal)] γ ); Figure 3 is the enhanced image after Retinex processing and gamma transformation;
[0100] S3: Use the image morphological processing method and connected component analysis method to extract the text region in the grayscale image H(x, y), and calculate the average height of the characters in the image;
[0101] In the step S3, the morphological method and connected component analysis method are used to extract the text region in the image, and the average height value of the characters in the image is calculated. The specific steps are as follows:
[0102] S31: Use the Otsu algorithm to obtain the binarization threshold of the enhanced grayscale image H(x, y), convert the grayscale image H(x, y) into a binary image by obtaining the binarization threshold, and invert the binary image, that is, perform amplitude inversion on the binary image to obtain a binary image B(x, y) with a background of 0 and tables and characters of 1; then perform closing operation on the binary image B(x, y), and the size of the closing operation kernel can be selected as 5×5 or 7×7;
[0103] S32: Use the image connected component extraction method to extract the connected regions in the binary image, and calculate the width w and height h of each connected region; set thresholds according to the scale size (width W, height H) of the text in the image relative to the image, set the thresholds Tw = W / 10, Th = H / 20; filter out the connected regions with w > Tw or h > Th to obtain the height of the connected regions that meet the conditions;
[0104] S33: Calculate the average value of the heights of all connected regions that meet the conditions, and use the average value as the average height value H_char of the characters in the image.
[0105] S4: Use the Maximally Stable Extremal Regions (MSERs) algorithm to detect the overall table region in the image, and determine the minimum detection area of the MSER algorithm according to the average height of the characters in the image obtained in the step S3, so as to obtain the contour of the overall table region;
[0106] In step S4, the MSER method for region detection is used to detect the overall table region in the image. The specific steps are as follows:
[0107] S41: Set the parameters for initializing the MSER object. Among them:
[0108] _min_area is the minimum table area detected by the MSER algorithm. This value is determined according to the average character height H_char in the image obtained in step S3. Set _min_area = k1×H_char×H_char, indicating that the minimum area of the table to be detected should be greater than k1 times the average character area in the image; the value range of k1 is 50 - 300;
[0109] _max_area is the maximum table area detected by the MSER algorithm. This value is determined according to the maximum ratio k2 of the area of the table to be detected to the area of the image. Set it as _max_area = k2×W×H, where W and H are the width and height of the image respectively, and 0 < k2 < 1;
[0110] S42: Use the MSER algorithm to obtain the contour coordinate values [Cent1, Cent2,..., CentN] of each region in the image and the coordinate values [Box1, Box2,..., BoxN] of the circumscribed rectangles corresponding to each contour;
[0111] S43: Apply the non - maximum suppression (NMS) algorithm to all the circumscribed rectangles of the contours obtained in step S42, filter out all the inscribed rectangles within the circumscribed rectangle of the overall table contour, and obtain the overall contour Contour_Table of the table in the image;
[0112] S5: Use the minimum circumscribed rectangle and the circumscribed rectangle of the overall table region contour in step S4 to calculate the perspective transformation matrix, perform a correction operation on the image, and calculate the coordinate values of each corrected table region in the image. The specific steps in step S5 are as follows:
[0113] S51: Obtain the minimum circumscribed rectangle RectminBox and the circumscribed rectangle RectBox of the table contour Contour_Table. The corresponding four corner points are [SrcPoint1, SrcPoint2, SrcPoint3, SrcPoint4] and the four corner points [DstPoint1, DstPoint2, DstPoint3, DstPoint4] of the minimum circumscribed rectangle; As Figure 4 shown, Figure 4 the rectangle with 4 hollow circles as the corner points in is the minimum circumscribed rectangle of the table contour to be recognized, and the rectangle with 4 small squares as the corner points is the circumscribed rectangle of the table contour to be recognized;
[0114] S52: [SrcPoint1, SrcPoint2, SrcPoint3, SrcPoint4] and [DstPoint1, DstPoint2, DstPoint3, DstPoint4] form four pairs of points with a perspective mapping relationship. Substitute the coordinates of the four pairs of points into the perspective transformation formula to obtain the perspective transformation coefficient matrix M of the image.
[0115]
[0116] S53: Use the perspective transformation matrix M to perform distortion correction on the image H(x, y). The mapping relationship between the pixel coordinate points (X, Y) in the corrected image I(x, y) and the corresponding coordinate points (x, y) in the original image is:
[0117]
[0118]
[0119] S6: Use the MSERS algorithm to detect each table cell after image distortion correction, merge the rows and columns of each table cell, and determine the row coordinates and column coordinates of the table cell; as Figure 5 shown, Figure 5 is the table image after distortion correction;
[0120] The specific steps of the said step S6 are:
[0121] S61: Use the MSER algorithm to detect each table cell in the distorted corrected image I(x, y);
[0122] S611: Set each parameter of the MSER algorithm, where:
[0123] _min_area: The minimum cell area of the table detected by MSER, set to k3 × H_char × H_char, indicating that the minimum area of the table cell to be detected should be greater than k3 times the average area of the characters in the image; the value of k3 is 5 - 10;
[0124] _max_area: The maximum cell area of the table detected by MSER, set to k4 × H_char × H_char, indicating that the maximum area of the table cell to be detected should be less than k4 times the average area of the characters in the image; the value of k4 is 100 - 200;
[0125] S612: Use the MSER algorithm to obtain the contour coordinate values of each region in the image and the circumscribed rectangles corresponding to each contour;
[0126] S613: Apply the non-maximum suppression (NMS) algorithm to all the bounding rectangles of the contours, filter out all the inscribed rectangles within the bounding rectangles of the table cell contours, and obtain the bounding rectangles of the N cell contours of the table in the image. The top-left corner points of the bounding rectangles are [Top_left_x_i, Top_left_y_i], where i = 1, 2, …, N;
[0127] S62: Let the sequence [Label_x_1, …, Label_x_N] be the column labels of each cell in the table, and set Label_x_1 = 1; sort the vertical coordinates Top_left_x_i (i = 1, 2, …, N) of the top-left corner points of the bounding rectangles in ascending order to obtain the sequence Top_left_sort_x_i (i = 1, 2, …, N), and calculate the difference Delt_x_i+1 between the points i+1 and i in the sorted sequence. The formula is:
[0128] Delt_x_i+1 = Top_left_sort_x_i+1 - Top_left_sort_x_i;
[0129] If Delt_x_i+1 > H_char, the column coordinate Label_x_i+1 of the (i+1)-th cell = Label_x_i+1, otherwise Label_x_i+1 = Label_x_i; perform iterative calculation to obtain the column labels [Label_x_1, …, Label_x_N] of each cell in the table;
[0130] S63: Let the sequence [Label_y_1, …, Label_y_N] be the row labels of each cell in the table, and set Label_y_1 = 1; sort the vertical coordinates Top_left_y_i (i = 1, 2, …, N) of the top-left corner points of the bounding rectangles in ascending order to obtain the sequence Top_left_sort_y_i (i = 1, 2, …, N), and calculate the formula for Delt_y_i+1 as:
[0131] Delt_y_i+1 = Top_left_sort_y_i+1 - Top_left_sort_y_i;
[0132] If Delt_y_i+1 > H_char, the row coordinate Label_y_i+1 of the (i+1)-th cell = Label_y_i+1, otherwise Label_y_i+1 = Label_y_i; perform iterative calculation to obtain the row labels [Label_y_1, …, Label_y_N] of each cell in the table; as Figure 6 shown, Figure 6 in which the coordinates of each cell in the corrected table of the image are represented by different gray levels.
[0133] For those of ordinary skill in the art, the specific embodiments only exemplarily describe the present invention. Obviously, the specific implementation of the present invention is not limited by the above-mentioned manner. As long as various non-substantive improvements are made by adopting the method concept and technical solution of the present invention, or the concept and technical solution of the present invention are directly applied to other occasions without improvement, they are all within the protection scope of the present invention.
Claims
1. A method for detecting image tables based on MSER under natural photographing conditions, characterized in that, It includes the following steps: S1: Convert the collected image into a grayscale image, eliminate the influence of uneven illumination in the image, and obtain the grayscale image R(x, y); S2: Adopt image gamma transformation to perform contrast enhancement processing on the grayscale image R(x, y) obtained in the step S1, and obtain the enhanced grayscale image H(x, y); S3: Extract the text area in the grayscale image H(x, y), and calculate the average height of the characters in the image; S4: Detect the overall table area in the image H(x, y), determine the minimum detection area according to the average height of the characters in the image H(x, y) obtained in the step S3, so as to obtain the contour of the overall table area; S5: Use the minimum circumscribed rectangle and the circumscribed rectangle of the overall table area contour in the step S4 to calculate the perspective transformation matrix, perform a correction operation on the image, and calculate the coordinate values of each corrected table area in the image; S6: Detect each table cell after image distortion correction, merge the rows and columns of each table cell, and determine the row coordinates and column coordinates of the table cell; In the step S4, the MSER method is adopted to detect the overall table area in the image, and the specific steps are as follows: S41: Set the parameters for initializing the MSER object, where: _min_area is the minimum table area detected by the MSER algorithm, and this value is determined according to the average character height H_char in the image obtained in the step S3, and is set to k1×H_char×H_char, indicating that the minimum area of the table to be detected should be greater than k1 times the average character area in the image; _max_area is the maximum table area detected by the MSER algorithm, and this value is determined according to the maximum ratio k2 of the area of the table to be detected to the area of the image, and is set to k2×W×H, where W and H are the width and height of the image respectively, and 0 < k2 < 1; S42: Use the MSER algorithm to obtain the contour coordinate values [Cent1, Cent2,..., CentN] of each area in the image and the circumscribed rectangle coordinate values [Box1, Box2,..., BoxN] corresponding to each contour; S43: Adopt the non-maximum suppression algorithm for all the circumscribed rectangles of the contours obtained in the step S42, filter out all the inscribed rectangles inside the circumscribed rectangle of the overall table contour, and obtain the overall contour Contour_Table of the table in the image.
2. The detection method of an image table based on MSER under natural photographing conditions according to claim 1, characterized in that, The specific steps for performing image enhancement processing on the grayscale image R(x, y) by using image gamma transformation in the step S2 include: S21: Perform histogram statistical analysis on the grayscale image R(x, y), and record the gray level corresponding to the histogram peak as maxVal, and record the minimum gray level of the grayscale image R(x, y) as minVal; S22: Perform gray level transformation on the pixel r(x, y) in the grayscale image R(x, y) to h(x, y) to obtain the enhanced image H(x, y), where the transformation formula is: 。 3. The detection method of an image table based on MSER under natural photographing conditions according to claim 1, characterized in that, In the step S3, the morphological method and the connected component analysis method are adopted to extract the text area in the image and calculate the average height value of the characters in the image. The specific steps are as follows: S31: Obtain the binarization threshold of the enhanced grayscale image H(x, y) using the Otsu algorithm. Convert the grayscale image H(x, y) into a binary image by obtaining the binarization threshold, invert the binary image, and perform a closing operation on the binary image. The size of the closing operation kernel can be selected as 5×5 or 7×7; S32: Use the image connected component extraction method to extract the connected regions in the binary image, and calculate the width w and height h of each connected region. According to the scale of the text in the image relative to the overall image, where the width is W and the height is H, set the thresholds Tw = W / 10 and Th = H / 20, and filter out the connected regions with w > Tw or h > Th to obtain the height of the connected regions that meet the conditions; S33: Calculate the average value of the heights of all connected regions that meet the conditions, and use the average value of the heights as the average height value H_char of the characters in the image.
4. The method for detecting an image table based on MSER under natural photographing conditions according to claim 1, characterized in that, The specific steps in step S5 are as follows: S51: Obtain the minimum bounding rectangle RectminBox and the bounding rectangle RectBox of the table contour Contour_Table. The four corner points of the corresponding bounding rectangle RectBox are [SrcPoint1, SrcPoint2, SrcPoint3, SrcPoint4], and the four corner points of the minimum bounding rectangle RectminBox are [DstPoint1, DstPoint2, DstPoint3, DstPoint4]; S52: [SrcPoint1, SrcPoint2, SrcPoint3, SrcPoint4] and [DstPoint1, DstPoint2, DstPoint3, DstPoint4] form four pairs of points with a perspective mapping relationship. Substitute the coordinates of the four pairs of points into the perspective transformation formula to obtain the perspective transformation coefficient matrix of the image. M , ; S53: Using a perspective transformation matrix M To implement distortion correction for the image H(x, y), the mapping relationship between the pixel coordinate points (X, Y) in the corrected image I(x, y) and the corresponding coordinate points (x, y) in the original image is as follows: ; 。 5. The detection method of an image table based on MSER under natural photographing conditions according to claim 4, characterized in that, The specific steps of step S6 are as follows: S61: Use the MSER algorithm to detect each table cell in the distortion-corrected image I(x, y); S611: Set each parameter of the MSER algorithm, where: _min_area: The minimum cell area of the table detected by MSER, set to k3×H_char×H_char, indicating that the minimum area of the table cell to be detected should be greater than k3 times the average area of the characters in the image; _max_area: The maximum cell area of the table detected by MSER, set to k4×H_char×H_char, indicating that the maximum area of the table cell to be detected should be less than k4 times the average area of the characters in the image; S612: Use the MSER algorithm to obtain the contour coordinate values of each region in the image and the bounding rectangle corresponding to each contour; S613: Use the non-maximum suppression algorithm for all the bounding rectangles of the contours to filter out all the inscribed rectangles inside the bounding rectangle of the table cell contour, and obtain the bounding rectangles of the N cell contours of the table in the image. The upper left corner point of the bounding rectangle is [Top_left_x_i, Top_left_y_i], i = 1, 2,..., N; S62: Let the sequence [Label_x_1,..., Label_x_N] be the column labels of each cell in the table, and set Label_x_1 = 1; sort the ordinates Top_left_x_i of the upper-left corner points of the circumscribed rectangles in ascending order for i = 1, 2,..., N to obtain the sequence Top_left_sort_x_i, i = 1, 2,..., N, and calculate the difference Delt_x_i+1 between the (i + 1)-th and i-th points of the sorted sequence. The formula is: Delt_x_i+1 = Top_left_sort_x_i+1 - Top_left_sort_x_i; If Delt_x_i+1 > H_char, the column coordinate Label_x_i+1 of the (i + 1)-th cell is Label_x_i + 1, otherwise Label_x_i+1 = Label_x_i; perform iterative calculation to obtain the column labels [Label_x_1,..., Label_x_N] of each cell in the table; S63: Let the sequence [Label_y_1,..., Label_y_N] be the row labels of each cell in the table, and set Label_y_1 = 1; sort the ordinates Top_left_y_i of the upper-left corner points of the circumscribed rectangles in ascending order for i = 1, 2,..., N to obtain the sequence Top_left_sort_y_i, i = 1, 2,..., N, and calculate the formula for Delt_y_i+1 as: Delt_y_i+1 = Top_left_sort_y_i+1 - Top_left_sort_y_i; If Delt_y_i+1 > H_char, the row coordinate Label_y_i+1 of the (i + 1)-th cell is Label_y_i + 1, otherwise Label_y_i+1 = Label_y_i; perform iterative calculation to obtain the row labels [Label_y_1,..., Label_y_N] of each cell in the table.
6. The detection method of an image table based on MSER under natural photographing conditions according to claim 4, wherein In step S1, the Retinex image enhancement algorithm is used to process the image taken under natural photographing conditions to eliminate the influence of uneven distribution of illumination components in the image on the subsequent table structure extraction and OCR recognition performance; the specific steps are: convert the table image collected by the mobile terminal from the RGB space to the grayscale space to obtain a grayscale image; then use the single-scale Retinex algorithm to perform enhancement processing on the grayscale image to eliminate the influence of uneven illumination on the overall performance of the algorithm.
7. The method for detecting an image table based on MSER under natural photographing conditions according to claim 6, wherein The specific steps of using the single-scale Retinex algorithm to perform enhancement processing on the grayscale image in step S1 are: S11: Use the large-scale Gaussian blur method to calculate the estimated value i'(x, y) of the illumination component at the pixel point (x, y) of the grayscale image. The size of the Gaussian kernel template is size×size, and the calculation formula for size is: ; where H and W are the height and width of the grayscale image, and the value of k ranges from 9 to 13; S12: Take the natural logarithm ln[f(x, y)] of the gray value f(x, y) of the grayscale image at the pixel point (x, y). S13: In the logarithmic domain, subtract the illumination component i'(x, y) from f(x, y) to obtain the high-frequency reflection component R(x, y) of the grayscale image. R(x, y) is the single-scale Retinex output image R(x, y) of the grayscale image. The formula is: 。 8. The detection method of the image table based on MSER under natural photographing conditions according to claim 2, characterized in that, The specific steps for performing gray-scale transformation on the pixel r(x, y) in the grayscale image R(x, y) to h(x, y) in step S22 are as follows: S221: Compress the gray-scale dynamic range of the image R(x, y) from 0 to 255 to minVal to maxVal. ; S222: Normalize the gray value r'(x, y) to 0 to 1, perform gamma transformation on the normalized image with the gamma coefficient γ = 3, and then map the gamma-transformed normalized image to an integer type of 0 to 255 to obtain h(x, y). The mapping formula is: 。
Citation Information
Patent Citations
Table reconstruction method and device, electronic equipment and storage medium
CN110738030A
Method and device for automatically identifying paper table structure
CN112036294A