A method for identifying deformation table structure
Through the deformation table structure recognition method of steps such as image preprocessing, character removal, corner point positioning and cell positioning, the problem of the existing technology being difficult to identify deformation table images that are disturbed by background, lighting, and physical deformation is solved, and table structure recognition with high accuracy and strong anti-interference ability is achieved.
Patent Information
- Application Number
- CN202210573606.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-24
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-05-24
AI Technical Summary
Existing table structure recognition methods are difficult to effectively identify deformation table images with interference such as background, lighting, and physical deformation. Especially when the table lines of the image are bent, curled, and wrinkled, the accuracy is low.
A deformation table structure recognition method is proposed, and the structure of the table image is recognized through the steps of image preprocessing, character removal, corner point positioning, outline acquisition and cell positioning. The method includes image enhancement, binarization and skeleton extraction, character removal algorithm, corner point detection and clustering, contour acquisition and cell coordinate determination.
This method can effectively remove interference in the image, accurately obtain corner point information and position cell positions, significantly improving the accuracy and anti-interference ability of deformation table structure recognition.
Smart Images

Figure CN114973283B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the field of computer information technology, and in particular relates to a deformation table structure recognition method. Background Art
[0002] Table structure recognition is an important research topic in the field of image processing and document recognition. The styles and types of tables are diverse, making the structural recognition of table images a major problem. Nowadays, most mature table structure recognition methods are aimed at PDF, HTML or scanned table images, such as the patent document named "A PDF table structure recognition method based on image recognition" (publication number CN111144300A) and the patent document named "A PDF table structure recognition method based on graph attention mechanism" (publication number CN110751038A) respectively disclose the structural recognition methods for PDF table images. These methods are only aimed at PDF table images and have a relatively limited scope of application.
[0003] There are also patents that propose structural recognition methods for photographic table images. For example, a patent document named "A table structure extraction method" (publication number CN111368695A) discloses an image-based table structure extraction method, which obtains table cells through steps such as straight line detection, finding corner points, and connecting broken lines. Although this method can completely extract the table, it is no longer applicable when the table lines in the image are curved. A patent document named "A table structure completion algorithm based on table node recognition" (publication number CN109447007A) discloses a method for restoring the structural features of the table itself as much as possible by identifying table nodes. Although this method can correct images with perspective angles, it is still difficult to obtain high accuracy for images with curls and wrinkles. Summary of the invention
[0004] In order to solve the above technical problems, the present invention proposes a deformable table structure recognition method, which performs structure recognition on table images that are interfered by factors such as background, lighting, physical deformation, etc.; the structure of the table image is recognized through image preprocessing, character removal, corner point positioning, contour acquisition, cell positioning and other methods.
[0005] The present invention provides a method for identifying a deformation table structure, characterized in that the method comprises the following steps:
[0006] (1) Image preprocessing: performing image enhancement, binarization, and skeleton extraction on the input original image I containing the table to obtain a first binary image I1;
[0007] (2) Character removal: Most characters in the first binary image I1 are removed by using a character removal algorithm to obtain a second binary image I2; then the number of black pixels in four adjacent rectangular areas of the pixel is further determined, and all characters in the second binary image I2 are removed to obtain a third binary image I3;
[0008] (3) Corner point location: First, the corner points in the third binary image I3 are detected using a corner point detection algorithm to obtain a first corner point set P1; then, the corner points in the first corner point set P1 are clustered to obtain a second corner point set P2; finally, the corner points in the second corner point set P2 that do not meet the conditions are screened to obtain a corner point set P3 of the original image I;
[0009] (4) Contour acquisition: Delete pixels with a width of 1 in the horizontal direction of the third binary image I3 to obtain a fourth binary image I4 that retains only the horizontal lines; Obtain all contours Con1, Con2, ..., Con in the fourth binary image I4 β , where β is the total number of contours in the fourth binary image I4;
[0010] (5) Cell positioning: Classify all corner points in the corner point set P3 and classify those that belong to the contour in the corner point set P3. Add the corner points to the point set in Get the corner point set point1, point2, ..., point β ; Then according to the corner point set point1, point2, ..., point β The position of each corner point in the original image I is used to determine the coordinates of the upper left corner vertex, the upper right corner vertex, the lower right corner vertex, and the lower left corner vertex of each cell in the original image I, and the cell coordinate set CP is obtained.
[0011] In the above step (1), the input original image I containing the table is subjected to image enhancement, binarization and skeleton extraction, and the following methods are specifically used:
[0012] (1.1) The input original image I is enhanced using an image enhancement algorithm to obtain the image enhancement result image I 01 Then, the image binarization algorithm is used to enhance the image I 01 Perform image binarization to obtain the image enhancement result binary image I 02 ;
[0013] (1.2) Use skeleton extraction algorithm to enhance the binary image I 02 Perform skeleton extraction to obtain a first binary image I1.
[0014] In the above step (2), a character removal algorithm is used to remove most of the characters in the first binary image I1, and the following method is specifically used:
[0015] (2.1) Traverse each pixel in the first binary image I1 and extract the connected area Area in the first binary image I1 through the eight-neighborhood adjacency relationship of the pixel j , where j = 1, 2, ..., N, N is the total number of connected areas in the first binary image I1, and all the extracted connected areas are stored in a connected area set Area;
[0016] (2.2) For each connected area j , by solving the following determinant, calculate the connected area Area j The covariance matrix S j The eigenvalue of Where t = 1, 2,
[0017]
[0018]
[0019]
[0020]
[0021] in, Respectively represent the connected area Area j The average value of the horizontal coordinate and the average value of the vertical coordinate of all pixels in x j and j They represent the connected areas Are j The horizontal and vertical coordinates of the pixel points in a, j = 1, 2, ..., N, N is the total number of connected areas in the first binary image I1;
[0022] (2.3) Traverse each connected area in the connected area set Area j , for the connected area Area j The covariance matrix S j The eigenvalue of Make a judgment, where j = 1, 2, ..., N, N is the total number of connected areas in the first binary image I1, if the eigenvalue satisfy
[0023]
[0024]
[0025] Among them, th2, th3, and th4 are the set thresholds, Cj Area is the connected area j The number of pixels and connected area j The number of pixels C j satisfy
[0026] C j <th4 (19)
[0027] Then the connected area Area j By deleting it from the connected area set Area, the second binary image I2 can be obtained.
[0028] In the above step (2), the number of black pixels in the four adjacent rectangular areas of the pixel is further determined, and all characters in the second binary image I2 are removed to obtain the third binary image I3. Specifically, the following method is used:
[0029] (2.4) Traverse each pixel in the second binary image I2 and store the pixel with gray value 0 in the set WP; each pixel po in the set WP α , let its horizontal coordinate be x α , the vertical coordinate is y α , the horizontal coordinate of the upper left corner of the upper adjacent rectangular area Rect1 is x α -λ1, y is the vertical coordinate α -λ2, the horizontal coordinate of the lower right corner is x α +λ1, ordinate is y α ; The horizontal coordinate of the upper left corner of the lower adjacent rectangular area Rect2 is x α -λ1, y is the vertical coordinate α , the horizontal coordinate of the lower right corner is x α +λ1, ordinate is y α +λ2; the horizontal coordinate of the upper left corner of the left adjacent rectangular area Rect3 is x α -λ2, y is the vertical coordinate α -λ1, the horizontal coordinate of the lower right corner is x α , the vertical coordinate is y α +λ1; the horizontal coordinate of the upper left corner of the right adjacent rectangular area Rect4 is x α , the vertical coordinate is y α -λ1, the horizontal coordinate of the lower right corner is x α +λ2, ordinate is y α +λ1, where λ1 and λ2 are the set coordinate distance thresholds; calculate the pixel point po α The number of black pixels in the four adjacent rectangular areas Rect1, Rect2, Rect3, and Rect4 on the top, bottom, left, and right Where α=1,2,...,M, M is the total number of pixels in the set WP;
[0030] (2.5) For pixel po α The adjacent rectangular area Rect i , where i = 1, 2, 3, 4; if the number of black pixels If it is greater than 1, the adjacent rectangular area Rect i The black pixel value Assign value 1, otherwise 0:
[0031]
[0032] (2.6) Calculate pixel point po α The black pixel values in the four adjacent rectangular areas are The sum of C α , the calculation formula is as follows:
[0033]
[0034] (2.7) If the pixel point po α C of four adjacent rectangular areas α If the value is 1, the pixel point po α As an endpoint, the pixel point po α Add to the set EP; if the pixel point po α C of four adjacent rectangular areas α The value is 2, then the pixel point po α Consider it as a connection point and place the pixel point po α Add to the set LP; in other cases, the pixel po α Consider it as the intersection point and put the pixel point po α Add to the collection IP;
[0035] (2.8) The noise points in the sets EP and LP are filtered out to obtain the third binary image I3.
[0036] In the above step (2.8), the noise points in the sets EP and LP are filtered out by using the following method:
[0037] (2.8.1) Randomly select a pixel point P from the set EP;
[0038] (2.8.2) For each pixel point P in the set WP μ , where μ=1,2,...,M, M is the total number of pixels in the set WP; calculate the pixel point P μ If the distance between the pixel point P and the pixel point P is less than the set first distance threshold θ, the pixel point P μFurther judgment:
[0039] The calculation formula of the distance dis between point p1 and point p2 is as follows:
[0040]
[0041] Among them, x1 and y1 are the horizontal coordinate and vertical coordinate of point p1 respectively, and x2 and y2 are the horizontal coordinate and vertical coordinate of point p2 respectively.
[0042] a) If the pixel point P μ If it also exists in the set LP, then the pixel point P is deleted from the sets EP and WP, and the pixel point P is μ Store in the collection EP;
[0043] b) If the pixel point P μ If it also exists in the set IP, then the pixel point P is deleted from the sets EP and WP;
[0044] c) If the pixel point P μ If it also exists in the set EP, then the pixel point P is deleted from the sets EP and WP, and the pixel point P is μ Remove from collection WP;
[0045] (2.8.3) If there are still pixels in the set EP, go to step (2.8.1), otherwise the noise filtering algorithm ends.
[0046] In the above step (3), the corner points in the first corner point set P1 are clustered, and the following method is specifically used:
[0047] (3.1) Randomly select a point in the first corner point set P1 As the initial point p0, where t∈{1,2,...,n}, n is the total number of corner points in the first corner point set P1;
[0048] (3.2) Traverse all points in the first corner point set P1 except the initial point p0 where t'∈{1,2,...,n}; calculation point With point The distance between tt' ,point The horizontal axis is The vertical axis is point The horizontal axis is The vertical axis is The distance calculation is shown in formula (10), where the distance dis tt' Corner points less than the set second distance threshold k are stored in the set CLP;
[0049] (3.3) The corner points Stored in the second corner point set P2, and for each point p in the set CLP γ ,in is the total number of corner points in the set CLP, and determines whether point p exists in the first corner point set P1 γ , if there is a point p in the first corner point set P1 γ , then point p γ Delete from the first corner point set P1;
[0050] (3.4) Set the set CLP to empty;
[0051] (3.5) If the first corner point set P1 is not empty, go to step (3.1), otherwise the corner point clustering algorithm ends.
[0052] In the above step (3), the corner points that do not meet the conditions in the second corner point set P2 are screened, and the following method is specifically used:
[0053] (3.6) Calculate each corner point in the second corner point set P2 in turn The total number of pixels with grayscale value 0 in the four adjacent rectangular areas Rect1, Rect2, Rect3, and Rect4 above, below, left, and right Where c = 1, 2, ..., m, m is the total number of corner points in the second corner point set P2;
[0054] (3.7) For corner points The adjacent rectangular area Rect i , where i = 1, 2, 3, 4; if the number of black pixels If it is greater than 9, the adjacent rectangular area Rect i The black pixel value Assign value 1, otherwise 0:
[0055]
[0056] (3.8) If the corner point of One of the following six situations: a) b) c) d) e) f) Then the corner point Delete from the second corner point set P2.
[0057] In the above step (5), all corner points in the corner point set P3 are classified, and those belonging to the contour are classified into Add the corner points to the point set in Get the corner point set point1, point2, ..., point β , specifically the following methods were used:
[0058] (5.1) For each contour, calculate the horizontal coordinate of its contour centroid according to the following formula and the vertical coordinate
[0059]
[0060]
[0061] Among them, f(i,j) is the gray value of the contour at the horizontal coordinate i and the vertical coordinate j, and W and V are the gray values of the contour respectively. The width and height of the bounding rectangle;
[0062] (5.2) Each contour is divided into two parts according to its centroid ordinate value. Sort from small to large, and the sorted contours are recorded as Con'1, Con'2, ..., Con' β , Contour Con' i The horizontal coordinate of the center of mass is The vertical axis is Where i = 1, 2, ..., β, β is the total number of contours in the fourth binary image I4; for each contour Con' i , if the ordinate y of the corner point k in the corner point set P3 k satisfy Where k∈{1,2,...,Z}, Z is the total number of corner points in the corner point set P3, σ is the set vertical coordinate distance threshold, then the corner point k and the contour Con are calculated respectively. i Distance of the centroid ik , corner point k and contour Con i-1 Distance of the centroid (i-1)k , corner point k and contour Con i+1 Distance of the centroid (i+1)k , where the horizontal coordinate of the corner point k is x k , the vertical coordinate is y k , Contour i The horizontal coordinate of the center of mass is The vertical axis is Contour i-1 The horizontal coordinate of the center of mass is The vertical axis is Contour i+1 The horizontal coordinate of the center of mass is The vertical axis is The distance calculation is shown in formula (10); the comparison distance (i-1)k 、distance ik 、distance (i+1)k The size of the distance is selected with the smallest value. l, , then the corner point k is on the corresponding contour Con l , the calculation formula is as follows:
[0063]
[0064] (5.3) For each contour Con i , will be located on the same contour Con i The corner points on the i , and all the corner points in each point set are sorted by the horizontal coordinate value x of the corner point in the point set τ Sort from small to large, where i = 1, 2, ..., β, β is the total number of contours in the fourth binary image I4, τ = 1, 2, ..., φ, φ is the point set point i The total number of corner points in .
[0065] In the above step (5), according to the corner point set point1, point2, ..., point β The position of each corner point in the original image I is determined to determine the coordinates of the upper left corner vertex, the upper right corner vertex, the lower right corner vertex, and the lower left corner vertex of each cell in the original image I, and the cell coordinate set CP is obtained. The specific method is as follows:
[0066] (5.4) Traverse the corner point set point1, point2, ..., point β Each subset point k , where k = 1, 2, ..., β, β is the total number of contours in the fourth binary image I4, and the subset point k The first corner point in Find the horizontal coordinate value from the corner point set P3 With corner point The horizontal axis value The difference does not exceed the threshold th, the vertical coordinate value Greater than corner The vertical coordinate value of And away from the corner The distance of the corner point is less than the set distance third threshold δ The distance calculation is shown in formula (10);
[0067] (5.5) Determine the corner point With corner point Is there a connecting line between them? If the corner points With corner point If there is a connection between them, then take the subset point k The second corner point in Find the horizontal coordinate value from the corner point set P3 With corner point The horizontal axis value The difference does not exceed the threshold th, the vertical coordinate value Greater than corner The vertical coordinate value of And away from the corner The distance of the corner point is less than the set distance third threshold δ If the corner With corner point If there is no connection between them, let k = k + 1. If k = β + 1, the cell location algorithm ends, otherwise jump to step (5.4);
[0068] (5.6) Determine the corner point With corner point Is there a connecting line between them? If the corner points With corner point If there is a line between them, the upper left vertex, upper right vertex, lower right vertex, and lower left vertex of the cell are Store these four vertices in the cell coordinate set CP; if the corner point With corner point If there is no connection between them, then the subset point is calculated k The number of corner points in ψ, if the subset point k If the number of corner points in ψ ≥ 3, then take the subset point k The third corner point in is taken as the corner point Find the horizontal coordinate value from the corner point set P3 With corner point The horizontal axis value The difference does not exceed the threshold th, the vertical coordinate value Greater than corner The vertical coordinate value of And away from the corner The distance of the corner point is less than the set distance third threshold δ Jump to step (5.6), otherwise take k = k + 1, if k = β + 1, the cell positioning algorithm ends, otherwise jump to step (5.4);
[0069] (5.7) The corner points From the subset pointk If the subset point k There is only one corner point left in the , at this time take k = k + 1, if k = β + 1, the cell positioning algorithm ends, otherwise jump to step (5.4);
[0070] The algorithm for determining whether there is a connecting line between two corner points in the above steps (5.5) and (5.6) specifically adopts the following method:
[0071] (5.8) Let the two corner points be p 3h 、p 3g , corner point p 3h The horizontal axis is x p3h , the vertical coordinate is y p3h , corner point p 3g The horizontal axis is x p3g , the vertical coordinate is y p3g , where y p3g >y p3h ;
[0072] (5.9) Determine a rectangular area rect in the third binary image I3 based on the coordinates of the two corner points, and compare x p3h and x p3g The size of x p3h >x p3g , then the horizontal coordinate of the upper left corner of the rectangular area rect is x p3g , the vertical coordinate is y p3h , the horizontal coordinate of the upper right vertex is x p3h , the vertical coordinate is y p3h , the horizontal coordinate of the lower right vertex is x p3h , the vertical coordinate is y p3g , the horizontal coordinate of the lower left vertex is x p3g , the vertical coordinate is y p3g ; if x p3h =x p3g , then the horizontal coordinate of the upper left corner of the rectangular area rect is x p3g -ξ1, the vertical coordinate is y p3h , the horizontal coordinate of the upper right vertex is x p3h +ξ2, the vertical coordinate is y p3h , the horizontal coordinate of the lower right vertex is x p3h +ξ2, the vertical coordinate is y p3g , the horizontal coordinate of the lower left vertex is x p3g -ξ1, the vertical coordinate is y p3g ; if x p3h <x p3g , then the horizontal coordinate of the upper left corner of the rectangular area rect is x p3h , the vertical coordinate is y p3h, the horizontal coordinate of the upper right vertex is x p3g , the vertical coordinate is y p3h , the horizontal coordinate of the lower right vertex is x p3g , the vertical coordinate is y p3g , the horizontal coordinate of the lower left vertex is x p3h , the vertical coordinate is y p3g , where ξ1 and ξ2 are the set thresholds;
[0073] (5.10) Calculate the number of black pixels Cnt in the rectangular area rect; if It means that there is a line connecting the two corner points; otherwise, it means that there is no line connecting the two corner points, where φ is the set critical threshold and w is the width of the rectangular area rect.
[0074] The advantages of the present invention are: for the interference of background, lighting, physical deformation and the like in the deformation table, a deformation table structure recognition method is provided. The method can effectively remove characters in the image, accurately obtain the corner point information in the image, and locate the position of the cell at the same time. The method can be effectively applied to the structure recognition of the deformation table, and not only has strong anti-interference ability and high accuracy, but also has good application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0075] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0076] Figure 1 is a method flow chart of an embodiment of the present invention;
[0077] Figure 2 is an original image to be processed in an embodiment of the present invention;
[0078] Figure 3 It is the image enhancement result binary image after image enhancement and binarization;
[0079] Figure 4 is the first binary image after skeleton extraction;
[0080] Figure 5 It is a schematic diagram of four adjacent rectangular areas of pixels;
[0081] Figure 6 is the third binary image after character removal;
[0082] Figure 7 is the corner point positioning result image;
[0083] Figure 8 is the cell positioning result image. DETAILED DESCRIPTION
[0084] The specific implementation of the present invention will be further described in detail below in conjunction with the drawings in the embodiments of the present invention; it should be noted that the specific embodiment of the deformation table structure recognition method according to the present invention is only used as an example and is not used to limit the present invention.
[0085] This example describes the deformable table structure recognition algorithm by combining the input original image I containing a table. Figure 1 As shown in the method flow chart, the present invention adopts the following steps to identify the deformation table structure:
[0086] (1) Image preprocessing. Input the original image I containing the table, such as Figure 2 The input original image I containing the table is subjected to image enhancement, binarization and skeleton extraction to obtain a first binary image I1.
[0087] (2) Character removal. A character removal algorithm is used to remove most of the characters in the first binary image I1 to obtain a second binary image I2. Then, the number of black pixels in the four adjacent rectangular areas of the pixel is further determined, and all characters in the second binary image I2 are removed to obtain a third binary image I3, as shown in FIG. Figure 6 shown.
[0088] (3) Corner point location. First, the corner points in the third binary image I3 are detected using a corner point detection algorithm to obtain a first corner point set P1. Then, the corner points in the first corner point set P1 are clustered to obtain a second corner point set P2. Finally, the corner points that do not meet the conditions in the second corner point set P2 are screened to obtain a corner point set P3 of the original image I.
[0089] The Harris corner detection algorithm used in the above steps is a common image feature extraction method, see Harris CG, Stephens M J. "A Combined Corner and Edge Detector." Alvey Vision Conference, 1988.
[0090] (4) Contour acquisition. Delete the pixels with a width of 1 in the horizontal direction of the third binary image I3 to obtain a fourth binary image I4 that retains only the horizontal lines. Obtain all contours Con1, Con2, ..., Con β , where β is the total number of contours in the fourth binary image I4.
[0091] (5) Cell positioning. Classify all corner points in the corner point set P3 and classify those that belong to the contour. Add the corner points to the point set in Get the corner point set point1, point2, ..., point β ; Then according to the corner point set point1, point2, ..., point β The position of each corner point in the original image I is used to determine the coordinates of the upper left corner vertex, the upper right corner vertex, the lower right corner vertex, and the lower left corner vertex of each cell in the original image I, and the cell coordinate set CP is obtained.
[0092] In the above step (1), the input original image I containing the table is subjected to image enhancement, binarization and skeleton extraction, and the following methods are specifically used:
[0093] (1.1) The input original image I is enhanced using an image enhancement algorithm to obtain the image enhancement result image I 01 Then, the image binarization algorithm is used to enhance the image I 01 Perform image binarization to obtain the image enhancement result binary image I 02 ,like Figure 3 shown.
[0094] The MSRCR image enhancement algorithm used in the above steps is a commonly used method for improving image clarity, see Rahman Zu, Jobson DJ, Woodell G A. "Multi-scale retinex for color image enhancement" Proceedings of 3rd IEEE International Conference on Image Processing, Lausanne, Switzerland, 1996, 3: 1003-1006 vol. 3. The local adaptive binarization algorithm is a commonly used image analysis and processing method, see Zong Zehua, Zhang Haijun, Zhang Jinfeng. "ORB feature extraction and tracking based on adaptive threshold" Electro-Optics and Control: 1-7.
[0095] (1.2) Using Zhang-Suen skeleton extraction algorithm to enhance the binary image I 02 Perform skeleton extraction to obtain the first binary image I1, such as Figure 4 shown.
[0096] The Zhang-Suen skeleton extraction algorithm used in the above steps is a commonly used image thinning method, see Zhang TY, Suen CY. "A fast parallel algorithm for thinning digital patterns" Commun. ACM, 1984, 27: 236-239.
[0097] In the above step (2), a character removal algorithm is used to remove most of the characters in the first binary image I1, and the following method is specifically used:
[0098] (2.1) Traverse each pixel in the first binary image I1 and extract the connected area Area in the first binary image I1 through the eight-neighborhood adjacency relationship of the pixel j , where j=1, 2, …, N, N is the total number of connected areas in the first binary image I1, and all the extracted connected areas are stored in a connected area set Area.
[0099] In the above step, extracting the connected area in the image through the eight-neighborhood adjacency relationship of the pixel points is a commonly used image processing method, see Nakaigawa T, Mashiyama Y, Mitsukura Y, et al. A Robust Table Detection Method for Distortion in Image Acquired from Camera. IECON 2019-45th Annual Conference of the IEEE Industrial Electronics Society, Lisbon, Portugal, 2019, 1: 5347-5352.
[0100] (2.2) For each connected area j , by solving the following determinant, calculate the connected area Area j The covariance matrix S j The eigenvalue of Where t = 1, 2,
[0101]
[0102]
[0103]
[0104]
[0105] in, Respectively represent the connected area Area j The average value of the horizontal coordinate and the average value of the vertical coordinate of all pixels in x j and j They represent the connected areas Are jThe horizontal and vertical coordinates of the pixel points in a, j=1, 2,…, N, N is the total number of connected areas in the first binary image I1.
[0106] (2.3) Set thresholds th2 = 60, th3 = 100, and th4 = 40. Traverse each connected area in the connected area set Area j , for the connected area Area j The covariance matrix S j The eigenvalue of Make a judgment, where j = 1, 2, ..., N, N is the total number of connected areas in the first binary image I1, if the eigenvalue satisfy
[0107]
[0108]
[0109] Among them, C j Area is the connected area j The number of pixels and connected area j The number of pixels C j satisfy
[0110] C j <th4 (31)
[0111] Then the connected area Area j By deleting it from the connected area set Area, the second binary image I2 can be obtained.
[0112] In the above step (2), the number of black pixels in the four adjacent rectangular areas of the pixel is further determined, and all characters in the second binary image I2 are removed to obtain the third binary image I3. Specifically, the following method is used:
[0113] (2.4) Set the coordinate distance thresholds λ1 = 3 and λ2 = 7. Traverse each pixel in the second binary image I2 and store the pixels with grayscale value 0 in the set WP. Each pixel po in the set WP α , let its horizontal coordinate be x α , the vertical coordinate is y α , the horizontal coordinate of the upper left corner of the upper adjacent rectangular area Rect1 is x α -λ1, y is the vertical coordinate α -λ2, the horizontal coordinate of the lower right corner is x α +λ1, ordinate is y α ; The horizontal coordinate of the upper left corner of the lower adjacent rectangular area Rect2 is x α -λ1, y is the vertical coordinate α, the horizontal coordinate of the lower right corner is x α +λ1, ordinate is y α +λ2; the horizontal coordinate of the upper left corner of the left adjacent rectangular area Rect3 is x α -λ2, y is the vertical coordinate α -λ1, the horizontal coordinate of the lower right corner is x α , the vertical coordinate is y α +λ1; the horizontal coordinate of the upper left corner of the right adjacent rectangular area Rect4 is x α , the vertical coordinate is y α -λ1, the horizontal coordinate of the lower right corner is x α +λ2, ordinate is y α +λ1, pixel po α The four adjacent rectangular areas are as follows Figure 5 As shown. Calculate the pixel point po α The number of black pixels in the four adjacent rectangular areas Rec1t, Rect2, Rect3, and Rect4 on the top, bottom, left, and right Where α = 1, 2, ..., M, M is the total number of pixels in the set WP.
[0114] (2.5) For pixel po α The adjacent rectangular area Rect i , where i = 1, 2, 3, 4. If the number of black pixels If it is greater than 1, the adjacent rectangular area Rect i The black pixel value Assign value 1, otherwise 0:
[0115]
[0116] (2.6) Calculate pixel point po α The black pixel values in the four adjacent rectangular areas are The sum of C α , the calculation formula is as follows:
[0117]
[0118] (2.7) If the pixel point po α C of four adjacent rectangular areas α If the value is 1, the pixel point po α As an endpoint, the pixel point po α Add to the set EP. If the pixel point po α C of four adjacent rectangular areas α The value is 2, then the pixel point po αConsider it as a connection point and place the pixel point po α Add to the set LP. In other cases, the pixel po α Consider it as the intersection point and put the pixel point po α Stored in the collection IP.
[0119] (2.8) The noise points in the sets EP and LP are filtered out to obtain the third binary image I3.
[0120] In the above step (2.8), the noise points in the sets EP and LP are filtered out by using the following method:
[0121] (2.8.1) Randomly select a pixel point P from the set EP.
[0122] (2.8.2) Set the distance threshold θ = 3. For each pixel point P in the set WP μ , where μ=1,2,...,M, M is the total number of pixels in the set WP. Calculate the pixel P μ If the distance between the pixel point P and the pixel point P is less than the set first distance threshold θ, the pixel point P μ Further judgment:
[0123] The calculation formula of the distance dis between point p1 and point p2 is as follows:
[0124]
[0125] Among them, x1 and y1 are the horizontal coordinate and vertical coordinate of point p1 respectively, and x2 and y2 are the horizontal coordinate and vertical coordinate of point p2 respectively.
[0126] a) If the pixel point P μ If it also exists in the set LP, then the pixel point P is deleted from the sets EP and WP, and the pixel point P is μ Stored in the collection EP.
[0127] b) If the pixel point P μ If it also exists in the set IP, the pixel point P is deleted from the sets EP and WP.
[0128] c) If the pixel point P μ If it also exists in the set EP, then the pixel point P is deleted from the sets EP and WP, and the pixel point P is μ Removed from collection WP.
[0129] (2.8.3) If there are still pixels in the set EP, go to step (2.8.1), otherwise the noise filtering algorithm ends.
[0130] In the above step (3), the corner points in the first corner point set P1 are clustered, and the following method is specifically used:
[0131] (3.1) Randomly select a point in the first corner point set P1 As the initial point p0, where t∈{1,2,...,n}, n is the total number of corner points in the first corner point set P1.
[0132] (3.2) Set the distance threshold k = 3. Traverse all points in the first corner point set P1 except the initial point p0 Where t'∈{1,2,...,n}. Calculate the point With point The distance between tt' ,point The horizontal axis is The vertical axis is point The horizontal axis is The vertical axis is The distance calculation is shown in formula (10), where the distance dis tt' Corner points whose distance is less than the set second distance threshold k are stored in the set CLP.
[0133] (3.3) The corner points Stored in the second corner point set P2, and for each point p in the set CLP γ ,in is the total number of corner points in the set CLP, and determines whether point p exists in the first corner point set P1 γ , if there is a point p in the first corner point set P1 γ , then point p γ Delete from the first corner point set P1.
[0134] (3.4) Set the set CLP to empty.
[0135] (3.5) If the first corner point set P1 is not empty, go to step (3.1), otherwise the corner point clustering algorithm ends.
[0136] In the above step (3), the corner points that do not meet the conditions in the second corner point set P2 are screened, and the following method is specifically used:
[0137] (3.6) Calculate each corner point in the second corner point set P2 in turn The total number of pixels with grayscale value 0 in the four adjacent rectangular areas Rect1, Rect2, Rect3, and Rect4 above, below, left, and right Wherein c=1,2,...,m, m is the total number of corner points in the second corner point set P2.
[0138] (3.7) For corner points The adjacent rectangular area Rect i , where i = 1, 2, 3, 4. If the number of black pixels If it is greater than 9, the adjacent rectangular area Rect i The black pixel value Assign value 1, otherwise 0:
[0139]
[0140] (3.8) If the corner point of One of the following six situations: a) b) c) d) e) f) Then the corner point Delete from the second corner point set P2. Then we can get the corner point positioning result image, such as Figure 7 shown.
[0141] In the above step (5), all corner points in the corner point set P3 are classified, and those belonging to the contour are classified into Add the corner points to the point set in Get the corner point set point1, point2, ..., point β ,, the following methods were used:
[0142] (5.1) For each contour, calculate the horizontal coordinate of its contour centroid according to the following formula and the vertical coordinate
[0143]
[0144]
[0145] Among them, f(i,j) is the gray value of the contour at the horizontal coordinate i and the vertical coordinate j, and W and V are the gray values of the contour respectively. The width and height of the bounding rectangle.
[0146] (5.2) Each contour is divided into two parts according to its centroid ordinate value. Sort from small to large, and the sorted contours are recorded as Con'1, Con'2, ..., Con' β , Contour Con' i The horizontal coordinate of the center of mass is The vertical axis is Where i = 1, 2, ..., β, β is the total number of contours in the fourth binary image I4. i , if the ordinate y of the corner point k in the corner point set P3 k satisfy Where k∈{1,2,...,Z}, Z is the total number of corner points in the corner point set P3, σ is the set vertical coordinate distance threshold, then the corner point k and the contour Con are calculated respectively. i Distance of the centroid ik , corner point k and contour Con i-1 Distance of the centroid (i-1)k , corner point k and contour Con i+1 Distance of the centroid (i+1)k , where the horizontal coordinate of the corner point k is x k , the vertical coordinate is y k , Contour i The horizontal coordinate of the center of mass is The vertical axis is Contour i-1 The horizontal coordinate of the center of mass is The vertical axis is Contour i+1 The horizontal coordinate of the center of mass is The vertical axis is The distance calculation is shown in formula (10). (i-1)k 、distance ik 、distance (i+1)k The size of the distance is selected with the smallest value. l, , then the corner point k(x k ,y k ) on the corresponding contour Con l , the calculation formula is as follows:
[0147]
[0148] (5.3) For each contour Con i , will be located on the same contour Con i The corner points on the i , and all the corner points in each point set are sorted by the horizontal coordinate value x of the corner point in the point set τSort from small to large, where i = 1, 2, ..., β, β is the total number of contours in the fourth binary image I4, τ = 1, 2, ..., φ, φ is the point set point i The total number of corner points in .
[0149] In the above step (5), according to the corner point set point1, point2, ..., point β The position of each corner point in the original image I is determined to determine the coordinates of the upper left corner vertex, the upper right corner vertex, the lower right corner vertex, and the lower left corner vertex of each cell in the original image I, and the cell coordinate set CP is obtained. The specific method is as follows:
[0150] (5.4) Set the threshold th = 3. Traverse the corner point set point1, point2, ..., point β Each subset point k , where k = 1, 2, ..., β, β is the total number of contours in the fourth binary image I4, and the subset point k The first corner point in Find the horizontal coordinate value from the corner point set P3 With corner point The horizontal axis value The difference does not exceed the threshold th, the vertical coordinate value Greater than corner The vertical coordinate value of And away from the corner The distance of the corner point is less than the set distance third threshold δ The distance calculation is shown in formula (10).
[0151] (5.5) Determine the corner point With corner point Is there a connecting line between them? With corner point If there is a connection between them, then take the subset point k The second corner point in Find the horizontal coordinate value from the corner point set P3 With corner point The horizontal axis value The difference does not exceed the threshold th, the vertical coordinate value Greater than corner The vertical coordinate value of And away from the corner The distance of the corner point is less than the set distance third threshold δ If the corner With corner point If there is no connection between them, set k = k + 1. If k = β + 1, the cell positioning algorithm ends, otherwise jump to step (5.4).
[0152] (5.6) Determine the corner point With corner point Is there a connecting line between them? With corner point If there is a line between them, the upper left vertex, upper right vertex, lower right vertex, and lower left vertex of the cell are Store these four vertices in the cell coordinate set CP. With corner point If there is no connection between them, then the subset point is calculated k The number of corner points in ψ, if the subset point k If the number of corner points in ψ ≥ 3, then take the subset point k The third corner point in is taken as the corner point Find the horizontal coordinate value from the corner point set P3 With corner point The horizontal axis value The difference does not exceed the threshold th, the vertical coordinate value Greater than corner The vertical coordinate value of And away from the corner The distance of the corner point is less than the set distance third threshold δ Jump to step (5.6), otherwise take k=k+1, if k=β+1, the cell positioning algorithm ends, otherwise jump to step (5.4).
[0153] (5.7) The corner points From the subset point k If the subset point k There is only one corner point left in the image. At this time, k=k+1. If k=β+1, the cell positioning algorithm ends. Otherwise, jump to step (5.4). The cell positioning result image can be obtained, such as Figure 8 shown.
[0154] The algorithm for determining whether there is a connecting line between two corner points in the above steps (5.5) and (5.6) specifically adopts the following method:
[0155] (5.8) Let the two corner points be p 3h 、p 3g , corner point p 3h The horizontal axis is x p3h , the vertical coordinate is y p3h , corner point p 3g The horizontal axis is x p3g , the vertical coordinate is y p3g , where y p3g >y p3h .
[0156] (5.9) Set the thresholds ξ1 = 3 and ξ2 = 2. Determine a rectangular area rect in the third binary image I3 based on the coordinates of the two corner points and compare x p3h and x p3g The size of x p3h >x p3g , then the horizontal coordinate of the upper left corner of the rectangular area rect is x p3g , the vertical coordinate is y p3h , the horizontal coordinate of the upper right vertex is x p3h , the vertical coordinate is y p3h , the horizontal coordinate of the lower right vertex is x p3h , the vertical coordinate is y p3g , the horizontal coordinate of the lower left vertex is x p3g , the vertical coordinate is y p3g If x p3h =x p3g , then the horizontal coordinate of the upper left corner of the rectangular area rect is x p3g -ξ1, the vertical coordinate is y p3h , the horizontal coordinate of the upper right vertex is x p3h +ξ2, the vertical coordinate is y p3h , the horizontal coordinate of the lower right vertex is x p3h +ξ2, the vertical coordinate is y p3g , the horizontal coordinate of the lower left vertex is x p3g -ξ1, the vertical coordinate is y p3g If x p3h <x p3g , then the horizontal coordinate of the upper left corner of the rectangular area rect is x p3h , the vertical coordinate is y p3h , the horizontal coordinate of the upper right vertex is x p3g , the vertical coordinate is y p3h , the horizontal coordinate of the lower right vertex is x p3g , the vertical coordinate is y p3g , the horizontal coordinate of the lower left vertex is x p3h , the vertical coordinate is y p3g .
[0157] (5.10) Set the critical threshold φ = 0.8. Calculate the number of black pixels Cnt in the rectangular area rect. If , it means that there is a line connecting the two corner points; otherwise, it means that there is no line connecting the two corner points, where w is the width of the rectangular area rect.
[0158] The present invention provides a method for identifying the structure of a deformation table in response to interferences such as background, lighting, and physical deformation in the deformation table. The method can effectively locate the position of cells in an image, has strong anti-interference ability and high accuracy, and has good application prospects.
[0159] The above is a preferred embodiment of the present invention, but the present invention should not be limited to the contents disclosed in the embodiment and the drawings. Therefore, any equivalent or modification completed without departing from the spirit disclosed in the present invention shall fall within the scope of protection of the present invention.
Claims
1. A method for identifying a deformation table structure, characterized in that: The method comprises the following steps: (1) Image preprocessing: performing image enhancement, binarization, and skeleton extraction on the input original image i containing the table to obtain a first binary image I1; (2) Character removal: Most characters in the first binary image I1 are removed by using a character removal algorithm to obtain a second binary image I2; then the number of black pixels in four adjacent rectangular areas of the pixel is further determined, and all characters in the second binary image I2 are removed to obtain a third binary image I3; (3) Corner point location: First, the corner points in the third binary image I3 are detected using a corner point detection algorithm to obtain a first corner point set P1; then, the corner points in the first corner point set P1 are clustered to obtain a second corner point set P2; finally, the corner points in the second corner point set P2 that do not meet the conditions are screened to obtain a corner point set P3 of the original image I; (4) Contour acquisition: Delete pixels with a width of 1 in the horizontal direction of the third binary image I3 to obtain a fourth binary image I4 that retains only the horizontal lines; obtain all contours Con1, Con2, .., Con in the fourth binary image I4. β , where β is the total number of contours in the fourth binary image I4; (5) Cell positioning: Classify all corner points in the corner point set P3 and classify those that belong to the contour in the corner point set P3. Add the corner points to the point set in Get the corner point set point1, point2, ..., point β ; Then according to the corner point set point1, point2, ..., point β The position of each corner point in the original image i is used to determine the coordinates of the upper left corner vertex, the upper right corner vertex, the lower right corner vertex, and the lower left corner vertex of each cell in the original image i, and the cell coordinate set CP is obtained; In the above step (1), the input original image I containing the table is subjected to image enhancement, binarization and skeleton extraction, and the following methods are specifically used: (1.1) The input original image I is enhanced using an image enhancement algorithm to obtain the image enhancement result image I 01 Then, the image binarization algorithm is used to enhance the image I 01 Perform image binarization to obtain the image enhancement result binary image I 02 ; (1.2) Use skeleton extraction algorithm to enhance the binary image I 02 Perform skeleton extraction to obtain the first binary image I1: In the above step (2), a character removal algorithm is used to remove most of the characters in the first binary image I1, and the following method is specifically used: (2.1) Traverse each pixel in the first binary image I1 and extract the connected area Area in the first binary image I1 through the eight-neighborhood adjacency relationship of the pixel j , where j = 1, 2, ..., N, N is the total number of connected areas in the first binary image I1, and all the extracted connected areas are stored in a connected area set Area; (2.2) For each connected area j , by solving the following determinant, calculate the connected area Area j The covariance matrix S j The eigenvalue λ jt , where t = 1, 2, in, Respectively represent the connected area Area j The average value of the horizontal coordinate and the average value of the vertical coordinate of all pixels in x j and j Respectively represent the connected area Area j The horizontal and vertical coordinates of the pixel points in the binary image I1 are j=1, 2, ..., N, where N is the total number of connected regions in the first binary image I1; (2.3) Traverse each connected area in the connected area set Area j , for the connected area Area j The covariance matrix S j The eigenvalue of Make a judgment, where j = 1, 2, ..., N, N is the total number of connected areas in the first binary image I1, if the eigenvalue satisfy Among them, th2, th3, and th4 are the set thresholds, C j Area is the connected area j The number of pixels and connected area j The number of pixels C j satisfy C j <th4 (7) Then the connected area Area j By deleting it from the connected area set Area, the second binary image I2 can be obtained; In the above step (2), the number of black pixels in the four adjacent rectangular areas of the pixel is further determined, and all characters in the second binary image I2 are removed to obtain the third binary image I3. Specifically, the following method is used: (2.4) Traverse each pixel in the second binary image I2 and store the pixel with gray value 0 in the set WP; each pixel po in the set WP α , let its horizontal coordinate be x α , the vertical coordinate is y α , the horizontal coordinate of the upper left corner of the upper adjacent rectangular area Rect1 is x α -λ1, y is the vertical coordinate α -λ2, the horizontal coordinate of the lower right corner is x α +λ1, ordinate is y α ; The horizontal coordinate of the upper left corner of the lower adjacent rectangular area Rect2 is x α -λ1, y is the vertical coordinate α , the horizontal coordinate of the lower right corner is x α +λ1, ordinate is y α +λ2; the horizontal coordinate of the upper left corner of the left adjacent rectangular area Rect3 is x α -λ2, y is the vertical coordinate α -λ1, the horizontal coordinate of the lower right corner is x α , the vertical coordinate is y α +λ1; the horizontal coordinate of the upper left corner of the right adjacent rectangular area Rect4 is x α , the vertical coordinate is y α -λ1, the horizontal coordinate of the lower right corner is x α +λ2, ordinate is y α +λ1, where λ1 and λ2 are the set coordinate distance thresholds; calculate the pixel point po α The number of black pixels in the four adjacent rectangular areas Rect1, Rect2, Rect3, and Rect4 on the top, bottom, left, and right Where α=1,2,...,M, M is the total number of pixels in the set WP; (2.5) For pixel po α The adjacent rectangular area Rect i , where i = 1, 2, 3, 4; if the number of black pixels If it is greater than 1, the adjacent rectangular area Rect i The black pixel value Assign value 1, otherwise 0: (2.6) Calculate pixel point po α The black pixel values in the four adjacent rectangular areas are The sum of C α , the calculation formula is as follows: (2.7) If the pixel point po α C of four adjacent rectangular areas α If the value is 1, the pixel point po α As an endpoint, the pixel point po α Add to the set EP; if the pixel point po α C of four adjacent rectangular areas α The value is 2, then the pixel point po α Consider it as a connection point and place the pixel point po α Add to the set LP; in other cases, the pixel po α Consider it as the intersection point and put the pixel point po α Add to the collection IP; (2.8) The noise points in the sets EP and LP are filtered out to obtain the third binary image I3.
2. A deformation table structure recognition method according to claim 1, characterized in that: In the above step (2.8), the noise points in the sets EP and LP are filtered out by using the following method: (2.8.1) Randomly select a pixel point P from the set EP; (2.8.2) For each pixel point P in the set WP μ , where μ=1,2,...,M, M is the total number of pixels in the set WP; calculate the pixel point P μ If the distance between the pixel point P and the pixel point P is less than the set first distance threshold θ, the pixel point P μ Further judgment: The calculation formula of the distance dis between point p1 and point p2 is as follows: Among them, x1 and y1 are the horizontal and vertical coordinates of point p1, respectively, and x2 and y2 are the horizontal and vertical coordinates of point p2, respectively; a) If the pixel point P μ If it also exists in the set LP, then the pixel point P is deleted from the sets EP and WP, and the pixel point P is μ Store in the collection EP; b) If the pixel point P μ If it also exists in the set IP, then the pixel point P is deleted from the sets EP and WP; c) If the pixel point P μ If it also exists in the set EP, then the pixel point P is deleted from the sets EP and WP, and the pixel point P is μ Remove from collection WP; (2.8.3) If there are still pixels in the set EP, go to step (2.8.1), otherwise the noise filtering algorithm ends.
3. A deformation table structure recognition method according to claim 2, characterized in that: In the above step (3), the corner points in the first corner point set P1 are clustered, and the following method is specifically used: (3.1) Randomly select a point p in the first corner point set P1 1t As the initial point p0, where t∈{1,2,...,n}, n is the total number of corner points in the first corner point set P1; (3.2) Traverse all points in the first corner point set P1 except the initial point p0 where t'∈{1,2,...,n}; calculation point With point The distance between tt' ,point The horizontal axis is The vertical axis is point The horizontal axis is The vertical axis is The distance calculation is shown in formula (10), where the distance dis tt' Corner points less than the set second distance threshold k are stored in the set CLP; (3.3) The corner point p 1t' Stored in the second corner point set P2, and for each point p in the set CLP γ ,in is the total number of corner points in the set CLP, and determines whether point p exists in the first corner point set P1 γ , if there is a point p in the first corner point set P1 γ , then point p γ Delete from the first corner point set P1; (3.4) Set the set CLP to empty; (3.5) If the first corner point set P1 is not empty, go to step (3.1), otherwise the corner point clustering algorithm ends.
4. A deformation table structure recognition method according to claim 3, characterized in that: In the above step (3), the corner points in the second corner point set P2 that do not meet the conditions are screened, and the following method is specifically used: (3.6) Calculate each corner point in the second corner point set P2 in turn The total number of pixels with grayscale value 0 in the four adjacent rectangular areas Rect1, Rect2, Rect3, and Rect4 above, below, left, and right Where c = 1, 2, ..., m, m is the total number of corner points in the second corner point set P2; (3.7) For corner points The adjacent rectangular area Rect i , where i = 1, 2, 3, 4; if the number of black pixels If it is greater than 9, the adjacent rectangular area Rect i The black pixel value Assign value 1, otherwise 0: (3.8) If the corner point of One of the following six situations: a) b) c) d) e) f) Then the corner point Delete from the second corner point set P2.
5. A deformation table structure recognition method according to claim 4, characterized in that: In the above step (5), all corner points in the corner point set P3 are classified, and those belonging to the contour are classified into Add the corner points to the point set in Get the corner point set point1, point2, ..., point β , specifically the following methods were used: (5.1) For each contour, calculate the horizontal coordinate of its contour centroid according to the following formula and the vertical coordinate Among them, f(i,j) is the gray value of the contour at the horizontal coordinate i and the vertical coordinate j, and W and V are the gray values of the contour respectively. The width and height of the bounding rectangle; (5.2) Each contour is divided into two parts according to its centroid ordinate value. Sort from small to large, and the sorted contours are recorded as Con'1, Con'2, ..., Con' β , Contour Con' i The horizontal coordinate of the center of mass is The vertical axis is Where i = 1, 2, ..., β, β is the total number of contours in the fourth binary image I4; for each contour Con' i , if the ordinate y of the corner point k in the corner point set P3 k satisfy Where k∈{1,2,...,Z}, Z is the total number of corner points in the corner point set P3, σ is the set vertical coordinate distance threshold, then the corner point k and the contour Con are calculated respectively. i Distance of the centroid ik , corner point k and contour Con i-1 Distance of the centroid (i-1)k , corner point k and contour Con i+1 Distance of the centroid (i+1)k , where the horizontal coordinate of the corner point k is x k , the vertical coordinate is y k , Contour i The horizontal coordinate of the center of mass is The vertical axis is Contour i-1 The horizontal coordinate of the center of mass is The vertical axis is Contour i+1 The horizontal coordinate of the center of mass is The vertical axis is The distance calculation is shown in formula (10); the distance comparison distance (i-1)k 、distance ik 、distance (i+1)k The size of the distance is selected with the smallest value. l , then the corner point k is on the corresponding contour Con l , the calculation formula is as follows: (5.3) For each contour Con i , will be located on the same contour Con i The corner points on the i , and all the corner points in each point set are sorted by the horizontal coordinate value x of the corner point in the point set τ Sort from small to large, where i = 1, 2, ..., β, β is the total number of contours in the fourth binary image I4, τ = 1, 2, ..., φ, φ is the point set point i The total number of corner points in .
6. A deformation table structure recognition method according to claim 5, characterized in that: In the above step (5), according to the corner point set point1, point2, ..., point β The position of each corner point in the original image I is determined to determine the coordinates of the upper left corner vertex, the upper right corner vertex, the lower right corner vertex, and the lower left corner vertex of each cell in the original image I, and the cell coordinate set CP is obtained. The specific method is as follows: (5.4) Traverse the corner point set point1, point2, ..., point β Each subset point k , where k = 1, 2, ..., β, β is the total number of contours in the fourth binary image I4, and the subset point k The first corner point in Find the horizontal coordinate value from the corner point set P3 With corner point The horizontal axis value The difference does not exceed the threshold th, the vertical coordinate value Greater than corner The vertical coordinate value of And away from the corner The distance of the corner point is less than the set distance third threshold δ The distance calculation is shown in formula (10); (5.5) Determine the corner point With corner point Is there a connecting line between them? If the corner points With corner point If there is a line between them, then take the subset point k The second corner point in Find the horizontal coordinate value from the corner point set P3 With corner point The horizontal axis value The difference does not exceed the threshold th, the vertical coordinate value Greater than corner The vertical coordinate value of And away from the corner The distance of the corner point is less than the set distance third threshold δ If the corner With corner point If there is no connection between them, let k = k + 1. If k = β + 1, the cell location algorithm ends, otherwise jump to step (5.4); (5.6) Determine the corner point With corner point Is there a connecting line between them? If the corner points With corner point If there is a line between them, then the upper left vertex, upper right vertex, lower right vertex, and lower left vertex of the cell are p respectively. k1 、p k2 、p krd 、p kd , store these four vertices in the cell coordinate set CP; if the corner point p k2 With corner point p krd If there is no connection between them, then the subset point is calculated. k The number of corner points in ψ, if the subset point k If the number of corner points in ψ ≥ 3, then take the subset point k The third corner point in is taken as the corner point p k2 , find the horizontal coordinate value x from the corner point set P3 krd With corner point p k2 The horizontal coordinate value x k2 The difference does not exceed the threshold th, the ordinate value y krd Greater than corner point p k2 The vertical coordinate value y k2 And away from the corner point p k2 The corner point p whose distance is less than the set third threshold δ krd , jump to step (5.6), otherwise take k = k + 1, if k = β + 1, the cell positioning algorithm ends, otherwise jump to step (5.4); (5.7) The corner point p k1 From the subset point k If the subset point k There is only one corner point left in the , so take k = k + 1. If k = β + 1, the cell positioning algorithm ends, otherwise jump to step (5.4).
7. A deformation table structure recognition method according to claim 6, characterized in that: The algorithm for determining whether there is a connecting line between two corner points in the above steps (5.5) and (5.6) specifically adopts the following method: (5.8) Let the two corner points be p 3h 、p 3g , corner point p 3h The horizontal axis is x p3h , the vertical coordinate is y p3h , corner point p 3g The horizontal axis is x p3g , the vertical coordinate is y p3g , where y p3g >y p3h ; (5.9) Determine a rectangular area rect in the third binary image I3 based on the coordinates of the two corner points, and compare x p3h and x p3g The size of x p3h >x p3g , then the horizontal coordinate of the upper left corner of the rectangular area rect is x p3g , the vertical coordinate is y p3h , the horizontal coordinate of the upper right vertex is x p3h , the vertical coordinate is y p3h , the horizontal coordinate of the lower right vertex is x p3h , the vertical coordinate is y p3g , the horizontal coordinate of the lower left vertex is x p3g , the vertical coordinate is y p3g ; if x p3h =x p3g , then the horizontal coordinate of the upper left corner of the rectangular area rect is x p3g -ξ1, the vertical coordinate is y p3h , the horizontal coordinate of the upper right vertex is x p3h +ξ2, the vertical coordinate is y p3h , the horizontal coordinate of the lower right vertex is x p3h +ξ2, the vertical coordinate is y p3g , the horizontal coordinate of the lower left vertex is x p3g -ξ1, the vertical coordinate is y p3g ; if x p3h <x p3g , then the horizontal coordinate of the upper left corner of the rectangular area rect is x p3h , the vertical coordinate is y p3h , the horizontal coordinate of the upper right vertex is x p3g , the vertical coordinate is y p3h , the horizontal coordinate of the lower right vertex is x p3g , the vertical coordinate is y p3g , the horizontal coordinate of the lower left vertex is x p3h , the vertical coordinate is y p3g , where ξ1 and ξ2 are the set thresholds; (5.10) Calculate the number of black pixels Cnt in the rectangular area rect; if It means that there is a line connecting the two corner points; otherwise, it means that there is no line connecting the two corner points, where φ is the set critical threshold and w is the width of the rectangular area rect.
Citation Information
Patent Citations
A table structure complement algorithm based on table node recognition
CN109447007A
PDF table structure identification method based on graph attention mechanism
CN110751038A
Pdf table structure identification method based on image identification
CN111144300A
Table structure extraction method
CN111368695A