A table structure recognition method and system
By preprocessing and convolution kernel operations on the table images, combined with LSD line segment detection, identifying and adjusting table lines, the complex background and noise interference problems in roster table structure recognition are solved, and high-precision table structure recognition is achieved.
Patent Information
- Application Number
- CN202510180214.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-02-19
AI Technical Summary
In the digitalization of archives, the table structure identification of the roster is affected by factors such as complex background, noise interference, incomplete lines, resulting in low recognition accuracy.
Through grayscale, shadow removal, sharpening, Gaussian blur and binarization preprocessing, combined with convolution kernel operation and LSD line segment detection, text and noise are removed, table lines are identified and spliced, and the adjustment lines are compensated for missing lines, and the cell position is determined.
It improves the accuracy and efficiency of table recognition, is highly adaptable, can effectively deal with complex backgrounds and noise interference, and ensures the integrity of the table structure.
Smart Images

Figure CN119672745B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision and image processing, and in particular relates to a table structure recognition method and system. Background Art
[0002] In the process of digitalizing archives, it is necessary to process various rosters. These rosters record data in the form of tables. To accurately extract the data from the rosters, it is necessary to first identify the table structure.
[0003] However, due to the wide variety of rosters and the long time span, the processing may be disturbed by the following factors: complex background, the document may contain multiple colors, text interference or shading; incomplete lines, the lines of the table may be broken, blurred, distorted or tilted; noise problems, handwriting, creases or light influences will introduce noise; paper problems, aging of the paper makes the table unclear. Summary of the invention
[0004] Based on this, an embodiment of the present invention provides a table structure recognition method and system, which aims to accurately recognize the table structure in an image, extract cell boundaries, and construct complete table data.
[0005] A first aspect of an embodiment of the present invention provides a table structure recognition method, the method comprising:
[0006] Acquire a first image containing a table, and preprocess the first image to obtain a second image to highlight the line features of the table, wherein the preprocessing includes graying, shadow removal, sharpening, Gaussian blurring, and binarization;
[0007] By controlling the convolution kernel size, a convolution calculation is performed on the second image to remove the text in the second image and retain the lines, thereby obtaining a third image;
[0008] Performing noise removal processing on the third image, and performing LSD line segment detection on the third image after the noise removal processing to determine basic horizontal and vertical lines;
[0009] The basic lines belonging to the same straight line are spliced and the interference lines are removed to obtain the first line;
[0010] The first lines are adjusted by a compensation algorithm to obtain second lines. The adjustment process includes removing lines with non-tabular structures, supplementing unrecognized lines, and merging or splitting horizontal lines. Specifically, starting points of the first horizontal lines and the first vertical lines are obtained respectively, and the starting points of the first horizontal lines and the first vertical lines are compared with each other to determine the number of corresponding first horizontal lines or first vertical lines.
[0011] Determining whether the number of the first horizontal lines or the first vertical lines is greater than a corresponding preset number;
[0012] If it is determined that the number of the first horizontal lines or the first vertical lines is greater than the corresponding preset number, the corresponding first horizontal lines or the first vertical lines are removed;
[0013] Based on the positions of the first horizontal and vertical lines, the missing lines are inferred and added;
[0014] respectively obtaining the spacing between the first horizontal lines and the first vertical lines, and merging the adjacent first lines whose spacing is smaller than a threshold to obtain the second lines;
[0015] According to the second line, the coordinates of each intersection are determined, and according to the coordinates of each intersection, the corresponding cell position is determined to complete the table structure recognition.
[0016] Furthermore, the step of obtaining a first image containing a table and preprocessing the first image to obtain a second image includes:
[0017] Gray-scale the first image to obtain a first sub-image;
[0018] Perform 9 closing operations on the first sub-image using a convolution kernel of size 9×9 to generate a closeMat function;
[0019] Performing a difference operation between the closeMat function and the first sub-image to obtain a second sub-image after removing the shadow, so as to complete the shadow removal;
[0020] According to the second-order derivative operation, the second sub-image is sharpened to obtain a third sub-image;
[0021] Performing Gaussian blur processing on the third sub-image to obtain a fourth sub-image;
[0022] The OTSU algorithm is used to automatically calculate the threshold, and the fourth sub-image is binarized to obtain the second image.
[0023] Furthermore, in the step of performing convolution calculation on the second image by controlling the convolution kernel size, when identifying vertical lines, a convolution kernel with a size of 1×10 is used to perform morphological opening operation on the second image; when identifying horizontal lines, a convolution kernel with a size of 10×1 is used to perform morphological opening operation on the second image.
[0024] Furthermore, the step of performing noise removal processing on the third image includes:
[0025] When extracting vertical lines, a convolution kernel of size 10×1 is used to obtain an image that only retains horizontal lines, and then the convolution kernel of size 1×3 is expanded to amplify the noise to obtain the fifth sub-image;
[0026] performing a difference operation on the fifth sub-image and the third image to obtain a sixth sub-image, so as to remove the extracted horizontal lines from the third image and retain the vertical lines and other feature information;
[0027] Performing an opening operation on the sixth sub-image again using a convolution kernel of size 10×1 to obtain a seventh sub-image to further remove noise;
[0028] BlobCounter is used to analyze the connected domains in the seventh sub-image to obtain an eighth sub-image, wherein for the connected domains that meet the preset conditions, the corresponding connected domains are filled with black in the seventh sub-image, and the preset conditions are that the height range of the connected domains is set to 2≤h≤250, and at the same time, the noise with a height of h≤50 is removed;
[0029] Using a 3×3 Gaussian kernel to perform blur processing on the eighth sub-image to generate a first smoothed image;
[0030] Performing a closing operation on the first smoothed graph using a convolution kernel of size 1×10 to close the disconnected parts of the straight line;
[0031] When extracting horizontal lines, a convolution kernel of size 1×10 is used to obtain an image that only retains vertical lines, and then the convolution kernel of size 3×1 is expanded to amplify the noise to obtain the ninth sub-image;
[0032] performing a difference operation on the ninth sub-image and the third image to obtain a tenth sub-image, so as to remove the extracted longitudinal lines from the third image and retain the transverse lines and other characteristic information;
[0033] The tenth sub-image is again opened by a convolution kernel of size 1×10 to obtain an eleventh sub-image, so as to further remove noise;
[0034] BlobCounter is used to analyze the connected domains in the eleventh sub-image to obtain a twelfth sub-image, wherein for the connected domains that meet the preset conditions, the corresponding connected domains are filled with black in the eleventh sub-image, and the preset conditions are that the height range of the connected domains is set to 2≤h≤250, and at the same time, the noise with a height of h≤50 is removed;
[0035] Using a 3×3 Gaussian kernel to perform blur processing on the twelfth sub-image to generate a second smoothed image;
[0036] A closing operation is performed on the second smoothed image using a convolution kernel of size 10×1 to close the disconnected parts of the straight line.
[0037] Furthermore, in the step of performing LSD line segment detection on the third image after noise removal processing to determine the basic horizontal and vertical lines, the LSD algorithm is used to perform line segment detection on the third image after noise removal processing to extract line segment information, and the line segment information includes the starting point, end point and length of the line segment.
[0038] Furthermore, the step of splicing the basic lines belonging to the same straight line and removing the interfering lines to obtain the first line includes:
[0039] Analyze the start and end positions of adjacent basic lines, determine the distance between the start and end positions of adjacent basic lines, and determine whether the distance is less than a preset distance;
[0040] If it is determined that the distance is less than the preset distance, connecting the start and end points of the adjacent basic lines;
[0041] Obtaining the length of each basic line, and determining whether the length is less than a preset length;
[0042] If it is determined that the length is less than the preset length, the corresponding basic line is deleted;
[0043] According to the difference in coordinate points between the horizontal basic lines and the vertical basic lines, the adjacent basic lines are dynamically searched and merged, wherein the horizontal basic lines are merged and the vertical basic lines are merged;
[0044] Determine whether the horizontal basic lines and the vertical basic lines are aligned;
[0045] If it is determined that the horizontal basic lines and the vertical basic lines are not aligned, a virtual connecting line is inserted to connect the unaligned basic lines to obtain the first line.
[0046] A second aspect of an embodiment of the present invention provides a table structure recognition system, which is used to implement a table structure recognition method provided by the first aspect of an embodiment of the present invention, and the system includes:
[0047] A preprocessing module, used for acquiring a first image containing a table, and preprocessing the first image to obtain a second image to highlight the line features of the table, wherein the preprocessing includes graying, shadow removal, sharpening, Gaussian blurring, and binarization;
[0048] A convolution module, configured to perform convolution calculation on the second image by controlling the size of a convolution kernel, so as to remove text from the second image and retain lines, thereby obtaining a third image;
[0049] A denoising module, used for performing a denoising process on the third image, and performing LSD line segment detection on the third image after the denoising process to determine basic horizontal and vertical lines;
[0050] A splicing module, used for splicing basic lines belonging to the same straight line and removing interfering lines to obtain a first line;
[0051] An adjustment module, configured to adjust the first lines by a compensation algorithm to obtain second lines, wherein the adjustment process includes removing lines with non-tabular structures, supplementing unrecognized lines, and merging or splitting horizontal lines;
[0052] The determination module is used to determine the coordinates of each intersection point according to the second line, and determine the corresponding cell position according to the coordinates of each intersection point to complete the table structure recognition.
[0053] A third aspect of an embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the table structure recognition method provided in the first aspect.
[0054] A fourth aspect of an embodiment of the present invention provides an electronic device, including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor implements the table structure recognition method provided in the first aspect when executing the program.
[0055] A table structure recognition method and system are provided in an embodiment of the present invention. The method obtains a second image by preprocessing a first image containing a table; performs convolution calculation on the second image by controlling the convolution kernel size to remove text in the second image and retain lines to obtain a third image; performs noise removal processing on the third image, and performs LSD line segment detection on the third image after noise removal to determine basic horizontal and vertical lines; splices basic lines belonging to the same straight line and removes interference lines to obtain first lines; adjusts the first lines through a compensation algorithm to obtain second lines; determines the coordinates of each intersection based on the second line, and determines the corresponding cell position based on the coordinates of each intersection to complete table structure recognition. The above method has obvious advantages in accuracy, adaptability and efficiency of table recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 A flowchart of a table structure recognition method provided in the first embodiment of the present invention;
[0057] Figure 2 A structural block diagram of a table structure recognition system provided in Embodiment 2 of the present invention;
[0058] Figure 3 This is a structural block diagram of an electronic device provided in Embodiment 3 of the present invention. DETAILED DESCRIPTION
[0059] In order to facilitate the understanding of the present invention, the present invention will be described more fully below with reference to the relevant drawings. Several embodiments of the present invention are given in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, the purpose of providing these embodiments is to make the disclosure of the present invention more thorough and comprehensive.
[0060] It should be noted that when an element is referred to as being "fixed to" another element, it may be directly on the other element or there may be a central element. When an element is considered to be "connected to" another element, it may be directly connected to the other element or there may be a central element at the same time. The terms "vertical", "horizontal", "left", "right" and similar expressions used herein are for illustrative purposes only.
[0061] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which the present invention belongs. The terms used herein in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.
[0062] Embodiment 1
[0063] See also Figure 1 , Figure 1 A table structure recognition method provided by the first embodiment of the present invention is shown. The table structure recognition method specifically includes steps S01 to S06.
[0064] Step S01, obtaining a first image containing a table, and preprocessing the first image to obtain a second image to highlight the line features of the table, wherein the preprocessing includes graying, shadow removal, sharpening, Gaussian blurring, and binarization.
[0065] Specifically, the first image containing the table can be captured by a shooting device. After the first image is obtained, preprocessing is performed. It should be noted that the first image is firstly grayscaled to obtain a first sub-image. Grayscale can convert the color information of the image into a single-channel brightness value, providing a simplified basic image for subsequent processing;
[0066] The convolution kernel of size 9×9 is used to perform 9 closing operations on the first sub-image to generate the closeMat function, which can be expressed as:
[0067] closeMat(x,y)=9Morph_Close(G,K)
[0068] Where K is the structural element of the closing operation. The closing operation can fill the line breaks and gaps and extract the shadow area;
[0069] The closeMat function is subtracted from the first sub-image to obtain the second sub-image after removing the shadow, so as to complete the shadow removal. The subtraction operation can be expressed as:
[0070] shadowFreeMat(x,y)=Subtract(closeMat,grayImage)
[0071] Among them, grayImage represents the first sub-image, shadowFreeMat(x,y) represents the second sub-image, and the difference operation can eliminate the large shadows retained in the closing operation, retaining only the details and line information in the image, and generating a black background and white text image;
[0072] According to the second-order derivative operation, the second sub-image is sharpened to obtain the third sub-image. Specifically, in order to enhance the edge and line features in the image, the second sub-image after removing the shadow is sharpened, which is expressed as:
[0073] sharpMat(x,y)=shadowFreeMat(x,y)+λLaplacian(shadowFreeMat(x,y))
[0074] Among them, Laplacian is the second-order derivative operation of the image, which is used to enhance the edge, and λ is the sharpening strength coefficient (such as setting λ=0.5);
[0075] The third sub-image is Gaussian blurred to reduce noise interference, and the fourth sub-image is obtained. The Gaussian blurred part can be expressed as:
[0076] blurredMat(x,y)=GaussianBlur(sharpMat(x,y),k=3)
[0077] blurredMat(x,y) represents the fourth sub-image;
[0078] The OTSU algorithm is used to automatically calculate the threshold value, and the fourth sub-image is binarized to obtain the second image. It can be understood that the OTSU algorithm is a global image binarization threshold selection algorithm, also known as the maximum inter-class variance method. Its purpose is to convert a grayscale image into a binary image, and automatically determine an optimal threshold value to divide the pixels in the image into foreground and background.
[0079] It can be understood that there are two second images, one for identifying horizontal lines and one for identifying vertical lines.
[0080] Step S02, performing convolution calculation on the second image by controlling the convolution kernel size to remove the text in the second image and retain the lines to obtain a third image.
[0081] When recognizing vertical lines, a convolution kernel of size 1×10 is used to perform a morphological opening operation on the second image; when recognizing horizontal lines, a convolution kernel of size 10×1 is used to perform a morphological opening operation on the second image, which is expressed as:
[0082] imageOpen=Morph_Open(image,K)
[0083] imageOpen represents the third image.
[0084] Step S03, performing noise removal processing on the third image, and performing LSD line segment detection on the third image after the noise removal processing to determine basic horizontal and vertical lines.
[0085] In this embodiment, in order to remove the noise of the third image, specifically, when extracting the vertical lines, a convolution kernel of size 10×1 is used to obtain an image retaining only the horizontal lines, and then the convolution kernel of size 1×3 is expanded to amplify the noise, thereby obtaining the fifth sub-image;
[0086] Performing a difference operation, i.e., a difference operation, on the fifth sub-image and the third image to obtain a sixth sub-image, so as to remove the extracted horizontal lines from the third image and retain the vertical lines and other feature information;
[0087] The sixth sub-image is opened again by a convolution kernel of size 10×1 to obtain the seventh sub-image to further remove noise;
[0088] BlobCounter is used to analyze the connected domains in the seventh sub-image to obtain the eighth sub-image, wherein for the connected domains that meet the preset conditions, the corresponding connected domains are filled with black in the seventh sub-image. The preset conditions are that the height range of the connected domains is set to 2≤h≤250, and at the same time, the noise with a height of h≤50 is removed;
[0089] The eighth sub-image is blurred using a 3×3 Gaussian kernel to generate a first smoothed image;
[0090] The first smoothed image is closed by using a convolution kernel of size 1×10 to close the disconnected parts of the straight line;
[0091] When extracting horizontal lines, a convolution kernel of size 1×10 is used to obtain an image that only retains vertical lines, and then the convolution kernel of size 3×1 is expanded to amplify the noise to obtain the ninth sub-image;
[0092] performing a difference operation on the ninth sub-image and the third image to obtain a tenth sub-image, so as to remove the extracted longitudinal lines from the third image and retain the transverse lines and other characteristic information;
[0093] The tenth sub-image is again opened by a convolution kernel of size 1×10 to obtain an eleventh sub-image, so as to further remove noise;
[0094] BlobCounter is used to analyze the connected domains in the eleventh sub-image to obtain a twelfth sub-image, wherein for the connected domains that meet the preset conditions, the corresponding connected domains are filled with black in the eleventh sub-image, and the preset conditions are that the height range of the connected domains is set to 2≤h≤250, and at the same time, the noise with a height of h≤50 is removed;
[0095] Using a 3×3 Gaussian kernel to perform blur processing on the twelfth sub-image to generate a second smoothed image;
[0096] A closing operation is performed on the second smoothed image using a convolution kernel of size 10×1 to close the disconnected parts of the straight line.
[0097] Furthermore, the LSD algorithm is used to perform line segment detection on the third image after the noise removal process to extract line segment information, which includes the starting point, end point and length of the line segment. The line segment detection can be expressed as:
[0098] lines2 = LSD_Detect(image).
[0099] Step S04, splicing the basic lines belonging to the same straight line and removing the interfering lines to obtain a first line.
[0100] Specifically, the starting and ending positions of adjacent basic lines are analyzed, and the scattered vertical and horizontal line segments are connected into a continuous vertical and horizontal line to ensure the vertical and horizontal structure of the table is complete, wherein the distance between the starting and ending positions of adjacent basic lines is determined, and it is judged whether the distance is less than a preset distance;
[0101] If the distance is judged to be less than the preset distance, the starting point and the end point of the adjacent basic lines are connected;
[0102] Obtain the length of each basic line and determine whether the length is less than a preset length;
[0103] If the length is judged to be less than the preset length, the corresponding basic line is deleted;
[0104] According to the difference in coordinate points between the horizontal basic lines and the vertical basic lines, the adjacent basic lines are dynamically searched and merged, wherein the horizontal basic lines are merged and the vertical basic lines are merged;
[0105] Determine whether the horizontal basic lines and the vertical basic lines are aligned;
[0106] If it is determined that the horizontal basic lines and the vertical basic lines are not aligned, a virtual connecting line is inserted to connect the unaligned basic lines to obtain a first line.
[0107] Step S05, adjusting the first lines by a compensation algorithm to obtain second lines, wherein the adjustment process includes removing lines without a tabular structure, supplementing unrecognized lines, and merging or splitting horizontal lines.
[0108] It should be noted that the starting points of the first horizontal and vertical lines are obtained respectively, and each of the first horizontal lines and the first vertical lines is taken as an object, and the starting points of the first horizontal lines and the first vertical lines are compared to determine the number of the corresponding first horizontal lines or the first vertical lines;
[0109] Determining whether the number of the first horizontal lines or the first vertical lines is greater than a corresponding preset number;
[0110] If it is determined that the number of the first horizontal lines or the first vertical lines is greater than the corresponding preset number, the corresponding first horizontal lines or the first vertical lines are removed. Specifically, taking the removal of the first vertical lines as an example, that is, removing the vertical lines, for each vertical line, the number of horizontal lines whose starting point X coordinates are less than the maximum starting point X coordinate of all line segments of the vertical line by a preset threshold ΔX is calculated. If the number of horizontal lines that meet the condition exceeds α (for example, 75%) of the total number of horizontal lines, it is inferred that the vertical line is close to the left edge and should be removed.
[0111] According to the position of the first horizontal and vertical lines, the missing lines are estimated and added. It is understandable that some vertical or horizontal lines are missing because the lines at the edge of the image may not be correctly detected. Therefore, based on the existing horizontal and vertical line positions, the missing lines are estimated and added. Taking the estimation of vertical lines as an example, the steps are as follows:
[0112] Step 1: Sort the vertical lines. Sort the vertical lines by the X coordinate of the starting point, making sure that the vertical lines are arranged from left to right.
[0113] Step 2: Initialize the inferred variables. Initialize the variables commonStartX and commonEndX to record the inferred leftmost and rightmost vertical line positions.
[0114] Step 3: Set the threshold. Set the distance threshold ΔX and the count threshold α (such as 0.7) to determine the degree of proximity between the horizontal line and the vertical line, and thus infer the position of the vertical line.
[0115] Step 4: Estimate the leftmost vertical line. For each horizontal line, calculate the number of horizontal lines whose X coordinate difference with the starting point of the current horizontal line is less than ΔX. If this number exceeds α of the total number of horizontal lines, it is estimated that the starting point X of the current horizontal line may be the correct position of the leftmost vertical line.
[0116] Step 5: Add missing vertical lines. If the position of the leftmost or rightmost vertical line is inferred and there are no vertical lines around it, add a new vertical line by copying the nearest vertical line and adjusting its position;
[0117] The spacing between the first horizontal lines and the first vertical lines is obtained respectively, and the adjacent first lines whose spacing is less than the threshold are merged to obtain the second lines. It should be noted that the purpose of this step is to make the table structure more concise and accurate by merging the horizontal lines (rows) with smaller heights in the process of table recognition when it is known that the heights of each row of the table are close. It merges the adjacent horizontal lines with smaller heights by judging the height difference between the horizontal lines, and inserts new horizontal lines when necessary to maintain the integrity of the table structure. The specific steps include the following:
[0118] Sort the vertical lines by the X coordinates of the starting points, ensuring that the vertical lines are arranged from left to right;
[0119] Calculate the middleX in the middle, which is used to determine whether the horizontal lines need to be merged later;
[0120] Sort the horizontal lines by the Y coordinate of the starting point, making sure to process the horizontal lines from top to bottom;
[0121] Calculate the height difference Height of each horizontal line, that is, the vertical distance between two adjacent horizontal lines;
[0122] By traversing all horizontal lines, find the most common horizontal line height commonHeight and set it as the standard height for merging. If no suitable height is found, skip the merging operation;
[0123] By traversing the horizontal lines, check whether each pair of adjacent horizontal lines needs to be merged. If the height difference between the two horizontal lines is less than commonHeight, and the height of one of the horizontal lines is less than half of the common height, try to merge them;
[0124] If the merge condition is met, merge the two horizontal lines.
[0125] Step S06, determining the coordinates of each intersection point according to the second line, and determining the corresponding cell position according to the coordinates of each intersection point to complete the table structure recognition.
[0126] In summary, the table structure recognition method in the above embodiment of the present invention obtains a second image by preprocessing a first image containing a table; performs convolution calculation on the second image by controlling the convolution kernel size to remove text from the second image and retain the lines to obtain a third image; performs noise removal processing on the third image, and performs LSD line segment detection on the third image after noise removal to determine basic horizontal and vertical lines; splices basic lines belonging to the same straight line, and removes interference lines to obtain first lines; adjusts the first lines through a compensation algorithm to obtain second lines; determines the coordinates of each intersection based on the second line, and determines the corresponding cell position based on the coordinates of each intersection to complete table structure recognition. The above method has obvious advantages in accuracy, adaptability and efficiency of table recognition.
[0127] Embodiment 2
[0128] See also Figure 2 , Figure 2 : is a structural block diagram of a table structure recognition system 200 provided in the second embodiment of the present invention. The table structure recognition system 200 specifically includes: a preprocessing module 21, a convolution module 22, a denoising module 23, a splicing module 24, an adjustment module 25 and a determination module 26, wherein:
[0129] A preprocessing module 21 is used to obtain a first image containing a table, and preprocess the first image to obtain a second image to highlight the line features of the table, wherein the preprocessing includes graying, shadow removal, sharpening, Gaussian blurring, and binarization;
[0130] A convolution module 22 is used to perform a convolution calculation on the second image by controlling the size of a convolution kernel to remove text from the second image and retain lines to obtain a third image, and perform a morphological opening operation on the second image using a convolution kernel of a size of 1×10;
[0131] A denoising module 23 is used to perform denoising on the third image, and perform LSD line segment detection on the third image after the denoising to determine basic horizontal and vertical lines, wherein the line segment detection is performed on the third image after the denoising using an LSD algorithm to extract line segment information, wherein the line segment information includes a starting point, an end point and a length of the line segment;
[0132] A splicing module 24 is used to splice the basic lines belonging to the same straight line and remove the interfering lines to obtain a first line;
[0133] An adjustment module 25, configured to adjust the first lines by a compensation algorithm to obtain second lines, wherein the adjustment process includes removing lines with non-tabular structures, supplementing unrecognized lines, and merging or splitting horizontal lines;
[0134] The determination module 26 is used to determine the coordinates of each intersection point according to the second line, and determine the corresponding cell position according to the coordinates of each intersection point to complete the table structure recognition.
[0135] Further, in some optional embodiments of the present invention, the preprocessing module 21 includes:
[0136] A grayscale processing unit, used for performing grayscale processing on the first image to obtain a first sub-image;
[0137] A first closing operation unit, used for performing nine closing operations on the first sub-image by using a convolution kernel of a size of 9×9 to generate a closeMat function;
[0138] A difference operation unit, used for performing a difference operation between the closeMat function and the first sub-image to obtain a second sub-image after removing the shadow, so as to complete the shadow removal;
[0139] a sharpening processing unit, configured to perform a sharpening process on the second sub-image according to a second-order derivative operation to obtain a third sub-image;
[0140] A Gaussian blur processing unit, configured to perform Gaussian blur processing on the third sub-image to obtain a fourth sub-image;
[0141] The binarization processing unit is used to automatically calculate the threshold value by using the OTSU algorithm, and perform binarization processing on the fourth sub-image to obtain the second image.
[0142] Furthermore, in some optional embodiments of the present invention, the denoising module 23 includes:
[0143] a first convolution processing unit, for obtaining an image retaining only horizontal lines by using a convolution kernel of size 10×1 when extracting vertical lines, and then expanding the image by using a convolution kernel of size 1×3 to amplify noise, thereby obtaining a fifth sub-image;
[0144] a first differential operation unit, configured to perform a differential operation on the fifth sub-image and the third image to obtain a sixth sub-image, so as to remove the extracted horizontal lines from the third image and retain the vertical lines and other characteristic information;
[0145] a first opening operation unit, configured to perform an opening operation on the sixth sub-image again through a convolution kernel having a size of 10×1 to obtain a seventh sub-image, so as to further remove noise;
[0146] a first analyzing unit, configured to analyze the connected domain in the seventh sub-image by using BlobCounter to obtain an eighth sub-image, wherein for a connected domain that meets a preset condition, the corresponding connected domain is filled with black in the seventh sub-image, and the preset condition is that a height range of the connected domain is set to 2≤h≤250, and at the same time, noise with a height of h≤50 is removed;
[0147] A first fuzzy processing unit, configured to perform fuzzy processing on the eighth sub-image by using a 3×3 Gaussian kernel to generate a first smoothed image;
[0148] A second closing operation unit, configured to perform a closing operation on the first smoothed image by using a convolution kernel of a size of 1×10, so as to close the disconnected parts of the straight line;
[0149] a second convolution processing unit, for obtaining an image retaining only the vertical lines by using a convolution kernel of size 1×10 when extracting the horizontal lines, and then dilating the image by using a convolution kernel of size 3×1 to amplify the noise, thereby obtaining a ninth sub-image;
[0150] a second differential operation unit, configured to perform a differential operation on the ninth sub-image and the third image to obtain a tenth sub-image, so as to remove the extracted longitudinal lines from the third image and retain the transverse lines and other characteristic information;
[0151] a second opening operation unit, configured to perform an opening operation on the tenth sub-image again through a convolution kernel of a size of 1×10 to obtain an eleventh sub-image, so as to further remove noise;
[0152] a second analysis unit, configured to analyze the connected domains in the eleventh sub-image by using BlobCounter to obtain a twelfth sub-image, wherein for connected domains that meet a preset condition, the corresponding connected domains are filled with black in the eleventh sub-image, and the preset condition is that the height range of the connected domains is set to 2≤h≤250, and at the same time, noise with a height of h≤50 is removed;
[0153] A second fuzzy processing unit, configured to perform fuzzy processing on the twelfth sub-image by using a 3×3 Gaussian kernel to generate a second smoothed image;
[0154] The third closing operation unit is used to perform a closing operation on the second smooth image through a convolution kernel with a size of 10×1, so as to close the disconnected parts of the straight line.
[0155] Furthermore, in some optional embodiments of the present invention, the splicing module 24 includes:
[0156] A first judging unit, configured to analyze the start and end positions of adjacent basic lines, determine the distance between the start and end positions of adjacent basic lines, and judge whether the distance is less than a preset distance;
[0157] A first connecting unit, configured to connect the start points and end points of adjacent basic lines if it is determined that the distance is less than a preset distance;
[0158] A second judging unit, used for obtaining the length of each basic line and judging whether the length is less than a preset length;
[0159] A deleting unit, configured to delete the corresponding basic line if it is determined that the length is less than a preset length;
[0160] A first merging unit, used for dynamically searching for adjacent basic lines and merging them according to the difference in coordinate points between the horizontal basic lines and the vertical basic lines, wherein the horizontal basic lines are merged with each other and the vertical basic lines are merged with each other;
[0161] A third judging unit is used to judge whether the horizontal basic lines and the vertical basic lines are aligned;
[0162] The second connecting unit is used for inserting a virtual connecting line to connect the misaligned basic lines if it is determined that the horizontal basic lines and the vertical basic lines are misaligned, so as to obtain the first lines.
[0163] Further, in some optional embodiments of the present invention, the adjustment module 25 includes:
[0164] a comparison unit, configured to obtain the starting points of the first horizontal and vertical lines respectively, and compare the starting points of the first horizontal and vertical lines with each other, so as to determine the number of the corresponding first horizontal lines or first vertical lines;
[0165] A fourth judging unit, used to judge whether the number of the first horizontal lines or the first vertical lines is greater than a corresponding preset number;
[0166] a removing unit, configured to remove the corresponding first horizontal lines or first vertical lines if it is determined that the number of the first horizontal lines or the first vertical lines is greater than the corresponding preset number;
[0167] An adding unit, used for inferring and adding missing lines according to positions of the first horizontal and vertical lines;
[0168] The second merging unit is used to respectively obtain the spacing between the horizontal first lines and the vertical first lines, and merge the adjacent first lines whose spacing is smaller than a threshold to obtain the second lines.
[0169] Embodiment 3
[0170] Another aspect of the present invention provides an electronic device, see Figure 3 , shown is an electronic device in Embodiment 3 of the present invention, comprising a memory 20, a processor 10, and a computer program 30 stored in the memory and executable on the processor, wherein the processor 10 implements the table structure recognition method as described above when executing the computer program 30.
[0171] In some embodiments, the processor 10 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor or other data processing chip, used to run program codes or process data stored in the memory 20, such as executing access restriction programs.
[0172] The memory 20 includes at least one type of readable storage medium, and the readable storage medium includes a flash memory, a hard disk, a multimedia card, a card-type memory (for example, an SD or DX memory, etc.), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the memory 20 may be an internal storage unit of an electronic device, such as a hard disk of the electronic device. In other embodiments, the memory 20 may also be an external storage device of an electronic device, such as a plug-in hard disk equipped on the electronic device, a smart memory card (SmartMediaCard, SMC), a secure digital (SecureDigital, SD) card, a flash card (FlashCard), etc. Further, the memory 20 may also include both an internal storage unit and an external storage device of the electronic device. The memory 20 may be used not only to store application software and various types of data of the electronic device, but also to temporarily store data that has been output or is to be output.
[0173] It should be pointed out that Figure 3 The structure shown does not constitute a limitation on the electronic device. In other embodiments, the electronic device may include fewer or more components than those shown in the figure, or combine certain components, or arrange the components differently.
[0174] The embodiment of the present invention further provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the table structure recognition method as described above is implemented.
[0175] Those skilled in the art will appreciate that the logic and / or steps represented in the flowchart or otherwise described herein, for example, may be considered as an ordered list of executable instructions for implementing logical functions, and may be specifically implemented in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in conjunction with such instruction execution systems, devices or apparatuses. For purposes of this specification, "computer-readable medium" may be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in conjunction with such instruction execution systems, devices or apparatuses.
[0176] More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic device), a portable computer disk case (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be a paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering or, if necessary, processing in another suitable manner, and then stored in a computer memory.
[0177] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or a combination thereof: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0178] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0179] The above embodiments only express several implementation methods of the present invention, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present invention. It should be pointed out that, for a person of ordinary skill in the art, several modifications and improvements can be made without departing from the concept of the present invention, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention patent shall be subject to the attached claims.
Claims
1. A table structure recognition method, characterized in that: The method comprises: Acquire a first image containing a table, and preprocess the first image to obtain a second image to highlight the line features of the table, wherein the preprocessing includes graying, shadow removal, sharpening, Gaussian blurring, and binarization; By controlling the convolution kernel size, a convolution calculation is performed on the second image to remove the text in the second image and retain the lines, thereby obtaining a third image; Performing noise removal processing on the third image, and performing LSD line segment detection on the third image after the noise removal processing to determine basic horizontal and vertical lines; The basic lines belonging to the same straight line are spliced and the interference lines are removed to obtain the first line; The first lines are adjusted by a compensation algorithm to obtain second lines. The adjustment process includes removing lines with non-tabular structures, supplementing unrecognized lines, and merging or splitting horizontal lines. Specifically, starting points of the first horizontal lines and the first vertical lines are obtained respectively, and the starting points of the first horizontal lines and the first vertical lines are compared with each other to determine the number of corresponding first horizontal lines or first vertical lines. Determining whether the number of the first horizontal lines or the first vertical lines is greater than a corresponding preset number; If it is determined that the number of the first horizontal lines or the first vertical lines is greater than the corresponding preset number, the corresponding first horizontal lines or the first vertical lines are removed; Based on the positions of the first horizontal and vertical lines, the missing lines are inferred and added; respectively obtaining the spacing between the first horizontal lines and the first vertical lines, and merging the adjacent first lines whose spacing is smaller than a threshold to obtain the second lines; According to the second line, the coordinates of each intersection are determined, and according to the coordinates of each intersection, the corresponding cell position is determined to complete the table structure recognition.
2. The table structure recognition method according to claim 1, characterized in that: The step of acquiring a first image containing a table and preprocessing the first image to obtain a second image comprises: Gray-scale the first image to obtain a first sub-image; Perform 9 closing operations on the first sub-image using a convolution kernel of size 9×9 to generate a closeMat function; Performing a difference operation between the closeMat function and the first sub-image to obtain a second sub-image after removing the shadow, so as to complete the shadow removal; According to the second-order derivative operation, the second sub-image is sharpened to obtain a third sub-image; Performing Gaussian blur processing on the third sub-image to obtain a fourth sub-image; The OTSU algorithm is used to automatically calculate the threshold, and the fourth sub-image is binarized to obtain the second image.
3. The table structure recognition method according to claim 2, characterized in that: In the step of performing convolution calculation on the second image by controlling the convolution kernel size, when identifying vertical lines, a convolution kernel with a size of 1×10 is used to perform morphological opening operation on the second image; when identifying horizontal lines, a convolution kernel with a size of 10×1 is used to perform morphological opening operation on the second image.
4. The table structure recognition method according to claim 3, characterized in that: The step of performing noise removal processing on the third image comprises: When extracting vertical lines, a convolution kernel of size 10×1 is used to obtain an image that only retains horizontal lines, and then the convolution kernel of size 1×3 is expanded to amplify the noise to obtain the fifth sub-image; Performing a difference operation on the fifth sub-image and the third image to obtain a sixth sub-image, so as to remove the extracted horizontal lines from the third image and retain the vertical lines and other feature information; Performing an opening operation on the sixth sub-image again using a convolution kernel of size 10×1 to obtain a seventh sub-image to further remove noise; BlobCounter is used to analyze the connected domains in the seventh sub-image to obtain an eighth sub-image, wherein for the connected domains that meet the preset conditions, the corresponding connected domains are filled with black in the seventh sub-image, and the preset conditions are that the height range of the connected domains is set to 2≤h≤250, and at the same time, the noise with a height of h≤50 is removed; Using a 3×3 Gaussian kernel to perform blur processing on the eighth sub-image to generate a first smoothed image; Performing a closing operation on the first smoothed graph using a convolution kernel of size 1×10 to close the disconnected parts of the straight line; When extracting horizontal lines, a convolution kernel of size 1×10 is used to obtain an image that only retains vertical lines, and then the convolution kernel of size 3×1 is expanded to amplify the noise to obtain the ninth sub-image; performing a difference operation on the ninth sub-image and the third image to obtain a tenth sub-image, so as to remove the extracted longitudinal lines from the third image and retain the transverse lines and other characteristic information; The tenth sub-image is again opened by a convolution kernel of size 1×10 to obtain an eleventh sub-image, so as to further remove noise; BlobCounter is used to analyze the connected domains in the eleventh sub-image to obtain a twelfth sub-image, wherein for the connected domains that meet the preset conditions, the corresponding connected domains are filled with black in the eleventh sub-image, and the preset conditions are that the height range of the connected domains is set to 2≤h≤250, and at the same time, the noise with a height of h≤50 is removed; Using a 3×3 Gaussian kernel to perform blur processing on the twelfth sub-image to generate a second smoothed image; A closing operation is performed on the second smoothed image using a convolution kernel of size 10×1 to close the disconnected parts of the straight line.
5. The table structure recognition method according to claim 4, characterized in that: In the step of performing LSD line segment detection on the third image after noise removal to determine the basic horizontal and vertical lines, the LSD algorithm is used to perform line segment detection on the third image after noise removal to extract line segment information, wherein the line segment information includes the starting point, end point and length of the line segment.
6. The table structure recognition method according to claim 5, characterized in that: The step of splicing the basic lines belonging to the same straight line and removing the interfering lines to obtain the first line includes: Analyze the start and end positions of adjacent basic lines, determine the distance between the start and end positions of adjacent basic lines, and determine whether the distance is less than a preset distance; If it is determined that the distance is less than the preset distance, connecting the start and end points of the adjacent basic lines; Obtaining the length of each basic line, and determining whether the length is less than a preset length; If it is determined that the length is less than the preset length, the corresponding basic line is deleted; According to the difference in coordinate points between the horizontal basic lines and the vertical basic lines, the adjacent basic lines are dynamically searched and merged, wherein the horizontal basic lines are merged and the vertical basic lines are merged; Determine whether the horizontal basic lines and the vertical basic lines are aligned; If it is determined that the horizontal basic lines and the vertical basic lines are not aligned, a virtual connecting line is inserted to connect the unaligned basic lines to obtain the first line.
7. A table structure recognition system, characterized in that: Used to implement the table structure recognition method according to any one of claims 1 to 6, the system comprises: A preprocessing module, used for acquiring a first image containing a table, and preprocessing the first image to obtain a second image to highlight the line features of the table, wherein the preprocessing includes graying, shadow removal, sharpening, Gaussian blurring, and binarization; A convolution module, configured to perform convolution calculation on the second image by controlling the size of a convolution kernel, so as to remove text from the second image and retain lines, thereby obtaining a third image; A denoising module, used for performing a denoising process on the third image, and performing LSD line segment detection on the third image after the denoising process to determine basic horizontal and vertical lines; A splicing module, used for splicing basic lines belonging to the same straight line and removing interfering lines to obtain a first line; An adjustment module, configured to adjust the first lines by a compensation algorithm to obtain second lines, wherein the adjustment process includes removing lines with non-tabular structures, supplementing unrecognized lines, and merging or splitting horizontal lines; The determination module is used to determine the coordinates of each intersection point according to the second line, and determine the corresponding cell position according to the coordinates of each intersection point to complete the table structure recognition.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the table structure recognition method as described in any one of claims 1 to 6 is implemented.
9. An electronic device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and running on the processor, wherein when the processor executes the program, the table structure recognition method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Table detection method and device
CN109858325A
Topic content identification method and device, readable storage medium and computer equipment
CN110956173A