Table identification method and device

The table images are preprocessed and corrected through object detection and key point detection models, which solves the problem of cross-row and cross-column cell recognition in complex table recognition, improves the accuracy of table recognition and the robustness of the system, and is suitable for OCR content recognition and other business processing.

CN120236295APending Publication Date: 2025-07-01SINOPEC SHARED SERVICES CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510175758.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

Existing table recognition technologies have problems such as low recognition efficiency, poor adaptability and limited ability to recognize complex tables, especially in cross-row and cross-column cell recognition, and are susceptible to shooting angles.

Method used

The table image is preprocessed through the pre-trained object detection model, and the key point detection model is used to identify the outer contour of the table and perform correction, rotation and perspective transformation. Combined with row and column recognition conversion technology, cross-row and column cells are identified and merged to form a target table structure.

Benefits of technology

It improves the accuracy of table recognition, adapts to table inputs of different resolutions and quality, enhances the robustness of the system, and supports OCR content recognition and other business processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236295A_ABST
    Figure CN120236295A_ABST
Patent Text Reader

Abstract

The invention discloses a table recognition method and device, and the method comprises the steps: carrying out the preprocessing of a to-be-recognized table image through a pre-trained target detection model, and obtaining a target table image; processing the target table image through a pre-trained key point detection model to obtain a scanning body table; and performing row-column identification conversion on the scanning body table to obtain a target table structure. According to the method, the to-be-recognized table image is preprocessed through the target detection model; and processing the target table image through a pre-trained key point detection model to obtain a scanning body table. And carrying out row-column identification conversion on the scanning body table to obtain a target table structure. The image quality problem caused by an improper shooting angle is improved, so that the table image is more regular and clearer. By using a plurality of convolutional neural network models and an image preprocessing technology, the table recognition accuracy is remarkably improved when a complex or irregular table is processed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a table recognition method and device. Background Art

[0002] As an important form of information recording and display, tables are widely used in various documents and databases. Traditional table recognition technology mainly relies on a series of predefined rule logic or template matching, which has problems such as low recognition efficiency, poor adaptability and limited ability to recognize complex tables.

[0003] With the rapid development of deep learning technology and the widespread application of convolutional neural networks (CNN) and recurrent neural networks (RNN), the efficiency of table recognition has been significantly improved, but the accuracy of table recognition is still low.

[0004] At present, traditional table recognition is divided into two categories. One is to recognize each cell in the table. This recognition method requires a large sample size and has a low accuracy rate. The other is to recognize rows and columns separately and then combine them, but this method cannot recognize cells with cross-row and cross-column attributes. At the same time, the above methods all have the problem of poor recognition effect due to the shooting angle.

[0005] This section is intended to provide a background or context to the embodiments of the invention recited in the claims. No admission is made that the description herein is prior art by inclusion in this section. Summary of the invention

[0006] An embodiment of the present invention provides a table recognition method for improving the accuracy of table recognition and accurately identifying cells across rows and columns. The method includes:

[0007] Preprocessing the table image to be identified by using a pre-trained target detection model to obtain a target table image;

[0008] Processing the target table image by a pre-trained key point detection model to obtain a scanned volume table;

[0009] Perform row and column recognition conversion on the scanned body table to obtain a target table structure.

[0010] Furthermore, the target table image is processed by a pre-trained key point detection model to obtain a scanned body table, including:

[0011] Identify the outer contour endpoint coordinates of the target table image, and form the outer contour edge line of the target table image according to the outer contour endpoint coordinates; wherein the outer contour edge line includes a first contour edge line, a second contour edge line, a third contour edge line and a fourth contour edge line;

[0012] Rectify the target table image according to the relationships between the first contour line, the second contour line, the third contour line, and the fourth contour line and a preset reference line, to obtain a scanned body table; wherein, the preset reference line includes a preset horizontal line and a preset vertical line.

[0013] Further, the rectifying the target table image according to the relationships between the first contour line, the second contour line, the third contour line, and the fourth contour line and a preset reference line, to obtain a scanned body table includes:

[0014] Determine whether the first contour line is parallel to the preset horizontal line;

[0015] If the first contour line is not parallel to the horizontal line, determine the included angle between the first contour line and the horizontal line;

[0016] Rotate the target table image by the angle of the included angle;

[0017] Determine whether the third contour line and the fourth contour line are parallel to the preset vertical line;

[0018] If the third contour line and the fourth contour line are not parallel to the vertical line, perform a perspective transformation on the rotated target table image to obtain a scanned body table.

[0019] Further, the determining the included angle between the first contour line and the horizontal line includes:

[0020] Determine the slope of the first contour line according to the slope algorithm;

[0021] Calculate the included angle according to the arctangent function and the slope.

[0022] Further, the performing row and column recognition and conversion on the scanned body table to obtain a target table structure includes:

[0023] Recognize each row and each column of the scanned body table;

[0024] Determine the positions of each row and each column of the scanned body table and their corresponding intersection points;

[0025] Determine the coordinates of each cell according to the positions of each row and each column to obtain regular cells;

[0026] Merge the regular cells in the scanned body table to obtain a target table structure.

[0027] Further, the merging the regular cells in the scanned body table to obtain a target table structure includes:

[0028] Identify cells that span multiple rows or columns in the scanned body table to obtain cells that span rows and columns;

[0029] Compare the cells that span rows and columns with the regular cells;

[0030] If the cells that span rows and columns contain multiple regular cells, merge the multiple regular cells into one cell that spans rows and columns;

[0031] Obtain the target table structure based on the regular cells and the cells that span rows and columns.

[0032] Furthermore, the training process of the target detection model includes:

[0033] Collect an image dataset containing tables;

[0034] Perform data annotation on the table images in the image dataset;

[0035] Perform image preprocessing on the table images in the image dataset;

[0036] Set the training parameters of the initial target detection model;

[0037] Train the initial target detection model according to the image dataset and the training parameters to obtain the trained target detection model.

[0038] An embodiment of the present invention further provides a table recognition device for improving the accuracy of table recognition and accurately recognizing cells that span rows and columns. The device includes:

[0039] An image preprocessing module for preprocessing the table image to be recognized through a pre-trained target detection model to obtain a target table image;

[0040] A scanned body module for processing the target table image through a pre-trained key point detection model to obtain a scanned body table;

[0041] A row and column recognition module for performing row and column recognition conversion on the scanned body table to obtain a target table structure.

[0042] Furthermore, the scanned body module includes:

[0043] A contour processing unit for identifying the coordinates of the outer contour endpoints of the target table image and forming the outer contour border of the target table image according to the outer contour endpoint coordinates; wherein, the outer contour border includes a first contour border, a second contour border, a third contour border, and a fourth contour border;

[0044] An image correction unit, configured to correct the target table image according to the relationships between the first contour edge line, the second contour edge line, the third contour edge line, and the fourth contour edge line and a preset reference line, so as to obtain a scanned body table; wherein, the preset reference line includes a preset horizontal line and a preset vertical line.

[0045] Further, the image correction unit includes:

[0046] A first judgment sub-unit, configured to judge whether the first contour edge line is parallel to the preset horizontal line;

[0047] An included angle determination sub-unit, configured to determine the included angle between the first contour edge line and the horizontal line if the first contour edge line is not parallel to the horizontal line;

[0048] An image rotation sub-unit, configured to rotate the target table image by the angle of the included angle;

[0049] A second judgment sub-unit, configured to judge whether the third contour line and the fourth contour line are parallel to the preset vertical line;

[0050] A perspective transformation sub-unit, configured to perform perspective transformation on the rotated target table image to obtain a scanned body table if the third contour line and the fourth contour line are not parallel to the vertical line.

[0051] Further, the included angle determination sub-unit includes:

[0052] A slope determination sub-unit, configured to determine the slope of the first contour edge line according to a slope algorithm;

[0053] An included angle calculation sub-unit, configured to calculate the included angle according to the arctangent function and the slope.

[0054] Further, the row and column recognition module includes:

[0055] A row and column recognition sub-unit, configured to recognize each row and each column of the scanned body table;

[0056] A row and column position determination sub-unit, configured to determine the positions of each row and each column of the scanned body table and their corresponding intersection points;

[0057] A cell determination sub-unit, configured to determine the coordinates of each cell according to the positions of each row and each column to obtain regular cells;

[0058] A table structure determination sub-unit, configured to merge the regular cells in the scanned body table to obtain a target table structure.

[0059] Further, the table structure determination sub-unit includes:

[0060] A cell recognition subunit, configured to recognize cells that span multiple rows or columns in the scanned body table, so as to obtain cells that span rows and columns;

[0061] A cell comparison subunit, configured to compare the cells that span rows and columns with the regular cells;

[0062] A cell merging subunit, configured to, if the cells that span rows and columns contain multiple regular cells, merge the multiple regular cells into one cell that spans rows and columns;

[0063] A target table structure determination subunit, configured to obtain a target table structure according to the regular cells and the cells that span rows and columns.

[0064] Further, the image preprocessing module includes:

[0065] An image collection unit, configured to collect an image dataset containing a table;

[0066] An image processing unit, configured to perform image preprocessing on the table images in the image dataset;

[0067] A data annotation unit, configured to perform data annotation on the table images in the image dataset;

[0068] A training parameter setting unit, configured to set training parameters of an initial target detection model;

[0069] A model training unit, configured to train the initial target detection model according to the image dataset and the training parameters, so as to obtain a trained target detection model.

[0070] An embodiment of the present invention further provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the above-mentioned table recognition method is implemented.

[0071] An embodiment of the present invention further provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned table recognition method is implemented.

[0072] An embodiment of the present invention further provides a computer program product, where the computer program product includes a computer program, and when the computer program is executed by a processor, the above-mentioned table recognition method is implemented.

[0073] A table recognition method and device provided by an embodiment of the present invention preprocess a table image to be recognized through a pre-trained object detection model to obtain a target table image. Then, the target table image is processed through a pre-trained key point detection model to obtain a scanned table. And the scanned table is subjected to row and column recognition conversion to obtain a target table structure. Through image correction techniques such as rotation and perspective transformation, the image quality problem caused by improper shooting angles is improved, making the table image more regular and clear. By using a variety of convolutional neural network models and image preprocessing techniques, the table recognition accuracy in processing complex or irregular tables is significantly improved. The present invention can adapt to table image inputs with different resolutions and qualities, improving the robustness of the system. The recognized table structure can be directly used for OCR content recognition or other business processes, having good flexibility and scalability. BRIEF DESCRIPTION OF THE DRAWINGS

[0074] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. In the drawings:

[0075] Figure 1 It is a schematic flowchart of the table recognition method in an embodiment of the present invention;

[0076] Figure 2 It is a schematic diagram of table image cropping in an embodiment of the present invention;

[0077] Figure 3 It is a schematic flowchart of the table recognition method in another embodiment of the present invention;

[0078] Figure 4 It is a schematic diagram of table image rotation and perspective transformation in an embodiment of the present invention;

[0079] Figure 5 It is a schematic flowchart of the table recognition method in another embodiment of the present invention;

[0080] Figure 6 It is a schematic flowchart of the table recognition method in another embodiment of the present invention;

[0081] Figure 7 It is a schematic flowchart of the table recognition method in another embodiment of the present invention;

[0082] Figure 8 It is a schematic diagram of merging cells across rows and columns in a table image in an embodiment of the present invention;

[0083] Figure 9 Schematic flowchart of the table recognition method in another embodiment of the present invention;

[0084] Figure 10 Schematic flowchart of the table recognition method in another embodiment of the present invention;

[0085] Figure 11 Schematic structural diagram of the table recognition device in another embodiment of the present invention;

[0086] Figure 12 Schematic diagram of the physical structure of the electronic device provided by the embodiment of the present invention. Detailed implementation manners

[0087] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer and more understandable, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings. Herein, the illustrative embodiments of the present invention and their descriptions are used to explain the present invention, but do not limit the present invention.

[0088] In the technical solutions of the present application, the information collected is information and data authorized by the user or fully authorized by all parties. Moreover, the processing of the relevant data, such as collection, storage, use, processing, transmission, provision, disclosure, and application, all comply with the relevant laws, regulations, and standards of the relevant countries and regions, take necessary confidentiality measures, do not violate public order and good customs, and provide corresponding operation entrances for the user to choose to authorize or refuse.

[0089] It should be noted that in the embodiments of the present invention, some existing solutions in the industry, such as certain software, components, models, etc., may be mentioned. They should be regarded as exemplary, and their purpose is only to illustrate the feasibility in the implementation of the technical solutions of the present invention, but it does not mean that the applicant has already or necessarily used this solution.

[0090] In order to improve the accuracy of table recognition and recognize cells that span rows and columns, the present invention provides a table recognition method.

[0091] Figure 1 Schematic flowchart of the table recognition method in the embodiment of the present invention, as Figure 1 shown, the table recognition method includes steps 101 to 103.

[0092] Step 101: Preprocess the table image to be recognized through a pre-trained object detection model to obtain a target table image.

[0093] Step 102: Process the target table image through a pre-trained key point detection model to obtain a scanned table.

[0094] Step 103: Perform row and column recognition conversion on the scanned table to obtain a target table structure.

[0095] From Figure 1 As can be seen from the process of Figure 1 , in the embodiments of the present invention, the table image to be recognized is preprocessed by a pre-trained object detection model to obtain a target table image. Then, the target table image is processed by a pre-trained key point detection model to obtain a scanned table. And the rows and columns of the scanned table are recognized and converted to obtain a target table structure. Through image correction techniques such as rotation and perspective transformation, the image quality problem caused by improper shooting angles is improved, making the table image more regular and clear. By using a variety of convolutional neural network models and image preprocessing techniques, the table recognition accuracy in processing complex or irregular tables is significantly improved. The present invention can adapt to table image inputs with different resolutions and qualities, improving the robustness of the system. The recognized table structure can be directly used for OCR content recognition or other business processes, with good flexibility and scalability.

[0096] Such as Figure 1 shown below, each step will be explained in detail.

[0097] Such as Figure 1 and Figure 2 shown, Step 101: The table image to be recognized is preprocessed by a pre-trained object detection model to obtain a target table image.

[0098] Specifically, after obtaining the image containing the table, the non-table part of the image is removed, improving the accuracy of subsequent recognition of the table area. The table image to be recognized is cropped by a pre-trained object detection model. The object detection model identifies the boundary of the table in the table image and crops it out to generate a target table image containing only the table.

[0099] In one embodiment, the present invention trains the object detection model based on deep learning technology. The object detection model can identify specific targets in the table image, that is, the image of only the table part, by learning a large amount of table image data. During the training process of the object detection model, by identifying the preset table features (such as table shape, table edge, and table texture, etc.), the object detection model can accurately locate the table area in the table image.

[0100] In one embodiment, such as Figure 3 shown, the training process of the object detection model includes Step 301 to Step 305.

[0101] Step 301: Collect an image dataset containing tables.

[0102] In one embodiment, the image dataset contains tabular images of different types and complexities. These datasets can be sourced from open-source datasets or tabular images collected by those skilled in the art, and the present invention is not limited thereto.

[0103] Step 302: Perform image preprocessing on the tabular images in the image dataset.

[0104] Specifically, perform data preprocessing operations such as grayscale processing, binarization processing, filtering processing, or contour processing on the data in the image dataset.

[0105] Step 303: Perform data annotation on the tabular images in the image dataset.

[0106] Specifically, indicate the position and scope of the table by drawing a bounding box. Use the bounding box to label the table part of each image and set the corresponding class label for it, thereby forming the training set labels.

[0107] In one embodiment, the object detection model in the present invention is trained using the YOLOv5 model, and other open-source models such as the SSD model and the Faster R-CNN model can also be used, and the present invention is not limited thereto.

[0108] Divide the image dataset into a training set and a test set. The training set accounts for 70% of the total number of photos, and the test set accounts for 30% of the total number of photos.

[0109] In one embodiment, the image dataset can also be divided into a training set, a validation set, and a test set, usually divided according to a ratio coefficient of 7:2:1, and the present invention is not limited thereto.

[0110] In one embodiment, process the length, width, and grayscale of the images in the training set into a format of (42, 42, 1), that is, the image matrix is a two-dimensional matrix with 42 rows, 42 columns, and 1 layer, and other ratios are also possible, and the present invention is not limited thereto.

[0111] Step 304: Set the training parameters of the initial object detection model. That is, the parameters and hyperparameters of the YOLOv5 model. Among them, the model parameters include model weights and biases, etc., and the hyperparameters include learning rate, batch size, and training epochs, etc.

[0112] Specifically, for the model parameters, the parameters in the open-source pre-trained model can be used, or they can be randomly initialized. For the model hyperparameters, by setting the learning rate value at the initial stage of model training (i.e., the initial learning rate), usually 0.01 is used as the initial learning rate. It can also be adjusted according to the image dataset, and the present invention is not limited thereto.

[0113] In one embodiment, the learning rate at the end of model training (i.e., the final learning rate) is set, usually as a proportion of the initial learning rate. For example, setting the final learning rate to 0.2 means the final learning rate is 20% of the initial learning rate. Set the batch size of the model, which is used to determine the number of samples used in each model iteration. In the training command of YOLOv5, the batch size can be set through the batch hyperparameter. For example, batch64 means 64 samples are used in each training. Set the training epochs of the model, which is used to determine how many times the neural network model will be iterated. In the YOLOv5 command, the training epochs of the model are set through the epochs parameter. For example, epochs300 means the model training will be carried out for 300 rounds.

[0114] In one embodiment, the ReLU function is used as the activation function of the object detection model, the loss function is the cross-entropy function, and the optimizer of this model is Adam. The present invention is not limited thereto.

[0115] Step 305: Train the initial object detection model according to the image dataset and training parameters to obtain the trained object detection model.

[0116] Specifically, calculate the difference between the target value and the predicted value through the loss function, adjust the training parameters of the model for iterative training to obtain the object detection model. According to the prediction results of the object detection model, continuously adjust the model parameters and hyperparameters for model iterative training until the number of model training rounds is reached to obtain the trained object detection model.

[0117] Since the convolutional neural network (CNN) model usually needs to input in a fixed size when processing images. This is because the convolutional neural network model limits the size of the input layer to facilitate efficient calculation in subsequent convolutional layers and pooling layers.

[0118] Specifically, the input size of a convolutional neural network model is generally fixed to one or several types. In order to be able to predict images of any size, when the table image to be recognized is input, the table image is resized once to convert it into a matrix of the input size before input.

[0119] In one embodiment, when the resolution of the input table image is greater than the input size, the smaller the input size, the smaller the transformed size of the table image, and the smaller the degree of distortion of the table image. And when the resolution of the input table image is less than the input size and is enlarged, the greater the degree of enlargement, the more table features can be extracted.

[0120] Resample the table image to be recognized. The resampling process is used to adjust the resolution of the table image so that the size of the target table image matches the input size of the convolutional neural network model, that is, the size transformation (resize) in this embodiment.

[0121] First, read the table image to be recognized. Read the original table image data by using OpenCV or Pillow.

[0122] In one embodiment, in OpenCV, read the table image to be recognized through the cv2.imread function.

[0123] Then, adjust the image size of the read table image. First, determine the input size of the convolutional neural network model, such as 224x224 pixels or 256x256 pixels, etc. Adjust the size of the table image through the resize function provided by the image processing library.

[0124] In one embodiment, if using OpenCV, the table image can be adjusted to the preset target size through the cv2.resize function. For example, resized = cv2.resize(img, (224, 224)).

[0125] When adjusting the size of the table image, directly adjusting the table image to the target size through the resize function may cause image distortion. To ensure that the aspect ratio of the table image remains unchanged before and after adjustment, it is necessary to first calculate the aspect ratio of the table image. Adjust the size of the table image according to this aspect ratio. By first scaling the table image to a smaller size and then cropping or padding it to obtain the table image at the target size.

[0126] Specifically, when the aspect ratio of the table image does not match the target size, adjust the size of the table image by cropping or padding. Among them, cropping means removing the redundant parts other than the table itself in the table image. Padding means adding pixels around the table image to match the target size.

[0127] In one embodiment, in OpenCV, crop through the crop function. Padding through the copyMakeBorder function.

[0128] In one embodiment, when adjusting the size of the table image, an appropriate interpolation method can be selected. The present invention selects the nearest neighbor interpolation method to adjust the size of the target table image, but the present invention is not limited thereto. OpenCV provides a variety of interpolation methods, such as nearest neighbor interpolation, bilinear interpolation, and bicubic interpolation. For shrinking images, nearest neighbor interpolation is usually used, while for enlarging images, bilinear interpolation and bicubic interpolation are usually used.

[0129] In an embodiment of the present invention, after image preprocessing is performed on the table image to be recognized, a target table image is obtained. The image preprocessing directly affects the accuracy and efficiency of subsequent recognition of the table image.

[0130] As Figure 1 and Figure 4 shown, step 102: Process the target table image through a pre-trained key point detection model to obtain a scanned body table.

[0131] The training process of the key point detection model is as follows:

[0132] Specifically, collect an image dataset containing tables. And perform image preprocessing on the table images in the image dataset. Among them, the dataset contains table images of different types and complexities.

[0133] In one embodiment, these datasets can be sourced from open-source datasets or can be table images collected by those skilled in the art. The present invention is not limited thereto.

[0134] Perform data annotation on the table areas in the table images.

[0135] Specifically, draw a bounding box to indicate the position and scope of the table. Use the bounding box to annotate the table part of each image and its corresponding corner coordinates, and set corresponding category labels for these features. Among them, the corner coordinates are the vertices of the table, including the upper left point, upper right point, lower left point, and lower right point of the table.

[0136] Divide the image dataset into a training set and a test set. The training set accounts for 70% of the total number of photos, and the test set accounts for 30% of the total number of photos.

[0137] In one embodiment, the image dataset can also be divided into a training set, a validation set, and a test set, usually divided according to a ratio coefficient of 7:2:1. The present invention is not limited thereto.

[0138] Among them, the image data in the training set is used to train the YOLOv8 model. The image data in the validation set is used to evaluate the model performance and adjust the model hyperparameters during the model training process. The image data in the test set is used to finally evaluate the generalization ability of the trained model in actual applications, that is, the real prediction effect.

[0139] In one embodiment, the key point coordinate model in the present invention is trained using the YOLOv8 model, or other open-source models such as the SSD model and the Faster R-CNN model can also be used. The present invention is not limited thereto.

[0140] Set the training parameters of the key point detection model, and set the parameters and hyperparameters of the YOLOv8 model. Among them, the parameters include model weights and biases, etc., and the hyperparameters include learning rate, batch size, training epochs, etc. The specific settings of the parameters and hyperparameters can refer to the hyperparameter settings of the YOLOv5 model.

[0141] In one embodiment, the ReLU function is used as the activation function of the key point detection model, the loss function is the cross-entropy function, and the optimizer is Adam. The present invention is not limited thereto.

[0142] Calculate the difference between the target value and the predicted value through the loss function, and adjust the training parameters of the model for iterative training to obtain the key point detection model.

[0143] Specifically, according to the obtained model prediction results, continuously adjust the hyperparameters for model iterative training until the model training rounds are reached to obtain the trained key point detection model.

[0144] As Figure 5 shown, step 102 includes step 501 and step 502.

[0145] Step 501: Identify the outer contour endpoint coordinates of the target table image, and form the outer contour side lines of the target table image according to the outer contour endpoint coordinates. Among them, the outer contour side lines include the first contour side line, the second contour side line, the third contour side line, and the fourth contour side line.

[0146] Specifically, identify the four corner points of the target table image and their corresponding corner point coordinates through the key point coordinate model. Among them, the corner point coordinates are the outer contour endpoint coordinates. The corner point coordinates include the upper left point, the upper right point, the lower left point, and the lower right point of the table. Connect the upper left point and the upper right point to form the first contour side line, that is, the upper side line of the table. Connect the lower left point and the lower right point to form the second contour side line, that is, the lower side line of the table. Connect the upper left point and the lower left point to form the third contour side line, that is, the left side line of the table. Connect the upper right point and the lower right point to form the fourth contour side line, that is, the right side line of the table.

[0147] In the embodiment of the present invention, the four corner point coordinates in the target table image can be identified through the key point coordinate model, and the four sides of the table can be formed by connecting the four corner point coordinates. These four sides define the outer boundary of the table and are the key factors for judging whether the table is regular.

[0148] Step 502: Correct the target table image according to the relationship between the first contour side line, the second contour side line, the third contour side line, and the fourth contour side line and the preset reference lines to obtain the scanned body table. Among them, the preset reference lines include a preset horizontal line and a preset vertical line.

[0149] Specifically, the first contour line (i.e., the upper border line of the table) and the second contour line (i.e., the lower border line of the table) are compared with the horizontal line to determine whether the first contour line is parallel to the horizontal line respectively.

[0150] The third contour line (i.e., the left border line of the table) and the fourth contour line (i.e., the right border line of the table) are compared with the vertical line to determine whether the third contour line and the fourth contour line are perpendicular to the preset vertical line respectively. Thus, it is determined whether the table in the target table image is a regular image. If it is not a regular image, the target table image is corrected to obtain a scanned table.

[0151] As Figure 6 shown, step 502 includes steps 601 to 605.

[0152] Step 601: Determine whether the first contour line is parallel to the preset horizontal line.

[0153] Specifically, by calculating the slopes of the upper border line and the lower border line of the table, it can be determined whether the upper border line and the lower border line of the table are parallel to the horizontal line.

[0154] If the slopes of both the upper border line and the lower border line of the table approach zero, it indicates that both the upper border line and the lower border line of the table are parallel to the horizontal line.

[0155] Step 602: If the first contour line is not parallel to the horizontal line, determine the angle between the first contour line and the horizontal line.

[0156] Specifically, if the upper border line is not parallel to the horizontal line, connect the coordinates of the upper left point and the upper right point into a line segment to form the upper border line of the table. Then calculate the angle between the upper border line of the table and the horizontal line for the rotational correction of the target table image.

[0157] As Figure 7 shown, step 602 includes steps 701 to 702.

[0158] Step 701: Determine the slope of the first contour line according to the slope algorithm.

[0159] Specifically, assume that the coordinates of the upper left point of the recognized upper border line of the table are (x1, y1), and the coordinates of the upper right point are (x2, y2).

[0160] It is known that the slope m1 of the horizontal line is 0, and the slope m2 of the upper border line of the table can be calculated by formula (1).

[0161] m2 = (y2 - y1) / (x2 - x1) (1)

[0162] Among them, x1 is the abscissa of the upper left point, y1 is the ordinate of the upper left point, x2 is the abscissa of the upper right point, y2 is the ordinate of the upper right point, and m2 is the slope of the upper border line of the table.

[0163] Step 702: Calculate the included angle according to the arctangent function and the slope of the upper border line.

[0164] Specifically, the included angle θ can be calculated by the arctangent function formula (2).

[0165]

[0166] Among them, m2 is the slope of the upper border line of the table, m1 is the slope of the horizontal line, and θ is the included angle between the upper border line of the table and the horizontal line.

[0167] Since the slope m1 of the horizontal line is 0, formula (2) can be simplified to θ = arctan(m2).

[0168] Step 603: Rotate the target table image according to the angle of the included angle.

[0169] Specifically, according to the calculated included angle θ, rotate the target table image to rotate the target table image until the upper border line is parallel to the horizontal line.

[0170] Step 604: Determine whether the third contour line and the fourth contour line are parallel to the preset vertical line.

[0171] Connect the coordinates of the lower left point and the lower right point into a line segment to form the lower border line of the table.

[0172] Suppose the coordinates of the lower left point of the recognized lower border line of the table are (x3, y3), and the coordinates of the lower right point are (x4, y4).

[0173] Compare the x-axis coordinate x1 of the upper left point coordinate (x1, y1) with the x-axis coordinate x3 of the lower left point coordinate (x3, y3) to determine whether the x-axis coordinates of the left endpoints of the upper border line and the lower border line of the table are the same. If they are the same, it means that the left border line of the table is parallel to the vertical line.

[0174] Compare the x-axis coordinate x2 of the upper right point coordinate (x2, y2) with the x-axis coordinate x4 of the lower right point coordinate (x4, y4) to determine whether the x-axis coordinates of the right endpoints of the upper border line and the lower border line of the table are the same. If they are the same, it means that the right border line of the table is parallel to the vertical line.

[0175] In one embodiment, it is also possible to determine whether the left and right sides of the table are parallel to the vertical line by calculating the slopes of the left and right sides of the table. If the slopes of both the left and right sides of the table approach infinity (or a very large number in actual calculations), it indicates that both the left and right sides of the table are parallel to the vertical line.

[0176] Step 605: If the third contour line and / or the fourth contour line are not parallel to the vertical line, perform a perspective transformation on the rotated target table image to obtain a scanned body table.

[0177] Specifically, if the left side of the table is not parallel to the vertical line and / or the right side of the table is not parallel to the vertical line, it indicates that there is perspective distortion in the image, and it is necessary to correct the rotated target table image through perspective transformation.

[0178] In OpenCV, the perspective transformation matrix is calculated through the cv2.getPerspectiveTransform function. The perspective transformation matrix is calculated by inputting the original coordinate points of the table and the target coordinate points of the table to perform a perspective transformation on the target table image. Among them, the original coordinate points are the coordinates of the four corner points in the target table image, that is, the upper left point coordinate is (x1, y1), the upper right point coordinate is (x2, y2), the lower left point coordinate is (x3, y3), and the lower right point coordinate is (x4, y4). The target coordinate points are the coordinate points expected to be obtained after transformation, usually the four corner points of the perspective transformation matrix.

[0179] The cv2.getPerspectiveTransform function calculates the perspective transformation matrix based on the original coordinate points and the target coordinate points. According to the rotated target table image and the perspective transformation matrix, the cv2.warpPerspective function is used to obtain the corrected scanned body table.

[0180] In the embodiment of the present invention, after the above perspective transformation, the table in the target table image is converted into a standard rectangular table. The corrected scanned body table can be further subjected to table recognition and table content extraction. The OCR technology can be used to recognize the text content in the table, or the table parsing technology can be used to extract the table structure. Through the above steps, the non-regular table image can be effectively corrected to become a regular image, thereby improving the accuracy and efficiency of subsequent table recognition.

[0181] Such as Figure 1 、 Figure 8 and Figure 9 As shown in steps 103: Perform row and column recognition conversion on the scanned body table to obtain the target table structure.

[0182] Such as Figure 9As shown, step 103 includes steps 901 to 904.

[0183] Step 901: Identify each row and each column of the scanned body table.

[0184] Specifically, all rows and all columns in the scanned body table are identified through image processing techniques (such as edge detection or line detection, etc.). The rows in the table are usually composed of horizontal line segments and can be identified by detecting consecutive pixel points in the horizontal direction. The columns in the table are usually composed of vertical line segments and can be identified by detecting consecutive pixel points in the vertical direction.

[0185] In one embodiment, edge detection identifies the boundary between the table content and the table background in the scanned body table through the gradient of the image gray distribution. Commonly used edge detection operators include the Prewitt operator, Sobel operator, and Canny operator, etc. These edge detection operators determine the rows and columns of the table by calculating the local gradient of the pixels in the scanned body table. For example, the Canny operator identifies the table boundary through methods such as Gaussian filtering, gradient calculation, and non-maximum suppression.

[0186] In one embodiment, line detection identifies all rows and all columns in the table by identifying the line segments in the scanned body table. For example, through the Hough Transform, all coordinate points in the scanned body table are transformed into the parameter space, and lines are identified through an accumulator. By representing all the lines in the scanned body table in polar coordinate form, that is, r = xcos(θ1) + ysin(θ1), where r is the perpendicular distance from the origin to the line, and θ1 is the angle between the line and the x-axis. By using the accumulator value in the statistical parameter space, the lines in the image can be identified.

[0187] Step 902: Determine the positions of each row and each column of the scanned body table and their corresponding intersection points.

[0188] In the embodiment of the present invention, the above-mentioned edge detection and line detection are combined and used. First, the edges in the scanned body table are identified through the edge detection operator, and these edges correspond to the boundaries or lines of the table. Then, the line segments that form the rows and columns of the table are identified through the line detection algorithm (such as the Hough Transform) from these edges, so as to obtain the positions of all rows and all columns of the scanned body table.

[0189] The identified rows and columns are combined to obtain the intersection coordinates between all rows and all columns.

[0190] Step 903: Determine the coordinates of each cell according to the positions of each row and each column to obtain regular cells.

[0191] Combine the y - coordinates of each row with the x - coordinates of each column to obtain the positions of all regular cells in the scan body table.

[0192] Specifically, for example, in an 8 - row and 6 - column table, the result predicted by the object detection model is 14 objects, which are 8 row categories and 6 column categories respectively. Traverse these data, and divide the 14 objects into a row set with a length of 8 and a column category set with a length of 6 according to the category. Then perform nested traversal, either with the row set outside or the column set outside.

[0193] When the row set is outside and the column set is inside, the first row is combined with all columns. Obtain the two y - coordinates of the first row and the two x - coordinates of the first column, form the upper - left and lower - right corners of the first cell in the first row, and get the coordinates of the first cell. By traversing in sequence, the upper - left and lower - right coordinates of all cells can be obtained, thereby determining the coordinate positions of each cell and recording them.

[0194] In one embodiment, the intersection of each row and column can also be used as the upper - left corner of a cell. For each intersection, determine its corresponding cell boundary, thereby determining the four corners of the cell (i.e., the upper - left corner, the upper - right corner, the lower - left corner, and the lower - right corner). By the position difference between adjacent rows and columns, the width and height of each cell can be determined. For each cell, the coordinates of its upper - left corner can be determined by the intersection position of the row and column, and the coordinates of its lower - right corner can be determined by adding the width and height of the cell to the coordinates of the upper - left corner.

[0195] In one embodiment, a coordinate mapping can also be created for the y - coordinates of each row and the x - coordinates of each column, and the position of each cell is defined through the coordinate mapping. For each row, record the y - coordinates of its upper and lower boundaries. For each column, record the x - coordinates of its left and right boundaries. Combine the row coordinates and column coordinates to determine the four corners of each cell (the upper - left corner, the upper - right corner, the lower - left corner, and the lower - right corner). Among them, the upper - left corner is the x - coordinate of the first column and the y - coordinate of the first row, the upper - right corner is the x - coordinate of the second column and the y - coordinate of the first row, the lower - left corner is the x - coordinate of the first column and the y - coordinate of the second row, and the lower - right corner is the x - coordinate of the second column and the y - coordinate of the second row. By recursively expanding to all rows and all columns in the scan body table, the positions of all cells in the scan body table are obtained, that is, the regular cells are obtained.

[0196] All regular cells in the scan body table can be identified through the above - mentioned methods, but the present invention is not limited thereto.

[0197] Step 904: Merge the regular cells in the scan body table to obtain the target table structure.

[0198] Such as Figure 10As shown, step 904 includes steps 1001 to 1004.

[0199] Step 1001: Identify cells in the scanned body table that span multiple rows or columns to obtain cross-row and cross-column cells.

[0200] Specifically, identify all cells in the scanned body table. For each detected cell, perform horizontal and vertical matching to determine its adjacent cells. By judging the position of the cell center point and the length of the adjacent sides, determine whether the cell is a cross-row and cross-column cell.

[0201] In one embodiment, first perform horizontal matching. For a cell a, judge whether the center point b1 of another cell b is on the left or right side of cell a. If the center point b1 of cell b has a significant offset in the vertical direction from the center point a1 of cell a, that is, the y coordinate of the center point b1 differs greatly from the y coordinate of the center point a1. And if the vertical span of cell b is greater than a certain proportion of the vertical span of cell a, then it is considered that cell b spans the upper and lower boundaries of cell a, thereby determining that cell b is a cross-row and cross-column cell. Then perform vertical matching, and the judgment method of vertical matching is similar to that of horizontal matching.

[0202] Step 1002: Compare the cross-row and cross-column cells with the regular cells.

[0203] Specifically, calculate the area ratio of the regular cells covered by each cross-row and cross-column cell. If a regular cell is covered by a cross-row and cross-column cell by more than 80% of its area, then determine that the regular cell is part of the cross-row and cross-column cell.

[0204] Step 1003: If a cross-row and cross-column cell contains multiple regular cells, merge the multiple regular cells into one cross-row and cross-column cell.

[0205] Specifically, if a cross-row and cross-column cell contains multiple regular cells, then replace the multiple cells with the cross-row and cross-column cell to update the table structure.

[0206] Step 1004: Obtain the target table structure based on the regular cells and the cross-row and cross-column cells.

[0207] Specifically, use the table structure that contains multiple regular cells and multiple cross-row and cross-column cells after replacement as the target table structure.

[0208] In one embodiment, subsequent business processing can be performed according to the target table structure. For example, OCR technology can be used to identify the table content, or the table structure can be directly used. The present invention is not limited thereto.

[0209] In the embodiments of the present invention, a pre-trained object detection model is used to preprocess the table image to be recognized, and a target table image is obtained. Then, a pre-trained key point detection model is used to process the target table image to obtain a scanned table. And row and column recognition conversion is performed on the scanned table to obtain a target table structure. Through image correction techniques such as rotation and perspective transformation, the image quality problems caused by improper shooting angles are improved, making the table image more regular and clear. By using a variety of convolutional neural network models and image preprocessing techniques, the table recognition accuracy in processing complex or irregular tables is significantly improved. The present invention can adapt to table image inputs with different resolutions and qualities, improving the robustness of the system. The recognized table structure can be directly used for OCR content recognition or other business processes, having good flexibility and scalability.

[0210] In the embodiments of the present invention, a table recognition device is also provided as described in the following embodiments. Since the principle of the device for solving problems is similar to that of the table recognition method, the implementation of the device can refer to the implementation of the table recognition method, and the repeated parts will not be described again.

[0211] As Figure 11 shown, the table recognition device 110 includes an image preprocessing module 1101, a scanned body module 1102, and a row and column recognition module 1103.

[0212] The image preprocessing module 1101 is used to preprocess the table image to be recognized through a pre-trained object detection model to obtain a target table image.

[0213] The scanned body module 1102 is used to process the target table image through a pre-trained key point detection model to obtain a scanned table.

[0214] The row and column recognition module 1103 is used to perform row and column recognition conversion on the scanned table to obtain a target table structure.

[0215] The above-mentioned scanned body module 1102 includes a contour processing unit and an image correction unit.

[0216] The contour processing unit is used to identify the outer contour endpoint coordinates of the target table image and form the outer contour side lines of the target table image according to the outer contour endpoint coordinates; wherein, the outer contour side lines include a first contour side line, a second contour side line, a third contour side line, and a fourth contour side line.

[0217] The image correction unit is used to correct the target table image according to the relationship between the first contour side line, the second contour side line, the third contour side line, and the fourth contour side line and a preset reference line to obtain a scanned table; wherein, the preset reference line includes a preset horizontal line and a preset vertical line.

[0218] The above image correction unit includes a first judgment subunit, an included angle determination subunit, an image rotation subunit, a second judgment subunit, and a perspective transformation subunit.

[0219] The first judgment subunit is used to judge whether the first contour side line is parallel to a preset horizontal line.

[0220] The included angle determination subunit is used to determine the included angle between the first contour side line and the horizontal line if the first contour side line is not parallel to the horizontal line.

[0221] The image rotation subunit is used to rotate the target table image according to the angle of the included angle.

[0222] The second judgment subunit is used to judge whether the third contour line and the fourth contour line are parallel to a preset vertical line.

[0223] The perspective transformation subunit is used to perform perspective transformation on the rotated target table image to obtain a scanned body table if the third contour line and the fourth contour line are not parallel to the vertical line.

[0224] The above included angle determination subunit includes a slope determination subunit and an included angle calculation subunit.

[0225] The slope determination subunit is used to determine the slope of the first contour side line according to a slope algorithm.

[0226] The included angle calculation subunit is used to calculate the included angle according to the arctangent function and the slope.

[0227] The above row and column recognition module 1103 includes a row and column recognition subunit, a row and column position determination subunit, a cell determination subunit, and a table structure determination subunit.

[0228] The row and column recognition subunit is used to recognize each row and each column of the scanned body table.

[0229] The row and column position determination subunit is used to determine the positions of each row and each column of the scanned body table and their corresponding intersection points.

[0230] The cell determination subunit is used to determine the coordinates of each cell according to the positions of each row and each column to obtain regular cells.

[0231] The table structure determination subunit is used to merge the regular cells in the scanned body table to obtain a target table structure.

[0232] The above table structure determination subunit includes: a cell recognition subunit, a cell comparison subunit, a cell merging subunit, and a target table structure determination subunit.

[0233] The cell recognition subunit is used to recognize the cells that span multiple rows or columns in the scanned body table, and obtain the cells that span rows and columns.

[0234] The cell comparison subunit is used to compare the cells that span rows and columns with the regular cells.

[0235] The cell merging subunit is used to merge multiple regular cells into one cell that spans rows and columns if the cell that spans rows and columns contains multiple regular cells.

[0236] The target table structure determination subunit is used to obtain the target table structure according to the regular cells and the cells that span rows and columns.

[0237] The above image preprocessing module includes an image collection unit, an image processing unit, a data annotation unit, a training parameter setting unit, and a model training unit.

[0238] The image collection unit is used to collect an image data set containing tables.

[0239] The image processing unit is used to perform image preprocessing on the table images in the image data set.

[0240] The data annotation unit is used to perform data annotation on the table images in the image data set.

[0241] The training parameter setting unit is used to set the training parameters of the initial object detection model.

[0242] The model training unit is used to train the initial object detection model according to the image data set and the training parameters, and obtain the trained object detection model.

[0243] Figure 12 The following is a schematic diagram of the physical structure of the electronic device provided by the embodiment of the present invention. As Figure 12 shown, the electronic device 120 includes: a processor 1201, a memory 1202, and a bus 1203.

[0244] Among them, the processor 1201 and the memory 1202 communicate with each other through the bus 1203.

[0245] The processor 1201 is used to call the program instructions in the memory 1202 to execute the methods provided by the above method embodiments.

[0246] The embodiment of the present invention also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above table recognition method is implemented.

[0247] An embodiment of the present invention also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the above-mentioned table recognition method is implemented.

[0248] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.

[0249] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or a plurality of flows and / or blocks

[0250] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implement the functions specified in Figure 1 one or more of the flows Figure 1 or a plurality of flows and / or blocks

[0251] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or a plurality of flows and / or blocks

[0252] The specific embodiments described above further elaborate on the object, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only for the specific embodiments of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A table recognition method, characterized in that: include: Preprocessing the table image to be identified by using a pre-trained target detection model to obtain a target table image; Processing the target table image by a pre-trained key point detection model to obtain a scanned volume table; Perform row and column recognition conversion on the scanned body table to obtain a target table structure.

2. The method according to claim 1, characterized in that The target table image is processed by a pre-trained key point detection model to obtain a scanned volume table, including: Identify the outer contour endpoint coordinates of the target table image, and form the outer contour edge line of the target table image according to the outer contour endpoint coordinates; wherein the outer contour edge line includes a first contour edge line, a second contour edge line, a third contour edge line and a fourth contour edge line; According to the relationship between the first contour edge line, the second contour edge line, the third contour edge line and the fourth contour edge line and the preset reference line, the target table image is corrected to obtain a scanned volume table; wherein the preset reference line includes a preset horizontal line and a preset vertical line.

3. The method according to claim 2, characterized in that The method of correcting the target table image according to the relationship between the first contour edge line, the second contour edge line, the third contour edge line, the fourth contour edge line and the preset reference line to obtain a scanned volume table includes: Determining whether the first contour edge line is parallel to a preset horizontal line; If the first contour edge line is not parallel to the horizontal line, determining an angle between the first contour edge line and the horizontal line; Rotating the target table image according to the angle; Determining whether the third contour line and the fourth contour line are parallel to a preset vertical line; If the third contour line and the fourth contour line are not parallel to the vertical line, a perspective transformation is performed on the rotated target table image to obtain a scanned volume table.

4. The method according to claim 3, characterized in that Determining the angle between the first contour edge line and the horizontal line comprises: Determine the slope of the first contour edge line according to a slope algorithm; The angle is calculated according to the inverse tangent function and the slope.

5. The method according to claim 1, characterized in that: The performing row-column recognition conversion on the scanned body table to obtain a target table structure includes: Identifying each row and each column of the scan body table; Determine the position of each row and each column of the scanned volume table and their corresponding intersection points; Determine the coordinates of each cell according to the position of each row and column to obtain a regular cell; The regular cells in the scanned body table are merged to obtain a target table structure.

6. The method according to claim 5, characterized in that The step of merging the regular cells in the scanned body table to obtain a target table structure includes: Identifying cells that span multiple rows or columns of the scanned body table to obtain cells that span multiple rows or columns; Compare the cross-row and cross-column cells with the regular cells; If the cross-row and cross-column cell includes a plurality of the regular cells, the plurality of regular cells are merged into one cross-row and cross-column cell; A target table structure is obtained according to the regular cells and the cross-row and cross-column cells.

7. The method according to claim 1, characterized in that The training process of the target detection model includes: Collect an image dataset containing tables; Performing image preprocessing on the table images in the image data set; Performing data annotation on the table images in the image data set; Set the initial training parameters of the target detection model; The initial target detection model is trained according to the image data set and the training parameters to obtain a trained target detection model.

8. A table recognition device, characterized in that: include: An image preprocessing module is used to preprocess the table image to be identified by using a pre-trained target detection model to obtain a target table image; A scanning volume module, used for processing the target table image through a pre-trained key point detection model to obtain a scanning volume table; The row and column recognition module is used to perform row and column recognition conversion on the scanned table to obtain a target table structure.

9. The device according to claim 8, characterized in that The scanning body module comprises: A contour processing unit, used for identifying the outer contour endpoint coordinates of the target table image, and forming the outer contour edge line of the target table image according to the outer contour endpoint coordinates; wherein the outer contour edge line includes a first contour edge line, a second contour edge line, a third contour edge line and a fourth contour edge line; An image correction unit is used to correct the target table image according to the relationship between the first contour edge line, the second contour edge line, the third contour edge line and the fourth contour edge line and a preset reference line to obtain a scanned volume table; wherein the preset reference line includes a preset horizontal line and a preset vertical line.

10. The device according to claim 9, characterized in that The image correction unit comprises: A first judging subunit, used to judge whether the first contour edge line is parallel to a preset horizontal line; an angle determination subunit, configured to determine an angle between the first contour edge line and the horizontal line if the first contour edge line is not parallel to the horizontal line; An image rotation subunit, used for rotating the target table image according to the angle; A second judging subunit, used for judging whether the third contour line and the fourth contour line are parallel to a preset vertical line; The perspective transformation subunit is used for performing perspective transformation on the rotated target table image to obtain a scanned volume table if the third contour line and the fourth contour line are not parallel to the vertical line.

11. The device according to claim 10, characterized in that The angle determination subunit comprises: a slope determination subunit, configured to determine the slope of the first contour edge line according to a slope algorithm; The angle calculation subunit is used to calculate the angle according to the inverse tangent function and the slope.

12. The device according to claim 8, characterized in that The row and column identification module comprises: A row and column identification subunit, used to identify each row and each column of the scan body table; A row and column position determination subunit, used to determine the position of each row and column of the scanned volume table and their corresponding intersection points; The cell determination subunit is used to determine the coordinates of each cell according to the position of each row and each column to obtain a regular cell; The table structure determination subunit is used to merge the regular cells in the scan body table to obtain a target table structure.

13. The device according to claim 12, characterized in that The table structure determination subunit includes: A cell identification subunit, used for identifying cells that span multiple rows or columns of the scanned body table to obtain cells that span rows and columns; A cell comparison subunit, used for comparing the cross-row and cross-column cells with the regular cells; A cell merging subunit, used for merging the multiple regular cells into one cross-row and cross-column cell if the cross-row and cross-column cell includes multiple regular cells; The target table structure determination subunit is used to obtain the target table structure according to the regular cells and the cross-row and cross-column cells.

14. The device according to claim 8, characterized in that The image preprocessing module comprises: An image collection unit, used for collecting an image data set containing a table; An image processing unit, used for performing image preprocessing on the table image in the image data set; A data annotation unit, used for annotating data on the table image in the image data set; A training parameter setting unit, used to set the initial training parameters of the target detection model; The model training unit is used to train the initial target detection model according to the image data set and the training parameters to obtain a trained target detection model.

15. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.

16. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

17. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Table structure extraction method based on deep cascade network

    CN121121781A

  • Table structure identification method and system based on self-adaptive anchor frame, terminal and medium

    CN121214467A