Table identification method and related device

By extracting the visual characteristics of the table image and determining the key points of the row and column line, and predicting the row and column segmentation lines, the problem of low accuracy of complex table recognition in the prior art is solved, and more efficient table structured recognition is achieved.

CN119992578AActive Publication Date: 2025-05-13IFLYTEK CO LTD

Patent Information

Application Number
CN202510056895.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-13
Estimated Expiration
2045-01-14

Smart Images

  • Figure CN119992578A_ABST
    Figure CN119992578A_ABST
Patent Text Reader

Abstract

The invention provides a table recognition method and a related device, and relates to the technical field of image recognition. The table recognition method comprises the following steps: extracting visual features of a table image; wherein the visual features comprise table row features and table column features; determining row line key points based on the table row features, and determining column line key points based on the table column features; predicting a row segmentation line based on the row line key points, and predicting a column segmentation line based on the column line key points; and generating a structured table based on the row segmentation lines and the column segmentation lines obtained through prediction. The technical scheme provided by the invention is used for solving the problem of low table recognition accuracy in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image recognition technology, and in particular to a table recognition method and related devices. Background Art

[0002] As an efficient form of information organization and presentation, tables are widely used in various documents, such as financial statements, medical records, academic papers, etc. With the development of information technology and the acceleration of digital transformation, the need to recognize image-type tables and convert them into structured tables is becoming increasingly urgent.

[0003] Existing table structure recognition methods are mainly implemented through segmentation schemes, that is, first segmenting the row and column line areas of the table, then converting the row and column line areas into row and column lines through complex post-processing methods, and finally merging cells to restore the table structure.

[0004] However, faced with complex and ever-changing table structures, segmentation-based table recognition solutions have difficulty segmenting accurate row and column dividing lines when images contain noise, table tilt, table curvature, etc., which in turn affects the structuring of the table. Summary of the invention

[0005] Based on the above-mentioned defects and shortcomings of the prior art, the present application proposes a table recognition method and related devices to solve the problem of low table recognition accuracy in the prior art.

[0006] According to a first aspect of the present application, a table recognition method is provided, the method comprising:

[0007] Extracting visual features of a table image; wherein the visual features include: table row features and table column features;

[0008] Determine row line key points based on the table row features, and determine column line key points based on the table column features;

[0009] Predicting row segmentation lines based on the row line key points, and predicting column segmentation lines based on the column line key points;

[0010] A structured table is generated based on the predicted row dividing lines and the column dividing lines.

[0011] According to a second aspect of the present application, a table recognition device is provided, the device comprising:

[0012] A feature extraction module, used to extract visual features of a table image; wherein the visual features include: table row features and table column features;

[0013] A key point determination module, used to determine row line key points based on the table row characteristics, and to determine column line key points based on the table column characteristics;

[0014] A segmentation line prediction module, used for predicting row segmentation lines based on the row line key points, and predicting column segmentation lines based on the column line key points;

[0015] The table generation module is used to generate a structured table based on the predicted row dividing lines and the column dividing lines.

[0016] According to a third aspect of the present application, there is provided an electronic device, comprising: a memory and a processor;

[0017] The memory is connected to the processor and is used to store programs;

[0018] The processor is used to implement the table recognition method as described in the first aspect by running the program in the memory.

[0019] According to a fourth aspect of the present application, a storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the table recognition method as described in the first aspect is implemented.

[0020] According to a fifth aspect of the present application, a computer program product or a computer program is provided, the computer program product includes the computer program, and when the processor of a computer device executes the computer program, the steps in the table recognition method as described in the first aspect are implemented. Optionally, the computer program can be stored in a readable storage medium or a cloud of the computer device; the processor of the computer device reads the computer program from the readable storage medium or the cloud.

[0021] In the technical solution provided by the present application, in the process of table recognition, the key points of the row and column lines are obtained based on the visual row and column features of the table, and then the row and column dividing lines are predicted from the key points of the row and column lines, so that the predicted row and column dividing lines can be the same or similar to the row and column shapes of the table in the image, improving the continuity of the row and column dividing lines. Even in the face of complex table situations, such as tilted tables, curved tables, etc., the table in the image can be restored more realistically, reducing the probability of dividing a row and column dividing line into multiple row and column dividing lines, and improving the accuracy of table structured recognition. In addition, the technical solution provided by the present application does not require complex post-processing, so it can also improve the efficiency of table recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0023] Figure 1 A flowchart of a table recognition method provided in an embodiment of the present application;

[0024] Figure 2 A schematic diagram of a table recognition process provided in an embodiment of the present application;

[0025] Figure 3 A schematic diagram of key points of a line provided in an embodiment of the present application;

[0026] Figure 4 A schematic diagram of cell merging provided in an embodiment of the present application;

[0027] Figure 5 A block diagram of a table recognition device provided in an embodiment of the present application;

[0028] Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0029] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0030] Application Overview

[0031] With the rapid development of information technology, tables, as an important carrier of structured data, are widely used in various fields. However, many tables are scanned or photographed into image format during data processing, which means that the computer loses the structured information. Therefore, table structure recognition has become an important technology for current document understanding.

[0032] The table structured recognition technology aims to automatically extract the structure and content of the table from the document image to facilitate further processing and analysis of the data. In recent years, this field has received widespread attention from academia and industry, and a large number of research results have emerged. However, there are many styles of tables, including wired tables, tables with few lines, tables without lines, etc., as well as tables with a lot of noise, such as tables photographed under underexposure or overexposure conditions, and tables that are deformed by rotation, distortion, etc., such as tables that are rotated (that is, tilted) due to the shooting angle or scanning angle, and tables that are photographed when the written text has a certain curvature, and the rows and columns are curved. The existing technology still faces many challenges in processing complex tables, resulting in low accuracy in table recognition.

[0033] For example, for the table recognition method implemented by segmentation scheme, it first segments the row and column line area of ​​the table, then converts the row and column line area into row and column lines through complex post-processing means, and finally merges cells to restore the table structure. Since post-processing is implemented based on rules, these rules are difficult to cover various complex table situations, which may lead to errors in row and column line segmentation. For example, due to the tilt of the table, a row line may be divided into two rows during segmentation, resulting in errors in table structuring. This type of scheme has poor generalization ability and reduces the accuracy of table recognition.

[0034] In order to solve the above technical problems, the present application provides a table recognition technology. In the table recognition process, the row and column line key points are obtained based on the visual row and column features of the table, and the row and column dividing lines are predicted from the row and column key points. In this way, the predicted row and column dividing lines can adapt to the row and column shapes of the table in the image, improve the continuity of the row and column dividing lines, and even in the face of complex table situations, such as tilted tables, curved tables, etc., the table in the image can be restored more realistically, reducing the probability of splitting a row and column dividing line into multiple row and column dividing lines, and improving the accuracy of table structured recognition. In addition, the technical solution provided by the present application does not require complex post-processing, so it can also improve the efficiency of table recognition.

[0035] Exemplary Methods

[0036] The embodiment of the present application provides a table recognition method, and the execution subject may be an electronic device with a table recognition function. The electronic device may be a terminal device, such as a mobile phone, a tablet computer, a desktop computer, a personal digital assistant device, etc., or a server or other device.

[0037] The method is described in detail below through some embodiments. The following embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments. Figure 1 As shown, the table recognition method may include steps 101 to 104, which are specifically described as follows:

[0038] Step 101: extract visual features of the table image.

[0039] The visual features described here may include: table row features, table column features, and text features in the table, etc. The table row features described here are used to represent the row boundary features of the table, so they can also be called table row boundary features, or table row boundary features such as the boundary features between rows, and the row boundary features between the table and other surrounding information. The table column features described here are used to represent the column boundary features of the table, so they can also be called table column boundary features, or table column boundary features such as the boundary features between columns, and the column boundary features between the table and other surrounding information.

[0040] When the target position in the table has a table line, the table target feature is the table line feature; when the target position in the table does not have a table line, the table target feature is the region feature where the target position is located. The target position described here is a table row boundary position (such as a boundary between rows, a row boundary feature between a table and the outside world), or the target position is a table column boundary position (such as a boundary between columns, a column boundary feature between a table and the outside world). The table target feature described here is a table row feature or a table column feature.

[0041] That is, when there is a table row line at the table row boundary, the table row feature is the feature of the table row line; when there is no table row line at the table row boundary, the table row feature is the feature of the area where the table row boundary is located, such as the feature of the blank area between rows (such as the length, width, etc. of the area), or the feature of the blank area at the edge of the row. The blank area feature may include but is not limited to: the length feature, width feature, etc. of the area.

[0042] Similarly, when there is a table column line at the column boundary in the table, the table column feature is the feature of the table column line; when there is no table column line at the column boundary in the table, the table column feature is the feature of the area where the table column boundary is located, such as the feature of the blank area between columns (such as the length, width and other features of the area), or the feature of the blank area at the edge of the column. The blank area feature may include but is not limited to: the length feature, width feature, etc. of the area.

[0043] In an embodiment of the present application, a table image can be input into a visual encoder to extract multi-layer features of different resolutions of the image. The different resolution images mentioned here refer to table images of different resolutions obtained by compressing the table image to different degrees along the width direction and / or height direction of the table. Compressing the table image along the width direction of the table can make the row features more concentrated and obvious, which is conducive to the accurate extraction of row features; similarly, compressing the table image along the height direction of the table can make the column features more concentrated and obvious, which is conducive to the accurate extraction of column features.

[0044] Optionally, when performing visual feature extraction, a cross-attention module can also be used to generate enhanced row and column features from the height and width dimensions of the table. Among them, the cross-attention module is an attention mechanism widely used in deep learning, which allows the model to establish associations between two different input sequences and calculate attention weights, thereby promoting the exchange and integration of information. Therefore, the embodiment of the present application can extract enhanced table row features and table column features from the height and width dimensions of the table through the cross-attention module, which is conducive to improving the prediction effect of subsequent table row and column dividing lines.

[0045] The aforementioned visual encoder can also be called a backbone network. Figure 2 As shown, the table image is input into the backbone network, and the backbone network performs feature extraction to obtain the visual features of the table image. The backbone network is generally a convolutional neural network, and its task is to extract high-level visual features from the input image and generate feature maps. These feature maps contain various information of the input image, such as edges, textures, shapes, etc. Therefore, the visual features output by the backbone network may include table row features and table column features. Of course, in addition to this, the visual features may also include information features in the table, such as text content features specifically included in the table.

[0046] Step 102: Determine row line key points based on table row features, and determine column line key points based on table column features.

[0047] After obtaining the table row features and table column features in step 101 , row line key points for predicting row segmentation lines may be determined based on the table row features, and column line key points for predicting column segmentation lines may be determined based on the table column features.

[0048] For positions with rows in the table, multiple row key points can be determined based on the row characteristics. Figure 2 The prediction results of key points of the line in the wired table are shown (specifically Figure 2As shown in the upper right corner); for the position where there is no row line at the edge of a table row or between rows, multiple row line key points can be determined within the corresponding row boundary area, such as selecting multiple points on the center line of the table in the horizontal direction (i.e., the width direction) of the area as row line key points. It can be understood that the method for determining row line key points is not limited to this, and can be set according to actual needs.

[0049] Similarly, for positions with column lines in a table, multiple column line key points can be determined based on column line features; for positions without column lines at the edge of a table column or between columns, multiple column line key points can be determined within the corresponding column boundary region, such as selecting multiple points on the center line of the table in the length direction (i.e., height direction) of the region as column line key points. It is understandable that the method for determining column line key points is not limited to this, and can be set according to actual needs.

[0050] It should be noted that the key points of the row lines conform to the rules of the table rows, and the curve obtained based on their fitting can achieve the effect of row boundaries; similarly, the key points of the column lines conform to the rules of the table column lines, and the curve obtained based on their fitting can achieve the effect of column boundaries.

[0051] Step 103: predicting row segmentation lines based on row line key points, and predicting column segmentation lines based on column line key points.

[0052] After obtaining the row key points and column key points in step 102, the row segmentation lines can be predicted based on the coordinates of the row key points, and the column segmentation lines can be predicted based on the coordinates of the column key points. Specifically, the row segmentation lines can be obtained by fitting the row key points, and the column segmentation lines can be obtained by fitting the column key points.

[0053] like Figure 2 As shown, by combining the table row features and the table column features, the prediction results of the row dividing line and the column dividing line are obtained.

[0054] Step 104: Generate a structured table based on the predicted row partition lines and column partition lines.

[0055] In this step, based on the row dividing lines and column dividing lines obtained in step 103, and the text information in the table (such as Figure 2 "A, B, C, ..., K" in the table shown in the figure can generate a structured table. Optionally, for the text content in the table, the visual features can be recognized by optical character recognition (OCR) technology to obtain the text content.

[0056] After obtaining the row dividing lines and column dividing lines, for the cells formed by the intersection of the two, it can be determined whether there is a need to merge. If so, the cells can be merged, such as Figure 2 As shown in the figure, the table obtained based on the row and column dividing line prediction results is obviously inconsistent with the cell layout in the real table, so the subsequent merging process can be performed to restore the table structure. On the contrary, if the table obtained based on the row and column dividing line prediction results is consistent with the cell layout in the real table, the merging process can be omitted.

[0057] In the embodiment of the present application, in the process of table recognition, the row and column line key points are obtained based on the visual row and column features of the table, and the row and column dividing lines are predicted from the row and column key points, which can improve the continuity of the row and column dividing lines, and the predicted row and column dividing lines can be the same or similar to the row and column dividing shapes of the table in the image. Even in the face of complex table situations, such as tilted tables, curved tables, etc., the table in the image can be restored more realistically, reducing the probability of dividing a row and column dividing line into multiple row and column dividing lines, and improving the accuracy of table structured recognition. In addition, the technical solution provided by the present application does not require complex post-processing technology, so it can also improve the efficiency of table recognition.

[0058] In some optional embodiments, due to deformation of some tables, rows or columns may no longer be in a horizontal or vertical state. In this case, a row of row key points can be predicted for each row boundary. When predicting row dividing lines based on the row key points, the key points belonging to different row boundaries can be divided according to the distance between the key points in the vertical direction of the table, thereby realizing the prediction of row dividing lines for each row boundary.

[0059] Optionally, in order to facilitate the distinction of key points belonging to each row boundary, multiple rows of key points can also be predicted for each row boundary, and the multiple rows of key points can specifically include n×m key points. Wherein, n represents the number of horizontal key points in the table, that is, the number of key points in each row, and m represents the number of vertical key points in the table, which also represents the number of rows of key points. Both m and n are integers greater than or equal to 2, and the distance between each group of adjacent key points in the vertical direction of the table is small, at least less than the height of the smallest cell in the table, which makes these key points relatively concentrated, so that the key points of each row boundary can be clearly distinguished. Wherein, m and n are not fixed values, and can be flexibly adjusted based on the size of the table, wherein the number of key points and the width of the table can be positively correlated, that is, the higher the width of the table, the more key points there are, so as to fit a more realistic row dividing line.

[0060] like Figure 3 As shown, the row key points of two row boundaries 301 and 302 are illustrated. The row key points corresponding to each row boundary include 9×3 key points. Figure 3The row boundaries shown are the row boundaries after the table is bent and deformed.

[0061] Similarly, for each column boundary, a column line key point can be predicted. When predicting the column dividing line based on the column line key point, the key points belonging to different column boundaries can be divided according to the distance between the key points in the horizontal direction of the table, thereby realizing the prediction of the column dividing line for each column boundary.

[0062] In addition, in order to facilitate the distinction of key points belonging to each column boundary, multiple column line key points can be predicted for each column boundary, and the multiple column line key points include p×q key points, where p represents the number of longitudinal key points, i.e., the number of key points in each column, and q represents the number of transverse key points, i.e., the number of columns of key points. Both p and q are integers greater than or equal to 2, and the distance between each group of adjacent key points in the horizontal direction of the table is small, at least less than the width of the smallest cell in the table, which makes these key points relatively concentrated, so that the key points of each column boundary can be clearly distinguished. Among them, p and q are non-fixed values ​​and can be flexibly adjusted based on the size of the table. Among them, the number of key points and the height of the table can be positively correlated, i.e., the higher the height of the table, the more key points there are, so as to fit a more realistic column segmentation line.

[0063] In some optional embodiments, step 102: determining row line key points based on table row features, and determining column line key points based on table column features, may include:

[0064] Through the key point regression algorithm, the row line key points are determined based on the table row features, and the column line key points are determined based on the table column features.

[0065] In an embodiment of the present application, a key point regression algorithm can be used to regress some or all row and column line key points using a target detection model.

[0066] Among them, the key point regression algorithm is a technology used in the field of computer vision and image processing to predict the location of specific key points in an image. It can determine the key points from the image features through a machine learning model and output the coordinate values ​​of the key points.

[0067] Optionally, the key point regression algorithm described here can be a key point direct regression algorithm, which extracts features from the input image through a convolutional neural network (CNN) or other deep learning models, and directly outputs the coordinate values ​​of the key points through a fully connected layer or a convolutional layer, thereby quickly obtaining the required key points.

[0068] It is understandable that in addition to using the key point direct regression algorithm to determine the row and column key points, other algorithms can also be used to determine the key points, such as the cascade regression method, which gradually refines the coordinates of the key points through multiple stages of networks, and each stage of the network performs more precise positioning based on the previous stage; or deconstruction key point regression, which activates the pixels around the key point area through adaptive convolution, and uses these activated pixels to learn new features, thereby accurately regressing the key point coordinates, etc.

[0069] Optionally, in an embodiment of the present application, the model may be pre-trained to improve the model's prediction accuracy for key points.

[0070] During the training process, multiple key points can be manually annotated at the boundary of each row in the table image and multiple key points can be annotated at the boundary of each column in the table image for model learning. At the same time, the manually annotated key points can be used as true values ​​to compare with the key points output by the model, so as to adjust the model parameters based on the comparison results and improve the model accuracy.

[0071] Optionally, when multiple rows of row line key points need to be predicted for each row boundary and multiple columns of column line key points need to be predicted for each column boundary, a row of key points can be manually marked at each row boundary in the table image and a column of key points can be manually marked at each column boundary in the table image. Then, a computer can be used to shift each marked row line key point up and / or down by k pixels in the vertical direction of the table to obtain multiple rows of row line key points, and a computer can be used to shift each marked column line key point left and / or right by j pixels in the horizontal direction of the table to obtain multiple columns of column line key points.

[0072] Optionally, in an embodiment of the present application, all key points can be predicted by a key point regression algorithm. For example, when at least one row of key points needs to be predicted for each row boundary, the at least one row of key points can be predicted by a key point regression algorithm.

[0073] In the case where multiple rows of key points need to be predicted for each row boundary and multiple columns of key points need to be predicted for each column boundary, some key points can also be predicted by a key point regression algorithm, as described below:

[0074] The aforementioned step of determining row key points based on table row features and determining column key points based on table column features by using a key point regression algorithm may include steps A1 to A3 as follows:

[0075] Step A1: Determine multiple basic key points based on the tabular target features through a key point regression algorithm.

[0076] Wherein, the table target feature is the aforementioned table row feature or table column feature. When the table target feature is a table row feature, the multiple basic key points described here are row line key points; when the table target feature is a table column feature, the multiple basic key points described here are column line key points.

[0077] In the embodiment of the present application, a key point regression algorithm can be used to determine a row of basic key points for each row boundary based on the table row features, and the number of the row basic key points is multiple. Similarly, a key point regression algorithm can be used to determine a column basic key points for each column boundary based on the table column features, and the number of the column basic key points is multiple.

[0078] Step A2: Perform target operation on each target key point to obtain a new key point corresponding to the target key point.

[0079] The target key point described here is one of multiple row line basic key points, or one of multiple column line basic key points.

[0080] The target operation described here may include: for each target key point, moving k pixels in the target direction on the target coordinate axis; when the target key point is a row-based key point, the target direction includes at least one of the following: the positive direction of the y-axis and the negative direction of the y-axis, which can also be understood as the vertical direction of the table upward and the vertical direction of the table downward; when the target key point is a column-based key point, the target direction includes at least one of the following: the positive direction of the x-axis and the negative direction of the x-axis, which can also be understood as the horizontal direction of the table to the left and the horizontal direction of the table to the right. Wherein, the y-axis corresponds to the vertical direction of the table, and the x-axis corresponds to the horizontal direction of the table.

[0081] After obtaining the row-line basic key points, multiple pixels may be moved up and / or down on each row-line basic key point to obtain multiple row-line key points; and after obtaining the column-line basic key points, multiple pixels may be moved left and / or right on each column-line basic key point to obtain multiple column-line key points. It should be noted that the number of pixels moved up and down may be the same as or different from the number of pixels moved left and right.

[0082] Step A3: Based on multiple basic key points and corresponding new key points, all row key points or all column key points are obtained.

[0083] In an embodiment of the present application, multiple row line basic key points and corresponding new key points can be determined as final row line key points, and multiple column line basic key points and corresponding new key points can be determined as final column line key points.

[0084] It is understandable that after removing abnormal or redundant key points from multiple basic key points and corresponding new key points, the remaining key points can be determined as final key points. The abnormal key points mentioned here can be key points that are obviously deviated. The redundant key points mentioned here can be that the distance between two adjacent key points is smaller than the distance between other adjacent key points, or two key points overlap, in which case there are redundant key points.

[0085] In some optional embodiments, when each row boundary of the table corresponds to multiple row line key points and each column boundary of the table corresponds to multiple column line key points, step 103: predicting row segmentation lines based on row line key points and predicting column segmentation lines based on column line key points may include steps B1 to B2 as described below:

[0086] Step B1: Determine a final row line key point based on the coordinates of each group of longitudinal key points in the row line key points corresponding to each row boundary of the table; and determine a final column line key point based on the coordinates of each group of transverse key points in the column line key points corresponding to each column boundary of the table.

[0087] In the embodiment of the present application, when each row boundary of the table corresponds to multiple row key points, it is necessary to determine the final row key points of a row based on the multiple row key points, such as Figure 3 As shown, the final row of row key points can be determined based on the three rows of row key points in order to perform row prediction. The determination method can be to calculate the average coordinates of each group of longitudinal key points, and use the coordinate point corresponding to the average coordinate as a final row key point; or the key point at the target position in each group of longitudinal key points can be selected as a final row key point. For example, when each group of longitudinal key points includes three key points, the key point at the middle position (that is, the target position) can be preferably determined as a final row key point; of course, the first key point or the third key point from top to bottom can also be used as a final row key point. For another example, when each group of longitudinal key points includes two key points, a key point at any position can be used as a final row key point. For example, the first key point from top to bottom can be used as a final row key point, or the second key point from top to bottom can be used as a final row key point.

[0088] Similarly, when each column boundary of a table corresponds to multiple column line key points, it is necessary to determine the final column line key point of a column based on the multiple column line key points, such as Figure 3As shown, a final column of column line key points is determined based on three columns of key points in order to perform column line prediction. The determination method can be to calculate the average coordinates of each group of longitudinal key points, and use the coordinate point corresponding to the average coordinate as a final column line key point; or the key point at the target position in each group of longitudinal key points can be selected as a final column line key point. For example, when each group of transverse key points includes three key points, the key point at the middle position (that is, the target position) can be preferably determined as a final column line key point. Of course, the first key point or the third key point from left to right can also be used as a final column line key point. For another example, when each group of transverse key points includes two key points, a key point at any position can be used as a final column line key point. For example, the first key point from left to right can be used as a final column line key point, and the second key point from left to right can also be used as a final column line key point. Of course, this is just an example, and it can be set according to actual needs.

[0089] Step B2: predicting row segmentation lines based on the final row line key points, and predicting column segmentation lines based on each set of final column line key points.

[0090] After obtaining the final row line key points and the final column line key points, the row segmentation lines can be fitted based on the final row line key points, and the row segmentation lines can be fitted based on the final column line key points.

[0091] Since the key points in the embodiment of the present application are obtained based on the visual features of the rows and columns in the table image, the row and column dividing lines predicted based on the key points will be the same or similar to the row and column boundary shapes of the table in the image. Therefore, the structured table restored based on the row and column dividing lines can be more similar to the table in the image, that is, the table in the image can be restored more realistically, thereby improving the accuracy of the table recognition results.

[0092] Optionally, in the embodiment of the present application, in order to determine the correspondence between the row key points and the row dividing lines, for at least one row of row key points corresponding to each row boundary, the first key point of each row from left to right can be used as the row starting point, such as Figure 3 As shown, 303 in the figure represents the starting point of the row.

[0093] Similarly, in order to determine the correspondence between column line key points and column dividing lines, for at least one column of column line key points corresponding to each column boundary, the first key point from top to bottom of each column can be used as the column starting point.

[0094] In some optional embodiments, the aforementioned table obtained based on the prediction results of the row and column dividing lines is obviously inconsistent with the cell layout in the actual table and can be merged. In the prior art table recognition technology, additional modules can be relied on to detect basic elements such as text or cells, and then determine whether these elements belong to the same row, column or larger cells to parse the structure of the table. This method encounters challenges when processing tables containing many blank cells or empty rows and columns, because the lack of content makes it difficult to determine the relationship between basic elements.

[0095] To this end, the embodiment of the present application proposes a solution for merging cells based on the adjacency relationship between tables, which can quickly and accurately merge cells, as described below:

[0096] Wherein, step 104: generating a structured table based on the predicted row segmentation lines and column segmentation lines may include steps C1 to C2, as described below:

[0097] Step C1: Determine the correspondence between the target cell and the table cell by comparing the target cell with the table cell in the table image.

[0098] The target cell is obtained by the intersection of the row dividing line and the column dividing line.

[0099] Step C2: According to the corresponding relationship, the target cells are merged to generate a structured table.

[0100] The intersection of the predicted row and column split lines may destroy the same cell, for example, Figure 2 As shown, the cell where the text "A" is located is divided into 4 cells by the predicted row dividing line and column dividing line, and the cell where the text "E" is located is divided into 2 cells by the predicted row dividing line and column dividing line, etc. (see Figure 2 ), which is obviously inconsistent with the actual table in the image. Therefore, it is necessary to determine whether there is a need to merge cells after predicting the row and column dividing lines.

[0101] In an embodiment of the present application, the number of cells included in the table in the image and the positional relationship between the cells can be determined based on visual features; the number of target cells obtained by the intersection of row dividing lines and column dividing lines, the positional relationship between the target cells, etc. can also be determined.

[0102] By comparing the target cell with the table cell in the table image, the corresponding relationship between the target cell and the table cell can be determined.

[0103] Afterwards, based on the corresponding relationship, the target cell merging strategy can be determined, that is, the adjacency relationship between the target cells can be obtained, so as to quickly merge the target cells. The adjacency relationship mentioned here refers to the merging relationship between adjacent cells.

[0104] by Figure 2 Take the table shown as an example, see Figure 4 , the target cells obtained based on the predicted row and column segmentation lines can be regarded as a 4×4 matrix. After determining the merging strategy of the target cells based on the correspondence between the target cells and the table cells in the image, the processing strategy of each target cell can be represented by a number, where the number "0" can be used to indicate that the cell remains in the original position, the number "1" can be used to indicate that the cell needs to be merged to the left, and the number "2" can be used to indicate that the cell needs to be merged upward. The specific merging strategy of the target cells can be seen Figure 4 For cells that can be merged upward or left, the merge strategy can be randomly selected, such as Figure 4 The cell located in the 2nd row and 2nd column in the left picture can be merged upward or merged left. You can choose any one of the merge strategies for this cell.

[0105] Based on Figure 4 After merging the target cells using the merge strategy shown in the middle left figure, we can get Figure 4 The table shown in the middle right figure is a structured table that matches the table in the image.

[0106] Optionally, on the basis of merging cells based on the correspondence between the target cell and the table cell, the merging strategy can also be determined in combination with the text information in the table. For example, when continuous text content is divided into multiple cells, it indicates that at least these multiple cells need to be merged, such as Figure 2 As shown, in the row and column dividing line prediction results, the text content "B" is divided into two cells, so at least these two cells need to be merged, which helps to improve the accuracy of cell merging.

[0107] The above is a description of the table recognition method provided in the embodiment of the present application.

[0108] In summary, in the technical solution provided by the embodiment of the present application, in the process of table recognition, the row and column line key points are obtained based on the visual row and column features of the table, and then the row and column dividing lines are predicted by the row and column key points, so that the predicted row and column dividing lines can be adapted to the row and column shapes of the table in the image, and the continuity of the row and column dividing lines is improved. Even in the face of complex table situations, such as tilted tables, curved tables, etc., the table in the image can be restored more realistically, and the probability of dividing a row and column dividing line into multiple row and column dividing lines is reduced, thereby improving the accuracy of table structured recognition. And the technical solution provided by the present application does not require complex post-processing, so it can also improve the efficiency of table recognition. In addition, the embodiment of the present application can quickly merge cells based on the adjacency relationship between cells, which can further improve the efficiency of table structuring.

[0109] Exemplary Devices

[0110] Accordingly, the embodiment of the present application also provides a table recognition device, which can be applied to an electronic device with a table recognition function. The electronic device can be a terminal device, such as a mobile phone, a tablet computer, a desktop computer, a personal digital assistant device, etc., or a server.

[0111] like Figure 5 As shown, the device may include:

[0112] The feature extraction module 501 is used to extract visual features of a table image; wherein the visual features include: table row features and table column features.

[0113] The key point determination module 502 is used to determine the row line key points based on the table row characteristics, and determine the column line key points based on the table column characteristics.

[0114] The segmentation line prediction module 503 is used to predict the row segmentation lines based on the row line key points, and to predict the column segmentation lines based on the column line key points.

[0115] The table generating module 504 is configured to generate a structured table based on the predicted row dividing lines and the column dividing lines.

[0116] In some optional embodiments, the row line key points corresponding to each row boundary of the table include n×m key points, n represents the number of horizontal key points of the table, m represents the number of vertical key points of the table, and both m and n are integers greater than or equal to 2.

[0117] In some optional embodiments, the column line key points corresponding to each column boundary of the table include p×q key points, p represents the number of vertical key points of the table, q represents the number of horizontal key points of the table, and both p and q are integers greater than or equal to 2.

[0118] In some optional embodiments, the key point determination module 502 may be specifically used to:

[0119] The row line key points are determined based on the table row features, and the column line key points are determined based on the table column features through a key point regression algorithm.

[0120] In some optional embodiments, the key point determination module 502 may also be specifically used for:

[0121] Through the key point regression algorithm, multiple basic key points are determined based on the table target features; and target operations are performed on each target key point to obtain new key points corresponding to the target key points; and based on the multiple basic key points and the corresponding new key points, all the row line key points or all the column line key points are obtained.

[0122] Among them, the table target feature is the table row feature or the table column feature; when the table target feature is the table row feature, the multiple basic key points are row line key points; when the table target feature is the table column feature, the multiple basic key points are column line key points.

[0123] Among them, the target key point is one of the multiple basic key points; the target operation includes: for each of the target key points, moving k pixels in the target direction on the target coordinate axis; when the target key point is a row line key point, the target coordinate axis is the y-axis, and the target direction includes at least one of the following: the positive direction of the y-axis and the negative direction of the y-axis; when the target key point is a column line key point, the target coordinate axis is the x-axis, and the target direction includes at least one of the following: the positive direction of the x-axis and the negative direction of the x-axis; the y-axis corresponds to the vertical direction of the table, and the x-axis corresponds to the horizontal direction of the table.

[0124] In some optional embodiments, when each row boundary of a table corresponds to multiple row line key points and each column boundary of a table corresponds to multiple column line key points, the segmentation line prediction module 503 may be specifically used to:

[0125] Based on the coordinates of each group of longitudinal key points in the row line key points corresponding to each row boundary of the table, a final row line key point is determined; and based on the coordinates of each group of transverse key points in the column line key points corresponding to each column boundary of the table, a final column line key point is determined; based on the final row line key points, the row dividing lines are predicted, and based on each group of the final column line key points, the column dividing lines are predicted.

[0126] In some optional embodiments, the table generation module 504 may include:

[0127] The correspondence determination unit is used to determine the correspondence between the target cell and the table cell in the table image by comparing the target cell with the table cell; wherein the target cell is obtained by the intersection of the row dividing line and the column dividing line.

[0128] A table generating unit is used to merge the target cells according to the corresponding relationship to generate the structured table.

[0129] In some optional embodiments, when the target position in the table has a table line, the table target feature is a table line feature; when the target position in the table does not have a table line, the table target feature is a region feature where the target position is located.

[0130] The target position is a table row boundary position, or the target position is a table column boundary position, and the table target feature is the table row feature or the table column feature.

[0131] The table recognition device provided in this embodiment belongs to the same application concept as the table recognition method provided in the above embodiments of this application, and can execute the table recognition method provided in any of the above embodiments of this application, and has the corresponding functional modules and beneficial effects of the execution method. For technical details not fully described in this embodiment, please refer to the specific processing content of the table recognition method provided in the above embodiments of this application, and will not be repeated here.

[0132] It should be understood that the modules in the above table recognition device can be implemented in the form of a processor calling software. For example, the device includes a processor, the processor is connected to a memory, and instructions are stored in the memory. The processor calls the instructions stored in the memory to implement any of the above methods or realize the functions of each unit of the device, wherein the processor can be a general-purpose processor, such as a CPU or a microprocessor, etc., and the memory can be a memory in the device or a memory outside the device. Alternatively, the unit in the device can be implemented in the form of a hardware circuit, and the functions of some or all units can be realized by designing the hardware circuit. The hardware circuit can be understood as one or more processors; for example, in one implementation, the hardware circuit is an ASIC, and the functions of some or all units above are realized by designing the logical relationship of the components in the circuit; for another example, in another implementation, the hardware circuit can be implemented by PLD, taking FPGA as an example, which can include a large number of logic gate circuits, and the connection relationship between the logic gate circuits is configured by a configuration file, so as to realize the functions of some or all units above. All units of the above device can be implemented in the form of a processor calling software, or in the form of a hardware circuit, or in part by a processor calling software, and the remaining part is implemented in the form of a hardware circuit.

[0133] In an embodiment of the present application, a processor is a circuit with the ability to process signals. In one implementation, the processor may be a circuit with the ability to read and run instructions, such as a CPU, a microprocessor, a GPU, or a DSP; in another implementation, the processor may implement certain functions through the logical relationship of a hardware circuit, and the logical relationship of the hardware circuit is fixed or reconfigurable, such as a hardware circuit implemented by an ASIC or PLD, such as an FPGA. In a reconfigurable hardware circuit, the process of the processor loading a configuration document to implement the hardware circuit configuration can be understood as the process of the processor loading instructions to implement the functions of some or all of the above units. In addition, it can also be a hardware circuit designed for artificial intelligence, which can be understood as an ASIC, such as an NPU, TPU, DPU, etc.

[0134] It can be seen that each unit in the above device can be one or more processors (or processing circuits) configured to implement the above method, such as: CPU, GPU, NPU, TPU, DPU, microprocessor, DSP, ASIC, FPGA, or a combination of at least two of these processor forms.

[0135] In addition, all or part of the units in the above device can be integrated together, or can be implemented independently. In one implementation, these units are integrated together and implemented in the form of a SOC. The SOC may include at least one processor for implementing any of the above methods or implementing the functions of each unit of the device. The type of the at least one processor may be different, for example, including a CPU and an FPGA, a CPU and an artificial intelligence processor, a CPU and a GPU, etc.

[0136] Exemplary Electronic Devices

[0137] The present application also provides an electronic device, such as Figure 6 As shown, the electronic device includes: a memory 600 and a processor 610 .

[0138] The memory 600 is connected to the processor 610 and is used to store programs.

[0139] The processor 610 is used to implement the table recognition method in the above embodiment by running the program stored in the memory 600 .

[0140] Specifically, the electronic device may further include: a communication interface 620 , an input device 630 , an output device 640 and a bus 650 .

[0141] The processor 610, the memory 600, the communication interface 620, the input device 630 and the output device 640 are connected to each other via a bus.

[0142] Bus 650 may include a pathway for transferring information between the various components of the computer system.

[0143] The processor 610 may be a general-purpose processor, such as a general-purpose central processing unit (CPU), a microprocessor, etc., or an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the program of the scheme of the present invention. It may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0144] The processor 610 may include a main processor, and may also include a baseband chip, a modem, and the like.

[0145] The memory 600 stores a program for executing the technical solution of the present invention, and may also store an operating system and other key services. Specifically, the program may include a program code, and the program code includes a computer operation instruction. More specifically, the memory 600 may include a read-only memory (ROM), other types of static storage devices that can store static information and instructions, a random access memory (RAM), other types of dynamic storage devices that can store information and instructions, a disk storage, a flash, and the like.

[0146] The input device 630 may include a device for receiving data and information input by a user, such as a keyboard, a mouse, a camera, a scanner, a light pen, a voice input device, a touch screen, a pedometer, or a gravity sensor.

[0147] Output device 640 may include a device that allows information to be output to a user, such as a display screen, printer, speaker, etc.

[0148] The communication interface 620 may include any transceiver or the like to communicate with other devices or communication networks, such as Ethernet, a radio access network (RAN), a wireless local area network (WLAN), etc.

[0149] The processor 610 executes the program stored in the memory 600 and calls other devices, which can be used to implement the various steps of the table recognition method provided in the above embodiment of the present application.

[0150] Exemplary computer program products and storage media

[0151] In addition to the above-mentioned method and device, the embodiment of the present application may also be a computer program product, which includes computer program instructions, which, when executed by a processor, enable the processor to execute the steps in the table recognition method described in the embodiment of the present application.

[0152] The computer program product may be implemented in hardware, software or a combination thereof. In one optional embodiment, the computer program product is implemented as a computer storage medium. In another optional embodiment, the computer program product is implemented as a software product, such as a software development kit (SDK).

[0153] The computer program product may be written in any combination of one or more programming languages ​​to write program codes for performing the operations of the embodiments of the present application, including object-oriented programming languages, such as Java, C++, etc., and conventional procedural programming languages, such as "C" language or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as an independent software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0154] In addition, the embodiment of the present application may also be a storage medium on which a computer program is stored, and the computer program is executed by a processor to execute the steps of the table recognition method described in the embodiment of the present application.

[0155] For the aforementioned method embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the order of the actions described, because according to the present application, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application.

[0156] It should be noted that each embodiment in this specification is described in a progressive manner, and each embodiment focuses on the differences from other embodiments, and the same or similar parts between the embodiments can be referred to each other. For the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0157] The steps in the methods of each embodiment of the present application can be adjusted in order, combined and deleted according to actual needs, and the technical features recorded in each embodiment can be replaced or combined.

[0158] The modules and sub-modules in the devices and terminals in the various embodiments of the present application can be combined, divided and deleted according to actual needs.

[0159] In the several embodiments provided in the present application, it should be understood that the disclosed terminals, devices and methods can be implemented in other ways. For example, the terminal embodiments described above are only schematic, for example, the division of modules or submodules is only a logical function division, and there may be other division methods in actual implementation, for example, multiple submodules or modules can be combined or integrated into another module, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or modules, which can be electrical, mechanical or other forms.

[0160] The modules or submodules described as separate components may or may not be physically separated, and the components of the modules or submodules may or may not be physical modules or submodules, that is, they may be located in one place, or they may be distributed on multiple network modules or submodules. Some or all of the modules or submodules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0161] In addition, each functional module or submodule in each embodiment of the present application may be integrated into one processing module, or each module or submodule may exist physically separately, or two or more modules or submodules may be integrated into one module. The above-mentioned integrated modules or submodules may be implemented in the form of hardware or in the form of software functional modules or submodules.

[0162] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0163] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly by hardware, software units executed by a processor, or a combination of the two. The software units may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0164] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise one..." do not exclude the existence of other identical elements in the process, method, article or device including the elements. The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A table recognition method, characterized in that: The method comprises: Extracting visual features of a table image; wherein the visual features include: table row features and table column features; Determine row line key points based on the table row features, and determine column line key points based on the table column features; Predicting row segmentation lines based on the row line key points, and predicting column segmentation lines based on the column line key points; A structured table is generated based on the predicted row dividing lines and the column dividing lines.

2. The table recognition method according to claim 1, characterized in that: The row key points corresponding to each row boundary of the table include n×m key points, where n represents the number of horizontal key points of the table, and m represents the number of vertical key points of the table, and both m and n are integers greater than or equal to 2; The column line key points corresponding to each column boundary of the table include p×q key points, p represents the number of longitudinal key points of the table, q represents the number of transverse key points of the table, and both p and q are integers greater than or equal to 2.

3. The table recognition method according to claim 1 or 2, characterized in that: The determining of the row line key points based on the table row features, and the determining of the column line key points based on the table column features, comprises: The row line key points are determined based on the table row features, and the column line key points are determined based on the table column features through a key point regression algorithm.

4. The table recognition method according to claim 3, characterized in that: The determining the row line key points based on the table row features and the determining the column line key points based on the table column features by using a key point regression algorithm includes: Determine a plurality of basic key points based on the table target features by a key point regression algorithm; the plurality of basic key points are row line key points, or the plurality of basic key points are column line key points; Performing a target operation on each target key point to obtain a new key point corresponding to the target key point; wherein the target key point is one of the multiple basic key points; the target operation includes: for each target key point, moving k pixels in the target direction on the target coordinate axis; Based on the multiple basic key points and the corresponding new key points, all the row line key points or all the column line key points are obtained.

5. The table recognition method according to claim 2, characterized in that: In a case where each row boundary of a table corresponds to a plurality of row line key points and each column boundary of a table corresponds to a plurality of column line key points, predicting row segmentation lines based on the row line key points and predicting column segmentation lines based on the column line key points include: Determine a final row line key point based on the coordinates of each group of longitudinal key points in the row line key points corresponding to each row boundary of the table; and determine a final column line key point based on the coordinates of each group of transverse key points in the column line key points corresponding to each column boundary of the table; The row segmentation lines are predicted based on the final row line keypoints, and the column segmentation lines are predicted based on each set of the final column line keypoints.

6. The table recognition method according to claim 1, characterized in that: The step of generating a structured table based on the predicted row dividing line and the column dividing line includes: By comparing the target cell with the table cell in the table image, determining the correspondence between the target cell and the table cell; wherein the target cell is obtained by the intersection of the row dividing line and the column dividing line; According to the corresponding relationship, the target cells are merged to generate the structured table.

7. The table recognition method according to claim 1, characterized in that: When the target position in the table has a table line, the table target feature is the table line feature; when the target position in the table does not have a table line, the table target feature is the region feature where the target position is located; The target position is a table row boundary position, or the target position is a table column boundary position, and the table target feature is the table row feature or the table column feature.

8. A table recognition device, characterized in that: The device comprises: A feature extraction module, used to extract visual features of a table image; wherein the visual features include: table row features and table column features; A key point determination module, used to determine row line key points based on the table row characteristics, and to determine column line key points based on the table column characteristics; A segmentation line prediction module, used for predicting row segmentation lines based on the row line key points, and predicting column segmentation lines based on the column line key points; The table generation module is used to generate a structured table based on the predicted row dividing lines and the column dividing lines.

9. An electronic device, characterized in that: include: Memory and processor; The memory is connected to the processor and is used to store programs; The processor is used to implement the table recognition method according to any one of claims 1 to 7 by running the program in the memory.

10. A storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by the processor, the table recognition method according to any one of claims 1 to 7 is implemented.

11. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the table recognition method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Table extraction method and device, computer equipment and storage medium

    CN110287854A

  • Table reconstruction method and device, computer equipment and storage medium

    CN112200822A

  • Table row and column identification method and system based on depth image segmentation

    CN113743253A

  • Table recognition method and device, electronic equipment and storage medium

    CN115620321A

  • Table structure identification method, apparatus and device, and storage medium

    CN116343246A

Cited By

  • Table structuring method, related equipment and program product

    CN121033877A

  • A table structuring method, related device and program product

    CN121033877B