Table structure identification method based on Transform rank attention mechanism

By combining row and column attention mechanism (RCAM) with global attention to the table header, the problem of insufficient utilization of row and column relationships in table image recognition in existing technologies is solved, achieving more efficient and accurate table structure recognition.

CN120997856APending Publication Date: 2025-11-21FUJIAN FOXIT SOFTWARE DEV LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410634364.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-05-21
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing technologies fail to fully utilize the row and column relationships in table image recognition, resulting in a decrease in the model's ability to recognize non-table data, a lack of learning of cell text features, and high computational resource consumption.

Method used

A Transformer-based row and column attention mechanism (RCAM) is adopted, which uses a row and column mask matrix to ensure that feature positions only interact with features in the same row and column. Combined with global header attention, this improves the model's ability to recognize table structures.

Benefits of technology

It improves the efficiency and accuracy of tabular data processing, especially the accuracy of header information recognition, reduces the computational burden, and enhances the overall recognition effect of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120997856A_ABST
    Figure CN120997856A_ABST
Patent Text Reader

Abstract

The invention discloses a table structure recognition method based on a Transform row and column attention mechanism. The table structure recognition method comprises the steps that S1, a table picture is divided into image blocks in multiple rows and columns through a segmentation mode with the fixed width and height; s2, flattening the divided image blocks in two-dimensional arrangement into one-dimensional image blocks; s3, processing the one-dimensional image blocks into image vectors, performing linear projection, merging the one-dimensional position embedding and [CLS] vectors of the image blocks, and then taking the merged result as the input of the model; and S4, feature extraction based on an attention mechanism: excluding image blocks in different rows / columns through the row and column mask matrix to ensure that the feature of each position is only calculated with the features from the same row and column, obtaining a final matrix by using row and column attention and a global attention mechanism of a header, then extracting the feature vector of the image, and obtaining the feature vector of the image. Performing feature pyramid on the features to obtain multi-scale features; and S5, accessing a target detection model to detect a target on the feature map, and finally obtaining information of a table structure.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image recognition technology, and more specifically, to a table structure recognition method based on the Transformer row and column attention mechanism. Background Technology

[0002] The use of computer vision and deep learning technologies to recognize table images has become a hot research topic in this field in recent years. For example, the invention disclosed in CN116259064A, "A Method and Apparatus for Recognizing Table Structures and a Training Method for a Table Structure Recognition Model," obtains grid line information and image features through an attention mechanism to acquire the relative position and primary structural information of the grid lines. However, because this invention relies on table lines and lacks learning of cell text features, its ability to recognize tables without grid lines is somewhat reduced. Another example is the invention disclosed in CN117237968A, which discloses a Transformer-based row-by-row scanning method and system for recognizing table structures. It designs a decoding module according to the two-dimensional form of table logic and uses a semi-autoregressive approach to generate table sequences. However, this invention mainly relies on row features and lacks feature extraction from columns.

[0003] In addition, some scholars abroad have also researched methods for recognizing table structures. For example, in "An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale" by Dosovitskiy A et al., a global attention mechanism was proposed to learn global features. However, this method leads to unrelated regional features being treated as important features, and the difficulty in processing long sequences exacerbates the consumption of computational resources. Another example is "Longformer: The Long-Document Transformer" by Beltagy I et al., which mentions a local attention mechanism, but its fixed window mechanism cannot adapt to changes in the importance of different regions in a table image.

[0004] The main problem with the existing technology is that when processing table images, it cannot make full use of the row and column relationships of the table to improve the attention accuracy of the model, does not make full use of the text features of the cells, resulting in a decrease in the ability to recognize non-table data, and does not make full use of the structural features of the table to reduce the area of ​​attention. Summary of the Invention

[0005] To address the aforementioned issues, this invention provides a table structure recognition method based on the Transformer Row and Column Attention Mechanism (RCAM). By combining the row and column structure features of the table with cell information, the RCAM ensures through a row and column mask matrix that during the attention calculation process, the features at each position only interact with the features of the same row and column, thereby improving the processing efficiency and accuracy of table data.

[0006] The Row and Column Attention Mechanism (RCAM) attention calculation method provided in this invention is based on the design of a row and column mask matrix. This ensures that feature positions can only interact with features in the same row and column during attention calculation, thus enhancing the model's ability to recognize table structures. Simultaneously, to emphasize the influence of the table header on the overall table structure, global attention calculation for the table header information is added, improving the model's attention to and accuracy in recognizing table header information.

[0007] To achieve the above objectives, this invention provides a table structure recognition method based on the Transformer row and column attention mechanism, which includes the following steps:

[0008] Step S1: Divide the table image into two-dimensional image blocks with multiple rows and columns using a fixed width and height division method;

[0009] Step S2: Flatten the divided two-dimensional image blocks into one-dimensional image blocks;

[0010] Step S3: Process the one-dimensional image patch into an image vector and project it linearly. At the same time, merge the one-dimensional position embedding of the image patch and the [CLS] vector. The [CLS] vector is added at the beginning of the image vector so that the model can use the category information for classification. Then, use the merged result as the input of the model.

[0011] Step S4: Feature extraction based on attention mechanism. Image blocks that are not in the same row / column are excluded by row and column mask matrix to ensure that the feature at each position is only calculated with the feature from the same row and column. The final matrix is ​​obtained by using row and column attention and global attention mechanism of table header. Then the feature vector of the image is extracted and the feature pyramid is made to obtain multi-scale features.

[0012] Step S5: After extracting features in step S4, the target detection model is used to detect targets on the feature map. The target detection results include row detection boxes, column detection boxes, and cell detection boxes, ultimately obtaining information about the table structure.

[0013] In one embodiment of the present invention, step S1 specifically comprises:

[0014] Given the original image H, W, and C represent the height, width, and number of channels of the image, respectively. Image I is divided into P×P non-overlapping image blocks by linear projection, where P is the side length of the image block.

[0015] In one embodiment of the present invention, the one-dimensional image block flattened in step S2 is represented as follows: Where N×N represents the number of image patches, and All image blocks are numbered using numerical values, with a length of N×N, starting from 1.

[0016] In one embodiment of the present invention, the specific process of obtaining the final matrix in step S4 is as follows:

[0017] The row and column mask matrix M is Where N×N still represents the number of image patches, and the specific definition of the row and column mask matrix M is:

[0018]

[0019] Among them, M ij Let i be the image patch number to be calculated, j be the image patch number to be used for attention calculation with i, and i and j be in the range (1, N×N). w represents the width of the original image divided by the width of the image patch. The modulo operator indicates rounding down; mod indicates congruence.

[0020] 0 represents the row and column values ​​in the row and column mask matrix corresponding to the row and column positions of all image patches to be calculated; -1e 10 This represents a large negative number to eliminate the weight of numbers that are not in the same row or column;

[0021] Since the header occupies the image block of the first y rows, where y is a positive integer, global attention is calculated for the first y rows. Specifically, during the calculation, all row and column results of the first y rows in the row and column mask matrix M are set to 0.

[0022] Before all image patch vectors, there is a vector [CLS] used to calculate the attention for each row and column image patch to obtain global features. The final row and column attention mask matrix is ​​as follows:

[0023]

[0024] In the formula, M is the row and column mask matrix; Q, K, and V are the three input representation vectors of the self-attention mechanism, where Q represents the query vector, K represents the key vector, and V represents the numerical vector; d kThe dimension of the vector / hidden layer; This indicates element-wise addition.

[0025] In one embodiment of the present invention, the target detection model in step S5 includes Mask-RCNN or CascadeR-CNN.

[0026] The table structure recognition method based on the Transformer row and column attention mechanism provided by this invention, compared with the prior art, can effectively combine the row and column structure features of the table through the row and column attention mechanism (RCAM), thereby improving the model's performance in table structure recognition. The introduced global row and column attention can better capture the information of the table header, which plays a guiding role in the overall structure of the table. The combination of the two can complete the function of table structure recognition. Attached Figure Description

[0027] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0028] Figure 1 This is a flowchart illustrating an embodiment of the present invention;

[0029] Figure 2 This is a schematic diagram of image block division according to an embodiment of the present invention;

[0030] Figure 3 This is a schematic diagram of image block tiling according to an embodiment of the present invention;

[0031] Figure 4 This is a schematic diagram of image block embedding according to an embodiment of the present invention;

[0032] Figure 5 This is the final matrix obtained by using only row and column attention for feature extraction according to an embodiment of the present invention;

[0033] Figure 6 This is the final matrix obtained by feature extraction using a global attention mechanism with row and column attention tables in one embodiment of the present invention. Detailed Implementation

[0034] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0035] Figure 1 Therefore, this embodiment provides a table structure recognition method based on the Transformer row and column attention mechanism, which includes the following steps:

[0036] Step S1: Image block division. The table image is divided into two-dimensional image blocks with multiple rows and columns using a fixed width and height division method.

[0037] In this embodiment, step S1 specifically includes:

[0038] Given the original image H, W, and C represent the height, width, and number of channels of the image, respectively. Image I is divided into P×P non-overlapping image blocks by linear projection, where P is the side length of the image block.

[0039] For example, if the original image is resized to better support the model's input, the original image will be scaled to a size of 224×224, that is, H and W are both 224, and divided into 16×16 image blocks according to the number of 14×14 image blocks, that is, P is 16.

[0040] Step S2: Image block tiling, flattening the divided two-dimensional image blocks into one-dimensional image blocks; the image blocks obtained in step S1 can be tiled, and the original two-dimensional image blocks will become one-dimensional image blocks.

[0041] In this embodiment, the one-dimensional image block flattened in step S2 is represented as: Where N×N represents the number of image patches, and All image blocks are numbered using numerical values, with a length of N×N, starting from 1.

[0042] Continuing with the example image above, the number of image blocks is 14×14, that is, N is 14.

[0043] Step S3: Image patch embedding. The one-dimensional image patch is processed into an image vector and linearly projected. At the same time, the one-dimensional position embedding of the image patch and the [CLS] vector are merged. The [CLS] vector is added at the beginning of the image vector so that the model can use the category information for classification. Then the merged result is used as the input of the model. The [CLS] vector represents the category information of the entire image.

[0044] For example, one-dimensionally arranged image blocks can be represented by image embedding to obtain image vectors, and a [CLS] vector can be inserted at the beginning of the image vectors to obtain the category of global information. Then, the image vectors containing the [CLS] vectors are copied row by row to obtain a vector matrix.

[0045] Step S4: Feature extraction based on attention mechanism. Image blocks that are not in the same row / column are excluded by row and column mask matrix to ensure that the feature at each position is only calculated with the feature from the same row and column, thereby limiting the attention range and reducing the computational burden. The final matrix is ​​obtained by using row and column attention and global attention mechanism of table header. Then the feature vector of the image is extracted and the feature pyramid is made to obtain multi-scale features.

[0046] In this embodiment, the specific process of obtaining the final matrix in step S4 is as follows:

[0047] The row and column mask matrix M is Where N×N still represents the number of image patches, and the specific definition of the row and column mask matrix M is:

[0048]

[0049] Among them, M ij Let i be the image patch number to be calculated, j be the image patch number to be used for attention calculation with i, and i and j range from (1, N×N). w represents the width of the original image divided by the width of the image patch, which can be called the two-dimensional feature width; (symbol) The expression indicates rounding down, and mod indicates congruence. Therefore, i≠j(mod w) means that i and j have the same remainder after being divided by w.

[0050] 0 represents the row and column values ​​in the row and column mask matrix corresponding to the row and column positions of all image patches to be calculated; -1e 10 Represents larger negative numbers to eliminate weights from numbers that are not in the same row or column;

[0051] Considering that the table header hides the overall table structure, when introducing global attention at the header position, since the header may occupy the image blocks of the first y rows, global attention is calculated for the first y rows. Specifically, during the calculation, all row and column results in the first y rows of the row and column mask matrix M are set to 0, where y is a positive integer, which can be set according to the specific use case; this implementation does not restrict its specific value. Note that there is a vector [CLS] before all image block vectors, which is used to calculate the attention of each row and column image block to obtain global features. The final row and column attention mask matrix is ​​as follows:

[0052]

[0053] In the formula, M is the row and column mask matrix; Q, K, and V are the three input representation vectors of the self-attention mechanism, where Q represents the query vector, K represents the key vector, and V represents the numerical vector. These three vectors are obtained from the vectors of the original image patch through linear transformation; d k Choose the dimension of the vector / hidden layer. This allows Q×K to conform to the distribution of N(0,1), similar to normalization; This indicates element-wise addition. In this embodiment, the row and column mask matrix M needs to be used after multi-head attention calculation and before the Softmax operation.

[0054] Step S5: Predicting the detection head result. This involves using the features extracted in Step S4 and feeding them into the object detection model to detect targets on the feature map. The object detection result should output three types of results: row detection boxes 101, column detection boxes 102, and cell detection boxes 103, ultimately obtaining information about the table structure. The object detection model can be, for example, Mask-RCNN or Cascade R-CNN, but is not limited to these.

[0055] Figure 2 This is a schematic diagram of image block division according to an embodiment of the present invention. Figure 3 This is a schematic diagram of image block tiling according to an embodiment of the present invention. Figure 4 This is a schematic diagram of image block embedding according to an embodiment of the present invention. Figure 5 This is the final matrix obtained by using only row and column attention for feature extraction according to an embodiment of the present invention. Figure 6 This is the final matrix obtained by feature extraction using a global attention mechanism with row and column attention tables according to an embodiment of the present invention. For a clearer explanation of the aforementioned process, see [link to documentation]. Figures 2-6 and combined Figure 1 Taking an image patch of 3x3 as an example, the attention calculation in the second row and second column is explained in detail below:

[0056] First, the given original table image will be divided into 3x3 two-dimensional image blocks, such as Figure 2 ;

[0057] Then, the image patches are flattened into one-dimensional image patches, each with a corresponding number, such as 1 to 9, for a total of 9 patches. Afterwards, an embedding operation is performed to obtain an image vector, i.e., flattened into a single image vector. A [CLS] vector marker is added before this image vector, resulting in 3 x 3 + 1 = 10 vectors, as shown below. Figure 3 As shown;

[0058] Next, the merged vector is copied 3x3+1 times to form a 10x10 matrix, i.e., a vector matrix, as shown below. Figure 4Use this matrix as the input to the model, noting that each image patch will be processed into a vector;

[0059] Next, since the attention calculation is performed on the second row and second column, that is, for the image block with the sequence number 5, the row and column attention calculation needs to be combined with its row and column, referring to formula (1) and Figure 2 We can obtain the image patch numbers corresponding to the rows as 4 and 6, and the image patch numbers corresponding to the columns as 2 and 8, as well as the global marker CLS. We set the values ​​at position M of the row and column mask matrix corresponding to these numbers to 0, such as... Figure 5 Medium-light gray area;

[0060] Assuming the table header occupies the first 3 rows, meaning global row and column attention is introduced for the first 3 rows, then the 3 rows other than the first row will all participate in the calculation. The complete calculation position is as follows: Figure 6 The area shown in medium-light gray;

[0061] exist Figure 6 In the diagram, the dark gray area represents the self-attention calculation part, the light gray area represents the global row and column attention calculation, and the gray area represents the row and column attention calculation.

[0062] The table structure recognition method based on the Transformer row and column attention mechanism provided by this invention can effectively combine the row and column structure features of the table through the row and column attention mechanism (RCAM), thereby improving the model's performance in table structure recognition. The introduced global row and column attention can better capture the information of the table header, which plays a guiding role in the overall structure of the table. The combination of the two can complete the function of table structure recognition.

[0063] Table 1 shows the experiments conducted using the PubTables-1M dataset to evaluate the performance of the proposed method. Standard object detection metrics AP, AP50, AP75, AR, and table context exact matching accuracy Acc were used. Con To conduct an evaluation, among which Acc Con The metrics consider the proportion of tables that achieve an exact match for all cells (even blank cells). A comparative experiment was conducted between the Vit scheme (An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale) from the background technique and TATR (PubTables-1M: Towards comprehensive table extraction from unstructured documents) proposed for PubTables-1M. Clearly, RCAM performs better across all metrics.

[0064] Table 1

[0065] Model AP AP50 AP75 AR <![CDATA[Acc Con ]]> TATR 0.902 0.970 0.941 0.935 0.824 Vit 0.893 0.953 0.937 0.932 0.812 RCAM 0.906 0.976 0.943 0.938 0.825

[0066] In addition, to verify the impact of the table header on RCAM, a set of experiments was designed, as shown in Table 2, to examine the impact of the entire table structure recognition on performance when the image blocks in the first n rows are the table headers. It can be observed that for PubTables-1M data, the optimal value of n is 3. However, for real-world scenarios, the specific value of n needs to be carefully considered.

[0067] Table 2

[0068] AP AP50 AP75 AR <![CDATA[Acc Con ]]> First row 0.902 0.970 0.941 0.935 0.824 First 2 rows 0.905 0.973 0.942 0.938 0.826 First 3 rows 0.915 0.986 0.952 0.942 0.828 First 4 rows 0.904 0.981 0.950 0.940 0.822 First 5 rows 0.898 0.962 0.940 0.935 0.817

[0069] Those skilled in the art will understand that the accompanying drawings are merely schematic diagrams of one embodiment, and the modules or processes shown in the drawings are not necessarily essential for implementing the present invention.

[0070] Those skilled in the art will understand that the modules in the apparatus of the embodiments can be distributed in the apparatus of the embodiments as described in the embodiments, or they can be located in one or more devices different from this embodiment with corresponding changes. The modules of the above embodiments can be combined into one module, or they can be further divided into multiple sub-modules.

[0071] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A table structure recognition method based on Transformer row and column attention mechanism, characterized in that, Includes the following steps: Step S1: Divide the table image into two-dimensional image blocks with multiple rows and columns using a fixed width and height division method; Step S2: Flatten the divided two-dimensional image blocks into one-dimensional image blocks; Step S3: Process the one-dimensional image patch into an image vector and project it linearly. At the same time, merge the one-dimensional position embedding of the image patch and the [CLS] vector. The [CLS] vector is added at the beginning of the image vector so that the model can use the category information for classification. Then, use the merged result as the input of the model. Step S4: Feature extraction based on attention mechanism. Image blocks that are not in the same row / column are excluded by row and column mask matrix to ensure that the feature at each position is only calculated with the feature from the same row and column. The final matrix is ​​obtained by using row and column attention and global attention mechanism of table header. Then the feature vector of the image is extracted and the feature pyramid is made to obtain multi-scale features. Step S5: After extracting features in step S4, the target detection model is used to detect targets on the feature map. The target detection results include row detection boxes, column detection boxes, and cell detection boxes, ultimately obtaining information about the table structure.

2. The table structure recognition method based on Transformer row and column attention mechanism according to claim 1, characterized in that, Step S1 is as follows: Given the original image H, W, and C represent the height, width, and number of channels of the image, respectively. Image I is divided into P×P non-overlapping image blocks by linear projection, where P is the side length of the image block.

3. The table structure recognition method based on Transformer row and column attention mechanism according to claim 1, characterized in that, The one-dimensional image patch flattened in step S2 is represented as follows: Where N×N represents the number of image patches, and All image blocks are numbered using numerical values, with a length of N×N, starting from 1.

4. The table structure recognition method based on Transformer row and column attention mechanism according to claim 1, characterized in that, The specific process for obtaining the final matrix in step S4 is as follows: The row and column mask matrix M is Where N×N still represents the number of image patches, and the specific definition of the row and column mask matrix M is: Among them, M ij Let i be the image patch number to be calculated, j be the image patch number to be used for attention calculation with i, and i and j be in the range (1, N×N). w represents the width of the original image divided by the width of the image patch. The modulo operator indicates rounding down; mod indicates congruence. 0 represents the row and column values ​​in the row and column mask matrix corresponding to the row and column positions of all image patches to be calculated; -1e 10 This represents a large negative number to eliminate the weight of numbers that are not in the same row or column; Since the header occupies the image block of the first y rows, where y is a positive integer, global attention is calculated for the first y rows. Specifically, during the calculation, all row and column results of the first y rows in the row and column mask matrix M are set to 0. Before all image patch vectors, there is a vector [CLS] used to calculate the attention for each row and column image patch to obtain global features. The final row and column attention mask matrix is ​​as follows: In the formula, M is the row and column mask matrix; Q, K, and V are the three input representation vectors of the self-attention mechanism, where Q represents the query vector, K represents the key vector, and V represents the numerical vector; d k The dimension of the vector / hidden layer; This indicates element-wise addition.

5. The table structure recognition method based on Transformer row and column attention mechanism according to claim 1, characterized in that, The object detection model in step S5 includes Mask-RCNN or Cascade R-CNN.

Citation Information

Patent Citations

  • Table structure recognition method and training method and device of table structure recognition model

    CN116259064A

  • Transform-based progressive scanning table structure identification method and system

    CN117237968A