Table structure recognition model training method, electronic equipment and storage medium

By combining the model training methods of image encoder, layout encoder and logical structure decoder, the text area misalignment and complex table recognition problems in table structure recognition are solved, and a more efficient and general table structure recognition effect is achieved.

CN120599644AInactive Publication Date: 2025-09-05TIANJIN ANXIN DIGITAL TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510746677.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-09-05
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art has problems such as text area misalignment, difficulty in complex table recognition and poor universality in table structure recognition, especially when dealing with watermarked documents or non-English fields.

Method used

The combined model of image encoder, layout encoder, feature weighting module and logical structure decoder is adopted to train the model through the target loss function, including label classification loss, layout pointer loss and span perceived comparison loss, to enhance the ability to identify table structures.

Benefits of technology

Effectively avoiding text area misalignment, improve the accuracy and efficiency of identification of complex table structures, and improve the universality and competitiveness of the model in industrial document recognition scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120599644A_ABST
    Figure CN120599644A_ABST
Patent Text Reader

Abstract

According to the training method of the table structure recognition model, in the training process of the model, the spatial features and the visual features of the bounding box are fused, a traditional text region prediction and matching problem can be converted into a direct text region pointing problem, and therefore the training efficiency of the table structure recognition model is improved in the actual table structure recognition task. Text region dislocation can be avoided, and post-processing requirements are eliminated. Besides, for a complex table structure, in the training process, span perception comparison loss is added into target loss, and the recognition capability of the model on the table structure containing row span or column span can be enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of table structure recognition, and in particular to a training method, electronic equipment and storage medium for a table structure recognition model. Background Art

[0002] Tables are ubiquitous in various documents, but direct machine parsing is challenging. Table structure recognition aims to digitize table images into a machine-readable format to support downstream applications such as information retrieval and table question answering. Current table structure recognition methods face the following challenges in processing the logical and physical structures of tables: 1. Table Applications and Challenges: Tables are ubiquitous in various documents, including business documents and academic papers, and are highly favored for their compact and efficient information presentation. However, this compactness poses significant challenges for direct machine parsing. To enable machines to understand and process tabular information, the table structure recognition task has emerged. Its goal is to convert table images into machine-readable formats such as HTML, thereby supporting various downstream applications such as information retrieval and table question answering. 2. Task Components of Table Structure Recognition: Table structure recognition typically requires predicting both logical and physical structures before constructing the complete table structure. The logical structure reflects the semantic organization and relationships between table cells, commonly represented in HTML or LaTeX. The physical structure represents the layout information of table cells, such as their bounding boxes.

[0003] Currently, image-to-text methods are often used for table structure recognition. This method first predicts the logical structure and then predicts the physical structure based on this. However, when mapping the predicted cell bounding box to the text area, these frameworks often have alignment errors, resulting in incorrect text matching. This is because the alignment between the table text area and the predicted cell bounding box is not perfect, so this type of method requires fine calibration post-processing during the text area matching process to obtain satisfactory results. In addition, for complex tables containing row spans or column spans, the recognition ability of existing methods is insufficient, and it is difficult to accurately parse their structure. In addition, in industrial document table recognition scenarios, such as when processing watermarked documents or tables in non-English fields, existing methods have poor versatility and cannot effectively deal with these complex situations. Summary of the Invention

[0004] In view of the above technical problems, the technical solution adopted by the present invention is: According to a first aspect of the present invention, a method for training a table structure recognition model is provided. The table structure recognition model includes an image encoder, a layout encoder, a feature weighting module, and a logical structure decoder. The method includes the following steps: S100 , obtaining a set of sample table images with annotation information; the annotation information includes the annotation position and annotation category of each cell in the sample table image.

[0005] S200 , using the sample table image set, training an initial table structure recognition model to obtain a target table structure recognition model.

[0006] Among them, the image encoder is used to obtain the visual features of the sample table image; the layout encoder is used to obtain the spatial features of the sample table image and fuse the visual features and spatial features to obtain the layout embedding features of the sample table image; the feature weighting module is used to interactively process the visual features and layout embedding feature sequences of the sample table image based on the interactive attention mechanism to obtain the interactive features of the sample table image; the logical structure decoder is used to decode the interactive features to obtain the prediction results of the sample table image; the prediction results include the predicted position and predicted label category of each cell in the sample table image.

[0007] Among them, during the training process of the initial table structure recognition model, the target loss is used to update the parameters of the initial table structure recognition model. The target loss includes label classification loss, layout pointer loss and span-aware contrast loss. Among them, the label classification loss is obtained based on the prediction results and annotation information, the layout pointer loss is obtained based on the output features of the logical structure decoder, and the span-aware contrast loss is obtained based on the annotation position of each cell.

[0008] According to a second aspect of the present invention, an electronic device is provided, comprising a processor and a memory; the processor is configured to execute the steps of the method according to the first aspect of the present invention by calling a program or instruction stored in the memory.

[0009] According to a third aspect of the present invention, there is provided a computer-readable storage medium storing a program or instructions, wherein the program or instructions enable a computer to execute the steps of the method according to the first aspect of the present invention.

[0010] The present invention has at least the following beneficial effects: The table structure recognition model training method provided by the embodiments of the present invention integrates the spatial and visual features of bounding boxes during model training, transforming the traditional text region prediction and matching problem into a direct text region pointing problem. This, in turn, avoids text region misalignment and eliminates the need for post-processing in actual table structure recognition tasks. Furthermore, for complex table structures, a span-aware contrast loss is added to the target loss during training, enhancing the model's ability to recognize table structures containing row or column spans.

[0011] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0013] Figure 1 A flowchart of a method for training a table structure recognition model provided by an embodiment of the present invention; Figure 2 Schematic diagram for calculating span coefficient. DETAILED DESCRIPTION

[0014] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present invention.

[0015] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this invention pertains. The terms used herein in the specification of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0016] It should be noted that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the steps as sequential processes, many of the steps can be performed in parallel, concurrently, or simultaneously. In addition, the order of the steps can be rearranged. A process can be terminated when its operation is completed, but can also have additional steps not included in the accompanying drawings. A process can correspond to a method, function, procedure, subroutine, subprogram, etc.

[0017] This embodiment of the present invention provides a training method for a table structure recognition model, aiming to address existing issues such as text region misalignment, difficulty recognizing complex tables, and poor versatility. Specific goals include: eliminating misalignment in text region matching through a layout pointer mechanism without the need for post-processing; enhancing the recognition of complex table structures through span-aware contrast supervision; achieving state-of-the-art performance across multiple benchmarks, and demonstrating strong versatility and competitiveness in industrial document table recognition scenarios. This approach will improve the accuracy and efficiency of table structure recognition and promote development in related fields.

[0018] In an embodiment of the present invention, the table structure recognition model may include an image encoder, a layout encoder, a feature weighting module, and a logical structure decoder.

[0019] Furthermore, a training method for a table structure recognition model provided by an embodiment of the present invention may include: Figure 1 The following steps are shown: S100 , obtaining a set of sample table images with annotation information; the annotation information includes the annotation position and annotation category of each cell in the sample table image.

[0020] In an embodiment of the present invention, the sample table image can be a sample table image obtained by directly converting the format of a table (such as an Excel table document), or a sample table image obtained by photographing the table through a camera. In an optional embodiment, the sample table image can also be a sample table image obtained by identifying and cutting out a portion of the table area from a candidate image. The candidate image can be an image document containing a table, such as a PDF document; the candidate image can also be a candidate image obtained by converting the format of a document containing a table, or a candidate image obtained by photographing a document containing a table. The document containing a table can be a word document, a bill document, a web page document, etc., but is not limited thereto. In an optional embodiment, a pre-trained table extraction model can also be used to identify the candidate image, identify the target position of the table in the candidate image, and then extract the image area at the target position of the table from the candidate image to obtain a sample table image. When training the table extraction model, an image with a table is used as a sample image, and the position of the table in the sample image is marked, and supervised training is performed until convergence to obtain a trained table extraction model.

[0021] In the embodiment of the present invention, each sample table image is marked with corresponding annotation information, and the annotation information includes the annotation position of each cell, the annotation category of each cell, and the annotation connection relationship between each cell.

[0022] In an embodiment of the present invention, the position of a cell may be represented by (x, y, w, h), where x and y may represent the coordinate position of the center point, the coordinate position of the upper left corner, or the coordinate position of the lower left corner of the cell. w represents the width of the table cell, and h represents the height of the table cell.

[0023] In an embodiment of the present invention, the categories of cells can be divided as needed. For example, the categories of cells may include a first category, a second category, a third category, a fourth category, and a fifth category, wherein the first category characterizes that no processing is required, the second category characterizes merging to the left, the third category characterizes merging upward, the fourth category characterizes merging from the left to the top, and the fifth type characterizes the background. The categories of cells can be represented by different characters, etc. In an exemplary embodiment, the categories of cells can be represented by different numeric characters, such as 1 for the first category, 2 for the second category, 3 for the third category, 4 for the fourth category, and 0 for the fifth category, without limitation thereto. In another exemplary embodiment, the categories of cells can be represented by different alphabetical characters, such as a for the first category, b for the second category, c for the third category, d for the fourth category, and t for the fifth category, without limitation thereto.

[0024] It is understandable that the label position and label category of the cell can be manually labeled and used as the actual value for reference comparison.

[0025] Optionally, the computer device may obtain multiple sample table images from a local storage space, or obtain multiple sample table images from a network or cloud or other computer device or external storage device, and obtain the annotation information corresponding to each sample table image.

[0026] S200 , using the sample table image set, training an initial table structure recognition model to obtain a target table structure recognition model.

[0027] Furthermore, in an embodiment of the present invention, an image encoder is used to obtain visual features of a sample table image. A layout encoder is used to obtain spatial features of a sample table image and fuse the visual features and spatial features to obtain layout embedding features of the sample table image; a feature weighting module is used to interactively process the visual features and layout embedding feature sequences of the sample table image based on an interactive attention mechanism to obtain interactive features of the sample table image; a logical structure decoder is used to decode the interactive features to obtain a prediction result of the sample table image; the prediction result includes the predicted position and predicted label category of each cell in the sample table image. In the process of training the initial table structure recognition model, the target loss is used to update the parameters of the initial table structure recognition model. The target loss includes label classification loss, layout pointer loss, and span-perceived contrast loss. The label classification loss is obtained based on the prediction result and annotation information, the layout pointer loss is obtained based on the output features of the logical structure decoder, and the span-perceived contrast loss is obtained based on the labeled position of each cell.

[0028] Furthermore, S200 may specifically include: S210, using the image encoder of the current table structure recognition model to extract visual features corresponding to the sample table images of the current batch, and sending the obtained visual features to the layout encoder; the initial value of the current table structure recognition model is the original table structure recognition model.

[0029] In an embodiment of the present invention, the image encoder may adopt an existing image encoder structure. In a non-limiting exemplary embodiment, it may be a SwinTransformer model having a hierarchical feature extraction capability and suitable for processing multi-scale layout information of tabular images.

[0030] The image encoder can obtain the visual features of each sample table image through the following steps: S2101 , scaling the received sample table image to an image with a fixed resolution as an intermediate image. In an illustrative embodiment, the fixed resolution may be, for example, 768×768.

[0031] S1202: Split the intermediate image into multiple sub-images. The number of sub-images can be set according to actual needs.

[0032] S1203, performing feature extraction on each sub-image to capture the spatial semantic information of the image and obtain corresponding visual features; and then obtaining the visual features of the sample table image, that is, the features obtained by concatenating the visual features of all sub-images.

[0033] In an embodiment of the present invention, the feature dimension of the visual features of each sub-image can be set based on actual needs. In one exemplary embodiment, the feature dimension can be 1024. The visual features Z of each sample table image can be a p×d matrix, where p is the number of sub-images and d is the feature dimension.

[0034] S220 , using a layout encoder to obtain spatial features of the sample table images of the current batch, and fusing the obtained spatial features with visual features to obtain layout embedding features of the sample table images.

[0035] In an embodiment of the present invention, the layout encoder may include a spatial feature acquisition module, a feature alignment module, and a feature fusion module.

[0036] Furthermore, S220 may specifically include: The spatial feature acquisition module is used to acquire the bounding box coordinates of each cell in the sample table image as the spatial feature of the cell.

[0037] In an embodiment of the present invention, the spatial feature acquisition module may be an OCR engine or a PDF parsing tool. The bounding box coordinates may be the upper left corner coordinates or the lower right corner coordinates of the bounding box.

[0038] Using the feature alignment module, based on the visual features of the sample table image, a preset feature alignment method is used to obtain the visual features corresponding to each bounding box area, that is, to find the corresponding visual features for each bounding box area.

[0039] In an embodiment of the present invention, the preset feature alignment method may be a region of interest alignment method, and specifically a 2×2 region of interest alignment method may be used, that is, the original region of interest is divided into a uniform grid of 2 rows×2 columns.

[0040] The feature fusion module is used to fuse the spatial features and visual features corresponding to each bounding box area of ​​the sample table image to obtain the layout embedding corresponding to the bounding box area, and then obtain the layout embedding features of the sample table image.

[0041] In an embodiment of the present invention, the feature fusion module may be a multilayer perceptron. The multilayer perceptron may have an existing structure. For example, the multilayer perceptron may be composed of two layers of neurons, namely an input layer and an output layer. A two-layer perceptron module may be obtained by adding two fully connected hidden layers between the input layer and the output layer, and transforming the output of the hidden layers using an activation function.

[0042] Those skilled in the art know that any method of using a multi-layer perceptron to fuse the spatial features and visual features corresponding to each bounding box area of ​​a sample table image to obtain a layout embedding corresponding to the bounding box area falls within the scope of protection of the present invention.

[0043] In an embodiment of the present invention, the layout embedding feature A of each sample table image may be a matrix of d×Q1, where Q1 is the number of bounding boxes in the sample table image.

[0044] S230: Using a feature weighting module, interactively process the visual features and the layout embedding features based on a cross-attention mechanism to obtain interactive features.

[0045] In the embodiment of the present invention, the interaction feature FI of each sample table image may be equal to the product of the transposed matrix of the corresponding visual feature and the layout embedding feature, that is, FI=ZA T .

[0046] S240, using a logical structure decoder to decode the interactive features to obtain corresponding prediction results; the prediction results include the predicted position and predicted label category of each cell in the sample table image.

[0047] In an embodiment of the present invention, the logical structure decoder may be a model that implements Donut support autoregressive generation of table structure labels based on the BART model, for example, the Donut model. It is known to those skilled in the art that the Donut model structure may be prior art.

[0048] In the embodiment of the present invention, the output of the logical structure decoder is an OTSL-tag sequence (mapped 1-1 with HTML tags to reduce the sequence length), such as table header, row, column, cell, etc., corresponding to the logical structure (such as HTML 、 、 ).

[0049] S250, obtain the target loss corresponding to the current table structure recognition model. If the target loss converges, use the current table structure recognition model as the target table structure recognition model. Otherwise, adjust the parameters of the current table structure recognition model based on the target loss, and use the sample table images of the next batch as the sample table images of the current batch, and execute S210.

[0050] Furthermore, the label classification loss of each sample table image can satisfy the following conditions: .

[0051] Among them, L els The label classification loss for each sample table image, N is the number of annotation positions, y ik is the one-hot representation of the true label identification value of the i-th annotation position, that is, the true label. If the true label of the i-th annotation position is the label category k, then y ik =1, otherwise, y ik =0, p ik To predict the probability that the i-th annotation position is of label category k, i ranges from 1 to N, k ranges from 1 to M, and M is the total number of label categories. The probability of a label category can be calculated using the softmax function.

[0052] Furthermore, the layout pointer loss of each sample table image is obtained by the following steps: The output features of the logical structure decoder are split into bounding box features and table label features, that is, the last hidden state features of the logical structure decoder are split into hidden state features corresponding to the bounding box and hidden state features corresponding to the table label.

[0053] The bounding box features and the table label features are linearly projected respectively to obtain the projection features of the bounding box and the projection features of the table label.

[0054] In an embodiment of the present invention, linear projection can be performed using an existing linear projection method. For example, linear projection can be performed using the method x=σ(w•x0+b), where x0 is the feature before projection, w is the linear projection weight matrix, b is the bias term, x is the feature after projection, σ( ) is a nonlinear activation function, such as a tanh activation function, and • represents dot product.

[0055] Based on the projection features of the bounding box and the projection features of the table label, the association relationship between the bounding box and the table label is constructed to obtain the layout pointer loss.

[0056] Furthermore, the layout pointer loss satisfies the following conditions: ; Among them, L ptr is the layout pointer loss of each sample table image, exp() is the exponential function, is the projection feature of the jth bounding box, where j ranges from 1 to Q1, and Q1 is the number of bounding boxes in each sample table image. is the projection feature of the true label corresponding to the j-th bounding box, τ is the temperature hyperparameter, is the projection feature of the rth table label, r ranges from 1 to Q2, and r≠j, Q2 is the number of table labels, and • represents the dot product.

[0057] In this embodiment of the present invention, a layout pointer loss is added to the target loss, transforming the traditional bounding box prediction into a "pointer pointing" problem. This directly associates the logical label with the physical bounding box, thus avoiding the misalignment problem of heuristic matching.

[0058] Furthermore, the span-aware loss of each sample table image includes row span-aware loss and column span-aware loss, where the row span-aware loss satisfies the following conditions: ; in, is the row span perception loss, c u1v1 is the span coefficient between the u1th row-span bounding box in the sample table image and the v1th bounding box in the same row as the u1th row-span bounding box, u1 ranges from 1 to Q3, Q3 is the number of row-span bounding boxes, v1 ranges from 1 to f(u1), f(u1) is the number of bounding boxes in the same row as the u1th row-span bounding box, is the projection feature of the u1th bounding box, is the projection feature of the v1th bounding box, It is the projection feature of the a1th bounding box in the Q1 bounding boxes, the value of a1 ranges from 1 to Q1, and a1≠u1.

[0059] Among them, the column span perception loss satisfies the following conditions: .

[0060] in, is the row span perception loss, c u2v2 is the span coefficient between the u2th column-spanned bounding box in the sample table image and the v2th bounding box in the same column as the u2th column-spanned bounding box, u2 ranges from 1 to Q4, Q4 is the number of column-spanned bounding boxes, and v2 ranges from 1 to f(u2), f(u2) is the number of bounding boxes in the same column as the u2th column-spanned bounding box. is the projection feature of the u2th bounding box, is the projection feature of the v2th bounding box, It is the projection feature of the a2th bounding box in the Q1 bounding boxes, the value of a2 ranges from 1 to Q1, and a2≠u2.

[0061] Furthermore, c u1v1 The following conditions are met: c u1v1 =(overlap(u1,v1)) 2 / span(u1)×span(v1).

[0062] Where overlap(u1, v1) is the number of overlapping cells between the u1-th bounding box and the v1-th bounding box, span(u1) is the number of row cells occupied by the u1-th bounding box, and span(v1) is the number of row cells occupied by the v1-th bounding box.

[0063] Furthermore, c u2v2 The following conditions are met: c u2v2 =(overlap(u2,v2)) 2 / span(u2)×span(v2).

[0064] Where overlap(u2, v2) is the number of overlapping cells between the u2th bounding box and the v2th bounding box, span(u2) is the number of column cells occupied by the u2th bounding box, and span(v2) is the number of column cells occupied by the v2th bounding box.

[0065] like Figure 2 As shown, Figure 2 Take the text box T1 in the example, the text box T1 is a column span text box, and the number of overlapping cells between it and the text box T2 belonging to the same column is 0. The number of column cells occupied by text box T1 is 1, and the number of column cells occupied by text box T2 is 6. In this way, the overlapping coefficient between text box T1 and text box T2 is (0) 2 / 1×6=0.

[0066] In the embodiment of the present invention, by introducing span information, the model's understanding of complex layouts can be enhanced, aiming to cluster the embeddings of span cells into the same category.

[0067] Those skilled in the art know that, by knowing the marked position of each cell, it is possible to know which cells are row-spanning cells, which cells are column-spanning cells, and the number of spans of the spanning cells.

[0068] In an embodiment of the present invention, the target loss may be a weighted sum of the label classification loss, the layout pointer loss, and the span-aware contrast loss, i.e., the target loss, where λ1 to λ4 are learnable weights. In one exemplary embodiment, λ1 = λ2 = 1, and λ3 = λ4 = 0.5.

[0069] An embodiment of the present invention further provides a table structure recognition method, comprising the following steps: S10, obtaining a table image to be recognized as an image to be processed; S11, input the image to be processed into the target table structure recognition model to obtain the corresponding prediction result.

[0070] An embodiment of the present invention also provides an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are configured to execute the method described in the embodiment of the present invention.

[0071] An embodiment of the present invention further provides a computer-readable storage medium storing computer-executable instructions, wherein the computer instructions are used to execute the method described in the embodiment of the present invention.

[0072] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in the present invention can be achieved. This is not limited herein.

[0073] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.

Claims

1. A training method for a table structure recognition model, characterized in that: The table structure recognition model includes an image encoder, a layout encoder, a feature weighting module and a logical structure decoder; the method includes the following steps: S100, obtaining a sample table image set with annotation information; the annotation information includes the annotation position and annotation category of each cell in the sample table image; S200, using the sample table image set, training an initial table structure recognition model to obtain a target table structure recognition model; Among them, the image encoder is used to obtain the visual features of the sample table image; the layout encoder is used to obtain the spatial features of the sample table image and fuse the visual features and spatial features to obtain the layout embedding features of the sample table image; the feature weighting module is used to interactively process the visual features and layout embedding feature sequences of the sample table image based on the interactive attention mechanism to obtain the interactive features of the sample table image; the logical structure decoder is used to decode the interactive features to obtain the prediction results of the sample table image; the prediction results include the predicted position and predicted label category of each cell in the sample table image; Among them, during the training process of the initial table structure recognition model, the target loss is used to update the parameters of the initial table structure recognition model. The target loss includes label classification loss, layout pointer loss and span-aware contrast loss. Among them, the label classification loss is obtained based on the prediction results and annotation information, the layout pointer loss is obtained based on the output features of the logical structure decoder, and the span-aware contrast loss is obtained based on the annotation position of each cell.

2. The method according to claim 1, characterized in that The layout encoder includes a spatial feature acquisition module, a feature alignment module and a feature fusion module; The layout encoder is specifically used for: Using the spatial feature acquisition module to acquire the bounding box coordinates of each cell in the sample table image as the spatial feature of the cell; Using the feature alignment module, based on the visual features of the sample table image, a preset feature alignment method is used to obtain the visual features corresponding to each bounding box area; The feature fusion module is used to fuse the spatial features and visual features corresponding to each bounding box area of ​​the sample table image to obtain the layout embedding corresponding to the bounding box area, and then obtain the layout embedding features of the sample table image.

3. The method according to claim 2, characterized in that The preset feature alignment method is a region of interest alignment method; the feature fusion module is a multi-layer perceptron.

4. The method according to claim 1, wherein The label classification loss for each sample table image satisfies the following conditions: ; Among them, L els The label classification loss for each sample table image, N is the number of annotation positions, y ik is the true label identification value of the i-th annotation position. If the true label of the i-th annotation position is label category k, then y ik =1, otherwise, y ik =0, p ik To predict the probability that the i-th annotation position is the label category k, the value of i ranges from 1 to N, the value of k ranges from 1 to M, and M is the total number of label categories.

5. The method according to claim 2, characterized in that The layout pointer loss of each sample table image is obtained by the following steps: Split the output features of the logical structure decoder into bounding box features and table label features; Perform linear projection on the bounding box features and the table label features respectively to obtain the projection features of the bounding box and the projection features of the table label; Based on the projection features of the bounding box and the projection features of the table label, the association relationship between the bounding box and the table label is constructed to obtain the layout pointer loss.

6. The method according to claim 5, characterized in that in, The layout pointer loss meets the following conditions: ; Among them, L ptr is the layout pointer loss of each sample table image, exp() is the exponential function, is the projection feature of the jth bounding box, where j ranges from 1 to Q1, where Q1 is the number of bounding boxes in each sample table image. is the projection feature of the true label corresponding to the j-th bounding box, τ is the temperature hyperparameter, is the projection feature of the rth table label, r ranges from 1 to Q2, and r≠j, Q2 is the number of table labels, and • represents the dot product.

7. The method according to claim 3, characterized in that The span-aware loss of each sample table image includes row span-aware loss and column span-aware loss, where the row span-aware loss satisfies the following conditions: ; in, is the row span perception loss, c u1v1 is the span coefficient between the u1th row-span bounding box in the sample table image and the v1th bounding box in the same row as the u1th row-span bounding box, u1 ranges from 1 to Q3, Q3 is the number of row-span bounding boxes, v1 ranges from 1 to f(u1), f(u1) is the number of bounding boxes in the same row as the u1th row-span bounding box, is the projection feature of the u1th bounding box, is the projection feature of the v1th bounding box, is the projection feature of the a1th bounding box in Q1 bounding boxes, the value of a1 ranges from 1 to Q1, and a1≠u1; Among them, the column span perception loss satisfies the following conditions: ; in, is the row span perception loss, c u2v2 is the span coefficient between the u2th column-spanned bounding box in the sample table image and the v2th bounding box in the same column as the u2th column-spanned bounding box, u2 ranges from 1 to Q4, Q4 is the number of column-spanned bounding boxes, and v2 ranges from 1 to f(u2), f(u2) is the number of bounding boxes in the same column as the u2th column-spanned bounding box. is the projection feature of the u2th bounding box, is the projection feature of the v2th bounding box, It is the projection feature of the a2th bounding box in the Q1 bounding boxes, the value of a2 ranges from 1 to Q1, and a2≠u2.

8. The method according to claim 7, characterized in that c u1v1 The following conditions are met: c u1v1 =(overlap(u1,v1)) 2 / span(u1)×span(v1); Where overlap(u1, v1) is the number of overlapping cells between the u1th bounding box and the v1th bounding box, span(u1) is the number of row cells occupied by the u1th bounding box, and span(v1) is the number of row cells occupied by the v1th bounding box; c u2v2 The following conditions are met: c u2v2 =(overlap(u2,v2)) 2 / span(u2)×span(v2); Where overlap(u2, v2) is the number of overlapping cells between the u2th bounding box and the v2th bounding box, span(u2) is the number of column cells occupied by the u2th bounding box, and span(v2) is the number of column cells occupied by the v2th bounding box.

9. An electronic device, characterized in that: including processor and memory; The processor is configured to execute the steps of the method according to any one of claims 1 to 8 by calling the program or instructions stored in the memory.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium is used to store a program or instruction, and the program or instruction enables a computer to execute the steps of the method according to any one of claims 1 to 8.