Information extraction method for picture type table

Through technical means such as deep nested U-shaped structure networks and fine-grained jump connections, the problem of poor information extraction effect in the existing technology when processing complex picture-type tables is solved, and a more efficient and accurate information extraction effect is achieved.

CN120107987APending Publication Date: 2025-06-06NORTHEASTERN UNIV CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510175797.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

When the prior art processes complex and variable picture tables, the information extraction effect is poor, and the model training is difficult, time-consuming and requires more training samples.

Method used

Information extraction is performed using a network with deep nested U-shaped structure. By introducing fine-grained jump connections and recursive structures, the integration of different levels and detailed features is improved, and auxiliary predictors are added to improve model performance.

Benefits of technology

It realizes accurate extraction of information at different scales in the picture, improves the recognition ability of complex tables and the accuracy of information extraction, and reduces the difficulty and time-consuming of model training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107987A_ABST
    Figure CN120107987A_ABST
Patent Text Reader

Abstract

The invention provides an information extraction method for a picture type table, and relates to the technical field of table structure recognition and information extraction. Aiming at the defects existing when the table information in the picture is processed in the prior art, the table information extraction method which is more accurate and better in robustness is developed by deeply researching the characteristics and rules of the table, and the accuracy of information extraction is improved. According to the method, information of different scales in a picture is extracted through a network with a deep nested U-shaped structure; fusion of different levels and detail features is improved by introducing fine-grained jump connection and a recursive structure, so that the model recognizes cells of different scales, and recognition can be more effectively carried out in the face of complex and changeable tables; and the model performance is further improved by adding an auxiliary predictor, so that the accuracy of information extraction is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of table structure recognition and information extraction, and in particular to an information extraction method for a picture-type table. Background Art

[0002] Data in tabular format is commonly found in various documents and is used to summarize and present information. Tabular information storage is centralized and easy to use. It accounts for a large proportion in the transmission of business information and has become a major form of data storage. In addition, there is a certain degree of mutual connection between tabular data. A large amount of tabular data contains rich information and is the focus of research on data collection and information induction. Some tables exist in the form of pictures or PDF files and do not have structured information, which makes it very difficult to extract information from tables.

[0003] In the prior art, there is an end-to-end approach to table structure recognition, which uses a sequence decoder to generate an HTML sequence representing the table structure. However, this requires the model to learn the syntax of HTML, making these methods difficult to train, and the model requires more training samples. In addition, these models are very time-consuming.

[0004] Others have proposed a table structure recognition framework - Logical Location Regression Network (LORE), which first locates the table cells on the input image and then predicts the logical position along the spatial position of the cells. In order to better model the dependencies and constraints between logical positions, a cascade regression framework is adopted, combining inter-unit and intra-unit supervision. LORE's reasoning is a parallel network forward pass, which does not require complex decoding strategies, but there are still some challenges when performing table structure recognition. For example, the layout and structure of the table are often complex and changeable, and the differences between different tables are large, which makes it ineffective when dealing with complex tables. Summary of the invention

[0005] The technical problem to be solved by the present invention is to provide an information extraction method for picture-type tables in view of the deficiencies of the above-mentioned prior art, to accurately extract information from the tables in the pictures, and to develop a more accurate and robust table information extraction method through in-depth research on the characteristics and rules of the tables, thereby improving the accuracy of information extraction.

[0006] In order to solve the above technical problems, the technical solution adopted by the present invention is:

[0007] A method for extracting information from image tables, which extracts information of different scales in an image through a deeply nested U-shaped network; introduces fine-grained skip connections and recursive structures to improve the fusion of different levels and detail features, so that the model can recognize cells of different scales; and further improves model performance by adding auxiliary predictors. Specifically, the method includes the following steps:

[0008] Step 1: Obtain the table structure recognition dataset, preprocess it, and divide the table structure recognition dataset into a training set, a validation set, and a test set;

[0009] Step 2: Build a table structure recognition model and a loss function;

[0010] The table structure recognition model includes a multi-scale feature extraction module, a cell physical position prediction module and a cell logical position prediction module;

[0011] Step 3: Use batch stochastic gradient descent and Adam optimization algorithm to train the table structure recognition model;

[0012] Step 4: Input the test set into the trained table structure recognition model to obtain the prediction results, including the physical location information and logical location information of each cell;

[0013] Step 5: Segment the image based on the predicted physical location information of the cells and input it into the PaddleOCR model, which is an open source OCR model based on PaddlePaddle, to output the cell content; generate the table structure information based on the predicted logical location of the cells; then combine the cell content and structure information to generate a table, which is represented by an HTML sequence;

[0014] Step 6: Design prompt words and use the big model to extract information from the HTML sequence. The keywords are: Assuming you are an entity-relationship triple extraction model, I will give you the head entity type list subject_types, the tail entity type list object_types, and the relationship list relations. I will give you an HTML sequence. Please extract the triples in the sequence based on these three lists; enter subject_types, object_types, and relations to get the triples.

[0015] Furthermore, the step 1 specifically includes:

[0016] Step 1.1: Obtain a table structure recognition dataset; the table structure recognition dataset is a table detection and structure recognition dataset for natural scenes, which includes not only table images extracted from photos, scanned documents and web pages, but also tilted tables, curved tables, occluded or blurred tables, tables with extreme aspect ratios, superimposed tables, multi-color tables and irregular tables;

[0017] Step 1.2: Preprocess the table structure data set;

[0018] The preprocessing includes: scaling the image according to the set input size, and performing data enhancement by random cropping, translation and rotation;

[0019] Step 1.3: Divide the dataset into training set, test set and validation set;

[0020] The dataset is divided into a training set and a test set; the test set in the dataset is further divided into a test set and a validation set in a 1:1 ratio.

[0021] Furthermore, in step 1.1, the open source dataset WTW-Dataset is downloaded from the Internet; both the training set and the test set include training images and labels, the labels are in XML file format, each image corresponds to an XML, and the labels include the bbox coordinate information of the cell and the cell structure information; the structure information includes the starting column, the ending column, the starting row, the ending row and the table number.

[0022] Furthermore, the multi-scale feature extraction module in step 2 is modified on the basis of U2-Net as a feature extractor; first, a 7×7 convolution layer with a step size of 2 and a padding of 3 is used to extract shallow features, then a BN layer and a ReLU activation function are used, and then a maximum pooling layer with a step size of 2 is used, and then a deep nested U-shaped structure network is used to extract features of four different scales; the deep nested U-shaped structure network includes 7 modules, namely 4 Encoder layers and 3 Decoder layers, namely Encoder1, Encoder2, Encoder3, Decoder1, Decoder2, Decoder3, which are the same type of modules with different depths; in the Encoder4 module, all sampling layers are replaced with dilated convolutions, and no upsampling or downsampling is performed; an auxiliary predictor is connected to the Encoder4, Decoder1, Decoder2 and Decoder3 modules, and the auxiliary predictor is an auxiliary component used in a deep learning model to enhance the learning process of the main task by providing additional supervision signals; finally, the features of the four different scales are feature fused.

[0023] Furthermore, the cell physical position prediction module in step 2 has features extracted by the multi-scale feature extraction module as input, and the extracted features are input into multiple prediction heads to predict different information, and the first prediction head generates a heatmap to represent classification information; each category has a heatmap, and a peak point is used on the output heatmap to represent the center position of the target; the heatmap is a two-dimensional probability distribution, and the peak position represents the most likely position of the target center; the heatmap has two categories, namely the center point of the cell and the vertex of the cell; a prediction head is used to predict a center offset to fine-tune the center point position to make the prediction result more accurate; next, based on the distance from the cell center point to the four vertices, the cell bbox is decoded, and based on the cell vertex, the distance to the center point of the cell sharing the vertex is regressed to decode the gbox; based on the gbox, the discrete bboxes are spliced ​​together to obtain the physical position of the cell.

[0024] Furthermore, the cell logical position prediction module in step 2 inputs the features extracted by the multi-scale feature extraction module into a prediction head, and extracts the features of each cell in combination with the heatmap in the physical position prediction module; then the physical position of the cell is embedded and fused with the features of the cell, the features of the cell are interacted through the attention mechanism, and finally the logical position is regressed through the linear layer; each cell output is four fields, namely the start column, the end column, the start row and the end row.

[0025] Furthermore, the loss function of the cell physical location prediction module is:

[0026] L det =L k +L off +L cv +L vc ;

[0027] Among them, L k Represents the heatmap loss, and the heatmap output by the network is It is a three-dimensional matrix, where R represents the stride of the output relative to the original image, C represents the number of categories; W represents the width of the original image, and H represents the height of the original image;

[0028] The heatmap loss expression is as follows:

[0029]

[0030] Among them, α and β are hyperparameters, and N represents the number of key points in the image; represents a point on the heatmap, xyc is the coordinate information, and the corresponding point is found according to the coordinate information; xyc is a point on the heatmap generated based on the actual cell position for training; the weight parameter (1-Y xyc ) β The closer to the center, the smaller the value. This weight parameter is used to reduce the penalty.

[0031] L off Represents the bias loss, which is used to bias the prediction of each key point To supervise, the expression is as follows:

[0032]

[0033] Where p represents the center point of the target box, R represents the downsampling multiple, R = 4; represents the deviation value; The output of the center offset prediction head of the cell physical location prediction module. is a point on the output feature map;

[0034] L cv represents the loss function based on the regression of the cell center point to the vertex; L vc Represents the loss function based on cell vertices and regressing to the cell center.

[0035] Furthermore, the loss function of the cell logical position prediction module is:

[0036]

[0037] Among them, N is the number of key points in the image, which is also the number of cells; Indicates the use of L1 loss function, l k Represents the four real logical coordinates of each cell, namely the starting column, ending column, starting row, and ending row. are the four logical coordinates for prediction;

[0038] Introducing the cell span-based loss function L span and the cell conflict-based loss function l conf , which is used to make the cell logic position prediction model better understand the dependencies and constraints between the cell logic positions, L span and L conf The loss function is as follows:

[0039]

[0040] Among them, A is the set of cells combined in pairs, and M is the set of cells; is the logical position of the prediction; i and j represent different cells, r e Represents the end of the line, r s Represents the starting line, c e Represents the end column, c s Represents the starting column.

[0041] The beneficial effect of adopting the above technical solution is that the information extraction method for picture-type tables provided by the present invention can extract information of different scales in pictures through a deeply nested U-shaped network. By introducing fine-grained skip connections and recursive structures, the fusion of different levels and detail features is effectively improved, so that the model can recognize cells of different scales and can more effectively recognize complex and changeable tables. Then, the model performance is further improved by adding auxiliary predictors, thereby improving the accuracy of information extraction. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 A flowchart of extracting information from a table in a picture provided by an embodiment of the present invention;

[0043] Figure 2 A schematic diagram of the structure of a multi-scale feature extraction module provided in an embodiment of the present invention;

[0044] Figure 3 A schematic diagram of the network structure of the Encoder1 module provided in an embodiment of the present invention;

[0045] Figure 4 A schematic diagram of the network structure of the Encoder4 module provided in an embodiment of the present invention; DETAILED DESCRIPTION

[0046] The specific implementation of the present invention is further described in detail below in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present invention, but are not intended to limit the scope of the present invention.

[0047] like Figure 1 As shown, the method of this embodiment is described as follows.

[0048] Step 1: Obtain the table structure recognition dataset and preprocess it to divide the table structure recognition dataset into a training set, a validation set, and a test set.

[0049] Step 1.1: Obtain a table structure recognition dataset; the table structure recognition dataset is a table detection and structure recognition dataset for natural scenes. This dataset not only contains table images extracted from photos, scanned documents and web pages, but also covers 7 extremely challenging table types, including tilted tables, curved tables, occluded or blurred tables, extreme aspect ratio tables, superimposed tables, multi-color tables and irregular tables. In this embodiment, the open source dataset WTW-Dataset (Wild Table Dataset) is downloaded from the Internet. The training set of WTW contains 10,970 pictures, and the test set contains 3,611 pictures. At the same time, the test set is also divided into 7 subcategories in detail. Both the training set and the test set include training pictures and labels. The label is in XML file format, and each picture corresponds to an XML. The label includes the bbox coordinate information of the cell and the cell structure information. The structural information includes the starting column, the ending column, the starting row, the ending row and the table number. In addition, the test set also has a separate file to explain the subcategory category to which each picture belongs. These diverse table types make WTW-Dataset an ideal choice for evaluating and improving the performance of table structure recognition algorithms.

[0050] Step 1.2: Preprocess the table structure data set;

[0051] The preprocessing includes: scaling the image according to a set input size, and performing data enhancement by random cropping, translation, and rotation.

[0052] Step 1.3: Divide the dataset into training set, test set and validation set;

[0053] The data set is divided into a training set and a test set. In this embodiment, the test set in the data set is divided into a test set and a validation set at a ratio of 1:1. The final data set includes 10,970 training images, 1,805 test images, and 1,806 validation images.

[0054] Step 2: Build a table structure recognition model and construct a loss function.

[0055] The table structure recognition model includes a multi-scale feature extraction module, a cell physical position prediction module and a cell logical position prediction module.

[0056] The multi-scale feature extraction module is modified based on U2-Net as a feature extractor, such as Figure 2As shown in the figure, a 7×7 convolutional layer with a stride of 2 and a padding of 3 is first used to extract shallow features, then a BN layer and a ReLU activation function are used, and then a maximum pooling layer with a stride of 2 is used. Then, a deeply nested U-shaped network is used to extract features of four different scales. The deeply nested U-shaped network consists of 7 modules, which include 4 Encoder layers and 3 Decoder layers. The structures of Encoder1, Encoder2, Encoder3, Decoder1, Decoder2, and Decoder3 are similar to those of UNet. These six modules are of the same type, but with different depths. Figure 3 The structure diagram of Encoder1 is shown in Figure 1. Encoder2 is obtained by cutting the network structure based on Encoder1. Encoder3 is obtained by cutting the network structure based on Encoder2. The structure of Decoder1 is the same as that of Encoder1. Decoder2 is obtained based on Decoder1. Decoder3 is obtained based on Decoder2. There is no upsampling or downsampling in the Encoder4 module. Instead, all the sampling layers are replaced by dilated convolutions. Figure 4 As shown. An auxiliary predictor is connected to the Encoder4, Decoder1, Decoder2, and Decoder3 modules. The auxiliary predictor is an auxiliary component used in deep learning models to enhance the learning process of the main task by providing additional supervision signals. By introducing auxiliary tasks, the model can learn more information and features, thereby improving the prediction accuracy of the main task. By introducing additional supervision signals, the auxiliary predictor helps the model maintain better generalization ability during training. The auxiliary predictor provides more feedback signals to the model, which helps the model converge faster, especially when data is scarce. Finally, the features of the four different scales are fused.

[0057] The cell physical position prediction module has the features extracted by the multi-scale feature extraction module as input, and the extracted features are input into multiple prediction heads to predict different information. The first prediction head generates a heatmap, which represents the classification information. Each category has a heatmap, and a peak point is used on the output heatmap to represent the center position of the target. This heatmap is a two-dimensional probability distribution, and the peak position represents the most likely position of the target center. In this embodiment, there are two types of heatmaps, namely the center point of the cell and the vertex of the cell. Due to operations such as downsampling, there may be a certain deviation in the predicted center point position. Therefore, a prediction head is used to predict a center offset to fine-tune the center point position, so that the prediction result is more accurate. Next, based on the distance from the center point of the cell to the four vertices, the cell bbox is decoded; at the same time, based on the cell vertex, the distance from the center point of the cell sharing the vertex is regressed to decode the gbox. Based on the gbox, the discrete bboxes are spliced ​​together to obtain the physical position of the cell, and the output consists of n rows of data, where n is the number of cells in the table, and each row of data records the coordinates of the four vertices of the cell.

[0058] The cell logical position prediction module inputs the features extracted by the multi-scale feature extraction module into a prediction head, and extracts the features of each cell in combination with the heatmap in the physical position prediction module. Then the physical position of the cell is embedded and fused with the cell features, the cell features are interacted through the attention mechanism, and finally the logical position is regressed through the linear layer. Each cell is output as four fields, namely the start column, end column, start row and end row.

[0059] The loss function of the physical location prediction module is: L det =L k +L off +L cv +L vc .

[0060] L k Represents the heatmap loss, and the heatmap output by the network is Where R represents the stride of the output relative to the original image. C represents the number of categories, W represents the width of the original image, and H represents the height of the original image. The heatmap loss expression is as follows:

[0061]

[0062] α and β are hyperparameters, and N represents the number of key points in the image. Represents a point on the heatmap, is the predicted heatmap, xyc is the coordinate information, and the corresponding point is found according to the coordinate information. xyc is a point on the heatmap generated based on the actual cell location for training. (1-Y xyc ) β The closer this weight parameter is to the center, the smaller its value is. This weight parameter is used to reduce the severity of the penalty.

[0063] Because the network downsamples the image, remapping the feature map to the original image will produce errors, so a bias is predicted for each key point L off Indicates bias loss, used to bias To supervise, the expression is as follows:

[0064]

[0065] Where p represents the center point of the target box, R represents the downsampling multiple, R = 4; represents the deviation value, The output of the center offset prediction head of the cell physical location prediction module. is a point on the output feature map.

[0066] L cv represents the loss function based on the regression of the cell center point to the vertex; L vc Represents the loss function based on cell vertices and regressing to the cell center.

[0067] The loss of the cell logical position prediction module is:

[0068]

[0069] Among them, N is the number of key points in the image, which is also the number of cells; Indicates the use of L1 loss function, l k Represents the four real logical coordinates of each cell, namely the starting column, ending column, starting row, and ending row. The four logical coordinates for the prediction.

[0070] In order to enable the cell logical position prediction model to better understand the dependencies and constraints between the cell logical positions, a loss function L based on cell span is introduced. span and the cell conflict-based loss function L conf :

[0071]

[0072]

[0073] Among them, A is the set of cells combined in pairs, M is the set of cells, is the predicted logical position, i and j represent different cells, r e Represents the end of the line, r s Represents the starting line, c e Represents the end column, c s Represents the starting column.

[0074] The total loss function is: L = L det +L log .

[0075] Step 3: Use batch stochastic gradient descent and Adam optimization algorithm to train the table structure recognition model. The table structure recognition model is trained for 100 epochs with an initial learning rate of 1×10 -4 , which decays to 1×10 at the 70th and 90th epochs. -5 and 1×10 -6 .

[0076] Step 4: Input the test set into the trained table structure recognition model to obtain the prediction results, which are the physical location information and logical location information of each cell.

[0077] The table structure recognition model of this embodiment is compared with other models on the WTW-Dataset (Wild Table Dataset) data set. As shown in Table 1, the accuracy, recall rate, F1 score and logical position accuracy of the table structure recognition model in this embodiment are better than those of other models. As shown in Table 2, the table structure recognition model in this embodiment is compared with the model based on table adjacent relationship prediction, and it can be seen that the table structure recognition model in this embodiment is better than other models.

[0078] Table 1 Comparison results of table structure recognition methods

[0079] Precision(%) Recall(%) F1(%) Acc(%) TGRNet 64.7 24.3 Lore 96.4 82.9 Ours 97.5 96.0 96.7 84.4

[0080] Table 2 Comparison results of table cell neighbor prediction methods

[0081] Precision(%) Recall(%) F1(%) FLAGNet 91.6 89.5 90.5 NCGM 93.7 94.6 94.1 Ours 94.7 94.0 94.4

[0082] In order to study how the key components of the proposed model contribute to the regression of physical and logical locations, an in-depth ablation study is performed on the WTW dataset. First, two models are trained, one with an auxiliary trainer and the other without an auxiliary trainer, to evaluate their effectiveness. According to the results in Table 3, it can be seen that the auxiliary trainer effectively improves the F1 value of the model. Then, three models are trained, one using only the L1 loss function.

[0083] One model uses L1 loss function and intra-cell loss function, and the last model uses L1 loss function, inter-cell loss function, row span loss function, column span and cell span loss function to evaluate their effectiveness. According to the results in Table 4, it can be seen that span-based loss function and conflict-based loss function effectively improve the accuracy of logical position prediction.

[0084] Table 3 Comparison of ablation experiment results for physical location prediction

[0085] Precision(%) Recall(%) F1(%) No auxiliary trainer 94.7 96.4 95.6 With auxiliary training device 97.5 96.0 96.7

[0086] Table 4 Comparison of ablation experimental results of logical position prediction

[0087] L1 Inter Intra Row-span Col-span Acc √ 79.6 √ √ 80.8 √ √ √ √ √ 84.4

[0088] Step 5: Next, the image is segmented according to the cell physical location information predicted by the cell physical location prediction module and then input into the PaddleOCR model (an open source OCR model based on PaddlePaddle) to output the cell content. Then, the table structure information is generated according to the cell logical location information predicted by the cell logical location prediction module. The table is generated by combining the cell content and structure information, and the table is represented by an HTML sequence.

[0089] Step 6: Design prompt words and use the big model to extract information from HTML sequences. The key words are: Assume that you are an entity-relation triple extraction model. I will give you the head entity type list: subject_types, the tail entity type list: object_types, and the relationship list: relations. Give you an HTML sequence, and extract the triples in the sequence based on these three lists. After entering subject_types, object_types, and relations, we can get triples.

[0090] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope defined by the claims of the present invention.

Claims

1. A method for extracting information from a picture-type table, characterized in that: Extract information of different scales in the image through a deeply nested U-shaped network; By introducing fine-grained skip connections and recursive structures, the fusion of features at different levels and details is improved, allowing the model to recognize cells of different scales; the model performance is further improved by adding auxiliary predictors; The specific steps include: Step 1: Obtain the table structure recognition dataset, preprocess it, and divide the table structure recognition dataset into a training set, a validation set, and a test set; Step 2: Build a table structure recognition model and a loss function; The table structure recognition model includes a multi-scale feature extraction module, a cell physical position prediction module and a cell logical position prediction module; Step 3: Use batch stochastic gradient descent and Adam optimization algorithm to train the table structure recognition model; Step 4: Input the test set into the trained table structure recognition model to obtain the prediction results, including the physical location information and logical location information of each cell; Step 5: Segment the image based on the predicted physical location information of the cells and input it into the PaddleOCR model, which is an open source OCR model based on PaddlePaddle, to output the cell content; generate the table structure information based on the predicted logical location of the cells; then combine the cell content and structure information to generate a table, which is represented by an HTML sequence; Step 6: Design prompt words and use the big model to extract information from the HTML sequence. The keywords are: Assuming you are an entity-relationship triple extraction model, I will give you the head entity type list subject_types, the tail entity type list object_types, and the relationship list relations. I will give you an HTML sequence. Please extract the triples in the sequence based on these three lists; enter subject_types, object_types, and relations to get the triples.

2. The information extraction method for a picture-type table according to claim 1, characterized in that: The step 1 specifically includes: Step 1.1: Obtain a table structure recognition dataset; the table structure recognition dataset is a table detection and structure recognition dataset for natural scenes, which includes not only table images extracted from photos, scanned documents and web pages, but also tilted tables, curved tables, occluded or blurred tables, tables with extreme aspect ratios, superimposed tables, multi-color tables and irregular tables; Step 1.2: Preprocess the table structure data set; The preprocessing includes: scaling the image according to the set input size, and performing data enhancement by random cropping, translation and rotation; Step 1.3: Divide the dataset into training set, test set and validation set; The dataset is divided into a training set and a test set; the test set in the dataset is further divided into a test set and a validation set in a 1:1 ratio.

3. The information extraction method for a picture-type table according to claim 2, characterized in that: In step 1.1, the open source dataset WTW-Dataset is downloaded from the Internet; both the training set and the test set include training images and labels, the labels are in XML file format, each image corresponds to an XML, and the labels include the bbox coordinate information of the cell and the cell structure information; the structure information includes the starting column, the ending column, the starting row, the ending row and the table number.

4. The information extraction method for a picture-type table according to claim 3, characterized in that: The multi-scale feature extraction module in step 2 is modified on the basis of U2-Net as a feature extractor; first, a 7×7 convolution layer with a step size of 2 and a padding of 3 is used to extract shallow features, then a BN layer and a ReLU activation function are used, and then a maximum pooling layer with a step size of 2 is used, and then a deep nested U-shaped structure network is used to extract features of four different scales; the deep nested U-shaped structure network includes 7 modules, namely 4 Encoder layers and 3 Decoder layers, namely Encoder1, Encoder2, Encoder3, Decoder1, Decoder2, Decoder3, which are the same type of modules with different depths; in the Encoder4 module, all sampling layers are replaced with dilated convolutions, and no upsampling or downsampling is performed; an auxiliary predictor is connected to the Encoder4, Decoder1, Decoder2 and Decoder3 modules, and the auxiliary predictor is an auxiliary component used in a deep learning model to enhance the learning process of the main task by providing additional supervision signals; finally, the features of the four different scales are feature fused.

5. The method for extracting information from a picture-type table according to claim 4, characterized in that: The cell physical position prediction module in step 2 has features extracted by the multi-scale feature extraction module as input, and the extracted features are input into multiple prediction heads to predict different information. The first prediction head generates a heatmap to represent classification information. Each category has a heatmap, and a peak point is used on the output heatmap to represent the center position of the target. The heatmap is a two-dimensional probability distribution, and the peak position represents the most likely location of the target center. The heatmap has two types, namely the center point of the cell and the vertex of the cell. A prediction head is used to predict a center offset to fine-tune the center point position, so that the prediction result is more accurate. Next, the distance from the center point of the cell to the four vertices is regressed to decode the cell bbox, and at the same time, based on the cell vertex, the distance from the center point of the cell sharing the vertex is regressed to decode the gbox. Based on gbox, the discrete bboxes are spliced ​​together to obtain the physical position of the cell.

6. The method for extracting information from a picture-type table according to claim 5, characterized in that: The cell logical position prediction module in step 2 inputs the features extracted by the multi-scale feature extraction module into a prediction head, and extracts the features of each cell in combination with the heatmap in the physical position prediction module; then the physical position of the cell is embedded and fused with the features of the cell, the features of the cell are interacted through the attention mechanism, and finally the logical position is regressed through the linear layer; each cell output is four fields, namely the start column, the end column, the start row and the end row.

7. The method for extracting information from a picture-type table according to claim 6, characterized in that: The loss function of the cell physical location prediction module is: L det =L k +L off +L cv +L vc ; Among them, L k Represents the heatmap loss, and the heatmap output by the network is It is a three-dimensional matrix, where R represents the stride of the output relative to the original image, C represents the number of categories; W represents the width of the original image, and H represents the height of the original image; The heatmap loss expression is as follows: Among them, α and β are hyperparameters, and N represents the number of key points in the image; represents a point on the heatmap, xyc is the coordinate information, and the corresponding point is found according to the coordinate information; xyc is a point on the heatmap generated based on the actual cell position for training; the weight parameter (1-Y xyc ) β The closer to the center, the smaller the value. This weight parameter is used to reduce the penalty. L off Represents the bias loss, which is used to bias the prediction of each key point To supervise, the expression is as follows: Where p represents the center point of the target box, R represents the downsampling multiple, R = 4; represents the deviation value; The output of the center offset prediction head of the cell physical location prediction module. is a point on the output feature map; L cv represents the loss function based on the regression of the cell center point to the vertex; L vc Represents the loss function based on cell vertices and regressing to the cell center.

8. The method for extracting information from a picture-type table according to claim 7, characterized in that: The loss function of the cell logical position prediction module is: Among them, N is the number of key points in the image, which is also the number of cells; Indicates the use of L1 loss function, l k Represents the four real logical coordinates of each cell, namely the starting column, ending column, starting row, and ending row. are the four logical coordinates for prediction; Introducing the cell span-based loss function L span and the cell conflict-based loss function L conf , which is used to make the cell logic position prediction model better understand the dependencies and constraints between the cell logic positions, L span and L conf The loss function is as follows: Among them, A is the set of cells combined in pairs, and M is the set of cells; is the logical position of the prediction; i and j represent different cells, r e Represents the end of the line, r s Represents the starting line, c e Represents the end column, c s Represents the starting column.