Lightweight table structure identification method

By building a lightweight deep learning model, combining downsampling module, feature extraction module and feature fusion module, the problem of limited table recognition performance in the existing technology is solved, and accurate identification and efficient inference of complex table structures are achieved.

CN119919952AActive Publication Date: 2025-05-02NANCHANG HANGKONG UNIVERSITY
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510409433.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-05-02
Estimated Expiration
2045-04-02

AI Technical Summary

Technical Problem

When faced with complex and diverse table structures and styles, the performance of existing table recognition methods is limited by the complexity and generalization ability of manual feature design, making it difficult to accurately identify table structures in various scenarios.

Method used

Using a lightweight deep learning model structure, a table structure recognition model is constructed through the combination of five downsampling modules, one activation function ReLU6, sixteen tabular feature extraction modules, nine tabular feature fusion modules and six prediction heads to achieve accurate identification of table structure.

Benefits of technology

On the basis of ensuring the recognition accuracy, it effectively reduces the consumption of computing resources, improves the inference speed of the model, and realizes rapid detection and accurate identification of table structures in various scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119919952A_ABST
    Figure CN119919952A_ABST
Patent Text Reader

Abstract

The invention discloses a lightweight table structure recognition method, which comprises the following steps of: collecting a public table image and corresponding annotation data, and performing primary processing to obtain a first table structure recognition data set; preprocessing the first table structure identification data set to obtain a second table structure identification data set; dividing the second table structure identification data set into a training set, a verification set and a test set; designing the structure of the table structure recognition model, and initializing parameters of the table structure recognition model to obtain a first table structure recognition model; training the first table structure recognition model by using the second table structure recognition data set to obtain a second table structure recognition model; and inputting the to-be-detected image into the second table structure identification model for reasoning and decoding to obtain an identification result of the table structure. The method has the advantages that on the basis of ensuring the detection precision, lightweight construction of the table structure recognition model is completed, and the model reasoning speed is increased.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the fields of computer vision and table structure recognition, and in particular to a lightweight table structure recognition method. Background Art

[0002] In today's era, tables are important tools for recording and summarizing key data, facilitating the rapid acquisition and clear presentation of information. With the large-scale application of smartphones, it has become a common phenomenon to use camera devices to obtain table images for information transmission. Therefore, the ability to automatically extract and parse table structures from photos or images of different scenes has become an urgent need. Although traditional table recognition methods, such as rule-based methods and machine learning methods, have achieved certain results in specific scenarios, their performance is often limited by the complexity and generalization ability of manual feature design when faced with complex and diverse table structures and styles. The table structure recognition method based on deep learning technology can automatically learn and extract high-level features in table images, so that the model shows strong generalization ability on diverse table images.

[0003] Therefore, it is particularly important to design a new method that can accurately identify table structures while taking into account adaptability and robustness. This technology came into being in this context, and is committed to achieving accurate recognition of table structures through innovative deep learning algorithms and strategies, while effectively reducing computing resource consumption and greatly improving reasoning speed. Summary of the invention

[0004] The present invention aims to provide a lightweight table structure recognition method that can quickly detect photographed and scanned table images in various scenarios and provide accurate table structure recognition results. The method adopts a lightweight model structure to effectively reduce computing resource consumption while ensuring high recognition accuracy.

[0005] The technical solution proposed by the present invention is as follows: A lightweight table structure recognition method mainly includes the following steps: Step S1: Obtaining a table structure recognition data set: collecting public table images and corresponding annotation data, and extracting the physical coordinates and logical coordinates of cells representing the table structure in the annotation data to obtain a first table structure recognition data set T1; Step S2: Preprocess the table structure recognition data set: perform data enhancement and standardization on the first table structure recognition data set T1 to obtain a second table structure recognition data set T2; divide the second table structure recognition data set T2 into a training set data t , validation set data v And test set data e ; Step S3: construct a table structure recognition model: design the basic architecture of the table structure recognition model, and connect five downsampling modules, one activation function ReLU6, sixteen table feature extraction modules, nine table feature fusion modules and six prediction heads in sequence according to the basic architecture of the table structure recognition model to obtain an uninitialized table structure recognition model Mr; initialize the parameters of the uninitialized table structure recognition model Mr to obtain a first table structure recognition model M; Step S4: training the table structure recognition model: using the second table structure recognition data set T2 to train the first table structure recognition model M to obtain a second table structure recognition model Mt; Step S5: Identify the table structure: from the test set data e A table image is randomly selected and input into the second table structure recognition model Mt for inference to obtain six feature maps representing the table structure; the six feature maps representing the table structure are decoded to obtain a table structure recognition result including the physical coordinates and logical coordinates of the cells.

[0006] Furthermore, in step S1, publicly available table images and corresponding annotation data are collected, and the physical coordinates and logical coordinates of cells representing the table structure in the annotation data are extracted to obtain a first table structure recognition data set T1. The specific steps are as follows: Step S11, collecting photographed and scanned table images containing table structure annotation data in various scenarios through the Internet to obtain a first table image; Step S12, parsing the annotation data corresponding to the first table image to obtain physical coordinates and logical coordinates of cells representing the table structure; correcting and integrating the physical coordinates and logical coordinates of cells representing the table structure to obtain a first label; Step S13: Combine the first table image and the first label to obtain a first data sample, and add the first data sample to the first table structure recognition data set T1.

[0007] Furthermore, in step S2, data enhancement and standardization are performed on the first table structure recognition data set T1 to obtain a second table structure recognition data set T2. The specific steps are as follows: Step S21, set the input image height h, set the input image width w, let I f represents the first table image in the fth data sample of the first table structure recognition data set T1, let g f Indicates the first label in the fth data sample of the first table structure identification data set T1, f=1,2,…,n, n is the number of data samples in the first table structure identification data set T1; sets an iteration variable k, and the initial value of the iteration variable k is set to 1; Step S22: identifying the first table image I in the kth data sample of the first table structure data set T1 k Perform color enhancement to obtain a color enhanced table image Ic; divide each pixel value of the color enhanced table image Ic by 255, perform data normalization, and obtain a normalized image Iv; Step S23, adjust the height of the normalized image Iv to the input image height h, and adjust the width to the input image width w, to obtain the second table image Ip and the affine transformation matrix AM; use the affine transformation matrix AM to identify the first label g in the kth data sample of the first table structure data set T1 k The physical coordinates of the cells in are transformed to obtain the second label g p ; Step S24: the second table image Ip and the second label g p Combine to obtain a second data sample, and add the second data sample to the second table structure recognition data set T2; Step S25, increase the value of the iteration variable k by 1, and if k≤n, jump to step S22.

[0008] Furthermore, in step S3, the basic architecture of the table structure recognition model is designed, and five downsampling modules, one activation function ReLU6, sixteen table feature extraction modules, nine table feature fusion modules and six prediction heads are connected in sequence according to the basic architecture of the table structure recognition model to obtain an uninitialized table structure recognition model Mr. The specific steps are as follows: Step S311, the first down-sampling module receives the preprocessed second table image Ip, the output of the first down-sampling module is connected to the input of the first activation function ReLU6; the output of the first activation function ReLU6 is connected to the input of the second down-sampling module, the output of the second down-sampling module is connected to the input of the first table feature extraction module, the output of the first table feature extraction module is connected to the input of the second table feature extraction module; the output of the second table feature extraction module is connected to the input of the third down-sampling module, the output of the third down-sampling module is connected to the input of the third table feature extraction module, the output of the third table feature extraction module is connected to the input of the fourth table feature extraction module; the output of the fourth table feature extraction module is connected to the input of the fourth down-sampling module, the output of the fourth down-sampling module is connected to the input of the fifth table feature extraction module, the output of the fifth table feature extraction module is connected to the input of the sixth table feature extraction module, the output of the sixth table feature extraction module The output of the seventh table feature extraction module is connected to the input of the eighth table feature extraction module, the output of the eighth table feature extraction module is connected to the input of the ninth table feature extraction module, the output of the ninth table feature extraction module is connected to the input of the tenth table feature extraction module, the output of the tenth table feature extraction module is connected to the input of the eleventh table feature extraction module, and the output of the eleventh table feature extraction module is connected to the input of the twelfth table feature extraction module; the output of the twelfth table feature extraction module is connected to the input of the fifth down-sampling module, the output of the fifth down-sampling module is connected to the input of the thirteenth table feature extraction module, the output of the thirteenth table feature extraction module is connected to the input of the fourteenth table feature extraction module, the output of the fourteenth table feature extraction module is connected to the input of the fifteenth table feature extraction module, and the output of the fifteenth table feature extraction module is connected to the input of the sixteenth table feature extraction module; Step S312, the output of the twelfth table feature extraction module is connected to the first input of the first table feature fusion module and the output of the sixteenth table feature extraction module is connected to the second input of the first table feature fusion module; the output of the fourth table feature extraction module is connected to the first input of the second table feature fusion module and the output of the twelfth table feature extraction module is connected to the second input of the second table feature fusion module, the output of the second table feature fusion module is connected to the first input of the third table feature fusion module and the output of the first table feature fusion module is connected to the second input of the third table feature fusion module; the output of the second table feature extraction module is connected to the first input of the fourth table feature fusion module and the output of the fourth table feature extraction module is connected to the second input of the fourth table feature fusion module, and the output of the fourth table feature fusion module is connected to the fifth table feature fusion module The first input and the output of the second table feature fusion module are connected to the second input of the fifth table feature fusion module, the output of the fifth table feature fusion module is connected to the first input of the sixth table feature fusion module and the output of the third table feature fusion module is connected to the second input of the sixth table feature fusion module; the output of the third table feature fusion module is connected to the first input of the seventh table feature fusion module and the output of the first table feature fusion module is connected to the second input of the seventh table feature fusion module; the output of the sixth table feature fusion module is connected to the first input of the eighth table feature fusion module and the output of the third table feature fusion module is connected to the second input of the eighth table feature fusion module, the output of the eighth table feature fusion module is connected to the first input of the ninth table feature fusion module and the output of the seventh table feature fusion module is connected to the second input of the ninth table feature fusion module; Step S313, the output of the ninth table feature fusion module is simultaneously connected to the first prediction head, the second prediction head, the third prediction head, the fourth prediction head, the fifth prediction head and the sixth prediction head to obtain an uninitialized table structure recognition model Mr.

[0009] Furthermore, in step S3, any downsampling module among the five downsampling modules includes a convolution layer and a batch normalization layer; the connection order of the downsampling modules is as follows: the input of the downsampling module is connected to the input of the first convolution layer, the output of the first convolution layer is connected to the input of the first batch normalization layer, and the output of the first batch normalization layer is the output of the downsampling module.

[0010] Further, in step S3, any of the sixteen table feature extraction modules includes three convolutional layers, two depth-separable convolutional layers, an activation function ReLU6 and two batch normalization layers; specifically: The input of the table feature extraction module is connected to the input of the first depth-separable convolutional layer, the output of the first depth-separable convolutional layer is connected to the input of the second batch normalization layer, the output of the second batch normalization layer is connected to the input of the second convolutional layer and the input of the third convolutional layer at the same time, the output of the second convolutional layer is connected to the input of the second activation function ReLU6, the output of the second activation function ReLU6 and the output of the third convolutional layer are element-wise multiplied and connected to the input of the fourth convolutional layer, the output of the fourth convolutional layer is connected to the input of the third batch normalization layer, the output of the third batch normalization layer is connected to the input of the second depth-separable convolutional layer, and the output of the second depth-separable convolutional layer and the input of the table feature extraction module are added together to form the output of the table feature extraction module.

[0011] Further, in step S3, any of the nine table feature fusion modules includes two deformable convolution layers, one transposed convolution upsampling layer, two batch normalization layers and two activation functions ReLU; specifically: The second input of the table feature fusion module is connected to the input of the first deformable convolution layer, the output of the first deformable convolution layer is connected to the input of the fourth batch normalization layer, the output of the fourth batch normalization layer is connected to the input of the first activation function ReLU, the output of the first activation function ReLU is connected to the input of the first transposed convolution upsampling layer, the output of the first transposed convolution upsampling layer and the first input of the table feature fusion module are added to the output of the second deformable convolution layer, the output of the second deformable convolution layer is connected to the input of the fifth batch normalization layer, the output of the fifth batch normalization layer is connected to the input of the second activation function ReLU, and the output of the second activation function ReLU is the output of the table feature fusion module.

[0012] Furthermore, in step S3, the parameters of the uninitialized table structure recognition model Mr are initialized to obtain the first table structure recognition model M. The specific steps are as follows: Step S32, initialize the weights of the transposed convolution upsampling layer of any of the nine table feature fusion modules to bilinear interpolation weights, initialize the bias of any of the six prediction heads to 0, and randomly initialize the remaining uninitialized model parameters to obtain the first table structure recognition model M.

[0013] Furthermore, in step S4, the first table structure recognition model M is trained using the second table structure recognition data set T2 to obtain a second table structure recognition model Mt. The specific steps are as follows: Step S41, set the number of model training rounds epochs, set the training data batch size batch, set the optimizer learning rate lr, set the learning rate prime decay step set steps; set the optimal verification loss value loss o, set the optimal table structure recognition model Mo, the optimal verification loss value loss o The initial value of is set to positive infinity, and the parameters of the optimal table structure recognition model Mo are initialized to the parameters of the first table structure recognition model M; let epoch represent the current training round, and initialize the current training round epoch to 1; Step S42, let count represent the current iteration batch, and initialize the current iteration batch count to , where m is the training set data t The number of data samples in ; Step S43, from the training set data t Randomly extract a second data sample of batch size, input the second table image in the second data sample into the first table structure recognition model M for forward propagation, and obtain the prediction result of the current iteration batch model; use the second label in the second data sample to calculate the loss of the prediction result of the current iteration batch model, and obtain the loss of the current iteration batch; Step S44, based on the loss of the current iteration batch, an adaptive momentum estimation optimization algorithm is used to update the parameters of the first table structure recognition model M through a back propagation process; Step S45, if epoch∈steps, divide the value of the learning rate lr by 10; Step S46, subtract 1 from the value of the current iteration batch count, if count>1, jump to step S43; Step S47: Verification set data v The second data sample in is read in batches and forward propagation is performed sequentially to calculate the validation set data v The average loss loss in the current training round epoch a ; Step S48, if loss a <loss o , then the optimal verification loss value loss o The value is set to loss a , and update the parameters of the optimal table structure recognition model Mo to the parameters of the first table structure recognition model M under the current training round epoch; Step S49, increase the value of the current training round epoch by 1; if epoch<epochs, jump to step S42; otherwise, save the optimal table structure recognition model Mo as the second table structure recognition model Mt.

[0014] The beneficial effects of the present invention are as follows: the method adopts an efficient table feature extraction module to realize table feature capture, optimizes the multi-scale feature fusion process through a table feature fusion module, and accurately identifies the physical coordinates and logical coordinates of cells in the table structure; on the basis of ensuring detection accuracy, the lightweight construction of the table structure recognition model is completed, thereby improving the model reasoning speed. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 A method flow chart of an example of the present invention; Figure 2 It is a structural diagram of the table structure recognition model of the present invention; Figure 3 This is a structural diagram of the downsampling module of the present invention; Figure 4 This is a structural diagram of the table feature extraction module of the present invention; Figure 5 This is a structural diagram of the table feature fusion module of the present invention. DETAILED DESCRIPTION

[0016] The specific implementation of the present invention will be further described in detail below in conjunction with the accompanying drawings in the embodiments of the present invention. It should be noted that the specific embodiment of a lightweight table structure recognition method according to the present invention is only used as an example and is not intended to limit the present invention.

[0017] The present invention provides a lightweight table structure recognition method, such as Figure 1 As shown, the method comprises the following steps: Step S1: Obtaining a table structure recognition data set: collecting public table images and corresponding annotation data, and extracting the physical coordinates and logical coordinates of cells representing the table structure in the annotation data to obtain a first table structure recognition data set T1; Step S2: Preprocess the table structure recognition data set: perform data enhancement and standardization on the first table structure recognition data set T1 to obtain a second table structure recognition data set T2; divide the second table structure recognition data set T2 into training set data according to a ratio of 8:1:1 t , validation set data v And test set data e ; Step S3: construct a table structure recognition model: design the basic architecture of the table structure recognition model, and connect five downsampling modules, one activation function ReLU6, sixteen table feature extraction modules, nine table feature fusion modules and six prediction heads in sequence according to the basic architecture of the table structure recognition model to obtain an uninitialized table structure recognition model Mr; initialize the parameters of the uninitialized table structure recognition model Mr to obtain a first table structure recognition model M; Step S4: training the table structure recognition model: using the second table structure recognition data set T2 to train the first table structure recognition model M to obtain a second table structure recognition model Mt; Step S5: Identify the table structure: from the test set data e A table image is randomly selected and input into the second table structure recognition model Mt for inference to obtain six feature maps representing the table structure; the six feature maps representing the table structure are decoded to obtain a table structure recognition result including the physical coordinates and logical coordinates of the cells.

[0018] Furthermore, in step S1, publicly available table images and corresponding annotation data are collected, and the physical coordinates and logical coordinates of cells representing the table structure in the annotation data are extracted to obtain a first table structure recognition data set T1. The specific steps are as follows: Step S11, collecting photographed and scanned table images containing table structure annotation data in various scenarios through the Internet to obtain a first table image; Step S12, parsing the annotation data corresponding to the first table image to obtain physical coordinates and logical coordinates of cells representing the table structure; correcting and integrating the physical coordinates and logical coordinates of cells representing the table structure to obtain a first label; Step S13: Combine the first table image and the first label to obtain a first data sample, and add the first data sample to the first table structure recognition data set T1.

[0019] Furthermore, in step S2, data enhancement and standardization are performed on the first table structure recognition data set T1 to obtain a second table structure recognition data set T2. The specific steps are as follows: Step S21, set the input image height h=1024, set the input image width w=1024, let I f represents the first table image in the fth data sample of the first table structure recognition data set T1, let g f Indicates the first label in the fth data sample of the first table structure identification data set T1, f=1,2,…,n, n is the number of data samples in the first table structure identification data set T1; sets a first variable k, and the initial value of the first variable k is set to 1; Step S22: identifying the first table image I in the kth data sample of the first table structure data set T1 k Perform color enhancement to obtain a color enhanced table image Ic; divide each pixel value of the color enhanced table image Ic by 255, perform data normalization, and obtain a normalized image Iv; Step S23, adjust the height of the normalized image Iv to the input image height h, and adjust the width to the input image width w, to obtain the second table image Ip and the affine transformation matrix AM; use the affine transformation matrix AM to identify the first label g in the kth data sample of the first table structure data set T1 k The physical coordinates of the cells in are transformed to obtain the second label g p ; Step S24: the second table image Ip and the second label g p Combine to obtain a second data sample, and add the second data sample to the second table structure recognition data set T2; Step S25, increase the value of the iteration variable k by 1, and if k≤n, jump to step S22.

[0020] Furthermore, if Figure 2 As shown, the basic architecture of the table structure recognition model is designed, and five downsampling modules, one activation function ReLU6, sixteen table feature extraction modules, nine table feature fusion modules and six prediction heads are connected in sequence according to the basic architecture of the table structure recognition model to obtain an uninitialized table structure recognition model Mr. The specific steps are as follows: Step S311, the first down-sampling module receives the preprocessed second table image Ip, the output of the first down-sampling module is connected to the input of the first activation function ReLU6; the output of the first activation function ReLU6 is connected to the input of the second down-sampling module, the output of the second down-sampling module is connected to the input of the first table feature extraction module, the output of the first table feature extraction module is connected to the input of the second table feature extraction module; the output of the second table feature extraction module is connected to the input of the third down-sampling module, the output of the third down-sampling module is connected to the input of the third table feature extraction module, the output of the third table feature extraction module is connected to the input of the fourth table feature extraction module; the output of the fourth table feature extraction module is connected to the input of the fourth down-sampling module, the output of the fourth down-sampling module is connected to the input of the fifth table feature extraction module, the output of the fifth table feature extraction module is connected to the input of the sixth table feature extraction module, the output of the sixth table feature extraction module The output of the seventh table feature extraction module is connected to the input of the eighth table feature extraction module, the output of the eighth table feature extraction module is connected to the input of the ninth table feature extraction module, the output of the ninth table feature extraction module is connected to the input of the tenth table feature extraction module, the output of the tenth table feature extraction module is connected to the input of the eleventh table feature extraction module, and the output of the eleventh table feature extraction module is connected to the input of the twelfth table feature extraction module; the output of the twelfth table feature extraction module is connected to the input of the fifth down-sampling module, the output of the fifth down-sampling module is connected to the input of the thirteenth table feature extraction module, the output of the thirteenth table feature extraction module is connected to the input of the fourteenth table feature extraction module, the output of the fourteenth table feature extraction module is connected to the input of the fifteenth table feature extraction module, and the output of the fifteenth table feature extraction module is connected to the input of the sixteenth table feature extraction module; Step S312, the output of the twelfth table feature extraction module is connected to the first input of the first table feature fusion module and the output of the sixteenth table feature extraction module is connected to the second input of the first table feature fusion module; the output of the fourth table feature extraction module is connected to the first input of the second table feature fusion module and the output of the twelfth table feature extraction module is connected to the second input of the second table feature fusion module, the output of the second table feature fusion module is connected to the first input of the third table feature fusion module and the output of the first table feature fusion module is connected to the second input of the third table feature fusion module; the output of the second table feature extraction module is connected to the first input of the fourth table feature fusion module and the output of the fourth table feature extraction module is connected to the second input of the fourth table feature fusion module, and the output of the fourth table feature fusion module is connected to the fifth table feature fusion module The first input and the output of the second table feature fusion module are connected to the second input of the fifth table feature fusion module, the output of the fifth table feature fusion module is connected to the first input of the sixth table feature fusion module and the output of the third table feature fusion module is connected to the second input of the sixth table feature fusion module; the output of the third table feature fusion module is connected to the first input of the seventh table feature fusion module and the output of the first table feature fusion module is connected to the second input of the seventh table feature fusion module; the output of the sixth table feature fusion module is connected to the first input of the eighth table feature fusion module and the output of the third table feature fusion module is connected to the second input of the eighth table feature fusion module, the output of the eighth table feature fusion module is connected to the first input of the ninth table feature fusion module and the output of the seventh table feature fusion module is connected to the second input of the ninth table feature fusion module; Step S313, the output of the ninth table feature fusion module is simultaneously connected to the first prediction head, the second prediction head, the third prediction head, the fourth prediction head, the fifth prediction head and the sixth prediction head to obtain an uninitialized table structure recognition model Mr; Further, in step S3, if Figure 3 As shown in FIG. 1 , any of the five downsampling modules includes a convolution layer and a batch normalization layer; the connection order of the downsampling modules is as follows: the input of the downsampling module is connected to the input of the first convolution layer, the output of the first convolution layer is connected to the input of the first batch normalization layer, and the output of the first batch normalization layer is the output of the downsampling module; Further, in step S3, if Figure 4 As shown, any of the sixteen table feature extraction modules includes three convolutional layers, two depth-separable convolutional layers, an activation function ReLU6 and two batch normalization layers; specifically: The input of the table feature extraction module is connected to the input of the first depth-separable convolutional layer, the output of the first depth-separable convolutional layer is connected to the input of the second batch normalization layer, the output of the second batch normalization layer is connected to the input of the second convolutional layer and the input of the third convolutional layer at the same time, the output of the second convolutional layer is connected to the input of the second activation function ReLU6, the output of the second activation function ReLU6 and the output of the third convolutional layer are element-wise multiplied and connected to the input of the fourth convolutional layer, the output of the fourth convolutional layer is connected to the input of the third batch normalization layer, the output of the third batch normalization layer is connected to the input of the second depth-separable convolutional layer, and the output of the second depth-separable convolutional layer and the input of the table feature extraction module are added to form the output of the table feature extraction module; Further, in step S3, if Figure 5 As shown, any of the nine table feature fusion modules includes two deformable convolution layers, one transposed convolution upsampling layer, two batch normalization layers and two activation functions ReLU; specifically: The second input of the table feature fusion module is connected to the input of the first deformable convolution layer, the output of the first deformable convolution layer is connected to the input of the fourth batch normalization layer, the output of the fourth batch normalization layer is connected to the input of the first activation function ReLU, the output of the first activation function ReLU is connected to the input of the first transposed convolution upsampling layer, the output of the first transposed convolution upsampling layer and the first input of the table feature fusion module are added to the output of the second deformable convolution layer, the output of the second deformable convolution layer is connected to the input of the fifth batch normalization layer, the output of the fifth batch normalization layer is connected to the input of the second activation function ReLU, and the output of the second activation function ReLU is the output of the table feature fusion module; Furthermore, in step S3, the parameters of the uninitialized table structure recognition model Mr are initialized to obtain the first table structure recognition model M. The specific steps are as follows: Step S32, initializing the weight of the transposed convolution upsampling layer of any table feature fusion module among the nine table feature fusion modules to the bilinear interpolation weight, initializing the bias of any prediction head among the six prediction heads to 0, and randomly initializing the remaining uninitialized model parameters to obtain a first table structure recognition model M; Furthermore, in step S4, the first table structure recognition model M is trained using the second table structure recognition data set T2 to obtain a second table structure recognition model Mt. The specific steps are as follows: Step S41, set the number of model training epochs = 200, set the training data batch size batch = 22, and set the optimizer learning rate lr = 1.25×10 -4, set the learning rate attenuation step set steps={140, 180}; set the optimal verification loss value loss o , set the optimal table structure recognition model Mo, the optimal verification loss value loss o The initial value of is set to positive infinity, and the parameters of the optimal table structure recognition model Mo are initialized to the parameters of the first table structure recognition model M; let epoch represent the current training round, and initialize the current training round epoch to 1; Step S42, let count represent the current iteration batch, and initialize the current iteration batch count to , where m is the training set data t The number of data samples in ; Step S43, from the training set data t Randomly extract a second data sample of batch size, input the second table image in the second data sample into the first table structure recognition model M for forward propagation, and obtain the prediction result of the current iteration batch model; use the second label in the second data sample to calculate the loss of the prediction result of the current iteration batch model, and obtain the loss of the current iteration batch; Step S44, based on the loss of the current iteration batch, an adaptive momentum estimation optimization algorithm is used to update the parameters of the first table structure recognition model M through a back propagation process; Step S45, if epoch∈steps, divide the value of the learning rate lr by 10; Step S46, subtract 1 from the value of the current iteration batch count, if count>1, jump to step S43; Step S47: Verification set data v The second data sample in is read in batches and forward propagation is performed sequentially to calculate the validation set data v The average loss loss in the current training round epoch a ; Step S48, if loss a <loss o , then the optimal verification loss value loss o The value is set to loss a , and update the parameters of the optimal table structure recognition model Mo to the parameters of the first table structure recognition model M under the current training round epoch; Step S49, increase the value of the current training round epoch by 1; if epoch<epochs, jump to step S42; otherwise, save the optimal table structure recognition model Mo as the second table structure recognition model Mt.

[0021] The beneficial effects of the present invention are as follows: the method adopts an efficient table feature extraction module to realize table feature capture, optimizes the multi-scale feature fusion process through a table feature fusion module, and accurately identifies the physical coordinates and logical coordinates of cells in the table structure; on the basis of ensuring detection accuracy, the lightweight construction of the table structure recognition model is completed, thereby improving the model reasoning speed.

[0022] The above description is a preferred embodiment of the present invention and is not intended to limit the present invention. Any equivalent or modification that does not deviate from the spirit disclosed in the present invention should be included in the scope of protection of the present invention.

Claims

1. A lightweight table structure recognition method, characterized by: The main steps include: Step S1: Obtaining a table structure recognition data set: collecting public table images and corresponding annotation data, and extracting the physical coordinates and logical coordinates of cells representing the table structure in the annotation data to obtain a first table structure recognition data set T1; Step S2: Preprocess the table structure recognition data set: perform data enhancement and standardization on the first table structure recognition data set T1 to obtain a second table structure recognition data set T2; divide the second table structure recognition data set T2 into a training set data t , validation set data v And test set data e ; Step S3: construct a table structure recognition model: design the basic architecture of the table structure recognition model, and connect five downsampling modules, one activation function ReLU6, sixteen table feature extraction modules, nine table feature fusion modules and six prediction heads in sequence according to the basic architecture of the table structure recognition model to obtain an uninitialized table structure recognition model Mr; initialize the parameters of the uninitialized table structure recognition model Mr to obtain a first table structure recognition model M; Step S4: training the table structure recognition model: using the second table structure recognition data set T2 to train the first table structure recognition model M to obtain a second table structure recognition model Mt; Step S5; Identify table structure: from test set data e A table image is randomly selected and input into the second table structure recognition model Mt for inference to obtain six feature maps representing the table structure; the six feature maps representing the table structure are decoded to obtain a table structure recognition result including the physical coordinates and logical coordinates of the cells.

2. A lightweight table structure recognition method according to claim 1, characterized in that: In step S1, publicly available table images and corresponding annotation data are collected, and the physical coordinates and logical coordinates of cells representing the table structure in the annotation data are extracted to obtain a first table structure recognition data set T1. The specific steps are as follows: Step S11, collecting photographed and scanned table images containing table structure annotation data in various scenarios to obtain a first table image; Step S12, parsing the annotation data corresponding to the first table image to obtain physical coordinates and logical coordinates of cells representing the table structure; correcting and integrating the physical coordinates and logical coordinates of cells representing the table structure to obtain a first label; Step S13: Combine the first table image and the first label to obtain a first data sample, and add the first data sample to the first table structure recognition data set T1.

3. A lightweight table structure recognition method according to claim 2, characterized in that: In step S2, data enhancement and standardization are performed on the first table structure recognition data set T1 to obtain a second table structure recognition data set T2. The specific steps are as follows: Step S21, set the input image height h, set the input image width w, let I f represents the first table image in the fth data sample of the first table structure recognition data set T1, let g f represents the first label in the f-th data sample of the first table structure identification data set T1, f=1,2,…,n, and n is the number of data samples in the first table structure identification data set T1; Set the iteration variable k, and the initial value of the iteration variable k is set to 1; Step S22: identifying the first table image I in the kth data sample of the first table structure data set T1 k Perform color enhancement to obtain a color enhanced table image Ic; Divide each pixel value of the color enhanced table image Ic by 255 to perform data normalization to obtain a normalized image Iv; Step S23, adjust the height of the normalized image Iv to the input image height h, and adjust the width to the input image width w, to obtain the second table image Ip and the affine transformation matrix AM; use the affine transformation matrix AM to identify the first label g in the kth data sample of the first table structure data set T1 k The physical coordinates of the cells in are transformed to obtain the second label g p ; Step S24: the second table image Ip and the second label g p Combine to obtain a second data sample, and add the second data sample to the second table structure recognition data set T2; Step S25, increase the value of the iteration variable k by 1, if k≤n, jump to step S22.

4. A lightweight table structure recognition method according to claim 3, characterized in that: In step S3, the basic architecture of the table structure recognition model is designed, and five downsampling modules, one activation function ReLU6, sixteen table feature extraction modules, nine table feature fusion modules and six prediction heads are connected in sequence according to the basic architecture of the table structure recognition model to obtain an uninitialized table structure recognition model Mr. The specific steps are as follows: Step S311, the first down-sampling module receives the preprocessed second table image Ip, the output of the first down-sampling module is connected to the input of the first activation function ReLU6; the output of the first activation function ReLU6 is connected to the input of the second down-sampling module, the output of the second down-sampling module is connected to the input of the first table feature extraction module, the output of the first table feature extraction module is connected to the input of the second table feature extraction module; the output of the second table feature extraction module is connected to the input of the third down-sampling module, the output of the third down-sampling module is connected to the input of the third table feature extraction module, the output of the third table feature extraction module is connected to the input of the fourth table feature extraction module; the output of the fourth table feature extraction module is connected to the input of the fourth down-sampling module, the output of the fourth down-sampling module is connected to the input of the fifth table feature extraction module, the output of the fifth table feature extraction module is connected to the input of the sixth table feature extraction module, the output of the sixth table feature extraction module The output of the seventh table feature extraction module is connected to the input of the eighth table feature extraction module, the output of the eighth table feature extraction module is connected to the input of the ninth table feature extraction module, the output of the ninth table feature extraction module is connected to the input of the tenth table feature extraction module, the output of the tenth table feature extraction module is connected to the input of the eleventh table feature extraction module, and the output of the eleventh table feature extraction module is connected to the input of the twelfth table feature extraction module; the output of the twelfth table feature extraction module is connected to the input of the fifth down-sampling module, the output of the fifth down-sampling module is connected to the input of the thirteenth table feature extraction module, the output of the thirteenth table feature extraction module is connected to the input of the fourteenth table feature extraction module, the output of the fourteenth table feature extraction module is connected to the input of the fifteenth table feature extraction module, and the output of the fifteenth table feature extraction module is connected to the input of the sixteenth table feature extraction module; Step S312, the output of the twelfth table feature extraction module is connected to the first input of the first table feature fusion module and the output of the sixteenth table feature extraction module is connected to the second input of the first table feature fusion module; the output of the fourth table feature extraction module is connected to the first input of the second table feature fusion module and the output of the twelfth table feature extraction module is connected to the second input of the second table feature fusion module, the output of the second table feature fusion module is connected to the first input of the third table feature fusion module and the output of the first table feature fusion module is connected to the second input of the third table feature fusion module; the output of the second table feature extraction module is connected to the first input of the fourth table feature fusion module and the output of the fourth table feature extraction module is connected to the second input of the fourth table feature fusion module, and the output of the fourth table feature fusion module is connected to the fifth table feature fusion module The first input and the output of the second table feature fusion module are connected to the second input of the fifth table feature fusion module, the output of the fifth table feature fusion module is connected to the first input of the sixth table feature fusion module and the output of the third table feature fusion module is connected to the second input of the sixth table feature fusion module; the output of the third table feature fusion module is connected to the first input of the seventh table feature fusion module and the output of the first table feature fusion module is connected to the second input of the seventh table feature fusion module; the output of the sixth table feature fusion module is connected to the first input of the eighth table feature fusion module and the output of the third table feature fusion module is connected to the second input of the eighth table feature fusion module, the output of the eighth table feature fusion module is connected to the first input of the ninth table feature fusion module and the output of the seventh table feature fusion module is connected to the second input of the ninth table feature fusion module; Step S313, the output of the ninth table feature fusion module is simultaneously connected to the first prediction head, the second prediction head, the third prediction head, the fourth prediction head, the fifth prediction head and the sixth prediction head to obtain an uninitialized table structure recognition model Mr.

5. A lightweight table structure recognition method according to claim 4, characterized in that: In step S3, any downsampling module among the five downsampling modules includes a convolution layer and a batch normalization layer; the connection order of the downsampling modules is as follows: the input of the downsampling module is connected to the input of the first convolution layer, the output of the first convolution layer is connected to the input of the first batch normalization layer, and the output of the first batch normalization layer is the output of the downsampling module.

6. A lightweight table structure recognition method according to claim 5, characterized in that: In step S3, any of the sixteen table feature extraction modules includes three convolutional layers, two depth-separable convolutional layers, an activation function ReLU6 and two batch normalization layers; specifically: The input of the table feature extraction module is connected to the input of the first depth-separable convolutional layer, the output of the first depth-separable convolutional layer is connected to the input of the second batch normalization layer, the output of the second batch normalization layer is connected to the input of the second convolutional layer and the input of the third convolutional layer at the same time, the output of the second convolutional layer is connected to the input of the second activation function ReLU6, the output of the second activation function ReLU6 and the output of the third convolutional layer are element-wise multiplied and connected to the input of the fourth convolutional layer, the output of the fourth convolutional layer is connected to the input of the third batch normalization layer, the output of the third batch normalization layer is connected to the input of the second depth-separable convolutional layer, and the output of the second depth-separable convolutional layer and the input of the table feature extraction module are added together to form the output of the table feature extraction module.

7. A lightweight table structure recognition method according to claim 6, characterized in that: In step S3, any of the nine table feature fusion modules includes two deformable convolution layers, one transposed convolution upsampling layer, two batch normalization layers and two activation functions ReLU; specifically: The second input of the table feature fusion module is connected to the input of the first deformable convolution layer, the output of the first deformable convolution layer is connected to the input of the fourth batch normalization layer, the output of the fourth batch normalization layer is connected to the input of the first activation function ReLU, the output of the first activation function ReLU is connected to the input of the first transposed convolution upsampling layer, the output of the first transposed convolution upsampling layer and the first input of the table feature fusion module are added to the output of the second deformable convolution layer, the output of the second deformable convolution layer is connected to the input of the fifth batch normalization layer, the output of the fifth batch normalization layer is connected to the input of the second activation function ReLU, and the output of the second activation function ReLU is the output of the table feature fusion module.

8. A lightweight table structure recognition method according to claim 7, characterized in that: In step S3, the parameters of the uninitialized table structure recognition model Mr are initialized to obtain the first table structure recognition model M. The specific steps are as follows: Step S32, initialize the weights of the transposed convolution upsampling layer of any of the nine table feature fusion modules to bilinear interpolation weights, initialize the bias of any of the six prediction heads to 0, and randomly initialize the remaining uninitialized model parameters to obtain the first table structure recognition model M.

9. A lightweight table structure recognition method according to claim 8, characterized in that: In step S4, the first table structure recognition model M is trained using the second table structure recognition data set T2 to obtain a second table structure recognition model Mt. The specific steps are as follows: Step S41, set the number of model training rounds epochs, set the training data batch size batch, set the optimizer learning rate lr, set the learning rate prime decay step set steps; set the optimal verification loss value loss o , set the optimal table structure recognition model Mo, the optimal verification loss value loss o The initial value of is set to positive infinity, and the parameters of the optimal table structure recognition model Mo are initialized to the parameters of the first table structure recognition model M; let epoch represent the current training round, and initialize the current training round epoch to 1; Step S42, let count represent the current iteration batch, and initialize the current iteration batch count to , where m is the training set data t The number of data samples in ; Step S43, from the training set data t Randomly extract a batch size of second data samples, input the second table image in the second data sample into the first table structure recognition model M for forward propagation, and obtain the prediction result of the current iteration batch model; use the second label in the second data sample to calculate the loss of the prediction result of the current iteration batch model, and obtain the loss loss of the current iteration batch; Step S44, based on the loss of the current iteration batch, an adaptive momentum estimation optimization algorithm is used to update the parameters of the first table structure recognition model M through a back propagation process; Step S45, if epoch∈steps, divide the value of the learning rate lr by 10; Step S46, subtract 1 from the value of the current iteration batch count, if count>1, jump to step S43; Step S47: Verification set data v The second data sample in is read in batches and forward propagation is performed sequentially to calculate the validation set data v The average loss loss in the current training round epoch a ; Step S48, if loss a <loss o , then the optimal verification loss value loss o The value is set to loss a , and update the parameters of the optimal table structure recognition model Mo to the parameters of the first table structure recognition model M under the current training round epoch; Step S49, increase the value of the current training round epoch by 1; if epoch<epochs, jump to step S42; otherwise, save the optimal table structure recognition model Mo as the second table structure recognition model Mt.

Citation Information

Patent Citations

  • Table structure analysis method and device

    CN116092105A

  • Table analysis method and device, computer equipment and storage medium

    CN117115841A

  • Image processing method and device, electronic equipment and storage medium

    CN117456545A

  • Lightweight document layout analysis method

    CN118172787A

  • Picture type archive Chinese table identification method based on deep learning

    CN119049074A