Lightweight table structure identification method
By building a lightweight deep learning model, combining downsampling module, feature extraction module and feature fusion module, the problem of limited table recognition performance in the existing technology is solved, and accurate identification and efficient inference of complex table structures are achieved.
Patent Information
- Application Number
- CN202510409433.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2045-04-02
AI Technical Summary
When faced with complex and diverse table structures and styles, the performance of existing table recognition methods is limited by the complexity and generalization ability of manual feature design, making it difficult to accurately identify table structures in various scenarios.
Using a lightweight deep learning model structure, a table structure recognition model is constructed through the combination of five downsampling modules, one activation function ReLU6, sixteen tabular feature extraction modules, nine tabular feature fusion modules and six prediction heads to achieve accurate identification of table structure.
On the basis of ensuring the recognition accuracy, it effectively reduces the consumption of computing resources, improves the inference speed of the model, and realizes rapid detection and accurate identification of table structures in various scenarios.
Smart Images

Figure CN119919952A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the fields of computer vision and table structure recognition, and in particular to a lightweight table structure recognition method. Background Art
[0002] In today's era, tables are important tools for recording and summarizing key data, facilitating the rapid acquisition and clear presentation of information. With the large-scale application of smartphones, it has become a common phenomenon to use camera devices to obtain table images for information transmission. Therefore, the ability to automatically extract and parse table structures from photos or images of different scenes has become an urgent need. Although traditional table recognition methods, such as rule-based methods and machine learning methods, have achieved certain results in specific scenarios, their performance is often limited by the complexity and generalization ability of manual feature design when faced with complex and diverse table structures and styles. The table structure recognition method based on deep learning technology can automatically learn and extract high-level features in table images, so that the model shows strong generalization ability on diverse table images.
[0003] Therefore, it is particularly important to design a new method that can accurately identify table structures while taking into account adaptability and robustness. This technology came into being in this context, and is committed to achieving accurate recognition of table structures through innovative deep learning algorithms and strategies, while effectively reducing computing resource consumption and greatly improving reasoning speed. Summary of the invention
[0004] The present invention aims to provide a lightweight table structure recognition method that can quickly detect photographed and scanned table images in various scenarios and provide accurate table structure recognition results. The method adopts a lightweight model structure to effectively reduce computing resource consumption while ensuring high recognition accuracy.
[0005] The technical solution proposed by the present invention is as follows: A lightweight table structure recognition method mainly includes the following steps: Step S1: Obtaining a table structure recognition data set: collecting public table images and corresponding annotation data, and extracting the physical coordinates and logical coordinates of cells representing the table structure in the annotation data to obtain a first table structure recognition data set T1; Step S2: Preprocess the table structure recognition data set: perform data enhancement and standardization on the first table structure recognition data set T1 to obtain a second table structure recognition data set T2; divide the second table structure recognition data set T2 into a training set data t , validation set data v And test set data e ; Step S3: construct a table structure recognition model: design the basic architecture of the table structure recognition model, and connect five downsampling modules, one activation function ReLU6, sixteen table feature extraction modules, nine table feature fusion modules and six prediction heads in sequence according to the basic architecture of the table structure recognition model to obtain an uninitialized table structure recognition model Mr; initialize the parameters of the uninitialized table structure recognition model Mr to obtain a first table structure recognition model M; Step S4: training the table structure recognition model: using the second table structure recognition data set T2 to train the first table structure recognition model M to obtain a second table structure recognition model Mt; Step S5: Identify the table structure: from the test set data e A table image is randomly selected and input into the second table structure recognition model Mt for inference to obtain six feature maps representing the table structure; the six feature maps representing the table structure are decoded to obtain a table structure recognition result including the physical coordinates and logical coordinates of the cells.
[0006] Furthermore, in step S1, publicly available table images and corresponding annotation data are collected, and the physical coordinates and logical coordinates of cells representing the table structure in the annotation data are extracted to obtain a first table structure recognition data set T1. The specific steps are as follows: Step S11, collecting photographed and scanned table images containing table structure annotation data in various scenarios through the Internet to obtain a first table image; Step S12, parsing the annotation data corresponding to the first table image to obtain physical coordinates and logical coordinates of cells representing the table structure; correcting and integrating the physical coordinates and logical coordinates of cells representing the table structure to obtain a first label; Step S13: Combine the first table image and the first label to obtain a first data sample, and add the first data sample to the first table structure recognition data set T1.
[0007] Furthermore, in step S2, data enhancement and standardization are performed on the first table structure recognition data set T1 to obtain a second table structure recognition data set T2. The specific steps are as follows: Step S21, set the input image height h, set the input image width w, let I f represents the first table image in the fth data sample of the first table structure recognition data set T1, let g f Indicates the first label in the fth data sample of the first table structure identification data set T1, f=1,2,…,n, n is the number of data samples in the first table structure identification data set T1; sets an iteration variable k, and the initial value of the iteration variable k is set to 1; Step S22: identifying the first table image I in the kth data sample of the first table structure data set T1 k Perform color enhancement to obtain a color enhanced table image Ic; divide each pixel value of the color enhanced table image Ic by 255, perform data normalization, and obtain a normalized image Iv; Step S23, adjust the height of the normalized image Iv to the input image height h, and adjust the width to the input image width w, to obtain the second table image Ip and the affine transformation matrix AM; use the affine transformation matrix AM to identify the first label g in the kth data sample of the first table structure data set T1 k The physical coordinates of the cells in are transformed to obtain the second label g p ; Step S24: the second table image Ip and the second label g p Combine to obtain a second data sample, and add the second data sample to the second table structure recognition data set T2; Step S25, increase the value of the iteration variable k by 1, and if k≤n, jump to step S22.
[0008] Furthermore, in step S3, the basic architecture of the table structure recognition model is designed, and five downsampling modules, one activation function ReLU6, sixteen table feature extraction modules, nine table feature fusion modules and six prediction heads are connected in sequence according to the basic architecture of the table structure recognition model to obtain an uninitialized table structure recognition model Mr. The specific steps are as follows: Step S311, the first down-sampling module receives the preprocessed second table image Ip, the output of the first down-sampling module is connected to the input of the first activation function ReLU6; the output of the first activation function ReLU6 is connected to the input of the second down-sampling module, the output of the second down-sampling module is connected to the input of the first table feature extraction module, the output of the first table feature extraction module is connected to the input of the second table feature extraction module; the output of the second table feature extraction module is connected to the input of the third down-sampling module, the output of the third down-sampling module is connected to the input of the third table feature extraction module, the output of the third table feature extraction module is connected to the input of the fourth table feature extraction module; the output of the fourth table feature extraction module is connected to the input of the fourth down-sampling module, the output of the fourth down-sampling module is connected to the input of the fifth table feature extraction module, the output of the fifth table feature extraction module is connected to the input of the sixth table feature extraction module, the output of the sixth table feature extraction module The output of the seventh table feature extraction module is connected to the input of the eighth table feature extraction module, the output of the eighth table feature extraction module is connected to the input of the ninth table feature extraction module, the output of the ninth table feature extraction module is connected to the input of the tenth table feature extraction module, the output of the tenth table feature extraction module is connected to the input of the eleventh table feature extraction module, and the output of the eleventh table feature extraction module is connected to the input of the twelfth table feature extraction module; the output of the twelfth table feature extraction module is connected to the input of the fifth down-sampling module, the output of the fifth down-sampling module is connected to the input of the thirteenth table feature extraction module, the output of the thirteenth table feature extraction module is connected to the input of the fourteenth table feature extraction module, the output of the fourteenth table feature extraction module is connected to the input of the fifteenth table feature extraction module, and the output of the fifteenth table feature extraction module is connected to the input of the sixteenth table feature extraction module; Step S312, the output of the twelfth table feature extraction module is connected to the first input of the first table feature fusion module and the output of the sixteenth table feature extraction module is connected to the second input of the first table feature fusion module; the output of the fourth table feature extraction module is connected to the first input of the second table feature fusion module and the output of the twelfth table feature extraction module is connected to the second input of the second table feature fusion module, the output of the second table feature fusion module is connected to the first input of the third table feature fusion module and the output of the first table feature fusion module is connected to the second input of the third table feature fusion module; the output of the second table feature extraction module is connected to the first input of the fourth table feature fusion module and the output of the fourth table feature extraction module is connected to the second input of the fourth table feature fusion module, and the output of the fourth table feature fusion module is connected to the fifth table feature fusion module The first input and the output of the second table feature fusion module are connected to the second input of the fifth table feature fusion module, the output of the fifth table feature fusion module is connected to the first input of the sixth table feature fusion module and the output of the third table feature fusion module is connected to the second input of the sixth table feature fusion module; the output of the third table feature fusion module is connected to the first input of the seventh table feature fusion module and the output of the first table feature fusion module is connected to the second input of the seventh table feature fusion module; the output of the sixth table feature fusion module is connected to the first input of the eighth table feature fusion module and the output of the third table feature fusion module is connected to the second input of the eighth table feature fusion module, the output of the eighth table feature fusion module is connected to the first input of the ninth table feature fusion module and the output of the seventh table feature fusion module is connected to the second input of the ninth table feature fusion module; Step S313, the output of the ninth table feature fusion module is simultaneously connected to the first prediction head, the second prediction head, the third prediction head, the fourth prediction head, the fifth prediction head and the sixth prediction head to obtain an uninitialized table structure recognition model Mr.
[0009] Furthermore, in step S3, any downsampling module among the five downsampling modules includes a convolution layer and a batch normalization layer; the connection order of the downsampling modules is as follows: the input of the downsampling module is connected to the input of the first convolution layer, the output of the first convolution layer is connected to the input of the first batch normalization layer, and the output of the first batch normalization layer is the output of the downsampling module.
[0010] Further, in step S3, any of the sixteen table feature extraction modules includes three convolutional layers, two depth-separable convolutional layers, an activation function ReLU6 and two batch normalization layers; specifically: The input of the table feature extraction module is connected to the input of the first depth-separable convolutional layer, the output of the first depth-separable convolutional layer is connected to the input of the second batch normalization layer, the output of the second batch normalization layer is connected to the input of the second convolutional layer and the input of the third convolutional layer at the same time, the output of the second convolutional layer is connected to the input of the second activation function ReLU6, the output of the second activation function ReLU6 and the output of the third convolutional layer are element-wise multiplied and connected to the input of the fourth convolutional layer, the output of the fourth convolutional layer is connected to the input of the third batch normalization layer, the output of the third batch normalization layer is connected to the input of the second depth-separable convolutional layer, and the output of the second depth-separable convolutional layer and the input of the table feature extraction module are added together to form the output of the table feature extraction module.
[0011] Further, in step S3, any of the nine table feature fusion modules includes two deformable convolution layers, one transposed convolution upsampling layer, two batch normalization layers and two activation functions ReLU; specifically: The second input of the table feature fusion module is connected to the input of the first deformable convolution layer, the output of the first deformable convolution layer is connected to the input of the fourth batch normalization layer, the output of the fourth batch normalization layer is connected to the input of the first activation function ReLU, the output of the first activation function ReLU is connected to the input of the first transposed convolution upsampling layer, the output of the first transposed convolution upsampling layer and the first input of the table feature fusion module are added to the output of the second deformable convolution layer, the output of the second deformable convolution layer is connected to the input of the fifth batch normalization layer, the output of the fifth batch normalization layer is connected to the input of the second activation function ReLU, and the output of the second activation function ReLU is the output of the table feature fusion module.
[0012] Furthermore, in step S3, the parameters of the uninitialized table structure recognition model Mr are initialized to obtain the first table structure recognition model M. The specific steps are as follows: Step S32, initialize the weights of the transposed convolution upsampling layer of any of the nine table feature fusion modules to bilinear interpolation weights, initialize the bias of any of the six prediction heads to 0, and randomly initialize the remaining uninitialized model parameters to obtain the first table structure recognition model M.
[0013] Furthermore, in step S4, the first table structure recognition model M is trained using the second table structure recognition data set T2 to obtain a second table structure recognition model Mt. The specific steps are as follows: Step S41, set the number of model training rounds epochs, set the training data batch size batch, set the optimizer learning rate lr, set the learning rate prime decay step set steps; set the optimal verification loss value loss o, set the optimal table structure recognition model Mo, the optimal verification loss value loss o The initial value of is set to positive infinity, and the parameters of the optimal table structure recognition model Mo are initialized to the parameters of the first table structure recognition model M; let epoch represent the current training round, and initialize the current training round epoch to 1; Step S42, let count represent the current iteration batch, and initialize the current iteration batch count to , where m is the training set data t The number of data samples in ; Step S43, from the training set data t Randomly extract a second data sample of batch size, input the second table image in the second data sample into the first table structure recognition model M for forward propagation, and obtain the prediction result of the current iteration batch model; use the second label in the second data sample to calculate the loss of the prediction result of the current iteration batch model, and obtain the loss of the current iteration batch; Step S44, based on the loss of the current iteration batch, an adaptive momentum estimation optimization algorithm is used to update the parameters of the first table structure recognition model M through a back propagation process; Step S45, if epoch∈steps, divide the value of the learning rate lr by 10; Step S46, subtract 1 from the value of the current iteration batch count, if count>1, jump to step S43; Step S47: Verification set data v The second data sample in is read in batches and forward propagation is performed sequentially to calculate the validation set data v The average loss loss in the current training round epoch a ; Step S48, if loss a <loss o , then the optimal verification loss value loss o The value is set to loss a , and update the parameters of the optimal table structure recognition model Mo to the parameters of the first table structure recognition model M under the current training round epoch; Step S49, increase the value of the current training round epoch by 1; if epoch<epochs, jump to step S42; otherwise, save the optimal table structure recognition model Mo as the second table structure recognition model Mt.
[0014] The beneficial effects of the present invention are as follows: the method adopts an efficient table feature extraction module to realize table feature capture, optimizes the multi-scale feature fusion process through a table feature fusion module, and accurately identifies the physical coordinates and logical coordinates of cells in the table structure; on the basis of ensuring detection accuracy, the lightweight construction of the table structure recognition model is completed, thereby improving the model reasoning speed. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 A method flow chart of an example of the present invention; Figure 2 It is a structural diagram of the table structure recognition model of the present invention; Figure 3 This is a structural diagram of the downsampling module of the present invention; Figure 4 This is a structural diagram of the table feature extraction module of the present invention; Figure 5 This is a structural diagram of the table feature fusion module of the present invention. DETAILED DESCRIPTION
[0016] The specific implementation of the present invention will be further described in detail below in conjunction with the accompanying drawings in the embodiments of the present invention. It should be noted that the specific embodiment of a lightweight table structure recognition method according to the present invention is only used as an example and is not intended to limit the present invention.
[0017] The present invention provides a lightweight table structure recognition method, such as Figure 1 As shown, the method comprises the following steps: Step S1: Obtaining a table structure recognition data set: collecting public table images and corresponding annotation data, and extracting the physical coordinates and logical coordinates of cells representing the table structure in the annotation data to obtain a first table structure recognition data set T1; Step S2: Preprocess the table structure recognition data set: perform data enhancement and standardization on the first table structure recognition data set T1 to obtain a second table structure recognition data set T2; divide the second table structure recognition data set T2 into training set data according to a ratio of 8:1:1 t , validation set data v And test set data e ; Step S3: construct a table structure recognition model: design the basic architecture of the table structure recognition model, and connect five downsampling modules, one activation function ReLU6, sixteen table feature extraction modules, nine table feature fusion modules and six prediction heads in sequence according to the basic architecture of the table structure recognition model to obtain an uninitialized table structure recognition model Mr; initialize the parameters of the uninitialized table structure recognition model Mr to obtain a first table structure recognition model M; Step S4: training the table structure recognition model: using the second table structure recognition data set T2 to train the first table structure recognition model M to obtain a second table structure recognition model Mt; Step S5: Identify the table structure: from the test set data e A table image is randomly selected and input into the second table structure recognition model Mt for inference to obtain six feature maps representing the table structure; the six feature maps representing the table structure are decoded to obtain a table structure recognition result including the physical coordinates and logical coordinates of the cells.
[0018] Furthermore, in step S1, publicly available table images and corresponding annotation data are collected, and the physical coordinates and logical coordinates of cells representing the table structure in the annotation data are extracted to obtain a first table structure recognition data set T1. The specific steps are as follows: Step S11, collecting photographed and scanned table images containing table structure annotation data in various scenarios through the Internet to obtain a first table image; Step S12, parsing the annotation data corresponding to the first table image to obtain physical coordinates and logical coordinates of cells representing the table structure; correcting and integrating the physical coordinates and logical coordinates of cells representing the table structure to obtain a first label; Step S13: Combine the first table image and the first label to obtain a first data sample, and add the first data sample to the first table structure recognition data set T1.
[0019] Furthermore, in step S2, data enhancement and standardization are performed on the first table structure recognition data set T1 to obtain a second table structure recognition data set T2. The specific steps are as follows: Step S21, set the input image height h=1024, set the input image width w=1024, let I f represents the first table image in the fth data sample of the first table structure recognition data set T1, let g f Indicates the first label in the fth data sample of the first table structure identification data set T1, f=1,2,…,n, n is the number of data samples in the first table structure identification data set T1; sets a first variable k, and the initial value of the first variable k is set to 1; Step S22: identifying the first table image I in the kth data sample of the first table structure data set T1 k Perform color enhancement to obtain a color enhanced table image Ic; divide each pixel value of the color enhanced table image Ic by 255, perform data normalization, and obtain a normalized image Iv; Step S23, adjust the height of the normalized image Iv to the input image height h, and adjust the width to the input image width w, to obtain the second table image Ip and the affine transformation matrix AM; use the affine transformation matrix AM to identify the first label g in the kth data sample of the first table structure data set T1 k The physical coordinates of the cells in are transformed to obtain the second label g p ; Step S24: the second table image Ip and the second label g p Combine to obtain a second data sample, and add the second data sample to the second table structure recognition data set T2; Step S25, increase the value of the iteration variable k by 1, and if k≤n, jump to step S22.
[0020] Furthermore, if Figure 2 As shown, the basic architecture of the table structure recognition model is designed, and five downsampling modules, one activation function ReLU6, sixteen table feature extraction modules, nine table feature fusion modules and six prediction heads are connected in sequence according to the basic architecture of the table structure recognition model to obtain an uninitialized table structure recognition model Mr. The specific steps are as follows: Step S311, the first down-sampling module receives the preprocessed second table image Ip, the output of the first down-sampling module is connected to the input of the first activation function ReLU6; the output of the first activation function ReLU6 is connected to the input of the second down-sampling module, the output of the second down-sampling module is connected to the input of the first table feature extraction module, the output of the first table feature extraction module is connected to the input of the second table feature extraction module; the output of the second table feature extraction module is connected to the input of the third down-sampling module, the output of the third down-sampling module is connected to the input of the third table feature extraction module, the output of the third table feature extraction module is connected to the input of the fourth table feature extraction module; the output of the fourth table feature extraction module is connected to the input of the fourth down-sampling module, the output of the fourth down-sampling module is connected to the input of the fifth table feature extraction module, the output of the fifth table feature extraction module is connected to the input of the sixth table feature extraction module, the output of the sixth table feature extraction module The output of the seventh table feature extraction module is connected to the input of the eighth table feature extraction module, the output of the eighth table feature extraction module is connected to the input of the ninth table feature extraction module, the output of the ninth table feature extraction module is connected to the input of the tenth table feature extraction module, the output of the tenth table feature extraction module is connected to the input of the eleventh table feature extraction module, and the output of the eleventh table feature extraction module is connected to the input of the twelfth table feature extraction module; the output of the twelfth table feature extraction module is connected to the input of the fifth down-sampling module, the output of the fifth down-sampling module is connected to the input of the thirteenth table feature extraction module, the output of the thirteenth table feature extraction module is connected to the input of the fourteenth table feature extraction module, the output of the fourteenth table feature extraction module is connected to the input of the fifteenth table feature extraction module, and the output of the fifteenth table feature extraction module is connected to the input of the sixteenth table feature extraction module; Step S312, the output of the twelfth table feature extraction module is connected to the first input of the first table feature fusion module and the output of the sixteenth table feature extraction module is connected to the second input of the first table feature fusion module; the output of the fourth table feature extraction module is connected to the first input of the second table feature fusion module and the output of the twelfth table feature extraction module is connected to the second input of the second table feature fusion module, the output of the second table feature fusion module is connected to the first input of the third table feature fusion module and the output of the first table feature fusion module is connected to the second input of the third table feature fusion module; the output of the second table feature extraction module is connected to the first input of the fourth table feature fusion module and the output of the fourth table feature extraction module is connected to the second input of the fourth table feature fusion module, and the output of the fourth table feature fusion module is connected to the fifth table feature fusion module The first input and the output of the second table feature fusion module are connected to the second input of the fifth table feature fusion module, the output of the fifth table feature fusion module is connected to the first input of the sixth table feature fusion module and the output of the third table feature fusion module is connected to the second input of the sixth table feature fusion module; the output of the third table feature fusion module is connected to the first input of the seventh table feature fusion module and the output of the first table feature fusion module is connected to the second input of the seventh table feature fusion module; the output of the sixth table feature fusion module is connected to the first input of the eighth table feature fusion module and the output of the third table feature fusion module is connected to the second input of the eighth table feature fusion module, the output of the eighth table feature fusion module is connected to the first input of the ninth table feature fusion module and the output of the seventh table feature fusion module is connected to the second input of the ninth table feature fusion module; Step S313, the output of the ninth table feature fusion module is simultaneously connected to the first prediction head, the second prediction head, the third prediction head, the fourth prediction head, the fifth prediction head and the sixth prediction head to obtain an uninitialized table structure recognition model Mr; Further, in step S3, if Figure 3 As shown in FIG. 1 , any of the five downsampling modules includes a convolution layer and a batch normalization layer; the connection order of the downsampling modules is as follows: the input of the downsampling module is connected to the input of the first convolution layer, the output of the first convolution layer is connected to the input of the first batch normalization layer, and the output of the first batch normalization layer is the output of the downsampling module; Further, in step S3, if Figure 4 As shown, any of the sixteen table feature extraction modules includes three convolutional layers, two depth-separable convolutional layers, an activation function ReLU6 and two batch normalization layers; specifically: The input of the table feature extraction module is connected to the input of the first depth-separable convolutional layer, the output of the first depth-separable convolutional layer is connected to the input of the second batch normalization layer, the output of the second batch normalization layer is connected to the input of the second convolutional layer and the input of the third convolutional layer at the same time, the output of the second convolutional layer is connected to the input of the second activation function ReLU6, the output of the second activation function ReLU6 and the output of the third convolutional layer are element-wise multiplied and connected to the input of the fourth convolutional layer, the output of the fourth convolutional layer is connected to the input of the third batch normalization layer, the output of the third batch normalization layer is connected to the input of the second depth-separable convolutional layer, and the output of the second depth-separable convolutional layer and the input of the table feature extraction module are added to form the output of the table feature extraction module; Further, in step S3, if Figure 5 As shown, any of the nine table feature fusion modules includes two deformable convolution layers, one transposed convolution upsampling layer, two batch normalization layers and two activation functions ReLU; specifically: The second input of the table feature fusion module is connected to the input of the first deformable convolution layer, the output of the first deformable convolution layer is connected to the input of the fourth batch normalization layer, the output of the fourth batch normalization layer is connected to the input of the first activation function ReLU, the output of the first activation function ReLU is connected to the input of the first transposed convolution upsampling layer, the output of the first transposed convolution upsampling layer and the first input of the table feature fusion module are added to the output of the second deformable convolution layer, the output of the second deformable convolution layer is connected to the input of the fifth batch normalization layer, the output of the fifth batch normalization layer is connected to the input of the second activation function ReLU, and the output of the second activation function ReLU is the output of the table feature fusion module; Furthermore, in step S3, the parameters of the uninitialized table structure recognition model Mr are initialized to obtain the first table structure recognition model M. The specific steps are as follows: Step S32, initializing the weight of the transposed convolution upsampling layer of any table feature fusion module among the nine table feature fusion modules to the bilinear interpolation weight, initializing the bias of any prediction head among the six prediction heads to 0, and randomly initializing the remaining uninitialized model parameters to obtain a first table structure recognition model M; Furthermore, in step S4, the first table structure recognition model M is trained using the second table structure recognition data set T2 to obtain a second table structure recognition model Mt. The specific steps are as follows: Step S41, set the number of model training epochs = 200, set the training data batch size batch = 22, and set the optimizer learning rate lr = 1.25×10 -4, set the learning rate attenuation step set steps={140, 180}; set the optimal verification loss value loss o , set the optimal table structure recognition model Mo, the optimal verification loss value loss o The initial value of is set to positive infinity, and the parameters of the optimal table structure recognition model Mo are initialized to the parameters of the first table structure recognition model M; let epoch represent the current training round, and initialize the current training round epoch to 1; Step S42, let count represent the current iteration batch, and initialize the current iteration batch count to , where m is the training set data t The number of data samples in ; Step S43, from the training set data t Randomly extract a second data sample of batch size, input the second table image in the second data sample into the first table structure recognition model M for forward propagation, and obtain the prediction result of the current iteration batch model; use the second label in the second data sample to calculate the loss of the prediction result of the current iteration batch model, and obtain the loss of the current iteration batch; Step S44, based on the loss of the current iteration batch, an adaptive momentum estimation optimization algorithm is used to update the parameters of the first table structure recognition model M through a back propagation process; Step S45, if epoch∈steps, divide the value of the learning rate lr by 10; Step S46, subtract 1 from the value of the current iteration batch count, if count>1, jump to step S43; Step S47: Verification set data v The second data sample in is read in batches and forward propagation is performed sequentially to calculate the validation set data v The average loss loss in the current training round epoch a ; Step S48, if loss a <loss o , then the optimal verification loss value loss o The value is set to loss a , and update the parameters of the optimal table structure recognition model Mo to the parameters of the first table structure recognition model M under the current training round epoch; Step S49, increase the value of the current training round epoch by 1; if epoch<epochs, jump to step S42; otherwise, save the optimal table structure recognition model Mo as the second table structure recognition model Mt.
[0021] The beneficial effects of the present invention are as follows: the method adopts an efficient table feature extraction module to realize table feature capture, optimizes the multi-scale feature fusion process through a table feature fusion module, and accurately identifies the physical coordinates and logical coordinates of cells in the table structure; on the basis of ensuring detection accuracy, the lightweight construction of the table structure recognition model is completed, thereby improving the model reasoning speed.
[0022] The above description is a preferred embodiment of the present invention and is not intended to limit the present invention. Any equivalent or modification that does not deviate from the spirit disclosed in the present invention should be included in the scope of protection of the present invention.
Claims
1. A lightweight table structure recognition method, characterized by: The main steps include: Step S1: Obtaining a table structure recognition data set: collecting public table images and corresponding annotation data, and extracting the physical coordinates and logical coordinates of cells representing the table structure in the annotation data to obtain a first table structure recognition data set T1; Step S2: Preprocess the table structure recognition data set: perform data enhancement and standardization on the first table structure recognition data set T1 to obtain a second table structure recognition data set T2; divide the second table structure recognition data set T2 into a training set data t , validation set data v And test set data e ; Step S3: construct a table structure recognition model: design the basic architecture of the table structure recognition model, and connect five downsampling modules, one activation function ReLU6, sixteen table feature extraction modules, nine table feature fusion modules and six prediction heads in sequence according to the basic architecture of the table structure recognition model to obtain an uninitialized table structure recognition model Mr; initialize the parameters of the uninitialized table structure recognition model Mr to obtain a first table structure recognition model M; Step S4: training the table structure recognition model: using the second table structure recognition data set T2 to train the first table structure recognition model M to obtain a second table structure recognition model Mt; Step S5; Identify table structure: from test set data e A table image is randomly selected and input into the second table structure recognition model Mt for inference to obtain six feature maps representing the table structure; the six feature maps representing the table structure are decoded to obtain a table structure recognition result including the physical coordinates and logical coordinates of the cells.
2. A lightweight table structure recognition method according to claim 1, characterized in that: In step S1, publicly available table images and corresponding annotation data are collected, and the physical coordinates and logical coordinates of cells representing the table structure in the annotation data are extracted to obtain a first table structure recognition data set T1. The specific steps are as follows: Step S11, collecting photographed and scanned table images containing table structure annotation data in various scenarios to obtain a first table image; Step S12, parsing the annotation data corresponding to the first table image to obtain physical coordinates and logical coordinates of cells representing the table structure; correcting and integrating the physical coordinates and logical coordinates of cells representing the table structure to obtain a first label; Step S13: Combine the first table image and the first label to obtain a first data sample, and add the first data sample to the first table structure recognition data set T1.
3. A lightweight table structure recognition method according to claim 2, characterized in that: In step S2, data enhancement and standardization are performed on the first table structure recognition data set T1 to obtain a second table structure recognition data set T2. The specific steps are as follows: Step S21, set the input image height h, set the input image width w, let I f represents the first table image in the fth data sample of the first table structure recognition data set T1, let g f represents the first label in the f-th data sample of the first table structure identification data set T1, f=1,2,…,n, and n is the number of data samples in the first table structure identification data set T1; Set the iteration variable k, and the initial value of the iteration variable k is set to 1; Step S22: identifying the first table image I in the kth data sample of the first table structure data set T1 k Perform color enhancement to obtain a color enhanced table image Ic; Divide each pixel value of the color enhanced table image Ic by 255 to perform data normalization to obtain a normalized image Iv; Step S23, adjust the height of the normalized image Iv to the input image height h, and adjust the width to the input image width w, to obtain the second table image Ip and the affine transformation matrix AM; use the affine transformation matrix AM to identify the first label g in the kth data sample of the first table structure data set T1 k The physical coordinates of the cells in are transformed to obtain the second label g p ; Step S24: the second table image Ip and the second label g p Combine to obtain a second data sample, and add the second data sample to the second table structure recognition data set T2; Step S25, increase the value of the iteration variable k by 1, if k≤n, jump to step S22.
4. A lightweight table structure recognition method according to claim 3, characterized in that: In step S3, the basic architecture of the table structure recognition model is designed, and five downsampling modules, one activation function ReLU6, sixteen table feature extraction modules, nine table feature fusion modules and six prediction heads are connected in sequence according to the basic architecture of the table structure recognition model to obtain an uninitialized table structure recognition model Mr. The specific steps are as follows: Step S311, the first down-sampling module receives the preprocessed second table image Ip, the output of the first down-sampling module is connected to the input of the first activation function ReLU6; the output of the first activation function ReLU6 is connected to the input of the second down-sampling module, the output of the second down-sampling module is connected to the input of the first table feature extraction module, the output of the first table feature extraction module is connected to the input of the second table feature extraction module; the output of the second table feature extraction module is connected to the input of the third down-sampling module, the output of the third down-sampling module is connected to the input of the third table feature extraction module, the output of the third table feature extraction module is connected to the input of the fourth table feature extraction module; the output of the fourth table feature extraction module is connected to the input of the fourth down-sampling module, the output of the fourth down-sampling module is connected to the input of the fifth table feature extraction module, the output of the fifth table feature extraction module is connected to the input of the sixth table feature extraction module, the output of the sixth table feature extraction module The output of the seventh table feature extraction module is connected to the input of the eighth table feature extraction module, the output of the eighth table feature extraction module is connected to the input of the ninth table feature extraction module, the output of the ninth table feature extraction module is connected to the input of the tenth table feature extraction module, the output of the tenth table feature extraction module is connected to the input of the eleventh table feature extraction module, and the output of the eleventh table feature extraction module is connected to the input of the twelfth table feature extraction module; the output of the twelfth table feature extraction module is connected to the input of the fifth down-sampling module, the output of the fifth down-sampling module is connected to the input of the thirteenth table feature extraction module, the output of the thirteenth table feature extraction module is connected to the input of the fourteenth table feature extraction module, the output of the fourteenth table feature extraction module is connected to the input of the fifteenth table feature extraction module, and the output of the fifteenth table feature extraction module is connected to the input of the sixteenth table feature extraction module; Step S312, the output of the twelfth table feature extraction module is connected to the first input of the first table feature fusion module and the output of the sixteenth table feature extraction module is connected to the second input of the first table feature fusion module; the output of the fourth table feature extraction module is connected to the first input of the second table feature fusion module and the output of the twelfth table feature extraction module is connected to the second input of the second table feature fusion module, the output of the second table feature fusion module is connected to the first input of the third table feature fusion module and the output of the first table feature fusion module is connected to the second input of the third table feature fusion module; the output of the second table feature extraction module is connected to the first input of the fourth table feature fusion module and the output of the fourth table feature extraction module is connected to the second input of the fourth table feature fusion module, and the output of the fourth table feature fusion module is connected to the fifth table feature fusion module The first input and the output of the second table feature fusion module are connected to the second input of the fifth table feature fusion module, the output of the fifth table feature fusion module is connected to the first input of the sixth table feature fusion module and the output of the third table feature fusion module is connected to the second input of the sixth table feature fusion module; the output of the third table feature fusion module is connected to the first input of the seventh table feature fusion module and the output of the first table feature fusion module is connected to the second input of the seventh table feature fusion module; the output of the sixth table feature fusion module is connected to the first input of the eighth table feature fusion module and the output of the third table feature fusion module is connected to the second input of the eighth table feature fusion module, the output of the eighth table feature fusion module is connected to the first input of the ninth table feature fusion module and the output of the seventh table feature fusion module is connected to the second input of the ninth table feature fusion module; Step S313, the output of the ninth table feature fusion module is simultaneously connected to the first prediction head, the second prediction head, the third prediction head, the fourth prediction head, the fifth prediction head and the sixth prediction head to obtain an uninitialized table structure recognition model Mr.
5. A lightweight table structure recognition method according to claim 4, characterized in that: In step S3, any downsampling module among the five downsampling modules includes a convolution layer and a batch normalization layer; the connection order of the downsampling modules is as follows: the input of the downsampling module is connected to the input of the first convolution layer, the output of the first convolution layer is connected to the input of the first batch normalization layer, and the output of the first batch normalization layer is the output of the downsampling module.
6. A lightweight table structure recognition method according to claim 5, characterized in that: In step S3, any of the sixteen table feature extraction modules includes three convolutional layers, two depth-separable convolutional layers, an activation function ReLU6 and two batch normalization layers; specifically: The input of the table feature extraction module is connected to the input of the first depth-separable convolutional layer, the output of the first depth-separable convolutional layer is connected to the input of the second batch normalization layer, the output of the second batch normalization layer is connected to the input of the second convolutional layer and the input of the third convolutional layer at the same time, the output of the second convolutional layer is connected to the input of the second activation function ReLU6, the output of the second activation function ReLU6 and the output of the third convolutional layer are element-wise multiplied and connected to the input of the fourth convolutional layer, the output of the fourth convolutional layer is connected to the input of the third batch normalization layer, the output of the third batch normalization layer is connected to the input of the second depth-separable convolutional layer, and the output of the second depth-separable convolutional layer and the input of the table feature extraction module are added together to form the output of the table feature extraction module.
7. A lightweight table structure recognition method according to claim 6, characterized in that: In step S3, any of the nine table feature fusion modules includes two deformable convolution layers, one transposed convolution upsampling layer, two batch normalization layers and two activation functions ReLU; specifically: The second input of the table feature fusion module is connected to the input of the first deformable convolution layer, the output of the first deformable convolution layer is connected to the input of the fourth batch normalization layer, the output of the fourth batch normalization layer is connected to the input of the first activation function ReLU, the output of the first activation function ReLU is connected to the input of the first transposed convolution upsampling layer, the output of the first transposed convolution upsampling layer and the first input of the table feature fusion module are added to the output of the second deformable convolution layer, the output of the second deformable convolution layer is connected to the input of the fifth batch normalization layer, the output of the fifth batch normalization layer is connected to the input of the second activation function ReLU, and the output of the second activation function ReLU is the output of the table feature fusion module.
8. A lightweight table structure recognition method according to claim 7, characterized in that: In step S3, the parameters of the uninitialized table structure recognition model Mr are initialized to obtain the first table structure recognition model M. The specific steps are as follows: Step S32, initialize the weights of the transposed convolution upsampling layer of any of the nine table feature fusion modules to bilinear interpolation weights, initialize the bias of any of the six prediction heads to 0, and randomly initialize the remaining uninitialized model parameters to obtain the first table structure recognition model M.
9. A lightweight table structure recognition method according to claim 8, characterized in that: In step S4, the first table structure recognition model M is trained using the second table structure recognition data set T2 to obtain a second table structure recognition model Mt. The specific steps are as follows: Step S41, set the number of model training rounds epochs, set the training data batch size batch, set the optimizer learning rate lr, set the learning rate prime decay step set steps; set the optimal verification loss value loss o , set the optimal table structure recognition model Mo, the optimal verification loss value loss o The initial value of is set to positive infinity, and the parameters of the optimal table structure recognition model Mo are initialized to the parameters of the first table structure recognition model M; let epoch represent the current training round, and initialize the current training round epoch to 1; Step S42, let count represent the current iteration batch, and initialize the current iteration batch count to , where m is the training set data t The number of data samples in ; Step S43, from the training set data t Randomly extract a batch size of second data samples, input the second table image in the second data sample into the first table structure recognition model M for forward propagation, and obtain the prediction result of the current iteration batch model; use the second label in the second data sample to calculate the loss of the prediction result of the current iteration batch model, and obtain the loss loss of the current iteration batch; Step S44, based on the loss of the current iteration batch, an adaptive momentum estimation optimization algorithm is used to update the parameters of the first table structure recognition model M through a back propagation process; Step S45, if epoch∈steps, divide the value of the learning rate lr by 10; Step S46, subtract 1 from the value of the current iteration batch count, if count>1, jump to step S43; Step S47: Verification set data v The second data sample in is read in batches and forward propagation is performed sequentially to calculate the validation set data v The average loss loss in the current training round epoch a ; Step S48, if loss a <loss o , then the optimal verification loss value loss o The value is set to loss a , and update the parameters of the optimal table structure recognition model Mo to the parameters of the first table structure recognition model M under the current training round epoch; Step S49, increase the value of the current training round epoch by 1; if epoch<epochs, jump to step S42; otherwise, save the optimal table structure recognition model Mo as the second table structure recognition model Mt.
Citation Information
Patent Citations
Table structure analysis method and device
CN116092105A
Table analysis method and device, computer equipment and storage medium
CN117115841A
Image processing method and device, electronic equipment and storage medium
CN117456545A
Lightweight document layout analysis method
CN118172787A
Picture type archive Chinese table identification method based on deep learning
CN119049074A