Table Automatic Reconstruction Method, Storage Medium and Electronic Device for Coal Quality Inspection Report Images

Through the multi-model collaboration method, the table reconstruction of coal quality detection report images is realized, solving the problem of inefficient data aggregation and statistics in the existing technology, and improving the accuracy and safety of the data.

CN115906786BActive Publication Date: 2025-05-30JIANGSU FRONTIER ELECTRIC TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211352582.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-31
Publication Date
2025-05-30
Estimated Expiration
2042-10-31

AI Technical Summary

Technical Problem

The prior art is difficult to realize the automatic reconstruction of coal quality inspection report images, resulting in inefficient data summary and statistical work, and prone to statistical errors and data fraud.

Method used

Using multi-model collaboration method, the table automatic reconstruction of coal quality detection report images is realized through text detection model EAST, text recognition model CRNN, tabular structure prediction model RARE based on adaptive attention mechanism, and the table is automatically reconstructed in the form of an Excel table.

Benefits of technology

It improves the accuracy and efficiency of table reconstruction, reduces the error rate of manual data processing, reduces the phenomenon of data manipulation in the intermediate links of data processing, and enhances the reliability and security of data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115906786B_ABST
    Figure CN115906786B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for automatically reconstructing a table in a coal quality test report image, a storage medium, and an electronic device, including collection; performing table recognition, that is, through the processing of a text detection model EAST, a text recognition model CRNN, a table structure prediction model RARE based on an adaptive attention mechanism, and a cell aggregation module; performing data post-processing; storing the processed data in the database in the form of a string to form a coal quality test information database; repeatedly storing the obtained coal quality test report image with a table displayed in the database. The present invention realizes the end-to-end reconstruction of the table in the coal quality test report image by adding model pruning, adjustment of the feature extraction network, an adaptive attention mechanism, cell coordinate aggregation, and text aggregation; on the basis of the table recognition technology, it realizes the automatic reconstruction of the table in the coal quality test report image, and at the same time improves the accuracy and efficiency of table reconstruction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for table reconstruction, and particularly to a method for automatic table reconstruction of coal quality inspection report images, a storage medium, and an electronic device. Background Art

[0002] In the past, coal quality inspection reports were mainly presented in paper form, and data summarization and statistics of coal quality inspection reports mainly relied on manual work. On the one hand, the workload was large, and statistical errors were likely to occur. On the other hand, the data was easily tampered with, providing a large space for carbon emission data fraud. To reduce the manual input in the data summarization and statistics of coal quality inspection reports, reduce the error rate of manual data processing, improve work efficiency, and prevent the phenomenon of manipulating data in the intermediate link of data processing, it is of great practical value and significance to study a method for automating the data statistics of coal quality inspection reports.

[0003] With the rapid development of deep learning technology, using deep learning technology to automate and intelligentize the data statistics of coal quality inspection reports is a way. The key to automating the data statistics of coal quality inspection reports lies in the reconstruction of the image table, that is, effectively extracting the data and structure information in the table from the coal quality inspection report image, which is also called table recognition technology. Traditional methods usually use image preprocessing methods, such as edge detection, binarization, etc., to segment the table lines, and then combine post-processing to achieve table reconstruction. This method requires adjustment of thresholds for different structures and types of tables, is difficult to adapt to general scenarios, and has low accuracy and low efficiency.

[0004] Therefore, there is an urgent need for a method for reconstructing the table of coal quality inspection report images based on deep learning to achieve automatic table reconstruction of coal quality inspection report images, while improving the accuracy and efficiency of table reconstruction. Summary of the Invention

[0005] In view of the deficiencies in the prior art, the present invention provides a method for automatic table reconstruction of coal quality inspection report images, a storage medium, and an electronic device to solve the problem of automatic table reconstruction of coal quality inspection report images.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] A method for automatic table reconstruction of coal quality inspection report images, characterized in that:

[0008] Step 1, collect coal quality inspection report images showing tables;

[0009] Step 2: Perform table recognition on the coal quality inspection report image with a table shown obtained in the previous step. After processing by the text detection model EAST, the text recognition model CRNN, the table structure prediction model RARE based on the adaptive attention mechanism, and the cell aggregation module, obtain the reconstructed table result and output it in the form of an Excel table;

[0010] Step 3: Perform data post-processing on the reconstructed Excel table information obtained in the previous step to form json-format data describing the key indicators and data in the table of the coal quality inspection report image;

[0011] Step 4: Store the json-format data obtained in the previous step as a string in the database to form a coal quality inspection information database;

[0012] Step 5: Repeat the operations in Steps 2 to 4, and continue to store the obtained coal quality inspection report images with a table shown in the database.

[0013] To optimize the above technical solution, the specific measures taken also include:

[0014] Further, the cell aggregation module includes a cell coordinate aggregation module, a cell text aggregation module, and a cell aggregation post-processing module.

[0015] Further, Step 2 includes the following steps:

[0016] Step 21: Input a coal quality inspection report image M with a table shown, having a width of W and a height of H, into the text detection model EAST, and output a set of single-line text coordinate boxes Y1. The text coordinate boxes in Y1 are four-point coordinates, in the format (x1,y1,x2,y2,x3,y3,x4,y4), where x1,y1,x2,y2,x3,y3,x4,y4 represent the coordinates of the text coordinate box;

[0017] Step 22: Input the set of single-line text coordinate boxes Y1 into the text recognition model CRNN, and output a set of single-line text coordinate boxes Yz with text content. The format of Yz is (x1,y1,x2,y2,x3,y3,x4,y4,word), where word represents the corresponding text content in the text coordinate box;

[0018] Step 23: Input the same coal quality inspection report image M with a table shown, having a width of W and a height of H, into the table structure prediction model RARE based on the adaptive attention mechanism, and output a string S describing the table structure information, and a set of cell coordinates Y2. Each cell coordinate has the same format as the text coordinate box, denoted as (x1’,y1’,x2’,y2’,x3’,y3’,x4’,y4’);

[0019] Step 24: Input the set Y1 of single-line text coordinate boxes and the set Y2 of cell coordinates into the cell coordinate aggregation module, calculate the intersection over union between the single-line text coordinate boxes and the cell coordinates, as well as the vertex distance between the single-line text coordinate boxes and the cell coordinates, and output the relationship Y3 between the single-line text coordinate boxes and the cell coordinates.

[0020] Step 25: Input the relationship Y3 between the single-line text coordinate boxes and the cell coordinates and the set Yz of single-line text coordinate boxes with text content into the cell text aggregation module, and output the complete information Yk of a single cell, where k is a positive number representing the number of single cells.

[0021] Step 26: Input the complete information Yk of a single cell and the string S describing the table structure information into the post-processing module after cell aggregation to obtain the reconstructed table result and output it in the form of an Excel table.

[0022] Furthermore, the text detection model EAST includes a feature extraction network, a feature fusion module, and a decoding network; the feature extraction network is constructed based on the VGG16 model, and the feature fusion module adopts the Feature Pyramid Network (FPN) structure; the decoding network has three parts of output. The first part has an output channel of 1, which is obtained by connecting a 1×1 convolutional kernel after the feature fusion layer, and the output is a score map representing the probability that each pixel belongs to the text region. The second part is used to predict the text of the rotated rectangle, with a total of 5 output channels, which are generated by two 1×1 convolutional kernels to produce 4 and 1 channels respectively. Among them, the 4 channels respectively represent the 4 distances from the pixel position to the top, right, bottom, and left boundaries of the rectangle, and the 1 channel represents the rotation angle of the bounding box. The third part is used to predict the text of the irregular quadrilateral; the output of the third part of the text detection model EAST is removed by using the model pruning method.

[0023] Furthermore, the text recognition model CRNN includes a feature extraction layer, a sequential modeling layer, and a decoding layer; the feature extraction layer is constructed based on the VGG11 model. According to the max-pooling operation, the feature extraction layer is divided into five feature extraction stages: First, remove the max-pooling layer in the fifth feature extraction stage of this model; second, use a 2×2 convolutional kernel in the fifth feature extraction stage; third, set the convolutional kernel size of the max-pooling layers in the third and fourth feature extraction stages to 1×2; fourth, add a batch normalization layer after the convolutional layers in the third and fourth feature extraction stages; the sequential modeling layer adopts the Bidirectional Long Short-Term Memory (BiLSTM) neural network; the decoding layer adopts the Connectionist Temporal Classification (CTC) layer.

[0024] Furthermore, the table structure prediction model RARE based on the adaptive attention mechanism includes an encoder, an adaptive attention mechanism, a decoder, and a post-processing module; the encoder is divided into two parts, the first half is composed of a convolutional neural network CNN, and the second half is composed of a recurrent neural network RNN; the CNN is constructed based on the VGG11 network structure, the adaptive attention mechanism, and FPN, and the RNN uses BiLSTM; in the CNN structure, the encoder part is divided into five stages according to the max pooling layer. The max pooling layer in the fourth stage uses a 1×2 convolutional kernel, and in the fifth stage, only one convolutional layer with a convolutional kernel size of 3×3, a stride of 1, and a padding value of 0 is set. Moreover, the adaptive attention mechanism is used to calculate the correlation between features obtained from the feature maps output at different encoding stages; the decoder uses GRU, and each unit contains 256 memory blocks and 37 output units; the post-processing module further parses the output of the decoder into an HTML string format for output, and the string contains the four-point coordinates of each cell and the table structure information.

[0025] Furthermore, the calculation process of the intersection over union is as follows:

[0026] S 1 =(x2 - x1)*(y2 - y1)

[0027] S 2 =(x2' - x1')(y2' - y1')

[0028]

[0029] Among them, S 1 and S 2 respectively represent the area of the single-line text coordinate box and the area of the cell coordinate box, and IOU represents the intersection over union of the two;

[0030] The calculation process of the distance between vertices and sorting is as follows:

[0031]

[0032] Among them, d ij represents the distance between the vertex of the i-th single-line text coordinate box and the vertex at the same position of the j-th cell coordinate box, where i and j are integers; sort represents the sorting function.

[0033] Furthermore, for the json format data, the main indicators include moisture, carbon element, hydrogen element, and nitrogen element.

[0034] Furthermore, a computer-readable storage medium stores a computer program, characterized in that: the computer program causes the computer to execute a method for automatically reconstructing a table of a coal quality detection report image as described above.

[0035] Furthermore, an electronic device, comprising: a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the computer program, a method for automatically reconstructing a table in a coal quality detection report image as described above is implemented.

[0036] The beneficial effects of the present invention are as follows:

[0037] 1) The present invention proposes a method for automatically reconstructing a table in a coal quality detection report image. This method uses multi-model collaboration to achieve end-to-end table reconstruction, which has higher efficiency and accuracy compared to traditional table reconstruction methods;

[0038] 2) The table reconstruction method proposed by the present invention uses the text detection model EAST, and prunes the original irregular text prediction output at the decoding layer, improving the speed of model training and inference while ensuring model performance;

[0039] 3) The table reconstruction method proposed by the present invention uses the text recognition model CRNN with an optimized feature extraction network. It uses a rectangular max pooling layer to expand the width and sequence length of the feature map, adapting to the diversity of table text information such as Chinese, English, numbers, characters, etc., and enhancing the robustness of the model in recognizing different types of text;

[0040] 4) The table reconstruction method proposed by the present invention uses the table structure prediction model RARE based on an adaptive attention mechanism. Through the adaptive attention mechanism, the relevance between features, that is, global information, is calculated for the feature maps output at different encoding stages, and spatial information is mined using a deep network, effectively improving the accuracy of table structure prediction;

[0041] 5) The table reconstruction method proposed by the present invention aggregates the outputs of the above-mentioned text detection model EAST, text recognition model CRNN, and table structure prediction model RARE based on an adaptive attention mechanism through a cell aggregation module, and finally exports the reconstructed image table in the form of Excel;

[0042] 6) The method for automatically reconstructing a table proposed by the present invention can reduce the human input in the coal quality detection report data statistics work, reduce the error rate in the process of manual data processing, improve work efficiency, effectively prevent data fraud behavior in the data processing process, and is also the basis for establishing a coal quality detection information database for coal quality detection data analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 It is the original structure diagram of the EAST network of the text detection model for the method for automatically reconstructing a table in a coal quality detection report image proposed by the present invention;

[0044] Figure 2 The structural diagram after pruning the EAST network, which is the text detection model of a method for automatically reconstructing the table in the coal quality detection report image proposed by the present invention;

[0045] Figure 3 The structural diagram of the CRNN feature extraction layer of the text recognition model of a method for automatically reconstructing the table in the coal quality detection report image proposed by the present invention;

[0046] Figure 4 The structural diagram of the RARE encoder of the table structure prediction model based on the adaptive attention mechanism of a method for automatically reconstructing the table in the coal quality detection report image proposed by the present invention;

[0047] Figure 5 The structural diagram of the adaptive attention mechanism of a method for automatically reconstructing the table in the coal quality detection report image proposed by the present invention;

[0048] Figure 6 The output example diagram of the RARE of the table structure prediction model of a method for automatically reconstructing the table in the coal quality detection report image proposed by the present invention;

[0049] Figure 7 The flow chart of the implementation method of table recognition of a method for automatically reconstructing the table in the coal quality detection report image proposed by the present invention;

[0050] Figure 8 The example diagram of the table recognition process of a method for automatically reconstructing the table in the coal quality detection report image proposed by the present invention. Detailed implementation manners

[0051] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.

[0052] A method for automatically reconstructing the table in the coal quality detection report image according to an embodiment of the present invention is characterized in that:

[0053] Step 1, collect as many coal quality detection report images showing tables as possible as the basis for establishing a coal quality detection information database; among them, the collection can separately collect the coal quality detection report images of each power plant of the coal quality to be analyzed;

[0054] Step 2: Perform table recognition on the coal quality inspection report image showing the table obtained in the previous step, that is, through the processing of the text detection model EAST, the text recognition model CRNN, the table structure prediction model RARE based on the adaptive attention mechanism, and the cell aggregation module, obtain the reconstructed table result, and output it in the form of an Excel table;

[0055] Step 3: Perform data post-processing on the reconstructed Excel table information obtained in the previous step to form json format data describing the key indicators and data in the table of the coal quality inspection report image; wherein, the data post-processing refers to, according to the structure information of the reconstructed Excel table, using the main indicators to be extracted as keywords, comparing with the recognized cell text, so as to locate the position of the cell coordinate box where the data corresponding to the indicator is located, obtain the corresponding value, and combine the indicator and the value in the form of a key-value pair to get json format data, thereby improving the readability of the key information in the table of the coal quality inspection report image;

[0056] Step 4: Store the json format data obtained in the previous step in the database in the form of a string to form a coal quality inspection information database; which is used to analyze the coal quality situation of the power plant, monitor the operation process of the thermal boiler, especially the carbon emissions of coal combustion elements, and study the characteristics of coal quality to improve the economy of combustion;

[0057] Step 5: Repeat the operations in Steps 2 to 4, and continue to store as many coal quality inspection report images showing tables as possible in the database to enrich the content of the coal quality inspection information database.

[0058] Among them, the cell aggregation module includes a cell coordinate aggregation module, a cell text aggregation module, and a cell aggregation post-processing module; wherein, the cell coordinate aggregation module takes the text coordinate box coordinates and the cell coordinate box as inputs, outputs the relationship between the two, and judges which cell a single-line text belongs to by calculating the intersection over union (IOU) between the two, and determines the order between single-line text boxes by the vertex distance between the text coordinate box and the cell coordinate box; the cell text aggregation module takes the relationship between the text coordinate box and the cell coordinate box and the recognized text content as inputs, and outputs the complete information of a single cell; the cell aggregation post-processing module takes the complete information of a single cell and the table structure information as inputs, and outputs the complete table information after reconstruction, presented in the form of an Excel table.

[0059] Among them, Step 2 includes the following steps:

[0060] Step 21: Input a coal quality inspection report image M with a width of W and a height of H, which shows a table, into the text detection model EAST. Output a set Y1 of single-line text coordinate boxes. The text coordinate boxes in Y1 are four-point coordinates, in the format of (x1,y1,x2,y2,x3,y3,x4,y4), where x1, y1, x2, y2, x3, y3, x4, y4 represent the coordinates of the text coordinate box.

[0061] Step 22: Input the set Y1 of single-line text coordinate boxes into the text recognition model CRNN to recognize the text in the text coordinate boxes, obtain the corresponding text content, and output a set Yz of single-line text coordinate boxes with text content. The format of Yz is (x1,y1,x2,y2,x3,y3,x4,y4,word), where word represents the corresponding text content in this text coordinate box.

[0062] Step 23: Input the same coal quality inspection report image M with a width of W and a height of H, which shows a table, into the table structure prediction model RARE based on an adaptive attention mechanism. Output an HTML string S that describes the table structure information, and a set Y2 of cell coordinates. Each cell coordinate has the same format as the text coordinate box, denoted as (x1’,y1’,x2’,y2’,x3’,y3’,x4’,y4’). Specifically, the input image M here is the same as the image M input in Step 21.

[0063] Step 24: Input the set Y1 of single-line text coordinate boxes and the set Y2 of cell coordinates into the cell coordinate aggregation module. Calculate the intersection over union between the single-line text coordinate boxes and the cell coordinates, and the vertex distance between the single-line text coordinate boxes and the cell coordinates. Complete the aggregation of the cell coordinates, and output the relationship Y3 between the single-line text coordinate boxes and the cell coordinates.

[0064] Step 25: Input the relationship Y3 between the single-line text coordinate boxes and the cell coordinates and the set Yz of single-line text coordinate boxes with text content into the cell text aggregation module for text aggregation, and output the complete information Yk of a single cell, where k is a positive number representing the number of single cells. Specifically, according to the relationship Y3 between the single-line text coordinate boxes and the cell coordinates, in the order from top to bottom and from left to right, combine the set Yz of single-line text coordinate boxes with text content to complete the splicing. The cell information containing multiple lines of text will be spliced into a complete string, that is, complete the reconstruction of each cell's information.

[0065] Step 26: Input the complete information Yk of a single cell and the string S that describes the table structure information into the post-processing module for cell aggregation to obtain the complete reconstructed table result, and output it in the form of an Excel table.

[0066] Among them, the text detection model EAST includes a feature extraction network, a feature fusion module, and a decoding network; the feature extraction network is constructed based on the VGG16 model, and the feature fusion module adopts the Feature Pyramid Network (FPN) structure, that is, Feature Pyramid Network; the decoding network includes three parts of outputs, and its structure is shown in the appendix Figure 1 : The first part has an output channel of 1, which is obtained by connecting a 1×1 convolutional kernel after the feature fusion layer. The output is a score map, representing the probability that each pixel belongs to the text region; the second part is used to predict the text of the rotated rectangle, with a total of 5 output channels, which are generated by two 1×1 convolutional kernels to produce 4 and 1 channels respectively. Among them, the 4 channels respectively represent the 4 distances from the pixel position to the top, right, bottom, and left boundaries of the rectangle, and 1 channel represents the rotation angle of the bounding box; the third part is used to predict the text of the irregular quadrilateral, and 8 numbers are used to represent the coordinate offsets from the four corner vertices of the quadrilateral to the pixel position; the output of the decoding branch of the third part of the text detection model EAST is removed by using the model pruning method, and the structure of the pruned model is shown in the appendix Figure 2 。In the coal quality detection report images involved in the specific tasks of the present invention, there is no situation of irregular text. Therefore, removing the output of the third part of the text detection model EAST can improve the speed of model training and inference on the basis of ensuring the performance of the text detection model EAST.

[0067] At the same time, the Pytorch deep learning framework is used to train the text detection model EAST. The training process mainly includes data preprocessing and model training; specifically, the ICDAR2017 dataset is downloaded. This dataset is a multilingual dataset containing 7,200 training data and 1,800 test data, involving 9 languages such as Chinese, English, and Japanese; to realize the training of the text detection model EAST, the data containing Chinese and English are selected from the ICDAR2017 dataset for training; and to ensure that the trained model is for single-line text detection, the selected data is re-annotated; in addition, 100 pieces of real data, that is, coal quality detection report images, are prepared for the subsequent fine-tuning of the model. To ensure the detection accuracy of the model, real data is used for fine-tuning.

[0068] Among them, the text recognition model CRNN includes a feature extraction layer, a sequence modeling layer, and a decoding layer. The feature extraction layer is constructed based on the VGG11 model. According to the max-pooling operation, the feature extraction layer can be divided into five feature extraction stages. To ensure the recognition accuracy of the text in the coal quality inspection report image table, the following adjustments are made: First, considering the deficiency of the max-pooling layer in retaining local feature information, the max-pooling layer in the fifth feature extraction stage of this model is removed. Second, to reduce network parameters for easy sequence operation, the convolutional kernel in the fifth feature extraction stage uses a 2×2 convolutional kernel. Third, to generate a feature map with a larger width, that is, a longer feature sequence, and improve the model's ability to recognize longer text lines and narrower characters, the convolutional kernel size of the max-pooling layer in the third and fourth feature extraction stages is set to 1×2. Fourth, to prevent problems such as difficult model convergence and excessive training time when training a deep network, a batch normalization layer, that is, BatchNormalization, is added after the convolutional layers in the third and fourth feature extraction stages. The structure of the adjusted feature extraction layer is shown in the appendix Figure 3 ; The sequence modeling layer uses a bidirectional deep recurrent convolutional neural network BiLSTM, that is, Bi-directional Long Short-Term Memory; the decoding layer uses a connectionist temporal classification layer, that is, Connectionist Temporal Classfication, CTC's conditional probability, to complete obtaining the label sequence from the probability matrix output by the sequence modeling layer, that is, the translated text.

[0069] Meanwhile, the Pytorch deep learning framework is used to train the text recognition model CRNN, and the ADADELTA strategy is adopted in the training process to automatically calculate the learning rate of each dimension. The training process mainly includes data preprocessing and model training. Specifically, the ICDAR2019 dataset is downloaded. This dataset consists of 450,000 Chinese street view images, of which 50,000 are fully annotated data and 400,000 are weakly annotated data. To realize the training of the text recognition model CRNN, 3,000 pictures are selected from the ICDAR2019 dataset to train the model, and 1,500 are used for testing. The characters marked in this dataset include Chinese, numbers, and Latin characters, which are relatively consistent with the character situation of the real data. In addition, 100 pieces of real data are also annotated for model fine-tuning.

[0070] Among them, the table structure prediction model RARE based on the adaptive attention mechanism includes an encoder, an adaptive attention mechanism, a decoder, and a post-processing module. The encoder is divided into two parts. The first half is composed of a convolutional neural network CNN, and the second half is composed of a recurrent neural network RNN. Its structure is shown in the appendix Figure 4As shown; the CNN is constructed based on the VGG11 network structure, an adaptive attention mechanism, and FPN, and the RNN uses BiLSTM; in the CNN structure, the encoder part is divided into five stages according to the max pooling layer. Among them, the convolutional kernel size of the max pooling layer in the fourth stage is 1×2, and in the fifth stage, the original max pooling operation is discarded, and only one convolution with a convolutional kernel size of 3×3, a stride of 1, and a padding value of 0 is set. And the relevance between features is obtained by calculating the feature maps output in different encoding stages through the adaptive attention mechanism. The structure of the adaptive attention mechanism, that is, Self-Attention, is as shown in the appendix Figure 5 As shown; the decoder uses GRU, that is, the Gated RecurrentUnit cell, and each cell contains 256 memory blocks and 37 output units; the post-processing module further parses the output of the decoder into an HTML string format for output, and the string contains the four-point coordinates of each cell and the table structure information, as shown in the appendix Figure 6 shown in

[0071] At the same time, the Pytorch deep learning framework is used to construct and train the table structure prediction model RARE based on the adaptive attention mechanism. The training process mainly includes data preprocessing and model training; specifically, download the PubTabNet dataset. The PubTabNet dataset contains a variety of complex table structures, including a large number of three-row table structures with empty cells or spanning cells, among which there are 500,000 training images, 9,115 validation images, and 9,138 test images; to realize the training of the table structure prediction model RARE based on the adaptive attention mechanism, 3,000 images are selected from the PubTabNet dataset for training. The table structure prediction model RARE based on the adaptive attention mechanism is trained with the prepared image data above. To ensure the accuracy of the model's prediction of the table structure information of the coal quality detection report image, the adaptive attention mechanism module is used to enhance the learning of the deep table structure information by the feature extraction network, and the model is fine-tuned in combination with 150 pieces of labeled real data.

[0072] Among them, the cell coordinate aggregation module judges which cell the single-line text belongs to by the IOU between the single-line text coordinate box and the cell coordinate, and judges the order between single-line text boxes by the distance between the single-line text coordinate box and the vertex of the cell coordinate; the specific calculation process of the intersection over union is as follows:

[0073] S 1 =(x2 - x1)*(y2 - y1)

[0074] S 2 =(x2' - x1')(y2' - y1')

[0075]

[0076] Among them, S1 and S2 respectively represent the area of the single-line text coordinate box and the area of the cell coordinate box, and IOU represents the intersection over union of the two;

[0077] The specific calculation process of the distance between vertices and sorting is as follows:

[0078]

[0079] Among them, d ij represents the distance between the vertex of the i-th single-line text coordinate box and the vertex at the same position of the j-th cell coordinate box, where i and j are integers; according to the known subordination relationship between the single-line text and the cell in the cell coordinate aggregation module, the relationship between i and j can be obtained, that is, j and i are in a many-to-one or one-to-one relationship; sort represents the sorting function, which sorts the calculated distances dij in ascending order to determine the order between the text boxes.

[0080] Among them, the main indicators of the json format data include moisture, carbon element, hydrogen element and nitrogen element indicators; the indicators can be further divided into as-received basis, air-dried basis, dry basis and dry ash-free basis.

[0081] Among them, the fine-tuning of the text detection model EAST, the text recognition model CRNN and the table structure prediction model RARE based on the adaptive attention mechanism means that first, it is trained on a public dataset, and then tested and fine-tuned on the real coal quality detection report table data; because there are special symbols in the coal quality detection report, such as chemical representation symbols of elements such as carbon, hydrogen and nitrogen, combined with the test situation, some characters with low recognition rate are retrained, that is, the fine-tuning process. The specific retraining is to initialize the model parameters with the weights obtained on the public dataset, use the coal quality detection report table data as the training data, and adjust the weights. Because this adjustment amplitude is very small, the main features of the table remain unchanged, but some special detailed features are re-learned and adjusted. This approach is also called model pre-training, which means that usually, a weight trained on a large amount of data is used to initialize the weights of the model, and then secondary and tertiary training is carried out on a specific task, but the premise is that the data used before and after is similar, that is, similar tasks can be done in this way, which can not only improve the training effect of the model, but also shorten the training time. Here, the weights trained by ourselves are used as the pre-weights.

[0082] In another embodiment, the present invention proposes a computer-readable storage medium storing a computer program, and the computer program causes a computer to execute a method for automatically reconstructing a table of a coal quality detection report image as described above.

[0083] In another embodiment, the present invention provides an electronic device, comprising: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, a method for automatically reconstructing a table of a coal quality detection report image as described above is implemented.

[0084] In the embodiments disclosed in the present application, a computer storage medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The computer storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the computer storage medium would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0085] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in the present application can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0086] The above are only the preferred embodiments of the present invention, and the protection scope of the present invention is not limited to the above embodiments. Any technical solutions falling within the concept of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art of this technology, several improvements and refinements made without departing from the principle of the present invention should be regarded as within the protection scope of the present invention.

Claims

1. A method for automatically reconstructing a table in a coal quality inspection report image, characterized in that: Step 1, collect coal quality inspection report images showing tables; Step 2, perform table recognition on the coal quality inspection report images showing tables obtained in the previous step. After processing by the text detection model EAST, the text recognition model CRNN, the table structure prediction model RARE based on the adaptive attention mechanism, and the cell aggregation module, obtain the reconstructed table result and output it in the form of an Excel table; Step 3, perform data post-processing on the reconstructed Excel table information obtained in the previous step to form json-format data describing the key indicators and data in the table of the coal quality inspection report image; Step 4, store the json-format data obtained in the previous step in the database in the form of a string to form a coal quality inspection information database; Step 5, repeat the operations in Steps 2 to 4, and continue to store the obtained coal quality inspection report images showing tables in the database; The cell aggregation module includes a cell coordinate aggregation module, a cell text aggregation module, and a cell aggregation post-processing module; Step 2 includes the following steps: Step 21, input a coal quality inspection report image M with a width of W and a height of H showing a table into the text detection model EAST, and output a set Y1 of single-line text coordinate boxes. The text coordinate boxes in Y1 are four-point coordinates, in the format of (x1, y1, x2, y2, x3, y3, x4, y4), where x1, y1, x2, y2, x3, y3, x4, y4 represent the coordinates of the text coordinate box; Step 22, input the set Y1 of single-line text coordinate boxes into the text recognition model CRNN, and output a set Yz of single-line text coordinate boxes with text content. The format of Yz is (x1, y1, x2, y2, x3, y3, x4, y4, word), where word represents the corresponding text content in the text coordinate box; Step 23, input the same coal quality inspection report image M with a width of W and a height of H showing a table into the table structure prediction model RARE based on the adaptive attention mechanism, and output a string S describing the table structure information and a set Y2 of cell coordinates. Each cell coordinate has the same format as the text coordinate box, denoted as (x1’, y1’, x2’, y2’, x3’, y3’, x4’, y4’); Step 24, input the set Y1 of single-line text coordinate boxes and the set Y2 of cell coordinates into the cell coordinate aggregation module, calculate the intersection-over-union ratio between the single-line text coordinate box and the cell coordinate, and the vertex distance between the single-line text coordinate box and the cell coordinate, and output the relationship Y3 between the single-line text coordinate box and the cell coordinate; Step 25: Input the relationship Y3 between the single-line text coordinate box and the cell coordinates and the set Yz of single-line text coordinate boxes with text content into the cell text aggregation module, and output the complete information Yk of a single cell, where k is a positive number representing the number of single cells. Step 26: Input the complete information Yk of a single cell and the string S describing the table structure information into the post-processing module for cell aggregation to obtain the reconstructed table result, and output it in the form of an Excel table.

2. A method for automatic reconstruction of a table in a coal quality detection report image according to claim 1, characterized in that: The text detection model EAST includes a feature extraction network, a feature fusion module, and a decoding network; the feature extraction network is constructed based on the VGG16 model, and the feature fusion module adopts the Feature Pyramid Network (FPN) structure; the decoding network has three parts of output. The first part has an output channel of 1, which is obtained by connecting a 1×1 convolutional kernel after the feature fusion layer. The output is a score map, representing the probability that each pixel belongs to the text region. The second part is used to predict the text of the rotated rectangle, with a total of 5 output channels, which are generated by two 1×1 convolutional kernels to produce 4 and 1 channels respectively. Among them, the 4 channels respectively represent the 4 distances from the pixel position to the top, right, bottom, and left boundaries of the rectangle, and the 1 channel represents the rotation angle of the bounding box. The third part is used to predict the text of the irregular quadrilateral; the output of the third part of the text detection model EAST is removed by using the model pruning method.

3. A method for automatic reconstruction of a table in a coal quality detection report image according to claim 1, characterized in that: The text recognition model CRNN includes a feature extraction layer, a sequential modeling layer, and a decoding layer; the feature extraction layer is constructed based on the VGG11 model. According to the max-pooling operation, the feature extraction layer is divided into five feature extraction stages: First, remove the max-pooling layer in the fifth feature extraction stage of this model; Second, use a 2×2 convolutional kernel in the fifth feature extraction stage; Third, set the convolutional kernel size of the max-pooling layers in the third and fourth feature extraction stages to 1×2; Fourth, add a batch normalization layer after the convolutional layers in the third and fourth feature extraction stages; the sequential modeling layer uses a bidirectional deep recurrent convolutional neural network (BiLSTM); the decoding layer uses a connectionist temporal classification layer.

4. A method for automatic reconstruction of a table in a coal quality detection report image according to claim 1, characterized in that: The table structure prediction model RARE based on the adaptive attention mechanism includes an encoder, an adaptive attention mechanism, a decoder, and a post-processing module; the encoder is divided into two parts, the first half is composed of a convolutional neural network CNN, and the second half is composed of a recurrent neural network RNN; the CNN is constructed based on the VGG11 network structure, the adaptive attention mechanism, and FPN, and the RNN uses BiLSTM; in the CNN structure, the encoder part is divided into five stages according to the max pooling layer. The convolutional kernel of the max pooling layer in the fourth stage is 1×2, and in the fifth stage, only one convolution with a convolution kernel size of 3×3, a stride of 1, and a padding value of 0 is set. And the relevance between features is obtained by calculating the feature maps output at different encoding stages through the adaptive attention mechanism; the decoder uses GRU, and each unit contains 256 memory blocks and 37 output units; the post-processing module further parses the output of the decoder into an HTML string format for output, and the string contains the four-point coordinates of each cell and the table structure information.

5. A method for automatically reconstructing a table in a coal quality detection report image according to claim 1, wherein: The calculation process of the intersection over union is as follows: Among them, S 1 and S 2 respectively represent the area of the single-line text coordinate box and the area of the cell coordinate box, and IOU represents the intersection over union of the two; The calculation process of sorting the distances between the single-line text coordinate box and the cell coordinate vertices is as follows: Among them, represents the distance between the vertex of the i-th single-line text coordinate box and the vertex at the same position of the j-th cell coordinate box, where i and j are integers and i is not equal to j; sort represents the sorting function.

6. A method for automatically reconstructing a table in a coal quality detection report image according to claim 1, wherein: The indicators of the json format data include moisture, carbon element, hydrogen element, and nitrogen element.

7. A computer-readable storage medium storing a computer program, wherein: The computer program causes the computer to execute a method for automatically reconstructing a table in a coal quality detection report image according to any one of claims 1-6.

8. An electronic device, wherein, comprising: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements a method for automatically reconstructing a table in a coal quality detection report image according to any one of claims 1-6.

Citation Information

Patent Citations

  • Form analysis method based on character positioning recognition, medium and computer equipment

    CN109961008A

  • Multi-scene multi-font Chinese character detection and recognition method

    CN110516554A