A Method for Implementing Document Layout Analysis Based on Domestic CPUs

Through the official document layout analysis method based on domestic CPUs, using feature extraction and model integration technology, the official document types and the official document area are automatically identified, which solves the complexity and inefficiency of manual annotation and input in the existing technology, and realizes automatic identification and labeling with high accuracy.

CN114140809BActive Publication Date: 2025-05-27INSPUR QILU SOFTWARE IND
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111318867.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-09
Publication Date
2025-05-27
Estimated Expiration
2041-11-09

AI Technical Summary

Technical Problem

The prior art cannot automatically identify the types and official document fields in official document pictures, resulting in the need for manual annotation or input, which is complex, inefficient, and errors may occur.

Method used

Using a method based on domestic CPU, through feature extraction and model integration technology, the official document types are automatically identified and the official document domain is crawled. The specific steps include: uploading official document pictures and extracting features, forming a two-dimensional tensor input trained classification model to determine the type of document, and using object detection technology to locate the official document domain.

Benefits of technology

Automatic identification and labeling of official documents is realized, the accuracy of official document layout analysis is improved, the operation process is simplified, efficiency is improved, and human errors are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114140809B_ABST
    Figure CN114140809B_ABST
Patent Text Reader

Abstract

The present invention particularly relates to a method for realizing official document layout analysis based on domestic CPUs. The method for realizing official document layout analysis based on domestic CPUs extracts features from official document pictures, erases the feature information of official document texts, extracts three features, namely line height, horizontal starting point, and horizontal ending point, from the lines after erasing the feature information of official document texts to form a two-dimensional tensor, and inputs it into a trained classification model to determine the type of official document; after obtaining the type of official document, the picture is input into a detection neural network trained for this type of document, and the location information of the official document area is captured using object detection technology. The method for realizing official document layout analysis based on domestic CPUs can automatically identify the type of official document, capture and label and classify the official document area in the official document, and at the same time uses three models to adopt a voting mechanism to determine the type of official document, and inputs official document pictures of different types into different detection neural networks, greatly improving the accuracy of official document layout analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of domestic software and hardware adaptation, and in particular to a method for realizing document layout analysis based on a domestic CPU. Background Art

[0002] With the strong support of the state, domestically produced software and hardware with independent intellectual property rights have developed rapidly, especially in recent years, many basic software and hardware products with independent intellectual property rights have emerged in my country. High-end general-purpose chips with independent intellectual property rights such as Loongson, Feiteng, and Shenwei have flourished, and their technical level has reached the world's advanced level of similar products.

[0003] At the same time, the development of domestic basic software products is also thriving, with the emergence of domestic basic software products such as the Kylin operating system, Shentong database, Jincang database, Damang database, Oriental Communication middleware, Kingdee middleware, Puhua office software, Kingsoft office software, Foxit layout office suite, Shuke layout office suite, Zhonganyuan electronic signature, Xinan electronic signature, etc. These basic software and hardware products have reached or are close to the world's advanced level in terms of performance and ease of use.

[0004] With the vigorous development of domestic basic software and hardware, it has brought unprecedented opportunities for the promotion and use of domestic basic software and hardware. In addition, based on the security and reliability of domestic software and hardware, it is imperative to replace domestic software and hardware in important fields such as government and military industry.

[0005] At present, the document image cannot be identified by computer for the document type and document domain, which makes it impossible to proceed to the next step of the text recognition process. The document type and document domain need to be manually marked, or the document needs to be manually entered into the office system. For example, if a paper external document needs to be entered into the system, a staff member must check the document type, the document unit, and other document systems of the paper document, and then select or enter it in the office system to enter it into the system. This process is complicated, inefficient, and may lead to errors.

[0006] Based on the above situation, the present invention proposes a method for realizing document layout analysis based on a domestically produced CPU. Summary of the invention

[0007] In order to make up for the defects of the prior art, the present invention provides a simple and efficient method for realizing document layout analysis based on a domestically produced CPU.

[0008] The present invention is achieved through the following technical solutions:

[0009] A method for realizing document layout analysis based on a domestically produced CPU, characterized in that it comprises the following steps:

[0010] The first step is to upload the official document image to the server, extract the features of the official document image, and erase the feature information of the official document text. The three features of line height, horizontal starting point and horizontal end point are extracted from the line after erasing the feature information of the official document text to form a two-dimensional tensor;

[0011] The second step is to transform the three features of line height, horizontal starting point and horizontal end point into a two-dimensional tensor and input it into the trained classification model to determine the type of official document.

[0012] The classification model adopts a model ensembling mechanism, including an RNN (Recurrent Neural Network) model, a 1DCNN (1D Convolutional Neural Networks) model and a DNN (Deep Neural Networks) model;

[0013] The RNN model is responsible for extracting sequential features, the 1DCNN model is responsible for extracting block features, and the DNN model is responsible for extracting global features. Finally, a voting mechanism is used to determine the type of official document.

[0014] The third step is to obtain the type of official document and input the image into the detection neural network trained for that type of document, and use target detection technology to capture the location information of the official document domain.

[0015] In the first step, the file formats that support recognition are preset, and after the official document image that needs to be recognized is selected, the file format of the selected official document image is automatically recognized and filtered: if the file format of the selected official document image is a file format that supports recognition, the official document image file is uploaded to the server; otherwise, the user is prompted that the file format does not support recognition and please reselect the file.

[0016] In the first step, the method of erasing the characteristic information of the official document is as follows:

[0017] (1) Rotate and straighten official document images

[0018] Binarize the image, then use Canny and Laplace edge detection operators to obtain the image contour map, and then use the contour extraction algorithm to obtain the maximum contour. The maximum contour is the edge of the paper. Adjust the deflection angle of the maximum contour to straighten the image.

[0019] (2) Erasing text features of official document images

[0020] Use the projection method to count the number of pixels in each line of the image, and then calculate the set of rows with more than 10 consecutive rows of 0 pixels based on the statistics. This set is all horizontal blank lines, that is, the line spacing. The lines obtained by inverting the line spacing are each line of text. Remove the blank spaces at the beginning and end of each line to get the actual area occupied by each line of text, and then fill all text areas with black.

[0021] The RNN model uses a bidirectional LSTM (Long Short-Term Memory) layer to obtain the sequential features in the official document and determine the type of the official document based on the sequential features;

[0022] The 1DCNN model uses a 1DCONV (one-dimensional vector convolution operation) neural network to obtain image position features and channel features of the document while preserving the order;

[0023] The DNN model adopts a fully connected method to obtain the overall features of the official document and determine the type of the document in turn.

[0024] In the second step, the row height, horizontal row start position and horizontal row end position of each row obtained in the first step are arranged in order to form a tensor with a dimension of (20, 3). The rows less than 20 are filled with 0, and the rows greater than 20 are removed for input into the RNN model.

[0025] The RNN model obtains sequential features through two layers of LSTM, and then identifies the global layout through two layers of fully connected layers, obtains the output of one-hot encoding, and obtains the corresponding classification result according to the output.

[0026] In the second step, a two-dimensional tensor (20, 3) is formed with the number of rows as the vertical axis and the row height, the horizontal row starting point, and the horizontal row end point as the horizontal axis, and the tensor is input into a 1DCONV (one-dimensional vector convolution operation) neural network.

[0027] In the second step, a two-dimensional tensor (20, 3) is input into a fully connected neural network, with the number of rows as the vertical axis and the row height, the horizontal row starting point, and the horizontal row end point as the horizontal axis.

[0028] In the second step, the voting weights of the RNN model, 1DCNN model, and DNN model are 45%, 35%, and 20%, respectively.

[0029] The beneficial effects of the present invention are as follows: the method for realizing official document layout analysis based on the domestic CPU can automatically identify the type of official document, capture the official document domain in the official document and mark and classify it, and at the same time adopt three models to determine the type of official document using a voting mechanism, and input official document images of different types into different detection neural networks, which greatly improves the accuracy of official document layout analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0031] Attached Figure 1 It is a schematic diagram of the method for realizing document layout analysis based on domestic CPU of the present invention.

[0032] Attached Figure 2 Schematic diagram of the network structure of the RNN model of the present invention.

[0033] Attached Figure 3 This is a schematic diagram of the network structure of the 1DCNN model of the present invention.

[0034] Attached Figure 4 Schematic diagram of the network structure of the DNN model of the present invention. DETAILED DESCRIPTION

[0035] In order to enable those skilled in the art to better understand the technical solutions in the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0036] The method for realizing document layout analysis based on domestic CPU comprises the following steps:

[0037] The first step is to upload the official document image to the server, extract the features of the official document image, and erase the feature information of the official document text. The three features of line height, horizontal starting point and horizontal end point are extracted from the line after erasing the feature information of the official document text to form a two-dimensional tensor;

[0038] The second step is to transform the three features of line height, horizontal starting point and horizontal end point into a two-dimensional tensor and input it into the trained classification model to determine the type of official document.

[0039] The classification model adopts a model ensembling mechanism, including an RNN (Recurrent Neural Network) model, a 1DCNN (1D Convolutional Neural Networks) model and a DNN (Deep Neural Networks) model;

[0040] Model ensembling is based on the following assumption: for different good models trained independently, their good recognition rates may be due to different reasons and characteristics: each model observes the data from a slightly different perspective to make predictions, and obtains part of the "truth", but not the whole truth. By bringing their views together, a more accurate description of the data can be obtained. Therefore, the use of model ensembling can significantly improve the accuracy of document classification.

[0041] The RNN model is responsible for extracting sequential features, the 1DCNN model is responsible for extracting block features, and the DNN model is responsible for extracting global features. Finally, a voting mechanism is used to determine the type of official document.

[0042] In the third step, the model trained for different document types will show a stronger ability to detect official document domains when faced with that document type. After obtaining the official document type, the image is input into the detection neural network trained for that document type, and the target detection technology is used to capture the location information of the official document domain. At this point, all the layout information of the official document has been extracted.

[0043] In the first step, the file formats that support recognition are preset, and after the official document image that needs to be recognized is selected, the file format of the selected official document image is automatically recognized and filtered: if the selected official document image file format is a file format that supports recognition, the official document image file is uploaded to the server; otherwise, the user is prompted that the file format does not support recognition and please reselect the file.

[0044] Since the text in the official document is changeable and unpredictable, the text cannot be used as the feature of the official document for recognition. In order to make the neural network training more accurate, the method of erasing the feature information of the official document text in the first step is as follows:

[0045] (1) Rotate and straighten official document images

[0046] Binarize the image, then use Canny and Laplace edge detection operators to obtain the image contour map, and then use the contour extraction algorithm to obtain the maximum contour. The maximum contour is the edge of the paper. Adjust the deflection angle of the maximum contour to straighten the image.

[0047] (2) Erasing text features of official document images

[0048] Use the projection method to count the number of pixels in each line of the image, and then calculate the set of rows with more than 10 consecutive rows of 0 pixels based on the statistics. This set is all horizontal blank lines, that is, the line spacing. The lines obtained by inverting the line spacing are each line of text. Remove the blank spaces at the beginning and end of each line to get the actual area occupied by each line of text, and then fill all text areas with black.

[0049] The RNN model uses a bidirectional LSTM (Long Short-Term Memory) layer to obtain the sequential features in the official document and determine the type of the official document based on the sequential features;

[0050] The 1DCNN model uses a 1DCONV (one-dimensional vector convolution operation) neural network to obtain image position features and channel features of the document while preserving the order;

[0051] The DNN model adopts a fully connected method to obtain the overall features of the official document and determine the type of the document in turn.

[0052] In the second step, the row height, horizontal row start position and horizontal row end position of each row obtained in the first step are arranged in order to form a tensor with a dimension of (20, 3). The rows less than 20 are filled with 0, and the rows greater than 20 are removed for input into the RNN model.

[0053] The RNN model obtains sequential features through two layers of LSTM, and then identifies the global layout through two layers of fully connected layers, obtains the output of one-hot encoding, and obtains the corresponding classification result according to the output.

[0054] In the second step, a two-dimensional tensor (20, 3) is formed with the number of rows as the vertical axis and the row height, the horizontal row starting point, and the horizontal row end point as the horizontal axis, and the tensor is input into a 1DCONV (one-dimensional vector convolution operation) neural network.

[0055] In the second step, a two-dimensional tensor (20, 3) is input into a fully connected neural network, with the number of rows as the vertical axis and the row height, the horizontal row starting point, and the horizontal row end point as the horizontal axis.

[0056] In the second step, the voting weights of the RNN model, 1DCNN model, and DNN model are 45%, 35%, and 20%, respectively.

[0057] The embodiment described above is only one specific implementation of the present invention. Common changes and substitutions made by those skilled in the art within the scope of the technical solution of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for analyzing official document layout based on domestic CPU. It is characterized in that The following steps are involved: The first step is to upload the official document image to the server, extract the features of the official document image, and erase the feature information of the official document text. The three features of line height, horizontal starting point and horizontal end point are extracted from the line after erasing the feature information of the official document text to form a two-dimensional tensor; The second step is to transform the three features of line height, horizontal starting point and horizontal end point into a two-dimensional tensor and input it into the trained classification model to determine the type of official document. The classification model adopts a model integration mechanism, including an RNN model, a 1DCNN model and a DNN model; The RNN model is responsible for extracting sequential features, the 1DCNN model is responsible for extracting block features, and the DNN model is responsible for extracting global features. Finally, a voting mechanism is used to determine the type of official document. The RNN model uses a bidirectional LSTM layer to obtain the sequential features in the official document and determine the type of the official document based on the sequential features; The 1DCNN model uses a 1DCONV neural network to obtain image position features and channel features of the document while preserving the order; The DNN model adopts a fully connected method to obtain the overall features of the official document and determine the type of the official document in turn; Arrange the row height, horizontal row start position, and horizontal row end position of each row obtained in the first step in order to form a tensor with a dimension of (20, 3). Fill the rows with 0 if they are less than 20, and remove the excess part if they are more than 20, so as to input them into the RNN model; The RNN model obtains sequential features through two layers of LSTM, then identifies the global layout through two layers of fully connected layers, obtains the output of one-hot encoding, and obtains the corresponding classification result according to the output; The third step is to obtain the type of official document and input the image into the detection neural network trained for that type of document, and use target detection technology to capture the location information of the official document domain.

2. According to the method for realizing document layout analysis based on domestic CPU according to claim 1, Features: In the first step, the file formats that support recognition are preset, and after the official document image that needs to be recognized is selected, the file format of the selected official document image is automatically recognized and filtered: if the selected official document image file format is a file format that supports recognition, the official document image file is uploaded to the server; otherwise, the user is prompted that the file format does not support recognition and please reselect the file.

3. According to the method for realizing document layout analysis based on domestic CPU according to claim 1, Features: In the first step, the method of erasing the characteristic information of the official document is as follows: (1) Rotate and straighten official document images Binarize the image, then use Canny and Laplace edge detection operators to obtain the image contour map, and then use the contour extraction algorithm to obtain the maximum contour. The maximum contour is the edge of the paper. Adjust the deflection angle of the maximum contour to straighten the image. (2) Erasing text features in official document images Use the projection method to count the number of pixels in each row of the image, and then calculate the set of rows with 0 pixel count for more than 10 consecutive rows based on the statistics. This set is all the horizontal blank rows, that is, the line spacing. The rows obtained by inverse selection of the line spacing are the text of each row. Removing the leading and trailing blanks of each row gives the actual area occupied by the text of each row, and then filling all the text areas with black.

4. The method for realizing official document layout analysis based on domestic CPU according to claim 3, characterized in that: In the second step, a two-dimensional tensor of (20, 3) is formed with the number of rows as the vertical axis and the row height, the starting point of the horizontal row, and the ending point of the horizontal row as the horizontal axis, and the tensor is input into the 1D CONV neural network.

5. The method for realizing official document layout analysis based on domestic CPU according to claim 3, characterized in that: In the second step, a two-dimensional tensor of (20, 3) is formed with the number of rows as the vertical axis and the row height, the starting point of the horizontal row, and the ending point of the horizontal row as the horizontal axis, and is input into the fully connected neural network.

6. The method for realizing official document layout analysis based on domestic CPU according to claim 1, characterized in that: In the second step, the voting weights of the RNN model, the 1DCNN model, and the DNN model are 45%, 35%, and 20% respectively.

Citation Information

Patent Citations

  • A Document Image Classification Method

    CN109344815A

  • Electronic official document recognition and reproduction method and system based on domestic CPU

    CN112949471A