Document Content Classification Method, System, Device and Computer Readable Storage Medium

By converting documents into pictures and using document content classification model and layout analysis model for area division and sorting, the problem of inflexible and effective classification of document content in the existing technology is solved, and effective classification and layout sorting of long-length documents is realized, and fault tolerance is improved.

CN114863408BActive Publication Date: 2025-06-20SICHUAN MEDICAL SHUN TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202110648550.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-06-10
Publication Date
2025-06-20
Estimated Expiration
2041-06-10

AI Technical Summary

Technical Problem

Existing document content classification technology is difficult to achieve flexible and effective document content classification in the case of insufficient corpus coverage and limited rule updates, especially in long-form documents, the semantic characteristics between word vectors cannot be reflected.

Method used

By obtaining the target document and converting it into image format, using the preset document content classification model and document layout analysis model, extracting content features and text types, and dividing and sorting the area according to the preset classification standards and layout rules, and finally obtaining the reorganized document.

Benefits of technology

It realizes more flexible and effective document content classification, can process long-form documents, reflect the semantic characteristics between word vectors, and reduces the impact on the entire document layout through single area sorting, with a higher fault tolerance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114863408B_ABST
    Figure CN114863408B_ABST
Patent Text Reader

Abstract

The present application discloses a method, system, device and computer-readable storage medium for document content classification, including: converting the document into a picture format to obtain a target picture corresponding to the target document; using a preset document content classification model to extract content features from the target picture, and dividing the target picture into regions according to the content features to obtain a plurality of segmented regions to be sorted; using a preset document layout analysis model to extract the text types of each segmented region, and sorting according to the text types of each segmented region to obtain a plurality of text regions with correct text order; re-sorting each text region to obtain a reorganized document. In the present application, the document is divided into multiple regions according to categories through image recognition, and each region is typeset separately, making the typesetting more flexible. Errors between regions do not seriously affect the whole. Finally, overall sorting is performed to obtain a complete document.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information retrieval, and particularly to a method, system, device and computer-readable storage medium for classifying document content. Background Art

[0002] The technology of document content classification is to label and classify information content under a certain classification system, which belongs to a research field of information retrieval technology. Its function is to help people improve the efficiency of managing and processing text information and is widely used in fields such as document structured processing, document organization, and text filtering. After research, traditional document content classification technologies are implemented based on statistical and rule-based methods. The statistical-based method is a probabilistic inference method with uncertainty learned from a large-scale corpus. The disadvantage of this method is that the coverage of the corpus needs to be wide enough to achieve good results. The rule-based method is to formulate certain classification rules according to some rule constraints in linguistics. This method is a deterministic inference method. The disadvantage of this method is that the formulation of rules requires the participation of domain experts, which in turn limits the update of rules. With the development of deep neural network technology, in recent years, most document content classification tasks are implemented based on NLP-related tasks. The basic implementation method is to first perform word segmentation on the text and perform Embedding operations to extract the feature vectors of the words, then go through a series of convolution and pooling operations, and finally obtain the classification result through softmax (Softmax logical regression). The advantage of this method is that the model is simple and easy to train. The disadvantage is to adjust the model parameters targeted according to the training results, and at the same time, the semantic features between word vectors cannot be reflected for long documents. In short, the above-mentioned text content classification methods require a large amount of text content that conforms to the correct semantics and has the correct word order as the basic support, and certain preprocessing of the text data is required, such as word segmentation, word frequency cleaning, processing of special symbols and stop words, construction of word vectors, etc.

[0003] Order is the premise to ensure the correct semantics of the text. Whether it is the result after document classification or the detection and recognition of the words in each category, the returned results may be out of order. Failing to process these results in terms of order will directly seriously affect the effects of downstream NLP (Natural Language Processing) related tasks. Therefore, it is crucial to return the correct order. In the prior art, during the process of sorting text, it is easy to make incorrect judgments, resulting in a chaotic document layout.

[0004] Therefore, a more accurate, flexible and effective method for classifying document content is needed. Summary of the Invention

[0005] In view of this, the purpose of the present invention is to provide a method, system, device and computer-readable storage medium for classifying document content, which is more flexible and effective. The specific solutions are as follows:

[0006] A method for classifying document content includes:

[0007] Obtain a target document, convert the document into a picture format to obtain a target picture corresponding to the target document;

[0008] Use a preset document content classification model to extract content features from the target picture according to a preset classification standard, and divide the target picture into regions according to the content features to obtain multiple split regions to be sorted;

[0009] Use a preset document layout analysis model to extract the text types of each split region, and sort the text order within each split region according to the text types of each split region according to a preset layout rule to obtain multiple text regions with correct text order;

[0010] Use the content features and text order of each text region to re-sort each text region to obtain a reorganized document;

[0011] Wherein, the document content classification model is obtained by pre-segmenting and training historical pictures according to the preset classification standard; the document layout analysis model is obtained by pre-layout training historical pictures according to the preset layout rule.

[0012] Optionally, the document content classification model uses ResNet+FPN as the backbone network, and first fuses the channel attention model and then the spatial attention model for the Feature Map generated by each ResBlock structure in the ResNet network to obtain the Feature Map with the attention mechanism fused generated by the entire backbone network.

[0013] Optionally, the classification standard includes: text, title, table body, table title, table annotation, list, image, annotation, header and footer.

[0014] Optionally, the process of using a preset document layout analysis model to extract the text types of each split region, and sorting the text order within each split region according to the text types of each split region according to a preset layout rule to obtain multiple text regions with correct text order includes:

[0015] Use the document layout analysis model to analyze the text types of the split regions;

[0016] Calculate the BoundingBox coordinate area corresponding to the segmented area using the text type of the segmented area;

[0017] Determine the vertical sorting order of the segmented areas using the width of the BoundingBox coordinate area and the width of the corresponding segmented area;

[0018] Judge the text spacing in the segmented area using the height of the BoundingBox coordinate area.

[0019] The present invention also discloses a document content classification system, including:

[0020] An image conversion module, configured to obtain a target document, convert the document into an image format, and obtain a target image corresponding to the target document;

[0021] A region classification module, configured to use a preset document content classification model to extract content features from the target image according to a preset classification standard, and divide the target image into multiple segmented areas to be sorted according to the content features;

[0022] A document layout module, configured to use a preset document layout analysis model to extract the text type of each segmented area, and sort the text order within each segmented area according to the text type of each segmented area according to a preset layout rule to obtain multiple text areas with correct text order;

[0023] A document recombination module, configured to re-sort each text area using the content features and text order of each text area to obtain a recombined document;

[0024] Wherein, the document content classification model is obtained by pre-segmenting and training historical images according to the preset classification standard; the document layout analysis model is obtained by pre-layout training historical images according to the preset layout rule.

[0025] Optionally, the document layout module includes:

[0026] A text type analysis unit, configured to analyze the text type of the segmented area using the document layout analysis model;

[0027] A BoundingBox calculation unit, configured to calculate the BoundingBox coordinate area corresponding to the segmented area using the text type of the segmented area;

[0028] A vertical sorting unit, configured to determine the vertical sorting order of the segmented areas using the width of the BoundingBox coordinate area and the width of the corresponding segmented area;

[0029] A spacing sorting unit for determining the text spacing in the divided area by using the height of the BoundingBox coordinate area.

[0030] The present invention also discloses a document content classification device, including:

[0031] A memory for storing a computer program;

[0032] A processor for executing the computer program to implement the document content classification method as described above.

[0033] The present invention also discloses a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the document content classification method as described above is implemented.

[0034] In the present invention, the document content classification method includes: obtaining a target document, converting the document into a picture format to obtain a target picture corresponding to the target document; using a preset document content classification model, extracting content features from the target picture according to a preset classification standard, dividing the target picture into regions according to the content features to obtain multiple divided regions to be sorted; using a preset document layout analysis model, extracting the text types of each divided region, and sorting the text order in each divided region according to the text types of each divided region according to a preset layout rule to obtain multiple text regions with correct text order; re-sorting each text region by using the content features and text order of each text region to obtain a reorganized document; wherein, the document content classification model is obtained by pre-dividing and training historical pictures according to the preset classification standard; the document layout analysis model is obtained by pre-training the layout of historical pictures according to the preset layout rule.

[0035] In the present invention, the document is divided into multiple regions according to categories through image recognition, and each region is typeset separately, making the typesetting more flexible. Finally, the overall sorting is performed to obtain a complete document. By sorting individual regions, even if the sorting within an individual region is incorrect, it can also reduce the impact on the layout of the entire document, and the error tolerance rate is higher. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0037] Figure 1Schematic flowchart of a document content classification method disclosed in an embodiment of the present invention;

[0038] Figure 2 Schematic structural diagram of a document content classification system disclosed in an embodiment of the present invention. Detailed implementation manners

[0039] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0040] An embodiment of the present invention discloses a document content classification method. Refer to Figure 1 as shown, the method includes:

[0041] S11: Obtain a target document, convert the document into a picture format, and obtain a target picture corresponding to the target document.

[0042] Specifically, in order to classify the document content using image recognition technology, therefore, the document in a non-picture format is converted into a picture format. Of course, a document that is already in a picture format does not need to be converted again and can be directly used as the target picture.

[0043] S12: Use a preset document content classification model, according to a preset classification standard, extract content features from the target picture, and divide the target picture into regions according to the content features to obtain multiple segmented regions to be sorted.

[0044] Specifically, the classification standard may include: text, title, table body, table title, table note, list, image, note, header, and footer. The picture will be divided into regions according to the classification standard to obtain each picture region corresponding to the classification standard. For example, the title region, the table region, and the footer region, etc. In this process, only various types of content in the picture are recognized and not sorted, so each region is a segmented region to be sorted.

[0045] S13: Use a preset document layout analysis model to extract the text types of each segmented region, and according to the text types of each segmented region, sort the text order in each segmented region according to a preset layout rule to obtain multiple text regions with correct text order.

[0046] Specifically, by extracting the text types of each segmented area, the text types can include paragraph spacing, and the document layout can be in one-column, two-column, three-column, mixed multi-column and other layout modes. Then, judge the text order within each segmented area. For example, judge whether the corresponding text content within the segmented area is in upper and lower paragraphs, and whether the text content corresponds to pictures or tables. After the analysis, the content in the segmented area can be re-sorted according to the preset layout rules. For example, the two-column text can be re-sorted into one column to obtain multiple text areas with correct text order.

[0047] S14: Re-order each text area by using the content features and text order of each text area to obtain the reorganized document.

[0048] Specifically, the text area corresponds to the segmented area. When the text order within each text area is normal, by using the content features and text order between each text area, the text areas can be re-ordered, and finally the reorganized document can be obtained.

[0049] Among them, the document content classification model is obtained by pre-segmenting and training historical pictures according to the preset classification criteria; the document layout analysis model is obtained by pre-training the layout of historical pictures according to the preset layout rules.

[0050] It can be seen that in the embodiment of the present invention, the document is divided into multiple regions according to categories through image recognition, and each region is typeset separately, making the typesetting more flexible. Finally, the overall sorting is performed to obtain the complete document. By sorting individual regions, even if the sorting within individual regions is incorrect, the impact on the layout of the entire document can be reduced, and the error tolerance rate is higher.

[0051] The embodiment of the present invention discloses a specific document content classification method. Compared with the previous embodiment, this embodiment further explains and optimizes the technical solution. Specifically:

[0052] Further, the document content classification model can adopt ResNet+FPN as the backbone network. First, fuse the channel attention model for the Feature Map generated by each ResBlock structure in the ResNet network, and then fuse the spatial attention model to obtain the Feature Map integrated with the attention mechanism generated by the entire backbone network.

[0053] Specifically, the document content classification model can include a training dataset construction stage, a backbone network construction stage for feature extraction, and a model training stage; among them,

[0054] In the stage of constructing the training data set, on the one hand, an open-source data set is used for model training, and on the other hand, data labeled by data annotators using a developed annotation system is incorporated. The main categories of the current document mainly include text, title, table_body, table_title, table_annotation, list, figure, annotation, page_header, footer, etc., a total of 10 categories. The data annotation type uses the COCO data set format.

[0055] In the stage of constructing the backbone network for feature extraction, the idea of instance segmentation is adopted. Compared with the object detection series models, the instance segmentation model performs segmentation calculations on the basis of detection, achieving pixel-level recognition. Furthermore, the coordinate positions of the bounding boxes recognized are more accurate, providing guarantee for the accuracy of subsequent text detection and recognition. At the same time, only a relatively small data set is required to train a model with high generalization ability.

[0056] In order to better extract features, the embodiment of the present invention uses a Feature Pyramid Network (FPN) for multi-scale object detection method, which realizes the fusion of features of each hierarchical structure. On the other hand, there may be a large number of blank areas in some image pages, the layout of each category area in different images has diversity, and there is a certain positional relationship between some categories. For example, for the table category, it includes table title, table body, table annotation, etc., that is, there is a certain connection between the features of different categories in space. These factors may all lead to a decline in the recognition performance of the model. Therefore, in order to avoid the above disadvantages, first, fully explore the features between different categories. Combining the characteristics of the deep neural network structure, it can be improved by increasing the depth of the network, increasing the number of channels of the feature map, and using multi-scale feature fusion technology, etc. Then, during the model training process, spatial position information is incorporated to suppress the common features of each category and improve the representation ability of specific regions to improve the recognition accuracy of the model. Therefore, ResNet+FPN is used as the backbone network. At the same time, for the Feature Map generated by each ResBlock structure in the ResNet network, the channel attention model is first fused, and then the spatial attention model is fused. Furthermore, the attention mechanism is fused for the Feature Map generated by the entire backbone network, automatically learning the importance of different feature channels and each feature space, helping us extract features in the image that contain both spatial feature weights and feature weights between different channels, improving the feature extraction ability of the model.

[0057] Suppose the size of the Feature Map generated by the backbone network for the original image is c*w*h, where c represents the number of channels of the feature map, w represents the width of the feature map, and h represents the height of the feature map. Global max pooling and global average pooling are respectively performed on this Feature Map in the spatial dimension. The output values then pass through a fully connected layer and a softmax activation function. Finally, an addition operation is performed on the respectively output feature vectors to obtain the weights of the channel attention model, and a dot product operation is performed with the input features to obtain the feature map passing through the channel attention model.

[0058] Global max pooling and global average pooling are respectively performed on the feature map output by the channel attention model in the channel dimension, and the obtained Feature Maps are merged. Finally, through a convolutional layer and a sofxmax activation layer, the weights of the spatial attention model with a size of 1*w*h are obtained, and a dot product operation is performed with the output features of the channel attention model.

[0059] In the model training stage, the RPN module is responsible for generating a fixed number of candidate regions ROIs, and performing foreground and background classification and regression to detect the position and size of the target bounding box, obtaining the filtered ROI regions. The RoIAlign module uses the method of bilinear interpolation to correspond the feature map of the ROI region with the original image region, cancels the rounding operation, alleviates the problem of the position deviation between the feature map and the original image caused by RoIPooling, and improves the detection accuracy. The MASK module performs classification, regression calculation of the candidate box, and generates a Mask for the ROI region to complete the instance segmentation task.

[0060] Further, the above S13 process of using a preset document layout analysis model to extract the text types of each segmentation region, and sorting the text order within each segmentation region according to the text types of each segmentation region according to the preset layout rules to obtain multiple text regions with correct text order can include S131 to S134; where

[0061] S131: Analyze the text type of the segmentation region using the document layout analysis model;

[0062] S132: Calculate the BoundingBox coordinate region corresponding to the segmentation region using the text type of the segmentation region.

[0063] Specifically, according to the type of image text, calculate the BoundingBox coordinate region corresponding to the text type in the segmentation region. Here, further processing is performed on the returned BoundingBox coordinate region, and the intersection-over-union ratio (IoU = [0, 1]) between each BoundingBox is calculated. If there are multiple intersecting BoundingBoxes, and the IoU value of two or more BoundingBoxes is greater than a fixed threshold (e.g., 0.98), it is considered that there is complete overlap between these BoundingBoxes, and the BoundingBox coordinates and categories that are completely contained within are removed.

[0064] S133: Determine the vertical sorting order of the segmentation regions by using the width of the BoundingBox coordinate region and the width of the corresponding segmentation region.

[0065] Specifically, calculate the ratio of the width of each BoundingBox returned in the previous step to the width of the entire image, find all BoundingBox coordinate regions where the ratio of the width of the BoundingBox coordinates to the width of the entire image is greater than a fixed threshold (e.g., 0.5, that is, more than half of the width of this image), and sort the BoundingBoxes in this part in ascending order along the Y-axis.

[0066] S134: Use the height of the BoundingBox coordinate region to judge the text spacing in the segmentation region.

[0067] Specifically, based on the calculated BoundingBox coordinate region, the layout in the entire image is segmented into multiple regions. The remaining BoundingBox coordinates are classified according to the calculated multiple regions, and then it is judged how many columns of layout the category in this region belongs to. All BoundingBoxes within each region are sorted in ascending order from left to right first, and then in ascending order according to the coordinate values from top to bottom. Then, the corresponding BoundingBoxes are added to the corresponding layout list in turn; the BoundingBox coordinate regions whose calculated ratio to the width of the entire image is greater than the fixed threshold are inserted into the corresponding positions (sorted along the Y-axis) in turn.

[0068] Furthermore, the layout between lines is adjusted. Based on the Bounding Boxes recognized by the OCR module, all Bounding Boxes are first pre-sorted according to the Y-axis coordinates. Then, it is determined whether the difference between the center coordinates of the current Bounding Box and the center coordinates of the next Bounding Box is greater than half of the height of the current Bounding Box. If it is greater than half, it is determined that the current Bounding Box is at the line break position. After finding the line break position, the Bounding Boxes in each line are sorted according to the X-axis. At this time, the line-level sorting rule within the paragraph is completed. Finally, some detailed processing is done. For example, there is a problem with the '-' connector at the end of a line. If the character is directly deleted, it will be too direct to connect with the characters in the next line. For example, after deleting '50-60', it becomes '5060', which directly causes a semantic error. Currently, it is processed according to the rule of whether it is a letter or a number. This detailed problem can also be judged with the help of subtasks in NLP.

[0069] Correspondingly, an embodiment of the present invention also discloses a document content classification system. Refer to Figure 2 As shown, the system includes:

[0070] An image conversion module 11, configured to obtain a target document, convert the document into an image format, and obtain a target image corresponding to the target document;

[0071] A region classification module 12, configured to use a preset document content classification model, extract content features from the target image according to a preset classification standard, and divide the target image into regions according to the content features to obtain multiple segmentation regions to be sorted;

[0072] A document layout module 13, configured to use a preset document layout analysis model, extract the text types of each segmentation region, and sort the text order within each segmentation region according to a preset layout rule according to the text types of each segmentation region to obtain multiple text regions with correct text order;

[0073] A document recombination module 14, configured to re-sort each text region according to the content features and text order of each text region to obtain a recombined document;

[0074] Among them, the document content classification model is obtained by pre-segmenting and training historical images according to a preset classification standard; the document layout analysis model is obtained by pre-training the layout of historical images according to a preset layout rule.

[0075] It can be seen that in the embodiments of the present invention, the document is divided into multiple regions according to categories through image recognition, and each region is typeset separately, making the typesetting more flexible. Finally, an overall sorting is performed to obtain a complete document. By sorting individual regions, even if there are sorting errors within individual regions, the impact on the layout of the entire document can be reduced, and the error tolerance rate is higher.

[0076] Specifically, the above-mentioned document layout module 13 includes a text type analysis unit, a BoundingBox calculation unit, a vertical sorting unit, and a spacing sorting unit; among them,

[0077] The text type analysis unit is used to analyze the text type of the segmented region by using the document layout analysis model;

[0078] The BoundingBox calculation unit is used to calculate the BoundingBox coordinate region corresponding to the segmented region by using the text type of the segmented region;

[0079] The vertical sorting unit is used to determine the vertical sorting order of the segmented regions by using the width of the BoundingBox coordinate region and the width of the corresponding segmented region;

[0080] The spacing sorting unit is used to judge the text spacing in the segmented region by using the height of the BoundingBox coordinate region.

[0081] Among them, the document content classification model uses ResNet+FPN as the backbone network. First, the channel attention model is fused with the Feature Map generated by each ResBlock structure in the ResNet network, and then the spatial attention model is fused to obtain the Feature Map with the attention mechanism fused generated by the entire backbone network.

[0082] Among them, the classification criteria include: text, title, table body, table title, table note, list, image, note, header, and footer.

[0083] In addition, the embodiments of the present invention also disclose a document content classification device, including:

[0084] A memory for storing a computer program;

[0085] A processor for executing the computer program to implement the document content classification method as described above.

[0086] In addition, the embodiments of the present invention also disclose a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the document content classification method as described above is implemented.

[0087] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0088] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this text can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0089] The above has introduced in detail the technical content provided by the present invention. Specific examples are used in this text to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A method for classifying document content, characterized in that, Including: Obtain a target document, convert the document into a picture format to obtain a target picture corresponding to the target document; Using a preset document content classification model, extract content features from the target picture according to a preset classification standard, and divide the target picture into regions based on the content features to obtain multiple segmented regions to be sorted. The document content classification model uses ResNet+FPN as the backbone network. For the Feature Map generated by each ResBlock structure in the ResNet network, first fuse the channel attention model and then fuse the spatial attention model to obtain the Feature Map with the attention mechanism fused generated by the entire backbone network. The classification standard includes: text, title, table body, table title, table annotation, list, image, annotation, header, and footer; Using a preset document layout analysis model, extract the text types of each segmented region, and sort the text order within each segmented region according to the preset layout rules based on the text types of each segmented region to obtain multiple text regions with correct text order; Among them, the step of using a preset document layout analysis model to extract the text types of each segmented region, and sort the text order within each segmented region according to the text types of each segmented region according to the preset layout rules to obtain multiple text regions with correct text order includes: analyzing the text types of the segmented regions using the document layout analysis model; calculating the BoundingBox coordinate region corresponding to the segmented region using the text types of the segmented regions; determining the vertical sorting order of the segmented regions using the width of the BoundingBox coordinate region and the width of the corresponding segmented region; judging the text spacing in the segmented region using the height of the BoundingBox coordinate region; Re-sort each text region using the content features and text order of each text region to obtain a reorganized document; Among them, the document content classification model is obtained by pre-training the segmentation of historical pictures according to the preset classification standard; the document layout analysis model is obtained by pre-training the layout of historical pictures according to the preset layout rules.

2. A system for classifying document content, characterized in that, Including: A picture conversion module for obtaining a target document and converting the document into a picture format to obtain a target picture corresponding to the target document; The region classification module is used to extract content features from the target picture according to a preset classification standard by using a preset document content classification model, and divide the target picture into multiple segmented regions to be sorted according to the content features. The document content classification model uses ResNet+FPN as the backbone network. For the Feature Map generated by each ResBlock structure in the ResNet network, the channel attention model is first fused, and then the spatial attention model is fused to obtain the Feature Map with the attention mechanism fused generated by the entire backbone network. The classification standard includes: text, title, table body, table title, table annotation, list, image, annotation, header, and footer; The document layout module is used to extract the text types of each segmented region by using a preset document layout analysis model, and sort the text order in each segmented region according to the text types of each segmented region according to a preset layout rule to obtain multiple text regions with correct text order; The document layout module includes: a text type analysis unit for analyzing the text type of the segmented region by using a document layout analysis model; a BoundingBox calculation unit for calculating the BoundingBox coordinate region corresponding to the segmented region by using the text type of the segmented region; a vertical sorting unit for determining the vertical sorting order of the segmented region by using the width of the BoundingBox coordinate region and the width of the corresponding segmented region; a spacing sorting unit for judging the text spacing in the segmented region by using the height of the BoundingBox coordinate region; The document recombination module is used to re-sort each text region by using the content features and text order of each text region to obtain the recombined document; Among them, the document content classification model is obtained by pre-segmenting and training historical pictures according to the preset classification standard; the document layout analysis model is obtained by pre-layout training historical pictures according to the preset layout rule.

3. A device for classifying document content, characterized in that, Including: A memory for storing computer programs; A processor for executing the computer program to implement the document content classification method as described in claim 1.

4. A computer-readable storage medium, characterized in that, The computer program is stored on a computer-readable storage medium, and when the computer program is executed by a processor, it implements the document content classification method as described in claim 1.

Citation Information

Patent Citations

  • Document information extraction method and device and electronic equipment

    CN111680491A

  • Character recognition method and device, storage medium and electronic equipment

    CN112183250A

  • Image segmentation method, network training method, electronic equipment and storage medium

    CN112613519A