A Method, System, Device and Medium for Complex Multi-Scene Document Segmentation

By introducing a deformable attention mechanism and multi-predictive head structure into the document segmentation model, combined with the post-processing method of connectivity domain analysis, the problem of insufficient performance of complex document segmentation in the prior art is solved, and efficient and accurate document segmentation in various scenarios is achieved.

CN114627484BActive Publication Date: 2025-06-13SOUTH CHINA UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210180055.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-25
Publication Date
2025-06-13
Estimated Expiration
2042-02-25

AI Technical Summary

Technical Problem

The prior art has insufficient segmentation performance in complex documents, and different targeted designs are required for different types of documents, and the versatility is poor, making it difficult to achieve accurate and efficient document segmentation.

Method used

The document segmentation model based on the deformable attention mechanism is adopted, and multi-task learning is carried out by constructing a network structure with multiple prediction heads, and redundant foreground area filtering and area shape optimization post-processing is carried out in combination with connectivity domain analysis to achieve the final result of document segmentation.

Benefits of technology

It realizes document segmentation at multiple levels in a variety of complex scenarios, improves the accuracy and efficiency of segmentation, and has an important positive effect on document restoration and machine automation document understanding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114627484B_ABST
    Figure CN114627484B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, system, device and medium for complex multi-scenario document segmentation. The method includes: constructing a data set; constructing a document segmentation model based on a deformable attention mechanism, inputting the images in the data set into the document segmentation model to obtain high-level features; designing a network structure with multiple prediction heads for multi-task learning, using multiple prediction heads to predict the high-level features to obtain predicted segmentation maps, and the prediction results include at least one of text, tables or pictures; performing post-processing on the predicted segmentation maps for redundant foreground region filtering and region shape optimization based on connected component analysis to obtain the final result of document segmentation. The present invention adopts a unified framework to solve the document segmentation problems at multiple levels in multiple scenarios. The method is simple and general, and has an important positive effect on document restoration and automatic document understanding of machines. The present invention can be widely applied to the technical field of document image segmentation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of document image segmentation, and particularly to a complex multi-scenario document segmentation method, system, device and medium. Background Art

[0002] There are various forms of documents in life, such as magazines, ancient books, handwritten documents, etc. Documents contain rich semantic information, so the automatic understanding of documents is of great significance. As an important technology in document layout analysis tasks, document segmentation can better understand the physical structure of documents and is crucial for downstream tasks such as text detection and recognition and document understanding. For different types of documents, the segmentation tasks are also different: irregular documents need to segment the main page area, magazine documents need to segment and locate different document element areas such as text and charts, and ancient book documents need to accurately segment the handwriting of various attribute texts. Therefore, designing a unified document segmentation method for various documents in various complex scenarios is a challenging task.

[0003] Some existing methods for document segmentation mainly fall into two categories: heuristic rule-based methods and learning-based methods. Among them, early heuristic rule algorithms were mainly designed by considering some inherent features of document images and were rather restricted by document types, such as the projection method based on the X-Y direction, the blank analysis method, and the document spectrum method. In recent years, the emerging learning-based methods can better break through these limitations and mainly include two categories: object detection-based methods and semantic segmentation-based methods. The detection-based methods mainly draw on the ideas of general object detection, use detection models to locate the positions of different document elements and classify them. However, due to the limitations of the characteristics of rectangular detection frames, this method can only be adapted to the situation where document elements are close to rectangles. And for the segmentation-based methods, since they can classify each pixel point, they can be better used for complex documents with more complex document element shapes. However, the current existing segmentation-based methods need to be improved in terms of segmentation performance in relatively complex documents, and different targeted designs are required for different types of documents, so the generality is relatively poor.

[0004] Therefore, there is an urgent need for a document segmentation method that can be used in multiple complex scenarios, is applicable to different segmentation tasks of different types of documents, realizes accurate and efficient document segmentation, better understands the physical structure of documents, and helps to build a document automatic understanding system. Summary of the Invention

[0005] To at least partly solve one of the technical problems existing in the prior art, an object of the present invention is to provide a complex multi-scenario document segmentation method, system, device and medium.

[0006] The technical solution adopted by the present invention is:

[0007] A complex multi-scenario document segmentation method includes the following steps:

[0008] Construct a dataset;

[0009] Based on the deformable attention mechanism, construct a document segmentation model, input the images in the dataset into the document segmentation model, and obtain high-level features;

[0010] Design a network structure with multiple prediction heads for multi-task learning, use multiple prediction heads to predict the high-level features, obtain a predicted segmentation map, and the prediction results include at least one of text, tables, or pictures;

[0011] Perform post-processing on the predicted segmentation map, including redundant foreground region filtering based on connected component analysis and region shape optimization, to obtain the final result of document segmentation.

[0012] Furthermore, the construction of the dataset includes:

[0013] Obtain real datasets in multiple scenarios, where the real datasets include page-level datasets, region-level datasets, and handwriting-level datasets;

[0014] Merge and summarize the real datasets in various scenarios to obtain a dataset.

[0015] Furthermore, the document segmentation model uses the ResNet18 network as the backbone network,

[0016] When an image is input into the backbone network, the outputs of the four stages of the backbone network obtain four feature maps with gradually decreasing sizes;

[0017] Transform and splice the obtained feature maps to obtain a one-dimensional feature vector, and input the one-dimensional feature vector into a multi-layer encoder containing a deformable attention mechanism to further extract high-level features;

[0018] Among them, the deformable attention mechanism uses deformable convolution to sample a preset number of feature points in a learnable manner and calculate attention weights with the features at the current position.

[0019] Furthermore, the use of prediction heads to predict multiple high-level features includes:

[0020] Transform the high-level features back into four feature maps, and the sizes of these four feature maps correspond to the sizes of the four feature maps output by the backbone network;

[0021] Send the four transformed feature maps into the network structure with multiple prediction heads to achieve predictions for text, tables, pictures, and multiple categories.

[0022] Further, the document segmentation model is trained in the following manner:

[0023] For page-level and region-level segmentation tasks, considering the integrity of document elements, the entire image is used as the input for training and testing;

[0024] For handwriting-level segmentation tasks, due to the relatively fine segmentation granularity, the sliding window method is used, and a part of the image is sequentially cut out and fed into the model for training and testing.

[0025] Further, the post-processing of filtering redundant foreground regions and optimizing the region shape based on connected component analysis for the predicted segmentation map includes:

[0026] By performing connected component analysis on the predicted segmentation map, each independent segmentation region is obtained;

[0027] Calculate the average pixel value at the corresponding original image position within the segmentation region. If the pixel value is greater than the preset threshold, determine that the segmentation region is a background region and filter out the background region;

[0028] After filtering out the background regions, for each remaining text region, estimate the size of the text within the text region according to the characteristics of the connected component size;

[0029] Select a filter according to the estimated text size, perform sliding within the text region, and filter out redundant background regions according to the characteristics of the pixels within the filter.

[0030] Further, obtaining the final result of document segmentation includes:

[0031] For page-level document segmentation tasks, obtain the position of the main page region;

[0032] For region-level document segmentation tasks, obtain the positions of three types of document elements: paragraphs, illustrations, and tables;

[0033] For handwriting-level document segmentation tasks, obtain the positions of text handwritings with various attributes.

[0034] Another technical solution adopted by the present invention is:

[0035] A complex multi-scenario document segmentation system includes:

[0036] A data acquisition module for constructing a data set;

[0037] A feature extraction module for constructing a document segmentation model based on the deformable attention mechanism, inputting the images in the data set into the document segmentation model, and obtaining high-level features;

[0038] A feature prediction module, which is used to design a network structure with multiple prediction heads for multi-task learning, and uses multiple prediction heads to predict high-level features to obtain a predicted segmentation map, and the prediction results include at least one of text, tables, or pictures;

[0039] A connected component analysis module, which is used to perform post-processing on the predicted segmentation map for redundant foreground area filtering and region shape optimization based on connected component analysis to obtain the final result of document segmentation.

[0040] Another technical solution adopted by the present invention is:

[0041] A complex multi-scenario document segmentation device, including:

[0042] At least one processor;

[0043] At least one memory, which is used to store at least one program;

[0044] When the at least one program is executed by the at least one processor, the at least one processor implements the above method.

[0045] Another technical solution adopted by the present invention is:

[0046] A computer-readable storage medium, in which a program executable by a processor is stored, and the program executable by the processor is used to execute the above method when executed by the processor.

[0047] The beneficial effect of the present invention is that the present invention adopts a unified framework to solve the document segmentation problems at multiple levels in multiple scenarios. The method is simple and general, and has an important positive effect on document restoration and automatic document understanding of machines. Description of the Drawings

[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following introduces the accompanying drawings related to the technical solutions in the embodiments of the present invention or the prior art. It should be understood that the accompanying drawings introduced below are only for conveniently and clearly expressing some embodiments of the technical solutions in the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative efforts.

[0049] Figure 1 It is a step flow chart of a complex multi-scenario document segmentation method in an embodiment of the present invention;

[0050] Figure 2 It is a schematic diagram of the diversity examples of a document data set in a complex multi-scenario used in an embodiment of the present invention;

[0051] Figure 3It is a schematic diagram of the structure of multiple prediction heads in the document segmentation model in the embodiment of the present invention;

[0052] Figure 4 It is a schematic diagram of the principle of the post - processing method for segmentation regions based on connected - component analysis in the embodiment of the present invention;

[0053] Figure 5 It is a schematic diagram of an example of the final result obtained by a complex multi - scenario document segmentation method according to an embodiment of the present invention. Detailed implementation manners

[0054] The embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the drawings, where the same or similar reference numerals indicate the same or similar elements or elements with the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary and are only used to explain the present invention, and should not be construed as a limitation of the present invention. For the step numbers in the following embodiments, they are only set for convenience of description and explanation, and no limitation is imposed on the order between steps. The execution order of each step in the embodiments can be adjusted adaptively according to the understanding of those skilled in the art.

[0055] In the description of the present invention, it should be understood that for the orientation description, such as the orientation or positional relationship indicated by up, down, front, back, left, right, etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, so it should not be construed as a limitation of the present invention.

[0056] In the description of the present invention, the meaning of several is one or more, the meaning of multiple is two or more, greater than, less than, exceeding, etc. are understood as not including the present number, and above, below, within, etc. are understood as including the present number. If there is a description of first and second, it is only for the purpose of distinguishing technical features and should not be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or the sequence of the indicated technical features.

[0057] In the description of the present invention, unless otherwise clearly defined, words such as setting, installing, connecting, etc. should be understood in a broad sense. Those skilled in the art can reasonably determine the specific meaning of the above words in the present invention in combination with the specific content of the technical solution.

[0058] As Figure 1 shown, this embodiment provides a complex multi - scenario document segmentation method, including the following steps:

[0059] S1. Construct a data set.

[0060] First, collect and summarize document segmentation datasets in various scenarios, synthesize corresponding additional data as needed, and clarify the segmentation tasks in each scenario. The collected and summarized document scenarios include magazines, papers, photos, handwritten texts, ancient books, etc., and the covered segmentation tasks include page-level, region-level, and handwriting-level. Figure 2 are some examples of diverse document images in document data under complex multi-scenarios.

[0061] Specifically, several publicly available synthetic datasets such as PubLayNet and real datasets in various scenarios such as ICDAR RDCL are used, covering three-level segmentation tasks including page-level, region-level, and handwriting-level. The datasets in various scenarios are merged and summarized, and necessary data is synthesized based on latex as needed.

[0062] S2. Build a document segmentation model based on the deformable attention mechanism, input the images in the dataset into the document segmentation model, and obtain high-level features.

[0063] Design of feature extraction in the document segmentation model: Use ResNet18 as the backbone network. The outputs of the 4 stages of the backbone network obtain 4 feature maps with gradually decreasing sizes. Then, these feature maps are transformed and concatenated to obtain a one-dimensional feature vector, which is input into a multi-layer encoder containing the deformable attention mechanism to further extract high-level features. Among them, the deformable attention mechanism realizes sampling of a certain number of feature points in a learnable manner through deformable convolution, and calculates the attention weights with the features at the current position. Experiments prove that in relatively simple images such as documents, only using sampling and attention calculation of several points is sufficient to meet the needs of feature extraction and can greatly reduce the computational amount of the attention mechanism. In order to fully extract features, 6 stacked encoder modules are used in this embodiment.

[0064] Among them, the training method of the document segmentation model is as follows: For the page-level and region-level segmentation tasks, considering the integrity of document elements, the entire image is used as the input for training and testing, and the input size used is 800x800; for the handwriting-level, due to the relatively fine segmentation granularity, in order to ensure the resolution, the sliding window method is used, and a part of the image is cut out in turn and sent into the model for training and testing (size 512x512). Such a training method has a certain data augmentation effect and can also achieve more accurate prediction.

[0065] S3. Design a network structure with multiple prediction heads for multi-task learning, use multiple prediction heads to predict the high-level features, obtain the predicted segmentation map, and the prediction results include at least one of text, table, or picture.

[0066] Design of Feature Decoding Prediction in Document Segmentation Model: Through the feature extraction in step S2, the high-level semantic features of the picture can be obtained, which is a one-dimensional semantic vector. It is re-transformed into a feature map with the same size as the above 4 feature maps, and then sent into the subsequent network structure with multiple prediction heads to respectively realize the prediction of text, table, picture, and multiple categories. Here, multiple categories mean that it can realize the prediction of three categories of text, table, and picture at the same time, rather than just the prediction of a certain category. The structural schematic diagram is as Figure 3 shown.

[0067] After using the ResNet18 backbone network to extract high-level features, multiple layers of transformer encoders are used for feature learning. Among them, the attention module in each layer of the encoder uses a deformable attention mechanism. The finally obtained high-level features are sent into the prediction head for segmentation map prediction. The structure of the multiple prediction heads is used for multi-category prediction and per-category prediction of three categories such as text, picture, and table, making full use of the multi-level feature information output by the transformer encoder. Multi-task learning and fusion learning further improve the performance and have a certain regularization effect, enhancing the generalization of the model.

[0068] S4. Perform post-processing of redundant foreground region filtering and region shape optimization based on connected component analysis on the predicted segmentation map to obtain the final result of document segmentation.

[0069] The post-processing method for segmentation regions based on connected component analysis can make full use of the prior that the document background is relatively clean and single, helping to filter out some irrelevant background regions. By performing connected component analysis on the prediction map, each independent segmentation region can be obtained. Setting a certain threshold can filter out some region blocks that may be the background. For the remaining region blocks belonging to the text type, further connected component analysis can be performed to estimate the font size of the text contained therein, and based on this, an adaptive-size filter is used to slide on the corresponding region to filter out some adhered background regions and optimize the shape of the segmentation region.

[0070] Specifically, the method for redundant foreground region filtering and region shape optimization post-processing based on connected component analysis is as follows: By performing connected component analysis on the segmentation map, each independent segmentation region can be obtained, and the average pixel value at the corresponding original image position within the region is calculated. If the pixel value is greater than a certain threshold, it is considered a background region and is filtered out. At the same time, for each remaining text region, further connected component analysis can be performed to estimate the size of the text within the region according to the size characteristics of the connected component. Select a filter according to the estimated text size and perform sliding within the region. Redundant background regions can be filtered out according to the characteristics of the pixels within the filter. The steps are implemented as shown in the appendix Figure 4 shown. Further, through experiments, it is verified that the proposed post-processing method can improve the performance of the document segmentation result.

[0071] Finally, through the accurate prediction of the document segmentation model and the further optimization of the post-processing method, the final segmentation result of the document can be obtained: for the region-level segmentation task, the position of the main page can be extracted; for the region-level segmentation task, the positions of different document elements such as text, pictures, and tables can be obtained; for the handwriting-level segmentation task, the accurate positions of the text handwritings with different attributes can be segmented. The final effect diagram is as shown in the appendix Figure 5 as follows

[0072] As can be seen from the above, this embodiment systematically proposes a new method for complex multi-scene document segmentation based on the deformable attention mechanism, mainly including designing a unified document segmentation framework to solve the page-level, region-level, and handwriting-level segmentation tasks of documents in multi-scenes, and applying the deformable attention mechanism to feature extraction in the document segmentation task, which helps to extract more discriminative features while increasing the receptive field; also proposes a new method for optimizing the post-processing of document segmentation based on connected component processing, which further improves the accuracy and segmentation quality of the segmentation map on the basis of model prediction. The method is effective, general, and simple

[0073] In summary, the method of this embodiment has the following advantages and beneficial effects compared with the prior art

[0074] (1) In constructing the document segmentation model, the present invention uses a deformable attention mechanism module, which solves the problem that document elements require a large receptive field through a global attention module, and at the same time has a small computational complexity, and is also beneficial to extracting more discriminative features

[0075] (2) The multi-prediction head structure used in the segmentation model of the present invention makes full use of the features of each layer output by the attention encoder. At the same time, the training paradigm of multiple prediction heads for multiple tasks is beneficial to the learning of the model. Finally, the fusion output of the results of each prediction head further improves the model performance

[0076] (3) The post-processing method of redundant foreground region filtering and region shape optimization based on connected component analysis adopted by the present invention can further optimize the document segmentation map on the basis of the model. Since the background of the document is relatively consistent, the method based on threshold filtering can better filter out the wrong foreground regions. At the same time, by estimating the size of the text in a specific text region to set an adaptive size sliding filter, the redundant regions adhered in the text region can also be better optimized, and the region shape can be optimized

[0077] (4) The present invention solves the problem of layout segmentation of multi-scenario documents with various types and complex styles, and can be uniformly applied to the segmentation tasks of different documents. The segmentation model cleverly uses a transformer encoder based on a deformable attention mechanism to extract features, extracts more discriminative features, further improves the performance, and at the same time can make the backbone network more lightweight, accelerate the model speed and reduce the number of parameters. The method plays an important role in further realizing the restoration and reconstruction of documents and automatic understanding.

[0078] This embodiment also provides a complex multi-scenario document segmentation system, including:

[0079] A data acquisition module for constructing a data set;

[0080] A feature extraction module for constructing a document segmentation model based on a deformable attention mechanism, inputting the images in the data set into the document segmentation model, and obtaining high-level features;

[0081] A feature prediction module for designing a network structure with multiple prediction heads for multi-task learning, using multiple prediction heads to predict the high-level features, obtaining a predicted segmentation map, and the prediction result includes at least one of text, table or picture;

[0082] A connected component analysis module for performing post-processing of redundant foreground region filtering and region shape optimization on the predicted segmentation map based on connected component analysis to obtain the final result of document segmentation.

[0083] A complex multi-scenario document segmentation system according to this embodiment can execute a complex multi-scenario document segmentation method provided by the method embodiment of the present invention, can execute any combination of the implementation steps of the method embodiment, and has the corresponding functions and beneficial effects of the method.

[0084] This embodiment also provides a complex multi-scenario document segmentation device, including:

[0085] At least one processor;

[0086] At least one memory for storing at least one program;

[0087] When the at least one program is executed by the at least one processor, the at least one processor is caused to implement Figure 1 The method shown.

[0088] A complex multi-scenario document segmentation device according to this embodiment can execute a complex multi-scenario document segmentation method provided by the method embodiment of the present invention, can execute any combination of the implementation steps of the method embodiment, and has the corresponding functions and beneficial effects of the method.

[0089] The embodiments of the present application also disclose a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes Figure 1 the method shown.

[0090] This embodiment also provides a storage medium, which stores instructions or programs that can execute a complex multi-scenario document segmentation method provided by the method embodiment of the present invention. When the instructions or programs are run, any combination of implementation steps of the method embodiment can be executed, and the corresponding functions and beneficial effects of the method are possessed.

[0091] In some alternative embodiments, the functions / operations mentioned in the block diagrams may not occur in the order mentioned in the operation diagrams. For example, depending on the functions / operations involved, two consecutive blocks shown may actually be executed substantially simultaneously, or the blocks can sometimes be executed in the reverse order. In addition, the embodiments presented and described in the flowcharts of the present invention are provided by way of example for the purpose of providing a more comprehensive understanding of the technology. The disclosed methods are not limited to the operations and logical flows presented herein. Alternative embodiments are contemplated, in which the order of various operations is changed and the sub-operations described as part of a larger operation are executed independently.

[0092] In addition, although the present invention is described in the context of functional modules, it should be understood that, unless otherwise stated to the contrary, one or more of the functions and / or features may be integrated in a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It can also be understood that a detailed discussion of the actual implementation of each module is not necessary for understanding the present invention. Rather, considering the attributes, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the modules will be understood within the ordinary skills of an engineer. Therefore, those skilled in the art can implement the present invention set forth in the claims without undue experimentation using ordinary skills. It can also be understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present invention, and the scope of the present invention is determined by the entire scope of the appended claims and their equivalents.

[0093] If the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.

[0094] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.

[0095] More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection portion with one or more wirings (electronic device), a portable computer disk cartridge (magnetic device), random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, then editing, interpreting, or otherwise processing it as appropriate, and then storing it in a computer memory.

[0096] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), and the like.

[0097] In the above description of this specification, the description with reference to the terms "one embodiment / example", "another embodiment / example", or "certain embodiments / examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0098] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the claims and their equivalents.

[0099] The above is a specific description of the preferred embodiments of the present invention, but the present invention is not limited to the above embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of this application.

Claims

1. A method for segmenting complex multi-scenario documents, characterized in that, it includes the following steps: Construct a dataset; Based on the deformable attention mechanism, construct a document segmentation model, input the images in the dataset into the document segmentation model to obtain high-level features; Design a network structure with multiple prediction heads for multi-task learning, use multiple prediction heads to predict the high-level features, obtain a predicted segmentation map, and the prediction results include at least one of text, tables, or pictures; Perform post-processing of redundant foreground region filtering and region shape optimization based on connected component analysis on the predicted segmentation map to obtain the final result of document segmentation; The document segmentation model uses the ResNet18 network as the backbone network, When an image is input into the backbone network, the outputs of the four stages of the backbone network obtain four feature maps with gradually decreasing sizes; Transform and splice the obtained feature maps to get a one-dimensional feature vector, input the one-dimensional feature vector into a multi-layer encoder containing the deformable attention mechanism to extract high-level features; Among them, the deformable attention mechanism realizes sampling of a preset number of feature points through deformable convolution and calculates attention weights with the features at the current position; The post-processing of redundant foreground region filtering and region shape optimization based on connected component analysis on the predicted segmentation map includes: Obtain each independent segmentation region by performing connected component analysis on the predicted segmentation map; Calculate the average pixel value of the corresponding original image position within the segmentation region. If the pixel value is greater than a preset threshold, determine that the segmentation region is a background region and filter out the background region; After filtering out the background region, for each remaining text region, estimate the size of the text within the text region according to the characteristics of the connected component size; Select a filter according to the estimated text size, slide within the text region, and filter out redundant background regions according to the characteristics of the pixels within the filter.

2. A method for segmenting complex multi-scenario documents according to claim 1, characterized in that, the construction of the dataset includes: Obtain real datasets under multiple scenarios, and the real datasets include page-level datasets, region-level datasets, and handwriting-level datasets; Merge and summarize the real datasets under various scenarios to obtain a dataset.

3. A method for segmenting complex multi-scenario documents according to claim 1, characterized in that, the use of prediction heads to predict multiple high-level features includes: Transform the high-level features back into 4 feature maps, and the sizes of these 4 feature maps correspond to the sizes of the 4 feature maps output by the backbone network; Send the 4 transformed feature maps into the network structure with multiple prediction heads to realize the prediction of text, tables, pictures, and multiple categories.

4. A method for segmenting complex multi-scenario documents according to claim 1, characterized in that, the document segmentation model is trained in the following way: For page-level and region-level segmentation tasks, considering the integrity of document elements, use the entire image as the input for training and testing; For handwriting-level segmentation tasks, due to the relatively fine segmentation granularity, use the sliding window method to sequentially cut out a part of the image and send it into the model for training and testing.

5. A method for segmenting complex multi-scenario documents according to claim 1, characterized in that, obtaining the final result of document segmentation includes: For the page-level document segmentation task, obtaining the position of the main page area; For the region-level document segmentation task, obtaining the positions of three document elements: paragraphs, illustrations, and tables; For the handwriting-level document segmentation task, obtaining the positions of text handwritings with various attributes.

6. A complex multi-scenario document segmentation system, characterized in that, including: A data acquisition module for constructing a data set; A feature extraction module for constructing a document segmentation model based on the deformable attention mechanism, inputting the images in the data set into the document segmentation model, and obtaining high-level features; A feature prediction module for designing a network structure with multiple prediction heads for multi-task learning, using multiple prediction heads to predict the high-level features, obtaining a predicted segmentation map, and the prediction result includes at least one of text, tables, or pictures; A connected component analysis module for performing post-processing of redundant foreground region filtering and region shape optimization on the predicted segmentation map based on connected component analysis to obtain the final result of document segmentation; The document segmentation model uses the ResNet18 network as the backbone network, When an image is input into the backbone network, the outputs of the 4 stages of the backbone network obtain 4 feature maps with gradually decreasing sizes; Transform and splice the obtained feature maps to obtain a one-dimensional feature vector, input the one-dimensional feature vector into a multi-layer encoder containing the deformable attention mechanism to extract high-level features; Among them, the deformable attention mechanism realizes sampling of a preset number of feature points in a learnable manner through deformable convolution, and calculates the attention weight with the features at the current position; The post-processing of redundant foreground region filtering and region shape optimization on the predicted segmentation map based on connected component analysis includes: By performing connected component analysis on the predicted segmentation map, obtaining each independent segmentation region; Calculating the average pixel value of the corresponding original image position within the segmentation region, if the pixel value is greater than a preset threshold, determining that the segmentation region is a background region, and filtering out the background region; After filtering out the background region, for each remaining text region, estimating the size of the text within the text region according to the size characteristics of the connected component; Selecting a filter according to the estimated text size, performing sliding within the text region, and filtering out redundant background regions according to the characteristics of the pixels within the filter.

7. A complex multi-scenario document segmentation device, characterized in that, including: At least one processor; At least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the method according to any one of claims 1-5.

8. A computer-readable storage medium, in which a program executable by a processor is stored, characterized in that, The program executable by the processor is used to execute the method according to any one of claims 1-5 when executed by the processor.