Image-text separation method and device for document image, chip, equipment and storage medium
By inputting document images into the image processing network model, and separating the labels and mask sub-pictures of candidate areas, the problem of low image-text separation accuracy in the prior art is solved, and high-precision image-text separation in complex scenarios is achieved.
Patent Information
- Application Number
- CN202510218325.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art is limited to coarse-grained segmentation in the separation of graphic and text of document images, and cannot adapt to scenes such as complex backgrounds, interlaced graphic and text or blurred text areas, resulting in low accuracy of graphic and text separation.
By inputting the document image into the trained image processing network model, the category labels, bounding box positions and mask sub-pictures of the candidate areas are used to separate the image and text, and high-precision extraction of the image content and text content in the document image is achieved.
High-precision graphic separation in scenes such as complex backgrounds, interlaced graphic or blurred text areas is achieved, and the accuracy and efficiency of graphic separation are improved.
Smart Images

Figure CN120071371A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of computer technology, and in particular, to a method, apparatus, chip, device, and storage medium for separating text and images in a document image. Background Art
[0002] With the advent of the digital age, a large number of paper documents need to be converted into digital document images for storage and management. Among them, accurately separating the text area and the image area in the document image can provide high-quality input for subsequent tasks such as text recognition, document reconstruction, and image optimization, and can also facilitate the classified storage, retrieval, and editing of digital content, enrich the content of information retrieval and mining, and improve the accuracy and efficiency of retrieval.
[0003] In the related art, methods such as image threshold segmentation, edge detection, and region growing are provided for separating text and images in a document image. However, the foregoing methods are often limited to coarse-grained segmentation and cannot adapt to application scenarios such as complex backgrounds, text and image interleaving, or blurred text areas. The accuracy of the text and image separation results is relatively low and needs to be improved. Summary of the Invention
[0004] The embodiments of the present application provide a method, apparatus, chip, device, and storage medium for separating text and images in a document image, which solve the problem that the related technology is often limited to coarse-grained segmentation and cannot adapt to application scenarios such as complex backgrounds, text and image interleaving, or blurred text areas, and the accuracy of the text and image separation results is relatively low. It realizes the accurate extraction of the image content and text content in the document image, completes high-precision text and image separation, and can be effectively deployed in application scenarios such as complex backgrounds, text and image interleaving, or blurred text areas.
[0005] In a first aspect, the embodiments of the present application provide a method for separating text and images in a document image, the method comprising:
[0006] Obtain a document image;
[0007] Input the document image into a trained image processing network model to obtain an output result, where the output result includes class labels, bounding box positions, and mask sub-images corresponding to at least one candidate region in the document image, and the candidate region is a region containing text or an image;
[0008] Separate the images and text in the document image according to the class labels, bounding box positions, and mask sub-images corresponding to the at least one candidate region to obtain the separated image content and text content.
[0009] In a second aspect, the embodiments of the present application further provide a device for separating text and images in a document image, comprising:
[0010] An acquisition module, configured to acquire a document image;
[0011] An image processing module, configured to input the document image into a trained image processing network model to obtain an output result, where the output result includes class labels, bounding box positions, and mask sub - graphs corresponding to at least one candidate region in the document image, and the candidate region is a region containing text or an image;
[0012] A text - image separation module, configured to separate the image and text in the document image according to the class labels, bounding box positions, and mask sub - graphs corresponding to the at least one candidate region, to obtain separated image content and text content.
[0013] In a third aspect, an embodiment of the present application further provides a chip, which includes:
[0014] One or more processors;
[0015] A memory, configured to store one or more programs,
[0016] When the one or more programs are executed by the one or more processors, the one or more processors implement the method for separating text and image of the document image according to the embodiments of the present application.
[0017] In a fourth aspect, an embodiment of the present application further provides an electronic device, which includes the chip according to the embodiment of the present application.
[0018] In a fifth aspect, an embodiment of the present application further provides a non - volatile storage medium storing computer - executable instructions, where the computer - executable instructions are configured to execute the method for separating text and image of the document image according to the embodiments of the present application when executed by a computer processor.
[0019] In the embodiments of the present application, a document image is obtained, and the document image is input into a trained image processing network model to obtain an output result. The output result includes class labels, bounding box positions, and mask sub-images corresponding to at least one candidate region in the document image. The candidate region is a region containing text or an image. Based on the class labels, bounding box positions, and mask sub-images corresponding to at least one candidate region, the image and text in the document image are separated to obtain the separated image content and text content. In the above solution, by inputting the document image into a trained image processing network model, the processing ability of the image processing network model can be effectively utilized to obtain key reference information for separating text content and image content. By separating the image and text in the document image based on the class labels, bounding box positions, and mask sub-images corresponding to at least one candidate region, the image content and text content in the document image can be accurately extracted to complete high-precision graphic-text separation, and it can be effectively deployed in application scenarios such as complex backgrounds, graphic-text interleaving, or blurred text regions. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 is a flowchart of a method for separating graphic and text of a document image provided by an embodiment of the present application;
[0021] Figure 2 is a flowchart of a method for separating graphic and text including the specific process of processing a document image by an image processing network model provided by an embodiment of the present application;
[0022] Figure 3 is a flowchart of a specific training process of a classification network in step S203 provided by an embodiment of the present application;
[0023] Figure 4 is a flowchart of a specific training process of a regression network in step S204 provided by an embodiment of the present application;
[0024] Figure 5 is a flowchart of a specific training process of a mask generation network in step S205 provided by an embodiment of the present application;
[0025] Figure 6 is a flowchart of a joint training process of the corresponding classification network, regression network, and mask generation network in steps S203, S204, and S205 provided by an embodiment of the present application;
[0026] Figure 7 is a flowchart of a method for separating graphic and text including the process of content extraction and processing of a document image provided by an embodiment of the present application;
[0027] Figure 8A flowchart of a text-image separation method provided by an embodiment of the present application, which includes a process of processing text content and image content;
[0028] Figure 9 A flowchart of a text-image separation method provided by an embodiment of the present application, which includes a process of preprocessing a document image;
[0029] Figure 10 A structural block diagram of a text-image separation device for a document image provided by an embodiment of the present application;
[0030] Figure 11 A schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0031] The following further describes the embodiments of the present application in detail with reference to the drawings and examples. It can be understood that the specific embodiments described herein are only used to explain the embodiments of the present application, rather than limiting the embodiments of the present application. Additionally, it should be noted that for the sake of description, only parts related to the embodiments of the present application are shown in the drawings rather than all the structures.
[0032] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of the same category, and the number of objects is not limited. For example, the first object can be one or multiple. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally indicates an "or" relationship between the associated objects before and after.
[0033] The text-image separation method for a document image provided by an embodiment of the present application is used to accurately separate image content and text content from a document image, and can provide strong data support for application scenarios such as document digitization, optical character recognition, document management, and information extraction. The present application aims to provide a text-image separation method for a document image to solve the problem that related technologies are often limited to coarse-grained segmentation, cannot adapt to application scenarios such as complex backgrounds, interlaced text and images, or blurred text regions, and the accuracy of text-image separation results is relatively low.
[0034] The execution subject of each step of the text and image separation method for document images provided in the embodiments of this application can be a computer device, which refers to any electronic device with data calculation, processing, and storage capabilities, such as terminal devices like mobile phones, PCs (Personal Computers), and tablet computers, or devices such as servers. The embodiments of this application do not make any limitations in this regard.
[0035] Figure 1 It is a flowchart of a text and image separation method for document images provided in the embodiments of this application. As Figure 1 shown, the text and image separation method for this document image specifically includes the following steps:
[0036] Step S101: Obtain a document image.
[0037] Among them, the document image can be a digital image obtained by scanning or photographing a paper document, or it can be sourced from local storage of the device or other terminal devices. The document image contains text content and image content.
[0038] Step S102: Input the document image into a trained image processing network model to obtain an output result. Among them, the output result includes class labels, bounding box positions, and mask sub - images corresponding to at least one candidate region in the document image, and the candidate region is a region containing text or an image.
[0039] Among them, the image processing network model can be a Mask R-CNN (Mask Region-based Convolutional Neural Network) model, a Faster R-CNN (Faster Region-based Convolutional Neural Network) model, etc., which are not limited in this application. Taking the Mask R-CNN model as an example, the Mask R-CNN model can not only perform object detection, but also perform pixel-level segmentation on each object in the image and generate a corresponding mask. Therefore, the Mask R-CNN model can be effectively applied to the detection and high-quality segmentation of the image area and text area of the document image, and can separate the image content and text content with high precision. The at least one candidate region can be obtained by the image processing network model detecting and segmenting text and images. The candidate region can be a text region that only contains text content, or an image region that only contains image content. The category label can be used to identify whether the candidate region belongs to the image region or the text region. For example, a text label and an image label. The bounding box position can be used to identify the specific coordinate position and range size of the candidate region in the document image. The mask sub-image can be used to identify the distribution position of the specific text content or the specific image content in the candidate region. Specifically, the mask sub-image can be in binary form, and different pixel values can be used to represent the corresponding target content for candidate regions of different content types. For example, for a candidate region with a category label of text label, the image processing network model can generate a mask sub-image with a pixel value of 1 for the pixel points corresponding to the text content. Another example is that for a candidate region with a category label of image label, the image processing network model can generate a mask sub-image with a pixel value of 0 for the pixel points corresponding to the image content. The mask sub-image can be accurate to the pixel level of each candidate region, which is beneficial to the high-precision extraction of image content and text content.
[0040] Step S103: Separate the images and text in the document image according to the category label, bounding box position, and mask sub-image corresponding to at least one candidate region, to obtain the separated image content and text content.
[0041] Among them, based on the bounding box position corresponding to each candidate region, the candidate region can be accurately located, and based on its corresponding category label, it can be determined whether the candidate region contains text content or image content. Finally, based on the mask sub-image corresponding to the candidate region, the image content or text content can be correspondingly extracted, so as to complete the separation of text and images in the document image.
[0042] As described above, obtain a document image, input the document image into a trained image processing network model, and obtain an output result. The output result includes class labels, bounding box positions, and mask sub-images corresponding to at least one candidate region in the document image. The candidate region is a region containing text or an image. Separate the image and text in the document image according to the class labels, bounding box positions, and mask sub-images corresponding to at least one candidate region to obtain the separated image content and text content. In the above solution, by inputting the document image into a trained image processing network model, the processing capabilities of the image processing network model can be effectively utilized to obtain key reference information for separating text content and image content. By separating the image and text in the document image according to the class labels, bounding box positions, and mask sub-images corresponding to at least one candidate region, the image content and text content in the document image can be accurately extracted, high-precision image-text separation can be completed, and it can be effectively deployed in application scenarios such as complex backgrounds, interleaved images and texts, or blurred text regions.
[0043] Figure 2 FIG. is a flowchart of an image-text separation method including a specific process of processing a document image by an image processing network model provided in an embodiment of the present application. The trained image processing network model includes a trained feature extraction network, a trained region proposal network, a trained classification network, a trained regression network, and a trained mask generation network. As Figure 2 shown, the image-text separation method for the document image specifically includes the following steps:
[0044] Step S201, obtain a document image.
[0045] Step S202, input the document image into the trained feature extraction network to obtain a first feature image, and input the first feature image into the trained region proposal network to obtain a second feature image, where at least one candidate region is marked in the second feature image.
[0046] Among them, the feature extraction network can be ResNet (Residual Network), FPN (Feature Pyramid Network), etc., which are not limited in this application. The feature extraction network can be used to extract key information of the document image, such as edges, textures, shapes, etc., and remove a large amount of redundant information, which is beneficial to reducing the computational complexity and enhancing the generalization ability of the region proposal network. Specifically, the residual network can include an input layer, a convolutional layer, a residual block, a pooling layer, and an output layer. Specifically, the input layer can receive a standard document image as input, the convolutional layer can use a series of convolutional operations to extract the features of the standard document image, the residual block can be used to learn the residual between the input and the output, the pooling layer can reduce the spatial resolution of the feature map while increasing its receptive field, and the output layer can be used to output the first feature image. The feature pyramid network can include a backbone network, a bottom-up path, a top-down path, lateral connections, and an output layer. Among them, the backbone network can extract feature maps of different levels, the bottom-up path can extract feature maps layer by layer from the lower levels to the higher levels of the backbone network, the top-down path can restore the spatial resolution of the feature maps of the higher levels to match that of the lower levels through upsampling, and then fuse these upsampled feature maps with the feature maps of the lower levels. The lateral connections can use 1x1 convolutions in the top-down path to adjust the number of channels of the feature maps of the lower levels so as to fuse with the upsampled feature maps. The output layer can be used to output the first feature image. The region proposal network can be used to recommend at least one candidate region of a text region that may contain text content or an image region that contains image content in the document image. Specifically, the region proposal network can include an input layer, a sliding window and anchor box layer, a classification layer, a regression layer, and an output layer. Among them, the input layer can be used to receive the first feature image, the sliding window and anchor box layer can be used to apply a sliding window on the first feature image to generate a series of fixed-size anchor boxes, each anchor box corresponding to a window position on the first feature image and having different scales and aspect ratios to adapt to targets of different sizes and shapes. The classification layer can be used to output a target score for classifying whether the anchor box contains text or an image. The regression layer can be used to adjust the position and size of the anchor points to more precisely match the position of the target object. The output layer can use non-maximum suppression to filter out overlapping anchor boxes and retain some of the higher-scoring anchor boxes as the final candidate regions, thereby obtaining a second feature image marked with multiple candidate regions, which is beneficial to identifying potential text regions or image regions.Optionally, the region proposal network model can be trained based on a sample document image marked with text regions and image regions, and a multi-task loss function can be used to simultaneously optimize the classification and regression tasks corresponding to the region proposal network model. Among them, the loss function of the classification task can adopt the cross-entropy loss function, which is used to measure the difference between the predicted probability of the candidate region and the true label, while the loss function of the regression task can adopt the smooth L1 loss function, which is used to measure the difference between the predicted bounding box of the candidate region and the true bounding box.
[0047] Step S203: Input the second feature image into the trained classification network to obtain the class label corresponding to each candidate region in the second feature image.
[0048] Among them, the classification network can be used to determine whether the candidate region contains text content or image content and give the class label corresponding to the candidate region. The classification network can be a network containing one or more fully connected layers, or a convolutional layer followed by a fully connected layer, which is used to output the class label corresponding to each candidate region in the second feature image. For example, a text label or an image label. Therefore, through the classification function of the classification network model, each candidate region in the second feature image can be accurately identified as a text region or an image region.
[0049] Step S204: Input the second feature image into the trained regression network to adjust the bounding box of each candidate region in the second feature image to obtain the third feature image and the bounding box position of each candidate region.
[0050] Among them, the regression network can be a network containing one or more fully connected layers that share some parameters with the classification network, or an independent convolutional layer followed by a fully connected layer, which is used to adjust the bounding box of each candidate region in the second feature image and obtain a more accurate bounding box position. Therefore, through the regression function of the regression network model, the boundary position of each candidate region can be precisely adjusted to obtain the third feature image, so as to more accurately divide the text region and the image region.
[0051] Step S205: Input the third feature image and the class label corresponding to each candidate region into the trained mask generation network to obtain the mask sub-image corresponding to each candidate region in the third feature image.
[0052] Among them, the mask generation network can include one or more convolutional layers and an upsampling layer, which can generate a binary mask sub-image for each classified candidate region, and use pixel-level segmentation ability to distinguish text content and image content. The size of the output mask sub-image is the same as that of the candidate region, which can provide support for the subsequent extraction of text content and image content.
[0053] Step S206: Separate the images and text in the document image according to the class labels, bounding box positions, and mask sub-images corresponding to at least one candidate region, to obtain the separated image content and text content.
[0054] As described above, by inputting the document image into the trained feature extraction network, the key features of the document image can be effectively extracted. By inputting the first feature image into the trained region proposal network, the potential text regions and image regions in the document image can be effectively identified, providing reliable reference information for subsequent accurate classification and region separation. By inputting the second feature image into the trained classification network, it can be accurately identified whether each candidate region belongs to a text region or an image region. Moreover, by inputting the second feature image into the trained regression network, the bounding boxes of each candidate region can be effectively adjusted to accurately locate the text regions and image regions. By inputting the third feature image and the class labels corresponding to each candidate region into the trained mask generation network, the target content of each candidate region can be accurately pixel-level labeled, which is beneficial to subsequent accurate separation of text content and image content.
[0055] Figure 3 The figure is a flowchart of a specific training process of the classification network in step S203 provided by the embodiment of the present application. It should be noted that the training process of the classification network shown in Figure 3 occurs at least before step S203. As shown in Figure 3 the training process of the classification network in step S203 specifically includes the following steps:
[0056] Step S301: Obtain a sample document image and the class labels corresponding to multiple sample regions in the sample document image, where the class labels include text labels and image labels.
[0057] Among them, the sample document image may include multiple sample regions, and the multiple sample regions include text regions and image regions. For the text regions, text labels are correspondingly set, and for the image regions, image labels are correspondingly set.
[0058] Step S302: Input the sample document image into the pre-constructed classification network to obtain the classification probability results corresponding to each sample region in the sample document image.
[0059] Among them, the pre-constructed classification network can output the classification probability results for dividing each sample region into a text region or an image region, which is used to characterize the likelihood of the sample region belonging to a text region or an image region.
[0060] Step S303: Calculate the first loss function based on the classification probability results and class labels corresponding to each sample region, and perform iterative optimization of the parameters of the classification network based on the calculated first loss value until the classification network converges to obtain the optimal classification network. The optimal classification network is the trained classification network.
[0061] Among them, the first loss function can be a cross-entropy loss function, etc., and this application does not make a limitation here. Specifically, the calculation formula of the cross-entropy loss function is as follows:
[0062]
[0063] Among them, L cls is the first loss value, y i is the class label, is the classification probability result predicted by the classification network. Using the cross-entropy loss function can effectively measure the difference between the classification probability result and the actual class label, and perform reasonable iterative optimization of the parameters of the classification network until the optimal classification network is obtained.
[0064] This optimal classification network is used as the trained classification network in step S203 to process the second feature image and obtain the class labels corresponding to each candidate region in the second feature image.
[0065] Based on the sample document image and the class labels corresponding to multiple sample regions in the sample document image, training the classification network can enable the classification network to have the ability to distinguish text regions and image regions. Furthermore, in the actual application process, the purpose of accurately distinguishing text regions and image regions can be achieved.
[0066] Figure 4 This is a flowchart of a specific training process of the regression network in step S204 provided by the embodiment of this application. It should be noted that for Figure 4 the training process of the regression network shown in step S204 occurs at least before step S204. As Figure 4 shown, the training process of the regression network in step S204 specifically includes the following steps:
[0067] Step S401: Obtain the sample document image and the positions of the reference bounding boxes corresponding to multiple sample regions in the sample document image. The multiple sample regions include text regions and image regions.
[0068] Among them, the sample document image can include multiple sample regions, and each sample region is marked with a corresponding reference bounding box in the sample document image. The position of the reference bounding box can be the bounding box coordinates, for example, the coordinates of the upper left corner and the lower right corner.
[0069] Step S402: Input the sample document image into the pre-constructed regression network to obtain the predicted bounding box positions corresponding to each sample region in the sample document image.
[0070] Among them, the pre-constructed regression network can output the predicted bounding box positions corresponding to each sample region, and the predicted bounding box positions can also be the bounding box coordinates. For example, the coordinates of the upper left corner and the lower right corner, using the same position description as the reference bounding box positions.
[0071] Step S403: Calculate the second loss function based on the predicted bounding box positions and the reference bounding box positions corresponding to each sample region, and perform iterative optimization of the parameters of the regression network based on the calculated second loss value until the regression network converges to obtain the optimal regression network, and the optimal regression network is the trained regression network.
[0072] Among them, the second loss function can be the least absolute deviation loss function, etc., which is not limited in this application. Specifically, the calculation formula of the least absolute deviation loss function is as follows:
[0073] L bbox =∑smooth L1 (b,b gt )
[0074] Among them, L bbox is the second loss value, b is the predicted bounding box position, and b gt is the reference bounding box position. Using the least absolute deviation loss function can effectively measure the difference between the predicted bounding box position and the reference bounding box position, and perform reasonable iterative optimization of the parameters of the regression network until the optimal regression network is obtained.
[0075] This optimal regression network is used as the trained regression network in step S204 to process the second feature image to adjust the bounding boxes of each candidate region in the second feature image to obtain the third feature image and the bounding box positions of each candidate region.
[0076] As described above, training the regression network based on the sample document image and the reference bounding box positions corresponding to multiple sample regions in the sample document image can enable the regression network to have the ability to accurately locate the sample regions. Furthermore, in the actual application process, the bounding boxes of the candidate regions in the document image can be accurately adjusted to accurately locate the text regions and image regions.
[0077] Figure 5 This is a flowchart of a specific training process of the mask generation network in step S205 provided by the embodiments of the present application. It should be noted that the training process of the mask generation network shown in Figure 5 step S205 occurs at least before step S205. AsFigure 5 As shown, the training process of the mask generation network in step S205 specifically includes the following steps:
[0078] Step S501: Obtain a sample document image and reference mask sub - graphs corresponding to multiple sample regions in the sample document image.
[0079] Among them, the sample document image may include multiple sample regions. The multiple sample regions include text regions and image regions, and each sample region is set with a pixel - level reference mask sub - graph. Exemplarily, the reference mask sub - graph may be a binary image. For a text region, its corresponding mask value may be 1, and for an image region, its corresponding mask value may be 0.
[0080] Step S502: Input the sample document image into a pre - constructed mask generation network to obtain a predicted mask sub - graph corresponding to each sample region in the sample document image.
[0081] Among them, the pre - constructed mask generation network can output a predicted mask sub - graph corresponding to each sample region. Exemplarily, the predicted mask sub - graph may be a binary image. For a text region, its corresponding mask value may be 1, and for an image region, its corresponding mask value may be 0.
[0082] Step S503: Calculate a third loss function based on the predicted mask sub - graph and the reference mask sub - graph corresponding to each sample region, and perform parameter iterative optimization of the mask generation network based on the calculated third loss value until the mask generation network converges to obtain an optimal mask generation network. The optimal mask generation network is the trained mask generation network.
[0083] Among them, the third loss function may be a binary cross - entropy loss function, etc., which is not limited in this application. Specifically, the calculation formula of the binary cross - entropy loss function is as follows:
[0084]
[0085] where L mask is the third loss value, y is the reference mask sub - graph, is the predicted mask sub - graph. Using the binary cross - entropy loss function can effectively measure the difference between the reference mask sub - graph and the predicted mask sub - graph, and perform reasonable parameter iterative optimization on the mask generation network until an optimal mask generation network is obtained.
[0086] This optimal mask generation network is used as the trained mask generation network in step S205 to process the third feature image and the class label corresponding to each candidate region, and obtain a mask sub - graph corresponding to each candidate region in the third feature image.
[0087] As described above, by training the mask generation network based on the sample document image and the reference mask sub - graphs corresponding to multiple sample regions in the sample document image, the mask generation network can be enabled to accurately mark the text content or image content at the pixel level for the text region or image region. Furthermore, in the actual application process, it is possible to accurately distinguish whether the specific pixels in the candidate region of the document image belong to the text content or the image content, which is convenient for subsequent extraction and separation of the text content and the image content.
[0088] Figure 6 FIG. is a flowchart of a joint training process of the classification network, regression network, and mask generation network corresponding to steps S203, S204, and S205 provided in the embodiments of the present application. It should be noted that Figure 6 for the joint training process of the classification network, regression network, and mask generation network corresponding to steps S203, S204, and S205 shown at least occurs before step S203. As Figure 6 shown, the joint training process of the classification network, regression network, and mask generation network corresponding to steps S203, S204, and S205 specifically includes the following steps:
[0089] Step S601: Obtain a sample document image, and the class labels, reference bounding box positions, and reference mask sub - graphs corresponding to multiple sample regions in the sample document image.
[0090] Among them, the class labels corresponding to the sample document image and multiple sample regions in the sample document image can be used as the training data of the classification network, the reference bounding box positions corresponding to the sample document image and multiple sample regions in the sample document image can be used as the training data of the regression network, and the reference mask sub - graphs corresponding to the sample document image and multiple sample regions in the sample document image can be used as the training data of the mask generation network.
[0091] Step S602: Based on the sample document image, and the class labels, reference bounding box positions, and reference mask sub - graphs corresponding to multiple sample regions in the sample document image, perform joint training on the pre - constructed classification network, regression network, and mask generation network after weighted combination of loss functions until the optimal classification network, regression network, and mask generation network are obtained. The optimal classification network, regression network, and mask generation network are the trained classification network, regression network, and mask generation network respectively.
[0092] Among them, during the training process, the loss functions corresponding to the classification network, regression network, and mask generation network can be weighted and combined through multi-task learning, so as to achieve the joint optimization of the classification network, regression network, and mask generation network. Specifically, the loss functions of different networks can be weighted according to the importance and difficulty of the training tasks. The relevant formula is as follows:
[0093] L = L cls + λ 1 L bbox + λ 2 L mask ,
[0094] Among them, L is the joint loss value, L cls is the loss value calculated by using the cross-entropy loss function during the training of the classification network, L bbox is the loss value calculated by using the least absolute deviation loss function during the training of the regression network, L mask is the loss value calculated by using the binary cross-entropy loss function during the training of the mask generation network, λ 1 and λ 2 are hyperparameters used to balance the losses of different training tasks.
[0095] The optimal classification network, regression network, and mask generation network are used as the trained classification network, regression network, and mask generation network corresponding to steps S203, S204, and S205 respectively, and are used to process the second feature image to obtain the class labels, bounding box positions, and mask subgraphs corresponding to at least one candidate region.
[0096] As described above, by weighted-combining the loss functions of different training tasks, the overall performance of multiple networks can be improved, information sharing between tasks can be promoted, overfitting can be avoided, and the generalization ability can be improved. At the same time, flexibility and stability are provided, which is beneficial to better complete the separation of subsequent text content and image content.
[0097] Figure 7 The figure is a flowchart of a text-image separation method provided by an embodiment of the present application, which includes a process of content extraction and processing of a document image. As Figure 7 shown, the text-image separation method of the document image specifically includes the following steps:
[0098] Step S701, obtain a document image.
[0099] Step S702, input the document image into the trained image processing network model to obtain an output result, where the output result includes the class labels, bounding box positions, and mask subgraphs corresponding to at least one candidate region in the document image, and the candidate region is a region containing text or an image.
[0100] Step S703: Based on the class labels, bounding box positions, and mask sub - graphs corresponding to at least one candidate region.
[0101] Step S704: When the class label corresponding to the candidate region is a text label, perform mask extraction based on the mask sub - graph and bounding box position corresponding to the candidate region to obtain the text content.
[0102] Step S705: When the class label corresponding to the candidate region is an image label, perform inverse mask extraction based on the mask sub - graph and bounding box position corresponding to the candidate region to obtain the image content.
[0103] Wherein, in this embodiment, the mask sub - graph of the candidate region can be a binary mask. For the text region, the pixel points with pixel value 1 in the mask sub - graph belong to the text content. For the image region, the pixel points with pixel value 0 in the mask sub - graph belong to the image content. Thus, if the class label corresponding to the candidate region is a text label, mask extraction can be directly performed based on the mask sub - graph and bounding box position to obtain the text content. If the class label corresponding to the candidate region is an image label, inverse mask extraction can be performed based on the mask sub - graph and bounding box position to obtain the image content, thereby realizing the separation of text and images in the document image.
[0104] As described above, by performing corresponding content extraction processing adapted to the class label corresponding to each candidate region, the pixel region containing the text content can be effectively obtained through mask extraction, and the pixel region containing the image content can be effectively obtained through inverse mask extraction, realizing the accurate separation of text content and image content in the document image.
[0105] Figure 8 The figure is a flowchart of a text - image separation method provided by an embodiment of the present application, which includes a process of processing text content and image content. As Figure 8 shown, the text - image separation method for this document image specifically includes the following steps:
[0106] Step S801: Obtain the document image.
[0107] Step S802: Input the document image into the trained image - processing network model to obtain the output result. The output result includes the class labels, bounding box positions, and mask sub - graphs corresponding to at least one candidate region in the document image, and the candidate region is a region containing text or an image.
[0108] Step S803: Based on the class labels, bounding box positions, and mask sub - graphs corresponding to at least one candidate region, separate the images and text in the document image to obtain the separated image content and text content.
[0109] Step S804: Perform text enhancement processing on the text content to obtain the target text content.
[0110] Among them, the text enhancement process can be the histogram equalization algorithm or the homomorphic filtering algorithm, etc., which is not limited in this application.
[0111] Step S805: Perform image enhancement processing on the image content to obtain the target image content.
[0112] Among them, the image enhancement processing can be performed by a deep learning module based on the super-resolution algorithm to improve the resolution of the image content, or can be based on the generative adversarial network model to improve the clarity of the image content, which is not limited in this application.
[0113] Step S806: Combine the target image content and the target text content to obtain the target document image.
[0114] As described above, by separately performing enhancement processing on the text content and the image content, and recombining the target image content and the target text content to obtain the target document image, the document image reconstruction can be effectively performed, and the quality of the document image can be improved.
[0115] Figure 9 It is a flowchart of a text-image separation method provided by an embodiment of this application, which includes a process of preprocessing a document image. As Figure 9 shown, the text-image separation method of this document image specifically includes the following steps:
[0116] Step S901: Obtain a document image.
[0117] Step S902: Perform preprocessing on the document image to obtain a standard document image. The preprocessing includes size adjustment and denoising processing.
[0118] Among them, the size adjustment can be used to adjust document images with inconsistent sizes to a unified specification. The denoising processing can be Gaussian filtering, mean filtering, Laplace filtering, etc., which is not limited in this application.
[0119] Step S903: Input the standard document image into the trained image processing network model to obtain an output result. Among them, the output result includes the class label, bounding box position, and mask sub-image corresponding to at least one candidate region in the document image. The candidate region is a region containing text or an image.
[0120] Step S904: Separate the image and text in the document image according to the class label, bounding box position, and mask sub-image corresponding to at least one candidate region to obtain the separated image content and text content.
[0121] As described above, preprocessing the document image to obtain a standard document image can ensure the quality of the document image input into the image processing network model, thereby improving the accuracy and robustness of the image processing network model.
[0122] Figure 10 The following is a structural block diagram of a device for separating text and images from a document image provided by an embodiment of the present application. The device is configured to execute the method for separating text and images from a document image provided by the above embodiment, and has corresponding functional modules and beneficial effects for executing the method. As Figure 10 shown, the device specifically includes:
[0123] An acquisition module 101, configured to acquire a document image.
[0124] An image processing module 102, configured to input the document image into a trained image processing network model to obtain an output result, where the output result includes class labels, bounding box positions, and mask sub-images corresponding to at least one candidate region in the document image, and the candidate region is a region containing text or an image.
[0125] A text and image separation module 103, configured to separate the image and text in the document image according to the class labels, bounding box positions, and mask sub-images corresponding to at least one candidate region, to obtain the separated image content and text content.
[0126] As described above, acquire a document image, input the document image into a trained image processing network model to obtain an output result, where the output result includes class labels, bounding box positions, and mask sub-images corresponding to at least one candidate region in the document image, and the candidate region is a region containing text or an image. Separate the image and text in the document image according to the class labels, bounding box positions, and mask sub-images corresponding to at least one candidate region to obtain the separated image content and text content. In the above solution, by inputting the document image into a trained image processing network model, the processing ability of the image processing network model can be effectively utilized to obtain key reference information for separating text content and image content. By separating the image and text in the document image according to the class labels, bounding box positions, and mask sub-images corresponding to at least one candidate region, the image content and text content in the document image can be accurately extracted to complete high-precision text and image separation, and it can be effectively deployed in application scenarios such as complex backgrounds, text and image interleaving, or blurred text regions.
[0127] In a possible embodiment, the trained image processing network model includes a trained feature extraction network, a trained region proposal network, a trained classification network, a trained regression network, and a trained mask generation network;
[0128] The image processing module 102 is further configured to:
[0129] Input the document image into the trained feature extraction network to obtain the first feature image;
[0130] Input the first feature image into the trained region proposal network to obtain the second feature image, with at least one candidate region marked in the second feature image;
[0131] Input the second feature image into the trained classification network to obtain the class label corresponding to each candidate region in the second feature image;
[0132] Input the second feature image into the trained regression network to adjust the bounding boxes of each candidate region in the second feature image, obtaining the third feature image and the bounding box positions of each candidate region;
[0133] Input the third feature image and the class label corresponding to each candidate region into the trained mask generation network to obtain the mask sub - image corresponding to each candidate region in the third feature image.
[0134] In a possible embodiment, it further includes a classification network training module, configured to:
[0135] Obtain the sample document image and the class labels corresponding to multiple sample regions in the sample document image, where the class labels include text labels and image labels;
[0136] Input the sample document image into the pre - constructed classification network to obtain the classification probability result corresponding to each sample region in the sample document image;
[0137] Calculate the first loss function according to the classification probability result corresponding to each sample region and the class label, and perform parameter iterative optimization of the classification network based on the calculated first loss value until the classification network converges to obtain the optimal classification network, and the optimal classification network is the trained classification network.
[0138] In a possible embodiment, it further includes a regression network training module, configured to:
[0139] Obtain the sample document image and the reference bounding box positions corresponding to multiple sample regions in the sample document image, where the multiple sample regions include text regions and image regions;
[0140] Input the sample document image into the pre - constructed regression network to obtain the predicted bounding box position corresponding to each sample region in the sample document image;
[0141] Calculate the second loss function according to the predicted bounding box positions and reference bounding box positions corresponding to each sample region, and perform iterative optimization of the parameters of the regression network based on the calculated second loss value until the regression network converges to obtain the optimal regression network, where the optimal regression network is the trained regression network.
[0142] In a possible embodiment, it further includes a mask generation network training module, configured to:
[0143] Obtain the sample document image and the reference mask sub-images corresponding to multiple sample regions in the sample document image respectively;
[0144] Input the sample document image into the pre-constructed mask generation network to obtain the predicted mask sub-images corresponding to each sample region in the sample document image;
[0145] Calculate the third loss function according to the predicted mask sub-images and reference mask sub-images corresponding to each sample region, and perform iterative optimization of the parameters of the mask generation network based on the calculated third loss value until the mask generation network converges to obtain the optimal mask generation network, where the optimal mask generation network is the trained mask generation network.
[0146] In a possible embodiment, it further includes a network joint training module, configured to:
[0147] Obtain the sample document image, and the class labels, reference bounding box positions and reference mask sub-images corresponding to multiple sample regions in the sample document image respectively;
[0148] Based on the sample document image, and the class labels, reference bounding box positions and reference mask sub-images corresponding to multiple sample regions in the sample document image respectively, perform joint training on the pre-constructed classification network, regression network and mask generation network after weighted combination of the loss functions until the optimal classification network, regression network and mask generation network are obtained, where the optimal classification network, regression network and mask generation network are the trained classification network, regression network and mask generation network respectively.
[0149] In a possible embodiment, the graphic and text separation module 103 is further configured to:
[0150] When the class label corresponding to the candidate region is a text label, perform mask extraction based on the mask sub-image and bounding box position corresponding to the candidate region to obtain the text content;
[0151] When the class label corresponding to the candidate region is an image label, perform inverse mask extraction based on the mask sub-image and bounding box position corresponding to the candidate region to obtain the image content.
[0152] In a possible embodiment, it further includes an enhancement combination module, configured to:
[0153] Perform text enhancement processing on the text content to obtain the target text content;
[0154] Perform image enhancement processing on the image content to obtain the target image content;
[0155] Combine the target image content and the target text content to obtain the target document image.
[0156] In a possible embodiment, it further includes an image preprocessing module configured to:
[0157] Perform preprocessing on the document image to obtain a standard document image, and the preprocessing includes size adjustment and denoising processing;
[0158] The image processing module 102 is further configured to:
[0159] Input the standard document image into the trained image processing network model.
[0160] The embodiment of the present application further provides a chip, which includes a processor and a memory; the number of processors in the chip can be one or more. The memory, as a computer-readable storage medium, can be configured to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the text and image separation method of the document image in the embodiment of the present application. The processor executes various functional applications and data processing of the device by running the software programs, instructions, and modules stored in the memory, that is, implements the above-mentioned text and image separation method of the document image.
[0161] Figure 11 It is a schematic structural diagram of an electronic device provided by the embodiment of the present application, as Figure 11 shown, the device includes the chip 201, the input device 202, and the output device 203 provided in the foregoing embodiment; the number of processors 2011 in the chip 201 can be one or more, and a memory 2012 is further provided in the chip 201. Figure 11 Taking one processor 2011 as an example; the chip 201, the input device 202, and the output device 203 in the device can be connected through a bus or other means. Figure 11 Taking the connection through the bus as an example. The input device 202 can be configured to receive input digital or character information, and generate key signal inputs related to the user settings and function controls of the device. The output device 203 may include a display device such as a display screen.
[0162] In one embodiment, the electronic device can be an MFP (Multi-Functional Peripheral), a scanner, etc., to process the image obtained by scanning a paper document using the method for separating text and graphics of a document image provided in the foregoing embodiment.
[0163] In other embodiments, the electronic device can also be a terminal device such as a mobile phone, a PC (Personal Computer), a tablet computer, etc., to process the document image received from other terminal devices or the document image stored internally using the method for separating text and graphics of a document image provided in the foregoing embodiment.
[0164] It should be noted that the electronic device provided in the embodiments of the present application is not limited to the examples mentioned above.
[0165] The electronic device provided above can be used to execute the method for separating text and graphics of a document image provided in any of the foregoing embodiments, and has the corresponding functions and beneficial effects.
[0166] The embodiments of the present application also provide a non-volatile storage medium containing computer-executable instructions. The computer-executable instructions are configured to execute a method for separating text and graphics of a document image described in any of the foregoing embodiments when executed by a computer processor. The method includes: obtaining a document image; inputting the document image into a trained image processing network model to obtain an output result, where the output result includes class labels, bounding box positions, and mask subgraphs corresponding to at least one candidate region in the document image, and the candidate region is a region containing text or an image; separating the image and text in the document image according to the class labels, bounding box positions, and mask subgraphs corresponding to at least one candidate region to obtain the separated image content and text content.
[0167] Storage medium - Any of various types of memory devices or storage devices. The term "storage medium" is intended to include: installation media such as CD-ROMs, floppy disks, or magnetic tape devices; computer system memory or random access memory such as DRAM, DDR RAM, SRAM, EDO RAM, Rambus RAM, etc.; non-volatile memory such as flash memory, magnetic media, or optical storage; registers or other similar types of memory elements, etc. The storage medium may also include other types of memory or combinations thereof. Additionally, the storage medium may be located in the first computer system in which the program is executed, or may be located in a different second computer system that is connected to the first computer system via a network (such as the Internet). The second computer system may provide program instructions to the first computer for execution. The term "storage medium" may include two or more storage media residing in different locations (e.g., in different computer systems connected via a network). The storage medium may store program instructions executable by one or more processors (e.g., embodied as a computer program).
[0168] Of course, for a storage medium containing computer-executable instructions provided by an embodiment of the present application, the computer-executable instructions are not limited to the above-described method for separating text and graphics in a document image, and may also perform related operations in the method for separating text and graphics in a document image provided by any embodiment of the present application.
[0169] It should be noted that in the embodiments of the above-described apparatus for separating text and graphics in a document image, the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; additionally, the specific names of the functional units are only for the convenience of mutual distinction and are not configured to limit the protection scope of the embodiments of the present application.
[0170] It should be noted that the numbering of the steps in this solution is only used to describe the overall design framework of this solution and does not represent an inevitable sequential relationship between the steps. On the basis that the overall implementation process conforms to the overall design framework of this solution, it all falls within the protection scope of this solution. The sequential order in the text form during description is not an exclusive limitation on the specific implementation process of this solution. Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory. The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM), and / or non-volatile memory in the form of, for example, read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0171] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising the element.
[0172] Note that the above is only a preferred embodiment of the present invention and the technical principles applied. Those skilled in the art will understand that the present invention is not limited to the specific embodiments here, and various obvious changes, re-adjustments and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in more detail through the above embodiments, the present invention is not limited to the above embodiments only. Without departing from the concept of the present invention, more other equivalent embodiments can be included, and the scope of the present invention is determined by the scope of the appended claims.
Claims
1. A method for separating text and images from a document image, characterized in that: The method comprises: Acquire document images; Inputting the document image into a trained image processing network model to obtain an output result, wherein the output result includes a category label, a bounding box position, and a mask sub-image corresponding to at least one candidate region in the document image, wherein the candidate region is a region containing text or an image; The image and text in the document image are separated according to the category label, the bounding box position and the mask sub-image corresponding to the at least one candidate region to obtain separated image content and text content.
2. The image-text separation method according to claim 1, characterized in that: The trained image processing network model includes a trained feature extraction network, a trained region proposal network, a trained classification network, a trained regression network and a trained mask generation network; The document image is input into the trained image processing network model to obtain an output result, which includes: Inputting the document image into a trained feature extraction network to obtain a first feature image; Inputting the first feature image into a trained region proposal network to obtain a second feature image, wherein at least one candidate region is marked in the second feature image; Inputting the second feature image into the trained classification network to obtain a category label corresponding to each of the candidate regions in the second feature image; Inputting the second feature image into the trained regression network to adjust the bounding box of each candidate region in the second feature image to obtain a third feature image and a bounding box position of each candidate region; The third feature image and the category label corresponding to each of the candidate regions are input into the trained mask generation network to obtain a mask sub-image corresponding to each of the candidate regions in the third feature image.
3. The image-text separation method according to claim 2, characterized in that: Before inputting the second feature image into the trained classification network to obtain the category label corresponding to each candidate region in the second feature image, the method further includes: Obtaining a sample document image and category labels corresponding to a plurality of sample regions in the sample document image, wherein the category labels include text labels and image labels; Inputting the sample document image into a pre-built classification network to obtain a classification probability result corresponding to each sample area in the sample document image; A first loss function is calculated according to the classification probability results and category labels corresponding to each of the sample regions, and parameters of the classification network are iteratively optimized based on the calculated first loss value until the classification network converges to obtain an optimal classification network, and the optimal classification network is the trained classification network.
4. The image-text separation method according to claim 2, characterized in that: Before inputting the second feature image into the trained regression network to adjust the bounding box of each candidate region in the second feature image to obtain the third feature image and the bounding box position of each candidate region, the method further includes: Acquire a sample document image and reference bounding box positions respectively corresponding to a plurality of sample areas in the sample document image, wherein the plurality of sample areas include a text area and an image area; Inputting the sample document image into a pre-built regression network to obtain a predicted bounding box position corresponding to each sample area in the sample document image; A second loss function is calculated according to the predicted bounding box position and the reference bounding box position corresponding to each of the sample areas, and parameters of the regression network are iteratively optimized based on the calculated second loss value until the regression network converges to obtain an optimal regression network, and the optimal regression network is the trained regression network.
5. The image-text separation method according to claim 2, characterized in that: Before inputting the third feature image and the category label corresponding to each of the candidate regions into the trained mask generation network to obtain the mask sub-image corresponding to each of the candidate regions in the third feature image, the method further includes: Acquire a sample document image and reference mask sub-images respectively corresponding to a plurality of sample areas in the sample document image; Inputting the sample document image into a pre-built mask generation network to obtain a predicted mask sub-image corresponding to each sample area in the sample document image; A third loss function is calculated according to the predicted mask sub-image and the reference mask sub-image corresponding to each of the sample areas, and parameters of the mask generation network are iteratively optimized based on the calculated third loss value until the mask generation network converges to obtain an optimal mask generation network, and the optimal mask generation network is the trained mask generation network.
6. The image-text separation method according to claim 2, characterized in that: Before inputting the second feature image into the trained classification network to obtain the category label corresponding to each candidate region in the second feature image, the method further includes: Obtaining a sample document image, and category labels, reference bounding box positions, and reference mask sub-images respectively corresponding to a plurality of sample regions in the sample document image; Based on the sample document image, and the category labels, reference bounding box positions and reference mask sub-images respectively corresponding to multiple sample areas in the sample document image, the pre-constructed classification network, regression network and mask generation network are jointly trained after weighted merging of loss functions until the optimal classification network, regression network and mask generation network are obtained, and the optimal classification network, regression network and mask generation network are respectively the trained classification network, regression network and mask generation network.
7. The image-text separation method according to claim 1, characterized in that: The step of separating the image and text in the document image to obtain separated image content and text content includes: When the category label corresponding to the candidate region is a text label, performing mask extraction based on the mask sub-image and the bounding box position corresponding to the candidate region to obtain text content; In the case where the category label corresponding to the candidate region is an image label, reverse mask extraction is performed based on the mask sub-image and the bounding box position corresponding to the candidate region to obtain the image content.
8. The image-text separation method according to claim 1, characterized in that: After separating the image and text in the document image to obtain the separated image content and text content, the method further includes: Performing text enhancement processing on the text content to obtain target text content; Performing image enhancement processing on the image content to obtain target image content; The target image content and the target text content are combined to obtain a target document image.
9. The image-text separation method according to claim 1, characterized in that: Before inputting the document image into the trained image processing network model, the method further includes: Preprocessing the document image to obtain a standard document image, wherein the preprocessing includes resizing and denoising; The step of inputting the document image into a trained image processing network model comprises: The standard document image is input into the trained image processing network model.
10. A device for separating image and text from a document image, characterized in that: include: An acquisition module configured to acquire a document image; An image processing module, configured to input the document image into a trained image processing network model to obtain an output result, wherein the output result includes a category label, a bounding box position, and a mask sub-image corresponding to at least one candidate region in the document image, wherein the candidate region is a region containing text or an image; The image-text separation module is configured to separate the image and text in the document image according to the category label, bounding box position and mask sub-image corresponding to the at least one candidate area to obtain separated image content and text content.
11. A chip, comprising: one or more processors; The memory is configured to store one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the method for separating image and text of a document image according to any one of claims 1 to 9.
12. An electronic device comprising the chip according to claim 11.
13. A non-volatile storage medium storing computer executable instructions, wherein the computer executable instructions are configured to execute the method for separating image and text of a document image according to any one of claims 1 to 9 when executed by a computer processor.
Citation Information
Patent Citations
Mixed-pasting bill image processing method, device, computer equipment and storage medium
CN111931664A
Text detection model training method and device and text detection method and device
CN112818975A
Neural network training method, image processing method and device
CN112990211A
Image target detection and instance segmentation method and system, computing equipment and medium
CN115375901A
Pepper rust disease image recognition method and system based on Mask RCNN
CN117422921A