Document Image Shadow Removal Method Based on Foreground Text Guidance and Shadow Probability Learning

CN122841201APending Publication Date: 2026-09-29YUNNAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611037577.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-13
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0004]然而,上述现有技术仍存在以下技术缺陷:传统优化算法对光照条件做了均匀性、平面性等理想化假设,无法适配真实场景中

Benefits of technology

本申请提供了一种前景文本引导与阴影概率学习的文档图像阴影消除方法,通过形态学膨胀运算配合像素差值的初始特征提取框架,由于采用无参数的物理运算实现阴影与文本的初步分离,相较于直接通过网络预测背景与前景的方式,避免了网络对文本与阴影特征混淆导致的特征提取失真问题,同时不增加模型参数数量,为后续网络训练降低特征学习难度;通过阴影概率估计分支基于原始文档图像与含阴影背景图像的配对特征强化阴影区域的特征响应,以及通过前景文本掩码重建分支基于原始文档图像与初始前景图的配对特征保留文本结构信息,从而能够准确识别阴影区域并精准重建文本结构,有效缓解特征提取过程中阴影与文本的混淆问题;通过图像重建网络根据阴影概率图中各像素的概率值对阴影区域进行光照校正,以及通过根据前景文本掩码中文本区域的掩码值对文本区域进行特征保持,实现了阴影区域精准消除与文本区域结构完整保持的协同优化,避免了传统方法中背景重建过程对文本笔画的侵蚀以及文本结构破碎的问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122841201A_ABST
    Figure CN122841201A_ABST
Patent Text Reader

Abstract

This application discloses a document image shadow removal method guided by foreground text and based on shadow probability learning, relating to the fields of computer vision and image processing technology. The method includes: acquiring a document image with shadows; performing morphological dilation on the image to generate a shadowed background image; calculating the pixel-by-pixel difference between the document image and the shadowed background image to obtain an initial foreground image; inputting the document image and the shadowed background image into a shadow probability estimation branch to generate a shadow probability map, and inputting them into a foreground text mask reconstruction branch to generate a foreground text mask; concatenating the document image, the shadow probability map, and the foreground text mask, and inputting the concatenation into an image reconstruction network to generate a shadow-free document image. This application achieves synergistic optimization of accurate shadow region removal and preservation of text region structure in document images, effectively avoiding the erosion of text strokes and structural fragmentation problems caused by background reconstruction in traditional methods.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of computer vision and image processing technology, and in particular to a method for removing shadows in document images based on foreground text guidance and shadow probability learning. Background Technology

[0002] With the rapid development of digital office and intelligent information processing, portable devices have become the mainstream tools for document image acquisition. However, factors such as uneven ambient lighting, occlusion, and document surface deformation during the shooting process can lead to large areas of non-uniform shadows and penumbra transition areas in the image. Shadow areas will significantly reduce text contrast and blur character edges, directly affecting subsequent optical character recognition (OCR).

[0003] Existing methods for removing shadows from document images can be broadly categorized into two types: traditional optimization algorithms and deep learning methods. Traditional optimization algorithms are designed based on prior knowledge such as physical lighting models, gradient characteristics, and region clustering, correcting shadow areas by establishing idealized lighting assumptions. Deep learning methods, on the other hand, focus on estimating the color of the shadow-free background, using neural networks to predict the shadow-free background image, and then using this prediction to guide shadow removal. In addition, there are natural image shadow removal methods, which prioritize visual naturalness, focusing on restoring the lighting consistency and visual realism of shadow areas.

[0004] However, the aforementioned existing technologies still have the following technical shortcomings: Traditional optimization algorithms make idealized assumptions about lighting conditions, such as uniformity and planarity, which cannot be adapted to real-world scenarios. Existing deep learning methods rely excessively on background feature estimation, which can erode text strokes during background reconstruction, leading to fragmented text and structural distortion. Natural image shadow removal methods, with visual naturalness as their core objective, do not match the specific requirements of document image background uniformity, high-fidelity text, and structural integrity, and cannot be directly transferred for application.

[0005] Therefore, there is an urgent need for a shadow removal solution specifically designed for document images that can accurately remove complex shadows while effectively protecting fine-grained structural information such as text strokes and edge corners, thereby improving the readability of document images and the recognition performance of subsequent OCR and other applications. Summary of the Invention

[0006] The purpose of this application is to provide a document image shadow removal method based on foreground text guidance and shadow probability learning. The method extracts the shadowed background image and the initial foreground image through morphological dilation operation, and generates a shadow probability map and a foreground text mask through a dual-branch network. The two are used as dual guidance information to drive the image reconstruction network, achieving synergistic optimization of accurate removal of shadow areas in document images and preservation of text region structure. This effectively avoids the problem of background reconstruction eroding text strokes and structural fragmentation in traditional methods.

[0007] To achieve the above objectives, this application provides the following solution: In a first aspect, this application provides a document image shadow removal method based on foreground text guidance and shadow probability learning, comprising: acquiring a document image with shadows; performing morphological dilation on the document image to generate a shadowed background image; calculating the pixel-wise difference between the document image and the shadowed background image to obtain an initial foreground image; inputting the document image and the shadowed background image into a shadow probability estimation branch to generate a shadow probability map, and inputting the document image and the initial foreground image into a foreground text mask reconstruction branch to generate a foreground text mask; concatenating the document image, the shadow probability map, and the foreground text mask, and inputting the concatenation into an image reconstruction network to generate a shadow-free document image; wherein the image reconstruction network performs illumination correction on the shadowed region based on the probability value of each pixel in the shadow probability map, and performs feature preservation on the text region based on the mask value of the text region in the foreground text mask.

[0008] In one embodiment, the morphological dilation operation employs an adaptively adjusted dilated convolution kernel, the size of which is determined based on the average size of the document characters.

[0009] In one embodiment, generating the image with the shadowed background further includes performing median filtering on the image after the morphological dilation operation.

[0010] In one embodiment, both the shadow probability estimation branch and the foreground text mask reconstruction branch adopt an encoder-decoder network architecture. The encoder uses a ResNet-18 backbone network for multi-scale feature extraction, and the decoder performs upsampling through deconvolution and merges the multi-scale features of the encoder with the upsampled features of the decoder through attention-gated skip connections.

[0011] In one embodiment, the attention gating is implemented according to the following formula: ;in, For attention weights, This is the feature map of the l-th layer of the encoder. It is the Sigmoid activation function. , , For learnable weight matrix, For bias terms, For encoder features, These are decoder features.

[0012] In one embodiment, the image reconstruction network performs illumination correction on the shadow region based on the probability values ​​of each pixel in the shadow probability map, including: adjusting the correction weights at corresponding positions in the image reconstruction network based on the probability values ​​of each pixel in the shadow probability map.

[0013] In one embodiment, the step of preserving the features of the text region based on the mask value of the text region in the foreground text mask includes: locking the feature values ​​of the text region in the image reconstruction network based on the mask value of the text region in the foreground text mask.

[0014] In one embodiment, the method further includes: jointly training the shadow probability estimation branch, the foreground text mask reconstruction branch, and the image reconstruction network using a multi-constraint loss function; wherein the multi-constraint loss function includes reconstruction loss, adversarial loss, shadow probability regression loss, text mask segmentation loss, and text feature consistency loss, and the text mask segmentation loss includes Dice loss and binary focus loss.

[0015] Secondly, this application provides a document image shadow removal device based on foreground text guidance and shadow probability learning, comprising: an image acquisition module for acquiring a document image with shadows; an initial feature extraction module for performing morphological dilation on the document image to generate a shadowed background image; calculating the pixel-wise difference between the document image and the shadowed background image to obtain an initial foreground image; a dual-branch feature estimation module, including a shadow probability estimation branch and a foreground text mask reconstruction branch; the shadow probability estimation branch is used to input the document image and the shadowed background image into an encoder-decoder network to generate a shadow probability map; the foreground text mask reconstruction branch is used to input the document image and the initial foreground image into an encoder-decoder network to generate a foreground text mask; and an image reconstruction module for concatenating the document image, the shadow probability map, and the foreground text mask, and inputting the concatenation into an image reconstruction network to generate a shadow-free document image; wherein the image reconstruction network performs illumination correction on the shadow area based on the probability value of each pixel in the shadow probability map, and performs feature preservation on the text area based on the mask value of the text area in the foreground text mask.

[0016] In one embodiment, the apparatus further includes a joint training module for jointly training the shadow probability estimation branch, the foreground text mask reconstruction branch, and the image reconstruction network using a multi-constraint loss function; the multi-constraint loss function includes reconstruction loss, adversarial loss, shadow probability regression loss, text mask segmentation loss, and text feature consistency loss, and the text mask segmentation loss includes Dice loss and binary focus loss.

[0017] According to the specific embodiments provided in this application, the following technical effects are disclosed: This application provides a document image shadow removal method guided by foreground text and based on shadow probability learning. It utilizes a morphological dilation operation combined with an initial feature extraction framework based on pixel differences. By employing parameterless physical operations to achieve initial separation of shadows and text, it avoids feature extraction distortion caused by the network confusing text and shadow features, compared to directly predicting the background and foreground through the network. Simultaneously, it does not increase the number of model parameters, reducing the difficulty of feature learning for subsequent network training. The shadow probability estimation branch enhances the feature response of the shadow region based on paired features of the original document image and the shadowed background image, while the foreground text mask reconstruction branch preserves text structure information based on paired features of the original document image and the initial foreground image. This enables accurate identification of shadow regions and precise reconstruction of text structure, effectively mitigating the problem of shadow and text confusion during feature extraction. Furthermore, the image reconstruction network performs illumination correction on the shadow region based on the probability values ​​of each pixel in the shadow probability map, and preserves the text region's features based on the mask values ​​of the text region in the foreground text mask. This achieves synergistic optimization of accurate shadow region removal and preservation of text region structure integrity, avoiding the erosion of text strokes and text structure fragmentation problems during background reconstruction in traditional methods.

[0018] Furthermore, the proposed method was tested on public datasets such as RDD, SD7K, and Kligler. The PSNR (Peak Signal-to-Noise Ratio) and SSIM (Structural Similarity) metrics were both superior to existing methods, verifying the dual advantages of the proposed method in terms of shadow removal effect and text fidelity. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 A flowchart illustrating a document image shadow removal method based on foreground text guidance and shadow probability learning, provided as an embodiment of this application; Figure 2 for Figure 1 A detailed flowchart of step S2; Figure 3 for Figure 1 A detailed flowchart of step S6; Figure 4This is a schematic diagram of the functional modules of a document image shadow removal device that uses foreground text guidance and shadow probability learning, provided in an embodiment of this application. Detailed Implementation

[0021] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0022] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0023] In one exemplary embodiment, such as Figure 1 As shown, a document image shadow removal method based on foreground text guidance and shadow probability learning is provided. This method can be directly applied to document image preprocessing scenarios such as optical character recognition (OCR), visually impaired assisted reading, automated document archiving, and electronic case file digitization. The method includes the following steps: S1, Get the document image with shadow.

[0024] S2 performs morphological dilation on the document image to generate an image with a shadowed background.

[0025] S3, calculate the pixel-by-pixel difference between the document image and the image with the shadowed background to obtain the initial foreground image.

[0026] S4: Input the document image and the image with the shadowed background into the shadow probability estimation branch to generate a shadow probability map.

[0027] S5, input the document image and the initial foreground image into the foreground text mask reconstruction branch to generate the foreground text mask.

[0028] S6 concatenates the document image, shadow probability map, and foreground text mask, and the input image reconstruction network generates a shadow-free document image.

[0029] The image reconstruction network performs illumination correction on the shadow region based on the probability values ​​of each pixel in the shadow probability map, and preserves the features of the text region based on the mask values ​​of the text region in the foreground text mask.

[0030] Implementing steps S1 to S6 above, steps S1 to S3 achieve initial separation of shadows and text through parameterless physical operations. Compared to directly predicting the background and foreground through the network, this avoids the feature extraction distortion problem caused by the network confusing text and shadow features, while not increasing the number of model parameters, thus reducing the difficulty of feature learning for subsequent network training. Steps S4 to S5 accurately identify shadow areas and precisely reconstruct the text structure, effectively alleviating the problem of shadow and text confusion during feature extraction. Step S6 achieves synergistic optimization of accurate shadow area elimination and preservation of text area structure, avoiding the problems of text stroke erosion and text structure fragmentation during background reconstruction in traditional methods.

[0031] The implementation process of one embodiment of this application will be described in detail below with reference to specific examples.

[0032] For example, in step S1 above, the document image to be processed is obtained. The document image with shadow can be a document image captured by a portable device, such as a smartphone, tablet or scanner, under conditions of uneven ambient lighting or occlusion. The image contains a large area of ​​non-uniform shadow and penumbra transition area, which leads to reduced text contrast and blurred character edges.

[0033] For example, in step S2 above, morphological dilation is performed on the document image to generate an image with a shadowed background, such as... Figure 2 As shown, the specific steps include the following: S21, the text character region in the document image is expanded by filling it with neighboring background pixels through morphological dilation operation.

[0034] Specifically, the morphological dilation operation uses a square-structured dilation convolution kernel. Through the dilation operation, the text character region in the image is filled and expanded with neighboring background pixels, which can erase the text characters and retain only the spatial distribution features of the shadows and background texture information in the image.

[0035] S22, the size of the dilated convolution kernel is determined based on the average size of the document characters.

[0036] Specifically, for documents with dense small characters, such as small font sizes and dense tables, a 3×3 dilated convolution kernel is used; for documents with sparse large characters, such as title pages and large-font posters, a 5×7 dilated convolution kernel is used. This adaptive adjustment mechanism of the dilated convolution kernel size ensures that while effectively erasing text, the original spatial distribution and brightness attenuation characteristics of shadows are not excessively destroyed.

[0037] S23 performs median filtering on the image after dilation.

[0038] Specifically, the median filter has a 5×5 filter window. Median filtering is used to eliminate edge artifacts and jagged noise introduced by dilation operations in text edge regions, smooth isolated noise points in the image, and make background transitions more natural.

[0039] Therefore, after dilation and median filtering, a shadowed background image is obtained that retains only the spatial distribution of shadows and background features.

[0040] For example, in step S3 above, the pixel-by-pixel difference between the document image and the image with the shadow is calculated to obtain the initial foreground image; specifically: the pixel-by-pixel difference is obtained by subtracting the pixel values ​​at corresponding positions in the original shadowed document image and the image with the shadow, that is: Initial foreground image (pixel values) = Original document image (pixel values) - Background image with shadow (pixel values); Then, the pixel-by-pixel difference results are normalized to map the pixel values ​​to the range [0,1], thus obtaining the initial foreground image.

[0041] Therefore, the initial foreground image filters out most of the shadow interference information in the image through pixel-level subtraction, while retaining high-frequency structural details such as core text strokes and character corners. This physical operation method does not require network parameters, fundamentally avoiding artifacts and feature confusion problems introduced by network prediction, and providing a high-fidelity text structure prior for the subsequent foreground text mask reconstruction branch.

[0042] Steps S2 and S3 achieve initial decoupling between shadow features and text features: the shadowed background image carries the spatial distribution features of the shadow, which is used as input for the subsequent shadow probability map estimation branch; the initial foreground image carries the structural features of the text, which is used as input for the subsequent foreground text mask reconstruction branch. For example, in step S4 above, the document image and the background image with shadow are input into the shadow probability estimation branch to generate a shadow probability map; specifically, the input to the shadow probability estimation branch is a concatenated tensor of the original shadowed document image and the background image with shadow along the channel dimension. In a specific implementation, the original shadow image is a 3-channel RGB image, and the background image with shadow is also a 3-channel RGB image. The two are concatenated to form an input tensor with 6 channels.

[0043] The shadow probability estimation branch employs an encoder-decoder network architecture. In the implementation, the encoding stage uses a ResNet-18 backbone network, downsampling the 6-channel input tensor five times to extract multi-scale features at scales of 1 / 2, 1 / 4, 1 / 8, 1 / 16, and 1 / 32. Each convolutional layer is followed by a batch normalization layer and a ReLU activation function to stabilize the training process and prevent gradient vanishing. In the decoding stage, deconvolution (transposed convolution) is used for upsampling, gradually restoring the spatial resolution of the feature maps to the original image size.

[0044] Between the encoder and decoder, attention-gated skip connections are used to concatenate and fuse feature maps from different scales of the encoder with corresponding upsampled feature maps from the decoder. The attention gating is implemented using the following formula: ; in, For attention weights, This is the feature map of the l-th layer of the encoder. It is the Sigmoid activation function. , , For learnable weight matrix, For bias terms, For encoder features, These are decoder features.

[0045] The working mechanism of attention gating is as follows: encoder features With decoder features Substituting into the attention gating formula above, and through the learnable weight matrix... , Weighted and biased terms After addition, the result is processed by the Sigmoid function. Mapped to the [0,1] range, an attention weight graph is generated. Then the attention weight map With encoder features Element-wise multiplication enables dynamic weighting of encoder features. This mechanism allows the network to adaptively enhance feature response values ​​in shadow regions, suppress interference from redundant background features, and effectively mitigate the confusion between shadow and text features during feature extraction.

[0046] Furthermore, the shadow probability estimation branch introduces convolutional attention blocks during the decoding stage to simultaneously extract channel attention and spatial attention from the concatenated encoder features and decoder upsampled features, thereby further enhancing key region features. Channel attention fuses feature information through global average pooling and global max pooling to generate channel weight vectors; spatial attention generates a spatial weight map through convolution operations; finally, the channel attention weights and spatial attention weights are multiplied element-wise with the feature map to achieve feature enhancement.

[0047] The decoder outputs the upsampled and recovered feature map as a continuous shadow probability map with 1 channel. The value range of each pixel in the shadow probability map is [0,1], which represents the probability value that the corresponding position belongs to the shadow area. The closer the probability value is to 1, the higher the confidence that the position is a shadow. The closer the probability value is to 0, the higher the confidence that the position is a non-shadow area.

[0048] For example, in step S5 above, the document image and the initial foreground image are input into the foreground text mask reconstruction branch to generate a foreground text mask. Specifically, the foreground text mask reconstruction branch adopts the same network architecture as the shadow probability estimation branch, that is, the encoding stage uses a ResNet-18 backbone network for multi-scale feature extraction, and the decoding stage uses deconvolution upsampling and attention-gated skip connections for multi-scale feature fusion. The two branches share the ResNet-18 backbone encoder, but each has an independent decoder to ensure that shadow localization and text reconstruction tasks are optimized independently without interfering with each other.

[0049] The foreground text mask reconstruction branch also introduces convolutional attention blocks in the decoding stage to simultaneously extract channel attention and spatial attention from the concatenated features, achieving a feature fusion enhancement mechanism consistent with the shadow probability estimation branch.

[0050] The decoder outputs the upsampled and recovered feature map as a foreground text mask with 1 channel. The value range of each pixel in the foreground text mask is [0,1], where 1 represents the text region, 0 represents the background region, and the intermediate transition value represents the probability that the region belongs to the text.

[0051] Through steps S4 and S5, the document image is modeled from two dimensions: shadow spatial distribution and text structural features. The shadow probability map provides continuous probability distribution information of the shadow region, which is used to guide the spatial allocation of correction weights in the subsequent shadow removal process. The foreground text mask provides binarized position information of the text region, which is used to apply structural protection constraints to the text region in the subsequent process.

[0052] For example, in step S6 above, the document image, shadow probability map, and foreground text mask are concatenated, and the input image reconstruction network generates a shadow-free document image, such as... Figure 3 As shown, it includes the following sub-steps: S61, Multi-source feature input construction.

[0053] Specifically, the original shadowed document image, the shadow probability map, and the foreground text mask are concatenated along the channel dimension to form a multi-source feature input tensor. The original shadow image has 3 channels (RGB), the shadow probability map has 1 channel, and the foreground text mask has 1 channel; the concatenation results in a multi-source feature input tensor with 5 channels. This multi-source feature input tensor simultaneously contains original pixel information, shadow spatial distribution information, and text structural location information.

[0054] S62, Multi-scale Feature Extraction and Fusion.

[0055] Specifically, a 5-channel multi-source feature input tensor is fed into the image reconstruction network. The image reconstruction network adopts the same attention-gated U-Net encoder-decoder architecture as in step S4. In the encoding stage, a ResNet-18 backbone network is used for multi-scale feature extraction, and multi-scale feature maps of 1 / 2 to 1 / 32 scale are obtained through 5 downsampling operations. In the decoding stage, deconvolution is used for layer-by-layer upsampling. Each upsampling layer is spliced ​​and fused with the corresponding scale feature map of the encoder through attention-gated skip connections. The dynamic weighting mechanism of attention gating strengthens shadow correction-related features and suppresses redundant background features.

[0056] S63, dual-guided shadow removal and image output.

[0057] Specifically, in the feature processing of the image reconstruction network, dual-guided collaborative optimization is achieved through shadow probability maps and foreground text masks: On one hand, the image reconstruction network performs illumination correction on shadow areas based on the probability values ​​of each pixel in the shadow probability map. Specifically, the shadow probability map is multiplied element-wise with the feature maps output by each layer of the image reconstruction network decoder. The larger the pixel value in the shadow probability map, the higher the correction weight of the corresponding feature map location, and the greater the brightness and color correction applied to that location by the network. Conversely, the smaller the pixel value in the shadow probability map, the lower the correction weight of the corresponding feature map location, and the network makes only minor brightness adjustments to that location. Thus, the network focuses its computational resources on high-probability shadow areas, achieving targeted shadow removal.

[0058] On the other hand, the image reconstruction network preserves the features of the text region based on the mask value of the text region in the foreground text mask. Specifically, the foreground text mask is multiplied element-wise with the feature map output by the image reconstruction network. For text regions with a mask value of 1, the network feature value is forced to remain stable to avoid stroke blurring and erosion during the correction process, and to avoid excessive modification of text strokes and character structure during the correction process. For non-text regions with a mask value of 0, shadow correction processing is performed to eliminate shadows.

[0059] In the last layer of the decoder, the fused feature map is mapped to 3 channels (RGB) through a 1×1 convolutional layer, and the pixel values ​​are normalized to the range of [0,1] through the Sigmoid activation function, so that it conforms to the physical lighting characteristics and the laws of human vision. Finally, the output is a shadowless document image with the same resolution as the original image.

[0060] Through step S6, this embodiment uses the shadow probability map as a spatial attention guide and the foreground text mask as a structural protection constraint to achieve synergistic optimization of precise shadow region elimination and complete text region preservation. In the output shadow-free document image, shadows are completely removed, while fine-grained structures such as text strokes and character corners are completely preserved.

[0061] In an exemplary embodiment, the shadow probability estimation branch, the foreground text mask reconstruction branch, and the image reconstruction network need to be jointly trained to obtain optimal parameters.

[0062] Specifically, a multi-constraint loss function is used to jointly train the shadow probability estimation branch, the foreground text mask reconstruction branch, and the image reconstruction network. The multi-constraint loss function includes reconstruction loss, shadow probability regression loss, text mask segmentation loss, and text feature consistency loss.

[0063] 1. Reconstruction Loss: Constructed by calculating the pixel-wise L1 norm distance and multi-scale structural similarity between the predicted shadowless image and the true shadowless image, i.e.: ; in, Describing the L1 norm, For the predicted shadowless document image, For realistic, shadow-free document images, This represents a multi-scale structural similarity index.

[0064] 2. Adversarial Loss: Using the image of the document with shadows as a condition, the generator and discriminator of the generative adversarial network are optimized by comparing the output probabilities of the discriminator for the predicted image without shadows with those for the real image without shadows. ; in, The original image of the document with the shadow. For the predicted shadowless document image, Let D represent the discriminator, which is a real, shadow-free document image. The discriminator generates an N×N real / fake rating map by detecting 70×70 image patches. The generator approximates the texture and lighting features of the real image by minimizing this loss, while the discriminator improves its ability to distinguish details by maximizing this loss.

[0065] 3. Shadow Probability Regression Loss: To guide the network to accurately learn the spatial location and probability intensity of shadow regions, pixel-level L1 regression constraints are applied to the shadow probability map, i.e.: ; in, This is the shadow probability map predicted by the network. This is the true probability map of shadows.

[0066] 4. Text masking segmentation loss: including Dice loss and binary focus loss.

[0067] The formula for calculating Dice loss is: ; in, For the predicted foreground text mask, The foreground text mask is the actual image, and N is the total number of pixels in the image. Let i be the value of the i-th pixel in the predicted foreground text mask. This is the value of the i-th pixel in the real foreground text mask.

[0068] The formula for calculating the binary focus loss is: ; in, For the predicted foreground text mask, The foreground text mask is the actual image, and N is the total number of pixels in the image. , These represent the number of pixels in the foreground and background, respectively. , For category weights, To focus parameters, The value of the i-th pixel in the true foreground text mask. 5. Text feature consistency loss: including image-level text correlation loss and mask-level text correlation loss.

[0069] Image-level text correlation loss is used to constrain the consistency of predicted shadowless images with real shadowless images using high-level text feature representation. The formula for calculating image-level text correlation loss is as follows: ; in, For the predicted shadowless document image, For realistic, shadow-free document images, This represents the depth features of the c-th layer of the CTPN model. This represents the deep features of the d-th layer of the DenseNet model. This represents the L1 norm.

[0070] The mask-level text correlation loss is used to apply high-level text feature consistency constraints between the predicted foreground text mask and the ground truth foreground text mask. The formula for calculating the mask-level text correlation loss is as follows: ; in, For the predicted foreground text mask, For the true foreground text mask, This represents the depth features of the c-th layer of the CTPN model. This represents the deep features of the d-th layer of the DenseNet model. This represents the L1 norm.

[0071] 6. Total Loss Function: The total loss function is formed by combining the above losses through weighted combinations. ; in, To rebuild the losses, To combat the losses, For shadow probability regression loss, For Dice's loss, For binary focus loss, For image-level text correlation loss, For mask-level text correlation loss, , , , , These are the weight hyperparameters corresponding to each loss term.

[0072] The training process employs a hybrid labeled dataset for phased training, divided into a pre-training phase and a fine-tuning phase. The pre-training phase utilizes a large-scale, fully synthetic labeled dataset, enabling the network to quickly learn the fundamental mapping relationship between shadow removal and text reconstruction. The fine-tuning phase trains on a real-world document image dataset, adapting to real-world lighting variations and shadow features. During training, a random sampling strategy is used to dynamically select training images, with each training iteration randomly cropping the images to a fixed resolution. Training is terminated when the validation set loss function stabilizes and no longer significantly decreases over multiple iterations. The number of iterations is determined based on the actual training situation, typically 5 to 10 iterations, to save the optimal model parameters.

[0073] In one exemplary embodiment, such as Figure 4 As shown, a document image shadow removal device based on foreground text guidance and shadow probability learning is provided, comprising: Image acquisition module 101 is used to acquire document images with shadows; The initial feature extraction module 102 is used to perform morphological dilation operation on the document image to generate a shadowed background image; and to calculate the pixel-by-pixel difference between the document image and the shadowed background image to obtain an initial foreground image. The dual-branch feature estimation module 103 includes a shadow probability estimation branch and a foreground text mask reconstruction branch; the shadow probability estimation branch is used to input the document image and the shadowed background image into the encoder-decoder network to generate a shadow probability map; the foreground text mask reconstruction branch is used to input the document image and the initial foreground image into the encoder-decoder network to generate a foreground text mask; The image reconstruction module 104 is used to concatenate the document image, the shadow probability map, and the foreground text mask, and input them into the image reconstruction network to generate a shadow-free document image; wherein, the image reconstruction network performs illumination correction on the shadow area according to the probability value of each pixel in the shadow probability map, and performs feature preservation on the text area according to the mask value of the text area in the foreground text mask; The joint training module 105 is used to jointly train the shadow probability estimation branch, the foreground text mask reconstruction branch, and the image reconstruction network using a multi-constraint loss function; the multi-constraint loss function includes reconstruction loss, adversarial loss, shadow probability regression loss, text mask segmentation loss, and text feature consistency loss, and the text mask segmentation loss includes Dice loss and binary focus loss.

[0074] In one exemplary embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.

[0075] In one exemplary embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.

[0076] In one exemplary embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.

[0077] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0078] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Furthermore, any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory.

[0079] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0080] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A document image shadow removal method based on foreground text guidance and shadow probability learning, characterized in that, include: Get the document image with shadow; Perform morphological dilation on the document image to generate an image with a shadowed background; Calculate the pixel-by-pixel difference between the document image and the image with the shadowed background to obtain an initial foreground image; The document image and the shadowed background image are input into the shadow probability estimation branch to generate a shadow probability map, and the document image and the initial foreground image are input into the foreground text mask reconstruction branch to generate a foreground text mask; The document image, the shadow probability map, and the foreground text mask are concatenated and input into an image reconstruction network to generate a shadow-free document image. The image reconstruction network performs illumination correction on the shadow area based on the probability value of each pixel in the shadow probability map, and preserves the features of the text area based on the mask value of the text area in the foreground text mask.

2. The document image shadow removal method based on foreground text guidance and shadow probability learning according to claim 1, characterized in that, The morphological dilation operation employs an adaptively adjusted dilation kernel, the size of which is determined based on the average size of the document characters.

3. The document image shadow removal method based on foreground text guidance and shadow probability learning according to claim 1, characterized in that, The process of generating the image with the shadow background further includes performing median filtering on the image after the morphological dilation operation.

4. The document image shadow removal method based on foreground text guidance and shadow probability learning according to claim 1, characterized in that, Both the shadow probability estimation branch and the foreground text mask reconstruction branch adopt an encoder-decoder network architecture. The encoder uses a ResNet-18 backbone network for multi-scale feature extraction, and the decoder performs upsampling through deconvolution and merges the multi-scale features of the encoder with the upsampled features of the decoder through attention-gated skip connections.

5. The document image shadow removal method based on foreground text guidance and shadow probability learning according to claim 4, characterized in that, The attention gating is implemented according to the following formula: ; in, For attention weights, This is the feature map of the l-th layer of the encoder. It is the Sigmoid activation function. , , For learnable weight matrix, For bias terms, For encoder features, These are decoder features.

6. The document image shadow removal method based on foreground text guidance and shadow probability learning according to claim 1, characterized in that, The image reconstruction network performs illumination correction on the shadow area based on the probability values ​​of each pixel in the shadow probability map, including: adjusting the correction weights at corresponding positions in the image reconstruction network based on the probability values ​​of each pixel in the shadow probability map.

7. The document image shadow removal method based on foreground text guidance and shadow probability learning according to claim 1, characterized in that, The step of preserving the features of the text region based on the mask value of the text region in the foreground text mask includes: locking the feature values ​​of the text region in the image reconstruction network based on the mask value of the text region in the foreground text mask.

8. The document image shadow removal method based on foreground text guidance and shadow probability learning according to claim 1, characterized in that, Also includes: The shadow probability estimation branch, the foreground text mask reconstruction branch, and the image reconstruction network are jointly trained using a multi-constraint loss function. The multi-constraint loss function includes reconstruction loss, adversarial loss, shadow probability regression loss, text mask segmentation loss, and text feature consistency loss. The text mask segmentation loss includes Dice loss and binary focus loss.

9. A document image shadow removal device based on foreground text guidance and shadow probability learning, characterized in that, include: The image acquisition module is used to acquire document images with shadows. An initial feature extraction module is used to perform morphological dilation on the document image to generate a background image with shadows; and to calculate the pixel-by-pixel difference between the document image and the background image with shadows to obtain an initial foreground image. The dual-branch feature estimation module includes a shadow probability estimation branch and a foreground text mask reconstruction branch. The shadow probability estimation branch is used to input the document image and the shadowed background image into the encoder-decoder network to generate a shadow probability map. The foreground text mask reconstruction branch is used to input the document image and the initial foreground image into the encoder-decoder network to generate a foreground text mask. An image reconstruction module is used to concatenate the document image, the shadow probability map, and the foreground text mask, and input them into an image reconstruction network to generate a shadow-free document image; wherein, the image reconstruction network performs illumination correction on the shadow area based on the probability value of each pixel in the shadow probability map, and performs feature preservation on the text area based on the mask value of the text area in the foreground text mask.

10. The document image shadow removal device based on foreground text guidance and shadow probability learning according to claim 9, characterized in that, It also includes a joint training module for jointly training the shadow probability estimation branch, the foreground text mask reconstruction branch, and the image reconstruction network using a multi-constraint loss function; the multi-constraint loss function includes reconstruction loss, adversarial loss, shadow probability regression loss, text mask segmentation loss, and text feature consistency loss, and the text mask segmentation loss includes Dice loss and binary focus loss.