Document image shadow removing method and device and storage medium

By adding noise to the document image and mask, and constructing a diffusion model and mask refinement module, the problem of dependence on external masks in existing methods is solved, and efficient shadow removal and document detail preservation are achieved.

CN121169747APending Publication Date: 2025-12-19WUHAN INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511097143.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-06
Publication Date
2025-12-19

AI Technical Summary

Technical Problem

Existing methods for removing shadows from document images rely on external shadow masks. The quality of the mask affects the removal effect. Traditional methods are incomplete in complex structures and are prone to artifacts or color distortion.

Method used

By adding noise to the original shadowed document image and the mask, a training model is built. The diffusion model and mask refinement module are used to improve the accuracy and robustness of shadow removal without the need for precise masking, while preserving document details.

Benefits of technology

It effectively improves the accuracy and robustness of shadow removal, reduces the dependence on mask precision, maintains the clarity and detail of text areas in documents, and overcomes the limitations of traditional and deep learning methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121169747A_ABST
    Figure CN121169747A_ABST
Patent Text Reader

Abstract

The invention provides a document image shadow removing method and device and a storage medium, and belongs to the technical field of image processing, and the method comprises the steps: importing a plurality of original shadow document images and original shadow masks; performing noise addition processing on the original shadow document image and the original shadow mask to obtain a noise-added shadow document image and a noise-added shadow mask; and performing model analysis on the training model through all the original shadow document images, all the noise-added shadow document images and all the noise-added shadow masks to obtain an image shadow removal model. According to the method, under the condition that the mask does not need to be accurately preset, the shadow removal accuracy and robustness can be effectively improved, meanwhile, the definition and details of the text area of the document are reserved, the limitation of a traditional method and an existing deep learning method is overcome, the shadow removal effect and the image quality are improved, and dependence on mask accuracy is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application mainly relates to the technical field of image processing, and in particular to a document image shadow removal method and device and storage medium. BACKGROUND

[0002] 1. During the shooting process, the document image often produces shadows due to uneven lighting, shielding and other reasons, which seriously affects the image quality and text readability, and further affects the subsequent recognition and processing effect.

[0003] 2. The existing document shadow removal methods mainly include traditional image processing algorithms and deep learning-based models. Although deep learning-based models such as BEDSR-Net can well preserve document details, they usually rely on external shadow masks, and the mask quality has a great influence on the removal effect. In practical applications, it is difficult to obtain high-quality masks, resulting in incomplete removal or blurred text. Although traditional image processing algorithms such as ShadowRefiner can process images without masks, they still have problems such as incomplete removal and detail loss in scenes where shadows overlap with text or edges are blurred.

[0004] 3. Traditional methods rely on image brightness or neighborhood interpolation to remove shadows, but when dealing with strong shadows or complex structures, they are prone to produce artifacts or color distortion. SUMMARY

[0005] The technical problem to be solved by the present application is to provide a document image shadow removal method, device and storage medium to solve the problems of the prior art.

[0006] The technical solution for solving the above technical problem is as follows: a document image shadow removal method, comprising the following steps: Importing a plurality of original shadow document images and original shadow masks corresponding to each of the original shadow document images; Respectively adding noise to each of the original shadow document images and the original shadow masks corresponding to each of the original shadow document images to obtain a noise-added shadow document image corresponding to each of the original shadow document images and a noise-added shadow mask corresponding to each of the original shadow document images; Constructing a training model, performing model analysis on the training model through all the original shadow document images, all the noise-added shadow document images and all the noise-added shadow masks, and obtaining an image shadow removal model; Importing a to-be-processed shadow document image, performing shadow removal processing on the to-be-processed shadow document image through the image shadow removal model, and obtaining a document image shadow removal result.

[0007] Another technical solution of the present application to solve the above technical problems is as follows: A document image de-shadowing device comprises: An import module is configured to import a plurality of original shadow document images and original shadow masks corresponding to the original shadow document images. A noise adding processing module is configured to add noise to each of the original shadow document images and the original shadow masks corresponding to the original shadow document images, to obtain a plurality of noise-added shadow document images corresponding to the original shadow document images and noise-added shadow masks corresponding to the original shadow document images. A model analysis module is configured to construct a training model, analyze the training model based on all the original shadow document images, all the noise-added shadow document images and all the noise-added shadow masks, and obtain an image de-shadowing model. The import module is further configured to import a to-be-processed shadow document image. A shadow processing module is configured to perform de-shadowing processing on the to-be-processed shadow document image based on the image de-shadowing model, to obtain a document image de-shadowing result.

[0008] Based on the above-mentioned document image de-shadowing method, the present application further provides a document image de-shadowing system.

[0009] Another technical solution of the present application to solve the above technical problems is as follows: A document image de-shadowing system comprises a memory, a processor and a computer program stored in the memory and executable on the processor, and when the processor executes the computer program, the document image de-shadowing method described above is implemented.

[0010] Based on the above-mentioned document image de-shadowing method, the present application further provides a computer readable storage medium.

[0011] Another technical solution of the present application to solve the above technical problems is as follows: A computer readable storage medium stores a computer program, and when the computer program is executed by a processor, the document image de-shadowing method described above is implemented.

[0012] The beneficial effects of the present application are: the original shadow document image and the original shadow mask are subjected to noise adding processing to obtain a noise-added shadow document image and a noise-added shadow mask, the original shadow document image, the noise-added shadow document image and the noise-added shadow mask are subjected to model analysis of a training model to obtain an image shadow removal model, and the image shadow removal model is subjected to shadow removal processing of a to-be-processed shadow document image to obtain a document image shadow removal result, the present application can effectively improve the accuracy and robustness of shadow removal without accurately presetting a mask, while the clarity and details of the document text area are retained, the limitations of traditional methods and existing deep learning methods are overcome, the shadow removal effect and image quality are improved, and the dependence on mask accuracy is reduced. BRIEF DESCRIPTION OF DRAWINGS

[0013] Figure 1 A flowchart of a document image shadow removal method provided by an embodiment of the present application is shown. Figure 2 A training model structure diagram of the document image shadow removal method provided by the embodiment of the present application is shown. Figure 3 A first experimental effect comparison diagram of the document image shadow removal method provided by the embodiment of the present application is shown. Figure 4 A second experimental effect comparison diagram of the document image shadow removal method provided by the embodiment of the present application is shown. Figure 5 A third experimental effect comparison diagram of the document image shadow removal method provided by the embodiment of the present application is shown. Figure 6 A module block diagram of a document image shadow removal device provided by the embodiment of the present application is shown. DETAILED DESCRIPTION

[0014] The principles and characteristics of the present application are described below in combination with the drawings, and the examples are only used to explain the present application and are not used to limit the scope of the present application.

[0015] Figure 1 A flowchart of a document image shadow removal method provided by an embodiment of the present application is shown.

[0016] As shown in Figure 1 A document image shadow removal method includes the following steps: S1: importing a plurality of original shadow document images and original shadow masks corresponding to each of the original shadow document images; S2: respectively performing noise adding processing on each of the original shadow document images and the original shadow masks corresponding to each of the original shadow document images to obtain a noise-added shadow document image corresponding to each of the original shadow document images and a noise-added shadow mask corresponding to each of the original shadow document images; S3: constructing a training model, performing model analysis on the training model through all the original shadow document images, all the shadow document images after adding noise, and all the shadow masks after adding noise, to obtain an image de-shadowing model; S4: importing a to-be-processed shadow document image, performing de-shadowing processing on the to-be-processed shadow document image through the image de-shadowing model, to obtain a document image de-shadowing result.

[0017] It should be understood that the shadow document image (i.e., the original shadow document image) and the initial shadow mask (i.e., the original shadow mask) are subjected to Gaussian noise addition.

[0018] In the above embodiment, the original shadow document image and the original shadow mask are subjected to noise addition processing to obtain the shadow document image after adding noise and the shadow mask after adding noise, the training model is subjected to model analysis through the original shadow document image, the shadow document image after adding noise, and the shadow mask after adding noise to obtain the image de-shadowing model, and the to-be-processed shadow document image is subjected to de-shadowing processing through the image de-shadowing model to obtain the document image de-shadowing result. The present application can effectively improve the accuracy and robustness of shadow removal without the need for accurate preset masks, while retaining the clarity and details of the document text area, overcoming the limitations of traditional methods and existing deep learning methods, improving the shadow removal effect and image quality, and reducing the dependence on mask accuracy.

[0019] Optionally, as one embodiment of the present application, as shown in Figure 1 and 2 the training model comprises a down-sampling module, an up-sampling module, and a plurality of sequentially connected first DSE residual blocks; the process of performing model analysis on the training model through all the original shadow document images, all the shadow document images after adding noise, and all the shadow masks after adding noise to obtain the image de-shadowing model comprises: respectively splicing each original shadow document image, the shadow document image after adding noise corresponding to each original shadow document image, and the shadow mask after adding noise corresponding to each original shadow document image to obtain a spliced document image corresponding to each original shadow document image; performing down-sampling analysis on each spliced document image through the down-sampling module to obtain a to-be-processed document feature corresponding to each original shadow document image; performing feature analysis on each to-be-processed document feature through a plurality of first DSE residual blocks to obtain a first feature-analyzed document feature corresponding to each original shadow document image; The up-sampling module is used for up-sampling analysis on each first feature analysis post-document feature, to obtain an initial de-shadowing document image corresponding to each original shadow document image and a target shadow mask corresponding to each original shadow document image. The real document image corresponding to each original shadow document image and the real shadow mask corresponding to each original shadow document image are imported, and loss function analysis is performed on all the initial de-shadowing document images, all the real document images, all the target shadow masks and all the real shadow masks, to obtain a target loss function. According to the target loss function, the training model is updated in parameters, and after the parameter update, each spliced post-document image is re-analyzed by down-sampling, until the iteration number is reached, and then the training model after the parameter update is used as an image de-shadowing model.

[0020] It should be understood that the DSRN network (i.e., the training model) takes a noisy image (i.e., a noisy post-shadow document image), a rough mask (i.e., a noisy post-shadow mask), and an original image (i.e., an original shadow document image) as input, predicts a noise map through a U-Net encoding-decoding structure, and introduces an improved residual module DSE (i.e., a first DSE residual block).

[0021] Specifically, the noisy image (i.e., the noisy post-shadow document image), the initial shadow mask (i.e., the noisy post-shadow mask), and the original document image (i.e., the original shadow document image) are jointly input into the noise estimation network DSRN (i.e., the training model) to perform de-noising and de-shadowing operations, to obtain a preliminary de-shadowing image (i.e., an initial de-shadowing document image), and a corresponding refined shadow mask (i.e., a target shadow mask) is generated by the LFCN mask refinement module integrated in the DSRN. Subsequently, the refined mask (i.e., the target shadow mask) is used to replace the initial mask as guide information and is input into the DSRN again with the current image to perform the next round of de-shadowing and mask optimization. This process forms a closed-loop mechanism of mutual enhancement of image and mask, and through iterative rounds, the accuracy of shadow area recognition and removal is continuously improved.

[0022] It should be understood that in the model training stage, the method adopts a fixed number of iteration strategies, and the model basically converges when the iteration is 200,000 times, at which time the network can output a high-quality de-shadowing image (i.e., an initial de-shadowing document image) and a corresponding refined mask (i.e., a target shadow mask).

[0023] It should be understood that the data processing process of each of the to-be-processed document features by the plurality of first DSE residual blocks is the same as the data processing process of the subsequent second DSE residual block, the third DSE residual block, the fourth DSE residual block and the fifth DSE residual block, and only the processed data is different.

[0024] In the above embodiment, the image de-shadowing model is obtained by performing model analysis on the training model through all original shadow document images, all noisy shadow document images and all noisy shadow masks, which improves the accuracy of shadow area recognition and removal, overcomes the limitations of traditional methods and existing deep learning methods, and reduces the dependence on mask accuracy.

[0025] Optionally, as an embodiment of the present application, the downsampling module comprises a plurality of downsampling layers, a plurality of second DSE residual blocks same in number as the downsampling layers, and a plurality of third DSE residual blocks same in number as the downsampling layers. The process of performing downsampling analysis on each of the spliced document images by the downsampling module to obtain the to-be-processed document features corresponding to each of the original shadow document images comprises: The first and last of the plurality of downsampling groups are connected. The second DSE residual block in the first downsampling group performs feature analysis on each of the spliced document images to obtain the second feature-analyzed document features corresponding to each of the original shadow document images. The third DSE residual block in the first downsampling group performs feature analysis on each of the second feature-analyzed document features to obtain the third feature-analyzed document features corresponding to each of the original shadow document images. The downsampling layer in the first downsampling group performs downsampling processing on each of the third feature-analyzed document features to obtain the downsampling document features corresponding to each of the original shadow document images, and each of the downsampling document features is input into the next downsampling group until the last downsampling group, and the output of the last downsampling group is taken as the to-be-processed document features, thereby obtaining the to-be-processed document features corresponding to each of the original shadow document images.

[0026] It should be understood that the data processing process of the second DSE residual block and the third DSE residual block is the same as that of the first DSE residual block, the fourth DSE residual block and the fifth DSE residual block, and only the processed data is different.

[0027] In the above embodiment, the feature of each spliced document image is analyzed by the downsampling module to obtain the to-be-processed document feature, which overcomes the limitations of traditional methods and existing deep learning methods and reduces the dependence on mask accuracy.

[0028] Optionally, as an embodiment of the present application, the second DSE residual block comprises a channel attention mechanism layer, a multi-scale pooling layer and a self-attention layer. The process of performing feature analysis on each of the spliced document images by the second DSE residual block in the first downsampling group to obtain the second feature-analyzed document feature corresponding to each of the original shadow document images comprises: The feature of each spliced document image is enhanced by the channel attention mechanism layer to obtain the original document feature corresponding to each of the original shadow document images; The original document feature is pooled by the multi-scale pooling layer to obtain the pooled document feature corresponding to each of the original shadow document images; The feature of each document feature is extracted by the self-attention layer to obtain the second feature-analyzed document feature corresponding to each of the original shadow document images.

[0029] It should be understood that the residual module DSE (i.e., the second DSE residual block) adopts a channel attention mechanism (i.e., a channel attention mechanism layer), multi-scale pooling (i.e., a multi-scale pooling layer) and a self-attention structure (i.e., a self-attention layer), effectively extracts shadow area and character edge information, and improves the adaptability of the model to complex document structures.

[0030] In the above embodiment, the feature of each spliced document image is analyzed by the second DSE residual block in the first downsampling group to obtain the second feature-analyzed document feature, effectively extracting shadow area and character edge information and improving the adaptability of the model to complex document structures.

[0031] Optionally, as an embodiment of the present application, the upsampling module comprises a plurality of fourth DSE residual blocks, a plurality of fifth DSE residual blocks identical in number to the fourth DSE residual blocks, a plurality of upsampling layers identical in number to the fourth DSE residual blocks, and a mask refinement module. The process of performing upsampling analysis on each of the first feature-analyzed document features by the upsampling module to obtain the initial de-shadowed document image corresponding to each of the original shadow document images and the target shadow mask corresponding to each of the original shadow document images comprises: One of the fourth DSE residual blocks, one of the fifth DSE residual blocks and one of the upsampling layers are taken as an upsampling group, thereby obtaining a plurality of first and last connected upsampling groups. the fourth DSE residual block in the first up-sampling group to obtain fourth feature-analyzed document features corresponding to each of the original shadow document images; the fifth DSE residual block in the first up-sampling group to obtain fifth feature-analyzed document features corresponding to each of the original shadow document images; the up-sampling layer in the first up-sampling group to obtain up-sampled document images corresponding to each of the original shadow document images, and input each of the up-sampled document images into the next up-sampling group until the last up-sampling group, and take the output of the last up-sampling group as an initial de-shadow document image, thereby obtaining an initial de-shadow document image corresponding to each of the original shadow document images; the mask refinement module to obtain a target shadow mask corresponding to each of the original shadow document images.

[0032] It should be understood that the data processing process of the fourth DSE residual block and the fifth DSE residual block is the same as that of the first DSE residual block, the second DSE residual block and the third DSE residual block, and only the data processed is different.

[0033] In the above embodiment, the up-sampling module is used to perform up-sampling analysis on each of the first feature-analyzed document features to obtain the initial de-shadow document image and the target shadow mask, thereby improving the accuracy of shadow area recognition and removal, overcoming the limitations of traditional methods and existing deep learning methods, and reducing the dependence on mask accuracy.

[0034] Optionally, as an embodiment of the present application, the mask refinement module comprises a plurality of convolution layers, a Sigmoid activation function layer and a DECA dynamic channel attention module, The process of performing mask analysis on each of the initial de-shadow document images by the mask refinement module to obtain a target shadow mask corresponding to each of the original shadow document images comprises: extracting features of each of the initial de-shadow document images by the plurality of convolution layers to obtain original de-shadow document features corresponding to each of the original shadow document images; mapping each of the original de-shadow document features by the Sigmoid activation function layer to obtain mapped de-shadow document features corresponding to each of the original shadow document images; The DECA dynamic channel attention module respectively performs mask processing on each of the mapped post-shadow removal document features, to obtain a target shadow mask corresponding to each of the original shadow document images.

[0035] It should be understood that, in order to solve the problem of reduced shadow removal effect caused by inaccurate mask, a DSRN output end is integrated with an LFCN module (i.e., a mask refinement module) for refining the initial shadow mask. The LFCN module (i.e., the mask refinement module) is composed of a multi-layer convolution structure (i.e., a multi-layer convolution layer) and a Sigmoid activation function (i.e., a Sigmoid activation function layer), and outputs a probability map of each pixel belonging to a shadow. In order to enhance the adaptability to different shadow intensities, a DECA dynamic channel attention module is introduced to adaptively adjust the channel response according to the shadow intensity, thereby improving the discrimination ability of the shadow area and the background text.

[0036] Specifically, the key role of the LFCN module (i.e., the mask refinement module) is to refine the initial shadow mask so that its boundary is closer to the real shadow area, thereby significantly reducing the risk of artifacts and misrecognition, and effectively improving the shadow removal accuracy of the final image while maintaining the clarity of the text.

[0037] Specifically, the LFCN module (i.e., the mask refinement module) not only has a lightweight structure and high computational efficiency, but also can accurately model shadows of different intensities through an adaptive mechanism (such as DECA dynamic channel attention), and has stronger stability and robustness in document images with blurred shadow boundaries and severely overlapped text.

[0038] In the above embodiment, the mask refinement module is used to analyze the mask of each initial shadow removal document image to obtain a target shadow mask, which significantly reduces the risk of artifacts and misrecognition, thereby effectively improving the shadow removal accuracy of the final image while maintaining the clarity of the text, and showing stronger stability and robustness in document images with blurred shadow boundaries and severely overlapped text.

[0039] Optionally, as an embodiment of the present application, the process of performing loss function analysis on all the initial shadow removal document images, all the real document images, all the target shadow masks, and all the real shadow masks to obtain a target loss function includes: performing loss function calculation on all the initial shadow removal document images and all the real document images to obtain a noise prediction loss function; performing loss function calculation on all the target shadow masks and all the real shadow masks to obtain a mask refinement loss function; performing calculation on the noise prediction loss function and the mask refinement loss function by a first formula to obtain a target loss function, wherein the first formula is: , wherein, is a target loss function, is a noise prediction loss function, is a hyperparameter, is a mask refinement loss function.

[0040] Specifically, the present application adopts a hybrid loss function, including: a noise prediction loss (i.e., a noise prediction loss function): guiding the model to learn the de-noising ability, driving the diffusion model to remove the shadow noise and restore the image content; a mask refinement loss (i.e., a mask refinement loss function): used to optimize the positioning accuracy of the shadow area, so that the generated mask boundary is more consistent with the real shadow (such as the half-shadow transition area), and the misidentification is reduced; both linearly weighted as a total loss function (i.e., a target loss function), and the weight is adjusted to balance the mask accuracy and image quality.

[0041] .

[0042] In the above embodiment, the loss function analysis is performed on all initial de-shadow document images, all real document images, all target shadow masks, and all real shadow masks to obtain the target loss function, which guides the model to learn the de-noising ability, drives the diffusion model to remove the shadow noise, restores the image content, optimizes the positioning accuracy of the shadow area, makes the generated mask boundary more consistent with the real shadow, and reduces the misidentification.

[0043] Optionally, as another embodiment of the present application, the existing document image de-shadowing method has obvious deficiencies: the traditional image processing method relies on simple rules and brightness adjustment, and it is difficult to deal with complex shadow structures and text areas, and it is almost impossible to achieve effective removal; the method based on deep learning has certain removal ability, but usually relies on external shadow masks, and if the masks are missing, the effect will be significantly reduced, and the masks themselves are prone to errors, which will further affect the final de-shadowing quality.

[0044] Based on the above technical problems, the present application introduces a dynamic mask optimization mechanism, which gradually refines the shadow mask in the reverse generation process of the diffusion model, and optimizes it in cooperation with the generation process of the shadow-free image. Without the need for accurate preset masks, the present application can effectively improve the accuracy and robustness of shadow removal, while preserving the clarity and details of the document text area, overcoming the limitations of traditional methods and existing deep learning methods.

[0045] Optionally, as another embodiment of the present application, the present application adopts a diffusion model-based image generation framework, including a forward noise adding process and an inverse noise removing process. In terms of model structure, an improved U-Net is adopted as the noise estimation backbone network DSRN, and an optimized shadow mask is generated while denoising. The entire network is composed of an input processing module, a down-sampling encoder, a DSE residual module, an up-sampling decoder, and a mask refinement module LFCN.

[0046] Optionally, as another embodiment of the present application, the present application adopts a diffusion modeling process of forward noise adding + inverse noise removing: Forward process: Gaussian noise is added to the shadow document image and the initial shadow mask; Inverse process: the noise-added image, the initial shadow mask, and the original document image are jointly input into the noise estimation network DSRN to perform denoising and shadow removal operations, to obtain a preliminary shadow-removed image, and a corresponding refined shadow mask is generated by the LFCN mask refinement module integrated in the DSRN. Subsequently, the refined mask is used to replace the initial mask as the guide information and is input into the DSRN again with the current image to perform the next round of shadow removal and mask optimization. This process forms a closed-loop mechanism of mutual enhancement of image and mask, and through iterative rounds, the accuracy of shadow area recognition and removal is continuously improved.

[0047] In the model training phase, the present application adopts a fixed number of iteration strategy, and the model basically converges when the iteration reaches 200,000 times, at which time the network can output high-quality shadow-removed images and corresponding refined masks.

[0048] Optionally, as another embodiment of the present application, the present application has the following beneficial effects: 1. Reducing dependence on mask accuracy: existing deep learning methods generally rely on external high-quality shadow masks, and the present application introduces a mask refinement module (LFCN) to dynamically optimize the mask boundary and improve the mask accuracy in the case of rough or erroneous initial mask, thereby enhancing the practicality and adaptability of the model.

[0049] 2. Improving shadow removal effect and image quality: the present application constructs a document image generation framework based on a diffusion model, and optimizes the shadow mask and image restoration process in the process of step-by-step denoising, effectively solving the problems of shadow residue, boundary artifacts, and fuzzy text in traditional methods, and generating more natural and clear results with significantly improved text retention effect.

[0050] 3. Enhancing robustness in complex shadow environment: by introducing the DSE residual block and the DECA dynamic attention mechanism, the modeling ability of the model for different shadow intensities, edge blur degrees, and text overlap situations is improved, and the model has stronger generalization ability and stable performance in various actual document scenarios.

[0051] 4. Training stability, inference efficiency: In the design of the diffusion model, the number of denoising steps is reasonably simplified, and the network efficiency is optimized with a lightweight structure to improve inference speed and training stability while ensuring effectiveness, facilitating deployment and application.

[0052] Optionally, as another embodiment of the present application, the input end of the present application is composed of three groups of elements, including: a shadow image (a document image with shadows), a corresponding initial shadow mask, and a clean image without shadows (as a true target). The three together constitute the training sample of the model. The core model is DSRN, which is responsible for extracting effective features from the input shadow image and its initial mask, and performing preliminary recovery of shadow removal. Subsequently, the output of DSRN is connected to LFCN, which focuses on refining the shadow mask to make the mask boundary more consistent with the actual shadow area, improving the accuracy and integrity of the mask. Through joint training, SdocDiff (i.e. the present application) can not only generate more accurate refined shadow masks, but also output high-quality shadow-free document images, significantly improving the clarity of the text and the quality of the image after shadow removal. The entire training process is optimized through multi-task optimization of mask refinement and image recovery, making the model more robust and effective in the document shadow removal task.

[0053] Optionally, as another embodiment of the present application, as shown in Figures 3 to 5 Optionally, as another embodiment of the present application, as shown in

[0054] Optionally, as another embodiment of the present application, the specific embodiments of the present application are: 1. Data preparation The present application uses three sets of data containing document shadow images, corresponding initial shadow masks, and shadow-free clean images in the training phase. The image data comes from the public datasets RDD and Kligler, covering various document types with different backgrounds, lighting, fonts, and language environments, ensuring that the training data is representative and diverse.

[0055] 2. Model structure The SdocDiff model proposed by the present application consists of two core modules: DSRN: used for preliminary recovery of shadow removal, based on diffusion model architecture design, introducing the DSE module to enhance the structural feature modeling capability, effectively improving the recovery quality of complex document regions.

[0056] LFCN: used for refining the initial shadow mask. By learning the spatial features of the shadow boundary, the generated mask is more accurate, effectively reducing the misjudgment area, and further optimizing the final output image.

[0057] 3. Model training and optimization The present application adopts an end-to-end joint training strategy to optimize the mask prediction and image restoration tasks. The loss function includes noise prediction error, image reconstruction error, and mask matching error, ensuring that the model can improve the structure recovery ability and boundary recognition accuracy during training.

[0058] 4. Output results After training, the SdocDiff model can receive the input shadow image and its initial mask in the inference stage, outputting two results: the refined shadow mask and the final shadow-free image. Experiments show that this method can effectively improve the clarity of the text area and the restoration effect of the shadow edge.

[0059] 5. Comparative experiment To verify the effectiveness of the present application, a comparative experiment with existing methods was constructed, and an ablation experiment of the mask refinement module (LFCN) and the DSE module was conducted. The results show that the complete SdocDiff model performs better than existing methods in various complex document scenarios, and the mask refinement and structure enhancement modules play an important role in improving the overall performance.

[0060] Optionally, as another embodiment of the present application, the present application protects the following points: 1. Document shadow removal method based on diffusion model The application first introduces a diffusion model into a document image shadow removal task, realizes higher quality image restoration through a step-by-step noise removal process, and significantly improves the shadow removal effect, especially in complex document structures and low-contrast areas.

[0061] 2. Document structure enhancement mechanism introducing a DSE module The DSE module is introduced in the diffusion modeling process, effectively enhancing the structure information modeling capability of the document, improving the reconstruction accuracy of the model in key areas such as text edges and table lines, and improving the shortcomings of traditional methods in structure detail recovery.

[0062] 3. Shadow mask refinement mechanism based on LFCN The application designs a lightweight fully convolutional network LFCN for refining and optimizing the initial shadow mask, so that its boundary is more consistent with the real shadow area. This mechanism effectively reduces artifacts, missed and false judgments, improves the accuracy of the mask, and further promotes the improvement of the final image quality.

[0063] Figure 6 A module block diagram of a document image shadow removal device provided by an embodiment of the application.

[0064] Optionally, as another embodiment of the application, as shown in Figure 6 A document image shadow removal device includes: An import module is configured to import a plurality of original shadow document images and original shadow masks corresponding to each of the original shadow document images. A noise adding processing module is configured to add noise to each of the original shadow document images and the original shadow masks corresponding to each of the original shadow document images, to obtain a noise-added shadow document image corresponding to each of the original shadow document images and a noise-added shadow mask corresponding to each of the original shadow document images. A model analysis module is configured to construct a training model, analyze the training model based on all of the original shadow document images, all of the noise-added shadow document images, and all of the noise-added shadow masks, and obtain an image shadow removal model. The import module is further configured to import a to-be-processed shadow document image. A shadow processing module is configured to perform shadow removal processing on the to-be-processed shadow document image based on the image shadow removal model, and obtain a document image shadow removal result.

[0065] Optionally, another embodiment of the present application provides a document image de-shadowing system, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, when the processor executes the computer program, the document image de-shadowing method as described above is implemented. The system can be a computer or the like.

[0066] Optionally, another embodiment of the present application provides a computer readable storage medium, the computer readable storage medium stores a computer program, when the computer program is executed by a processor, the document image de-shadowing method as described above is implemented.

[0067] It should be noted that, in this paper, the relationship terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between the entities or operations. Moreover, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or equipment including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or equipment.

[0068] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the above-described device and unit can refer to the corresponding process in the foregoing method embodiment, which will not be described here.

[0069] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed.

[0070] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, that is, they can be located in one place, or they can be distributed on a plurality of network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment of the present application.

[0071] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present alone, or two or more units can be integrated in one unit. The above integrated unit can be realized in the form of hardware or in the form of software functional unit.

[0072] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or say the part of the prior art that contributes, or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

[0073] The above description is only the preferred embodiment of the present application, and is not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A method for removing shadows from document images, characterized in that, Includes the following steps: Import multiple original shadow document images and the original shadow masks corresponding to each of the original shadow document images; Noise is added to each of the original shadow document images and the original shadow mask corresponding to each of the original shadow document images to obtain the noisy shadow document images and the noisy shadow mask corresponding to each of the original shadow document images. A training model is constructed, and model analysis is performed on the training model using all the original shadowed document images, all the noisy shadowed document images, and all the noisy shadow masks to obtain an image deshading model. Import the shadowed document image to be processed, and perform shadow removal processing on the shadowed document image using the image shadow removal model to obtain the shadow removal result of the document image.

2. The document image shadow removal method according to claim 1, characterized in that, The training model includes a downsampling module, an upsampling module, and multiple sequentially connected first DSE residual blocks; The process of performing model analysis on the trained model using all the original shadowed document images, all the noisy shadowed document images, and all the noisy shadow masks to obtain the image deshading model includes: Each of the original shadow document images, the noisy shadow document images corresponding to each of the original shadow document images, and the noisy shadow masks corresponding to each of the original shadow document images are concatenated to obtain the concatenated document images corresponding to each of the original shadow document images. The downsampling module performs downsampling analysis on each of the stitched document images to obtain the document features to be processed corresponding to each of the original shadow document images. By performing feature analysis on each of the document features to be processed through multiple first DSE residual blocks, the first feature analysis document features corresponding to each of the original shadow document images are obtained; The upsampling module performs upsampling analysis on each of the document features after the first feature analysis to obtain an initial de-shadowed document image corresponding to each of the original shadowed document images and a target shadow mask corresponding to each of the original shadowed document images; Import the real document images corresponding to each of the original shadowed document images and the real shadow masks corresponding to each of the original shadowed document images, and perform loss function analysis on all the initial unshaded document images, all the real document images, all the target shadow masks and all the real shadow masks to obtain the target loss function; The parameters of the training model are updated according to the target loss function. After the parameters are updated, each of the stitched document images is downsampled and analyzed again until the number of iterations is reached. Then, the training model with updated parameters is used as the image shadow removal model.

3. The document image shadow removal method according to claim 2, characterized in that, The downsampling module includes multiple downsampling layers, multiple second DSE residual blocks with the same number as the downsampling layers, and multiple third DSE residual blocks with the same number as the downsampling layers; The process of performing downsampling analysis on each of the stitched document images using the downsampling module to obtain the document features to be processed corresponding to each of the original shadowed document images includes: By taking a second DSE residual block, a third DSE residual block, and a downsampling layer as a downsampling group, multiple downsampling groups are obtained by connecting them end to end. By performing feature analysis on each of the stitched document images using the second DSE residual block in the first downsampling group, the second feature analysis document features corresponding to each of the original shadow document images are obtained. By performing feature analysis on each of the second feature-analyzed document features using the third DSE residual block in the first downsampling group, the third feature-analyzed document features corresponding to each of the original shadow document images are obtained. The document features analyzed by the third feature are downsampled in the first downsampling group to obtain downsampled document features corresponding to the original shadow document images. Each downsampled document feature is then input into the next downsampling group until the last downsampling group is passed. The output of the last downsampling group is used as the document features to be processed, thus obtaining the document features to be processed corresponding to the original shadow document images.

4. The document image shadow removal method according to claim 3, characterized in that, The second DSE residual block includes a channel attention mechanism layer, a multi-scale pooling layer, and a self-attention layer; The process of performing feature analysis on each of the stitched document images using the second DSE residual block in the first downsampling group to obtain the second feature-analyzed document features corresponding to each of the original shadowed document images includes: The channel attention mechanism layer is used to perform feature enhancement processing on each of the stitched document images to obtain the original document features corresponding to each of the original shadow document images; The multi-scale pooling layer is used to pool each of the original document features to obtain pooled document features corresponding to each of the original shadow document images. The self-attention layer extracts features from each of the document features to obtain the second feature-analyzed document features corresponding to each of the original shadow document images.

5. The document image shadow removal method according to claim 2, characterized in that, The upsampling module includes multiple fourth DSE residual blocks, multiple fifth DSE residual blocks of the same number as the fourth DSE residual blocks, multiple upsampling layers of the same number as the fourth DSE residual blocks, and a mask refinement module; The process of performing upsampling analysis on each of the document features after the first feature analysis through the upsampling module to obtain the initial de-shadowed document image corresponding to each of the original shadowed document images and the target shadow mask corresponding to each of the original shadowed document images includes: By taking one of the fourth DSE residual blocks, one of the fifth DSE residual blocks, and one of the upsampling layers as an upsampling group, multiple upsampling groups are obtained by connecting them end to end. By performing feature analysis on each of the first feature-analyzed document features using the fourth DSE residual block in the first upsampling group, the fourth feature-analyzed document features corresponding to each of the original shadow document images are obtained. The fifth DSE residual block in the first upsampling group is used to perform feature analysis on each of the document features after the fourth feature analysis, so as to obtain the document features after the fifth feature analysis corresponding to each of the original shadow document images. The document features analyzed by the fifth feature are upsampled by the upsampling layer in the first upsampling group to obtain the upsampled document images corresponding to the original shadow document images. The upsampled document images are then input into the next upsampling group until the last upsampling group is passed. The output of the last upsampling group is used as the initial de-shadowed document image, thus obtaining the initial de-shadowed document image corresponding to the original shadow document images. The mask refinement module performs mask analysis on each of the initial deshaded document images to obtain the target shadow mask corresponding to each of the original shadowed document images.

6. The document image shadow removal method according to claim 5, characterized in that, The masking refinement module includes multiple convolutional layers, a Sigmoid activation function layer, and a DECA dynamic channel attention module. The process of performing mask analysis on each of the initial deshaded document images through the mask refinement module to obtain the target shadow mask corresponding to each of the original shadowed document images includes: The initial unshaded document images are extracted by the multi-layer convolutional layers to obtain the original unshaded document features corresponding to each original unshaded document image. The Sigmoid activation function layer is used to map each of the original shadow-removed document features to obtain the mapped shadow-removed document features corresponding to each of the original shadow document images. The DECA dynamic channel attention module performs masking processing on each of the mapped and de-shadowed document features to obtain the target shadow mask corresponding to each of the original shadowed document images.

7. The document image shadow removal method according to claim 2, characterized in that, The process of performing loss function analysis on all the initial deshaded document images, all the real document images, all the target shadow masks, and all the real shadow masks to obtain the target loss function includes: The loss function is calculated for all the initial deshaded document images and all the real document images to obtain the noise prediction loss function; The loss function is calculated for all the target shadow masks and all the real shadow masks to obtain the mask refinement loss function; The target loss function is obtained by calculating the noise prediction loss function and the mask refinement loss function using the first equation, which is: , in, Let be the target loss function. Let the noise prediction loss function be... For hyperparameters, To refine the loss function for the mask.

8. A document image shadow removal device, characterized in that, include: The import module is used to import multiple original shadow document images and the original shadow mask corresponding to each of the original shadow document images; The noise-adding processing module is used to add noise to each of the original shadow document images and the original shadow mask corresponding to each of the original shadow document images, to obtain the noise-added shadow document image and the noise-added shadow mask corresponding to each of the original shadow document images. The model analysis module is used to construct a training model. It performs model analysis on the training model using all the original shadowed document images, all the noisy shadowed document images, and all the noisy shadow masks to obtain an image deshading model. The import module is also used to import the image of the shadow document to be processed; The shadow processing module is used to perform shadow removal processing on the shadow document image to be processed using the image shadow removal model, so as to obtain the shadow removal result of the document image.

9. A document image shadow removal apparatus, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the document image shadow removal method as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the document image shadow removal method as described in any one of claims 1 to 7.