Document image illumination recovery method and system based on diffusion model
By applying a diffusion model-based method in document image lighting recovery, using the conditional diffusion module and the light recovery module to predict and restore the lighting information, the problems of poor lighting recovery effect, low efficiency and inconsistent results in the prior art are solved, and efficient and accurate lighting recovery effect is achieved.
Patent Information
- Application Number
- CN202510237266.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-02
- Publication Date
- 2025-06-17
AI Technical Summary
The prior art has poor results in restoring document images lighting, low processing efficiency and inconsistent generation results.
Using a diffusion model-based method, the lighting information of the document image is gradually restored by constructing a conditional diffusion module and a lighting recovery module, and the U-Net model is used to predict the conditional scoring function and degradation mapping.
Effective recovery of document image lighting is achieved, the efficiency of processing high-resolution images is improved, and the consistency and quality of the generation results are significantly improved.
Smart Images

Figure CN120163753A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, specifically to a method for restoring the illumination of document images based on a diffusion model. The present invention also relates to a restoration system for implementing the method for restoring the illumination of document images based on a diffusion model. It aims to solve the problems of poor restoration effect, low processing efficiency, and inconsistent generation results existing in the prior art when restoring the illumination of document images. Background Art
[0002] With the advent of the digital age, the digital preservation of paper documents has become increasingly important, and the preservation and management of paper documents are gradually transformed into electronic forms. However, during the scanning or photographing process of paper documents, due to uneven illumination conditions, shadows often occur, resulting in a decline in the quality of the scanned document images, affecting the reading experience and the accuracy of OCR (Optical Character Recognition). Traditional illumination restoration methods mostly rely on image preprocessing and partition stitching strategies. However, when dealing with high-resolution document images, these methods often have problems such as high computational complexity, unsatisfactory restoration effects, and inconsistent generation results.
[0003] In recent years, with the rapid development of deep learning technology, especially the powerful performance of diffusion models (Diffusion Models) in image generation and enhancement tasks, it provides a new idea for the illumination restoration of document images. The diffusion model learns the data distribution by gradually adding and removing noise, and can effectively capture the internal features of the data, providing the possibility for achieving high-quality illumination restoration.
[0004] Related patent literature: CN104156916A discloses a light field projection method for scene illumination restoration, belonging to the technical field of virtual reality. This method designs a device for light field projection using a projector and a lens array, converts the illumination data into a light field for projection, and restores the scene illumination. The specific method includes: first, converting the sampled light measurement map sequence into a four-dimensional light field representation of the scene illumination; then, clustering the point light sources on the light source plane, generating projection sub-images for each clustered point light source, and performing distortion correction and image stitching; finally, after geometric calibration and brightness calibration, using the projector to project the projection image onto the lens array, and the light rays are refracted by the lens to form a planar light field on the other side to restore the illumination of the original scene.
[0005] The above technologies cannot solve the problems of poor restoration effect, low processing efficiency, and inconsistent generation results existing in the restoration of document image illumination. Summary of the Invention
[0006] The object of the present invention is to provide a method for restoring the illumination of a document image based on a diffusion model, which can effectively restore the illumination of the document image and improve the efficiency of processing high-resolution images, so as to solve the problems of poor restoration effect, low processing efficiency and inconsistent generation results in the prior art when restoring the illumination of document images.
[0007] To this end, another object of the present invention is to provide a restoration system for implementing the above-mentioned method for restoring the illumination of a document image based on a diffusion model.
[0008] To solve the above technical problems, the technical solutions adopted by the present invention are as follows:
[0009] A method for restoring the illumination of a document image based on a diffusion model, the technical solution thereof lies in that it includes the following steps:
[0010] S100: Construct a conditional diffusion module, which takes a shadow image as a conditional input and inputs a noise image at the same time to predict a conditional score function; wherein, the noise image is generated by adding Gaussian noise to an image without shadow, the conditional diffusion module uses a U-Net model to predict the noise level of the noise image, the U-Net model includes four encoder blocks and four decoder blocks, features are extracted by the encoder blocks, and the decoder blocks combine the shadow image and the extracted features to predict the score function, and the conditional score function contains the feature information between the shadow and the image without shadow, and is used to assist the following illumination restoration module to complete the image restoration task.
[0011] S200: Design a document image illumination restoration module, which adopts a coarse-to-fine processing strategy. First, a preliminary degradation map is predicted by a coarse network, and then a more accurate degradation map is predicted by a fine network in combination with the shadow image, the score function and the preliminary degradation mapping; the above-mentioned conditional diffusion module and illumination restoration module constitute the IlluRestDiff model (or framework, model framework, A model).
[0012] S300: Apply the accurate degradation mapping to the shadow image to achieve the illumination restoration of the document image; the finally output image is the document image after being repaired and shadow removed.
[0013] S400: Construct a high-resolution and complex illumination restoration test set with rich layout, that is, construct a high-resolution data set to evaluate the generalization ability of the model.
[0014] In some actual implementation processes, the light degradation model assumes that the shadowless image If is affected by the degradation mapping m, thus generating the shadow image is. The mathematical representation of this model is as follows: is = If + m, where + represents element-wise addition, and the degradation mapping m ∈ [-1, 1] is a constant. Since the distribution of shadow regions formed under natural light is usually uneven, our goal is to learn the difference distribution between the shadow image and the shadowless image through the model, so as to achieve the purpose of shadow removal. The diffusion model consists of two parameterized Markov chains, which respectively perform the forward diffusion process and the reverse generation process. The forward process gradually adds Gaussian noise to the original data according to a preset noise schedule, thus destroying the data distribution; the reverse process gradually removes the noise through the learned parameterized model to restore the original data distribution. In order to further improve the accuracy of light restoration, a conditional diffusion module (CMD) is designed. This module takes the shadow image as the conditional input, predicts the scoring function under the condition (i.e., the gradient of the data distribution), and uses this scoring function to assist the light restoration module to more accurately estimate the degradation mapping function m, so as to better restore the light information of the image.
[0015] In the above technical solution, a preferred technical solution may be that in step S100, the conditional diffusion module consists of four encoder blocks and four decoder blocks. The number of output channels of each encoder block is 64, 128, 256, and 256 in sequence. Each block contains normalization, rotation function, and ResNet block.
[0016] The light restoration module adopts a coarse-to-fine strategy. Through the stacked U-net network structure, first, the coarse U-net predicts the preliminary degradation map m p , and then the fine U-net combines the shadow image I s , the scoring function, and the preliminary degradation map m p , and predicts the fine degradation map m.
[0017] In step S100, the conditional diffusion module uses the generation ability of the diffusion model to process the input shadowed document image, predicts its scoring function, and makes the document image blurred in the time step by gradually adding Gaussian noise, and finally gradually restores the clarity of the image in the reverse process. The U-net network architecture is adopted, and the image features are used as conditions to combine the input noisy image to predict the feature information of the shadowless region, thereby assisting the light restoration module.
[0018] The process of using the diffusion model to generate the conditional scoring function includes the following steps:
[0019] First, Gaussian noise is introduced into the shadowless image to generate a noisy image, that is, Gaussian noise is gradually added to the original data through a parameterized Markov chain to achieve the forward process. Next, the U-net model is used to predict the noise level in the noisy image, that is, the original data distribution is gradually recovered from the noise through a Markov chain to achieve the reverse generation process. Using the basic principle of the diffusion model, that is, training is carried out by gradually adding noise to clean data (here it is the shadowless image), and learning how to recover from the noise to the original data, including:
[0020] Forward process: Gradually add noise to the noise-free data to generate a series of noisy images with different degrees of noise. The purpose of this step is to let the model learn how to transform the original shadowless image into an image containing different noise levels.
[0021] Reverse process: The model learns how to gradually recover from noisy images with different degrees of noise to the original noise-free image. Finally, what the model learns is how to generate a shadowless image that conforms to the data distribution from random noise.
[0022] By adding noise to the shadowless image, the model can learn how to remove the noise in the reverse process, helping the model master the ability to generate shadowless images from random noise, and then realizing the optimization of the conditional scoring function.
[0023] In step S200, both the coarse network and the fine network are of U-Net structure, where the coarse network is used to preliminarily predict the degradation map, and the fine network is used to further refine the prediction result.
[0024] In step S200, the implementation process of the illumination recovery module in the proposed IlluRestDiff model includes: The document illumination recovery module is designed to perform illumination recovery on high-resolution document images, avoiding complex preprocessing and postprocessing steps. Pixel-by-pixel illumination recovery is achieved by predicting the degradation map caused by shadows. A coarse-to-fine processing strategy is adopted to correct the document image, and a document illumination recovery framework is designed. As Figure 3 shown, a stacked network structure is adopted, which consists of two U-net networks sharing the same structural components. Based on the U-net structure therein, all skip connections are removed and the number of filters is reduced by half to reduce the computational complexity. The coarse U-net takes the shadow image as input, predicts the degradation map, and then inputs this map together with the shadow image and the scoring function into the fine U-net to predict the final degradation map. Finally, the fine degradation map is used for the illumination recovery of the shadow image.
[0025] The illumination restoration module adopts a coarse-to-fine restoration strategy, including two cascaded U-net networks. First, the coarse U-net predicts the degradation mapping of the document image, estimating the degradation information caused by uneven illumination in the document image. Second, the fine U-net further optimizes based on the coarse mapping, and uses the scoring function predicted by the degradation mapping and the conditional diffusion module to perform precise illumination restoration processing on the document image.
[0026] In the above technical solution, the preferred technical solution may be that the method for restoring the illumination of a document image based on a diffusion model further includes constructing a loss function, and the loss function includes a conditional diffusion module loss function, a coarse illumination restoration loss function, and a fine illumination restoration loss function, which are used to jointly train the conditional diffusion module and the illumination restoration module. The calculation formula of the total loss function is:
[0027]
[0028] where the total loss is the conditional diffusion model loss the coarse illumination restoration loss and the fine illumination restoration loss weighted sum.
[0029] The calculation formula of the conditional diffusion module loss function is:
[0030]
[0031] where G θ (I s , x t , t) represents the output of the conditional diffusion model CMD. It accepts the shadow image I s , the time step t, and the noisy image x t as inputs, and outputs the predicted value of the noise by the model. θ is the Gaussian noise actually added to the shadowless image, from the standard normal distribution ∈~N(0, I). x t is the noisy image after t steps of diffusion, and the calculation formula is where I f is the shadowless image, t represents the diffusion time step, and ||·||1 represents the L1 norm, which is used to measure the difference between the noise predicted by the model and the true noise.
[0032] The calculation formula of the coarse illumination restoration loss function is:
[0033]
[0034] where C θ (I s ) represents the output of the coarse U-net, that is, the coarsely predicted degradation mapping, m gtRepresents the true value of the degradation mapping, i.e., the shaded image I s and the shadowless image I f The difference m gt = I s - I f .
[0035] The above-mentioned document image illumination restoration method based on the diffusion model is characterized in that the calculation formula of the fine illumination restoration loss function is:
[0036]
[0037] Among them, R θ (I s , m p , score) represents the output of the fine U-net, that is, the fine predicted degradation mapping m, and m gt is the true degradation mapping. V(I f , I r ) is the VGG perceptual loss, which represents the L1 loss between the true shadowless image I f and the corrected image I r after extracting features using the pre-trained VGG-19 network. The VGG network is used to measure the similarity of high-level semantic features of images. SSIM(I f , I r ) is the structural similarity measure (SSIM), which is used to measure the structural similarity between the true shadowless image I f and the corrected image I r . The larger the SSIM value, the more similar the structures of the two images. γ is the weight of the VGG perceptual loss, which controls the trade-off between different loss terms.
[0038] In step S400, the data sources of the high-resolution dataset include: a self-made high-resolution document image dataset, a synthetic document image dataset, and a real-scene dataset for document image geometric correction and illumination restoration.
[0039] A document image illumination restoration system based on the diffusion model, which uses the above-mentioned document image illumination restoration method based on the diffusion model, is characterized in that the document image illumination restoration system based on the diffusion model includes:
[0040] A conditional diffusion module for predicting the conditional score function with the shaded image as the condition;
[0041] An illumination restoration module, including a coarse network and a fine network. The coarse network is used to predict the preliminary degradation map, and the fine network combines the shaded image, the score function, and the preliminary degradation map to predict the fine degradation map;
[0042] A restoration unit for applying a fine degradation map to a shadow image to achieve illumination restoration of a document image.
[0043] The conditional diffusion module includes at least one encoder block and at least one decoder block. The encoder block extracts features, and the decoder block combines the shadow image and the extracted features to predict a score function. Both the coarse network and the fine network are of U-Net structure. The coarse network is configured with a smaller number of convolutional layers and channels to reduce computational complexity, and the fine network is configured with more convolutional layers and channels to refine the prediction result.
[0044] The document image illumination restoration system based on the diffusion model further includes a loss calculation unit for calculating the conditional diffusion module loss, the coarse network loss, and the fine network loss, and jointly optimizing the parameters of the conditional diffusion module and the illumination restoration module.
[0045] Construct a high-resolution and complex illumination restoration test set with rich layouts to evaluate the generalization ability of the model.
[0046] The present invention provides a method and system for document image illumination restoration based on a diffusion model. The method realizes effective restoration of document image illumination by introducing a conditional diffusion model CMD and a document image illumination restoration module. The CMD uses the shadow image as a condition to predict a conditional score function, providing auxiliary information for the illumination restoration module. The illumination restoration module adopts a coarse-to-fine processing strategy. First, a preliminary degradation map is predicted through a coarse U-net network, and then, by combining the shadow image, the score function, and the preliminary degradation map, a fine degradation map is predicted through a fine U-net network, and finally, it is applied to the shadow image to achieve illumination restoration. Experimental results show that the IlluRestDiff model performs well in processing high-resolution document images, not only having a good restoration effect but also high processing efficiency and strong generalization ability. The present invention has broad application prospects and promotion value for improving the reading experience of document images and OCR performance.
[0047] Compared with the prior art, the present invention provides a method for efficiently processing high-resolution images. By directly predicting the global degradation map, it avoids complex preprocessing and postprocessing steps, improving the efficiency of processing high-resolution images. At the same time, the present invention utilizes the powerful learning ability of the diffusion model to more accurately restore the illumination of document images, retain document details and color information, and achieve a better restoration effect. Moreover, through experiments on multiple data sets, the results show that the method of the present invention performs well in restoring document images under different illumination conditions and has strong generalization ability.
[0048] In summary, the present invention provides a method and system for document image illumination restoration based on a diffusion model, which effectively restores the illumination of document images, improves the efficiency of processing high-resolution images, and solves the problems of poor restoration effect, low processing efficiency, and inconsistent generation results in the prior art when restoring the illumination of document images.
[0049] The following are the drawings for the IlluRestDiff model, which are intended to more intuitively show the system architecture, workflow, and core modules of the present invention. Description of the Drawings
[0050] Figure 1 It is a flowchart of the method for document image illumination restoration based on a diffusion model provided by an embodiment of the present invention.
[0051] Figure 2 It is a system architecture diagram of the IlluRestDiff model.
[0052] Figure 3 It is a workflow diagram of the illumination restoration module.
[0053] Figure 4 It is a model training flowchart of the method for document image illumination restoration based on a diffusion model provided by an embodiment of the present invention.
[0054] Figure 5 It is a schematic diagram of the functional modules of the method for document image illumination restoration based on a diffusion model.
[0055] Figure 6 It is a diagram showing experimental results. Detailed Embodiments
[0056] To make the invention objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on these embodiments, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of the present invention.
[0057] Embodiment 1: Refer to Figure 1 、 Figure 2 、 Figure 3 、 Figure 4 、 Figure 5 、 Figure 6 , Figure 2 It is a system architecture diagram of the IlluRestDiff model. Figure 2Shows the overall architecture of the IlluRestDiff model, including the Conditional Diffusion Module (CMD) and the Document Image Illumination Restoration Module. The Conditional Diffusion Module receives the shadow image as a condition and predicts the conditional score function; the Illumination Restoration Module then uses the score function output by the CMD to predict and remove the degraded image through a coarse-to-fine processing strategy to achieve illumination restoration. Figure 3 It is the workflow diagram of the Illumination Restoration Module. Figure 3 Details the workflow of the Illumination Restoration Module, including the coarse U-net predicting the initial degradation map, combining the conditional score function and the shadow image input to the fine U-net for refined prediction, and applying the predicted fine degradation map to the shadow image to achieve illumination restoration. Figure 4 It is the model training flowchart of the document image illumination restoration method based on the diffusion model. Figure 4 Describes the training process of the IlluRestDiff model, including key steps such as data preparation, model initialization, forward propagation, loss calculation, backpropagation, and parameter update. Figure 6 It is the experimental result display diagram. Figure 6 By comparing the document images before and after the experiment, intuitively shows the effect of the IlluRestDiff model in illumination restoration. It can include the original shadow image, the shadowless image after restoration, and the comparison results with other methods.
[0058] The document image illumination restoration method based on the diffusion model includes the following steps:
[0059] S100: Construct a Conditional Diffusion Module that takes the shadow image as a conditional input and simultaneously inputs a noise image, and predicts the conditional score function; wherein, the noise image is generated by adding Gaussian noise to the shadowless image, the Conditional Diffusion Module uses a U-Net model to predict the noise level of the noise image, the U-Net model includes four encoder blocks and four decoder blocks, extracts features through the encoder blocks, and the decoder blocks combine the shadow image and the extracted features to predict the score function. The conditional score function contains the feature information between the shadow and the shadowless image, and is used to assist the following Illumination Restoration Module to complete the image restoration task.
[0060] S200: Design a Document Image Illumination Restoration Module, adopting a coarse-to-fine processing strategy. First, predict the initial degradation map through a coarse network, and then combine the shadow image, the score function, and the initial degradation map to predict a more accurate degradation map through a fine network; the above Conditional Diffusion Module and Illumination Restoration Module constitute the IlluRestDiff model (or framework, model framework, Model A).
[0061] S300: Apply the precise degradation mapping to the shadow image to achieve the illumination restoration of the document image; the finally output image is the document image after repair and shadow removal.
[0062] S400: Construct a complex illumination restoration test set with high resolution and rich layout, that is, construct a high-resolution data set to evaluate the generalization ability of the model.
[0063] In some actual implementation processes, the illumination degradation model assumes that the shadow-free image If is affected by the degradation mapping m, thus generating the shadow image is. The mathematical representation of this model is as follows: is = If + m, where + represents element-wise addition, and the degradation mapping m ∈ [-1, 1] is a constant. Since the distribution of the shadow area formed under natural illumination is usually uneven, our goal is to learn the difference distribution between the shadow image and the shadow-free image through the model, so as to achieve the purpose of shadow removal. The diffusion model consists of two parameterized Markov chains, which respectively execute the forward diffusion process and the reverse generation process. The forward process gradually adds Gaussian noise to the original data according to a preset noise schedule, thus destroying the distribution of the data; the reverse process gradually removes the noise through the learned parameterized model to restore the original data distribution. To further improve the accuracy of illumination restoration, a conditional diffusion module (CMD) is designed. This module takes the shadow image as the conditional input, predicts the scoring function (i.e., the gradient of the data distribution) under the condition, and uses this scoring function to assist the illumination restoration module to more accurately estimate the degradation mapping function m, so as to better restore the illumination information of the image.
[0064] In step S100, the conditional diffusion module consists of four encoder blocks and four decoder blocks. The output channel numbers of each encoder block are 64, 128, 256, and 256 in sequence. Each block contains normalization, rotation function, and ResNet block.
[0065] The illumination restoration module adopts a coarse-to-fine strategy. Through the stacked U-net network structure, first, the coarse U-net predicts the preliminary degradation map m p , and then the fine U-net combines the shadow image I s , the scoring function, and the preliminary degradation map m p to predict the fine degradation map m.
[0066] In step S100, the conditional diffusion module uses the generative ability of the diffusion model to process the input shaded document image, predict its score function, make the document image blurred in the time steps by gradually adding Gaussian noise, and finally gradually restore the clarity of the image in the reverse process. The U-net network architecture is adopted, using image features as conditions and combining with the input noisy image to predict the feature information of the shadow-free area, thus assisting the illumination restoration module.
[0067] The process of using the diffusion model to generate the conditional score function includes the following steps:
[0068] First, Gaussian noise is introduced into the shadow-free image to generate a noisy image, that is, Gaussian noise is gradually added to the original data through a parameterized Markov chain to achieve the forward process. Next, the U-net model is used to predict the noise level in the noisy image, that is, the original data distribution is gradually restored from the noise through a Markov chain to achieve the reverse generation process. Using the basic principle of the diffusion model, that is, training by gradually adding noise to clean data (here is the shadow-free image), learning how to recover from noise to the original data, including:
[0069] Forward process: Gradually add noise to the noise-free data to generate a series of noisy images with different degrees. The purpose of this step is to let the model learn how to transform the original shadow-free image into an image containing different noise levels.
[0070] Reverse process: The model learns how to gradually restore from noisy images with different degrees to the original noise-free image. Finally, what the model learns is how to generate a shadow-free image that conforms to the data distribution from random noise.
[0071] By adding noise to the shadow-free image, the model can learn how to remove noise in the reverse process, helping the model master the ability to generate shadow-free images from random noise, and then realizing the optimization of the conditional score function.
[0072] In step S200, both the coarse network and the fine network are of U-Net structure, where the coarse network is used to preliminarily predict the degradation map, and the fine network is used to further refine the prediction result.
[0073] In step S200, the implementation process of the illumination restoration module in the proposed IlluRestDiff model includes: The document illumination restoration module is designed to perform illumination restoration on high-resolution document images, avoiding complex preprocessing and postprocessing steps. Pixel-by-pixel illumination restoration is achieved by predicting the degradation mapping caused by shadows. A coarse-to-fine processing strategy is adopted to correct the document image, and a document illumination restoration framework is designed. As Figure 3As shown, a stacked network structure is adopted, which consists of two U-net networks sharing the same structural components. Based on the U-net structure therein, all skip connections are removed and the number of filters is halved to reduce the computational complexity. The coarse U-net takes the shaded image as input, predicts the degradation map, and then inputs this map together with the shaded image and the scoring function into the fine U-net to predict the final degradation map. Finally, the fine degradation map is used for the illumination restoration of the shaded image.
[0074] The illumination restoration module adopts a coarse-to-fine restoration strategy, including two cascaded U-net networks. First, the coarse U-net predicts the degradation map of the document image, estimating the degradation information caused by uneven illumination in the document image. Second, the fine U-net further optimizes based on the coarse map, and uses the degradation map and the scoring function predicted by the conditional diffusion module to perform precise illumination restoration processing on the document image.
[0075] The described method for illumination restoration of document images based on the diffusion model further includes constructing a loss function, which includes a conditional diffusion module loss function, a coarse illumination restoration loss function, and a fine illumination restoration loss function, and is used to jointly train the conditional diffusion module and the illumination restoration module. The calculation formula of the total loss function is:
[0076]
[0077] where the total loss is the conditional diffusion model loss the coarse illumination restoration loss and the fine illumination restoration loss is the weighted sum of.
[0078] The calculation formula of the conditional diffusion module loss function is:
[0079]
[0080] where G θ (I s , x t , t) represents the output of the conditional diffusion model CMD. It takes the shaded image I s , the time step t, and the noisy image x t as inputs and outputs the predicted value of the noise by the model. θ is the Gaussian noise actually added to the shadowless image, coming from the standard normal distribution ∈~N(0, I). x t is the noisy image after t steps of diffusion, and the calculation formula is where I f is the shadowless image, t represents the diffusion time step, and ||·||1 represents the L1 norm, which is used to measure the difference between the noise predicted by the model and the real noise.
[0081] The calculation formula of the rough illumination recovery loss function is as follows:
[0082]
[0083] Among them, C θ (I s ) represents the output of the rough U-net, that is, the roughly predicted degradation map, and m gt represents the true value of the degradation map, that is, the difference m s between the shaded image I f and the shadowless image I gt = i s - i f .
[0084] The calculation formula of the fine illumination recovery loss function is as follows:
[0085]
[0086] Among them, R θ (I s , m p , score) represents the output of the fine U-net, that is, the finely predicted degradation map m, and m gt is the true degradation map. V(I f , I r ) is the VGG perceptual loss, which represents the L1 loss between the true shadowless image I f and the corrected image I r after extracting features using the pre-trained VGG-19 network. The VGG network is used to measure the similarity of high-level semantic features of images. SSIM(I f , I r ) is the structural similarity measure (SSIM), which is used to measure the structural similarity between the true shadowless image I f and the corrected image I r . The larger the SSIM value, the more similar the structures of the two images. γ is the weight of the VGG perceptual loss, which controls the trade-off between different loss terms.
[0087] In step S400, the data sources of the high-resolution dataset include: a self-made high-resolution document image dataset, a synthetic document image dataset, and a real-scene dataset for document image geometric correction and illumination recovery.
[0088] The IlluRestDiff model is trained on the DocProj dataset and tested on the DocUNet benchmark and the newly introduced DocIll dataset. The DocProj dataset contains 2700 synthetic images, while the DocUNet benchmark includes 130 real document images with various contents and formats.
[0089] As a preferred embodiment of the present invention, the DocIll dataset is created to evaluate the performance of the model on high-resolution images. DocIll contains 110 images captured in real scenes with different lighting conditions, most of which have a resolution of 3468×4624. Due to its high-resolution characteristics, DocIll places higher requirements on the illumination restoration model, especially in ensuring image details and overall illumination consistency. The test set contains rich document content, different types of documents such as tables, texts, and mixed images. Documents with different layouts pose a challenge to the generalization ability of the model, ensuring that the model can not only perform well on documents of a specific format, but also cope with different types of document illumination restoration tasks.
[0090] DocIll images are taken by different mobile devices in real-world scenarios, which leads to various degrees of illumination distortion in document images, including large shadows, uneven lighting, color distortion, etc. These illumination distortion phenomena add complexity to the test set, making it closer to actual application scenarios.
[0091] Due to different shooting angles and equipment, some DocIll test sets contain geometric distortions. These distortions affect the readability of the document and increase the difficulty of illumination restoration and image inpainting. To alleviate this problem, some images in the test set were geometrically corrected using the DocTR (Document Text Recognition) method, but the effects of illumination distortion were retained to test the ability of the illumination restoration model.
[0092] The document images in the DocIll test set include not only black and white documents, but also a large number of documents with rich color information. The color complexity in the documents increases the difficulty of model restoration, especially the need to preserve the original color and details of the document while performing lighting restoration.
[0093] Since the images in the DocIll test set are taken using mobile devices in different environments, the lighting conditions in the scenes vary greatly, such as natural light, indoor light, partial occlusion, etc. Such scene diversity makes DocIll an important test set for evaluating the performance of models in real environments.
[0094] The introduction of DocIll is mainly to make up for the lack of high-resolution, complex illumination distortion, and real-scene shooting data in the existing test sets, and to further verify the generalization ability and illumination restoration effect of the IlluRestDiff model in diverse scenarios. By testing on DocIll, the performance of the model in dealing with common illumination problems in practical applications can be effectively examined, and feedback can be provided for further improving the algorithm.
[0095] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following combines the Figure 2 of the present invention to provide a complete and clear description of the technical solutions in the embodiments of the present invention:
[0096] Three key datasets are used to evaluate the performance and generalization ability of the model, namely the DocProj dataset, the DocUNet Benchmark dataset, and the newly introduced DocIll test set. The design and application of these datasets verify the effectiveness and advantages of the IlluRestDiff model in the task of document image illumination restoration at different levels.
[0097] The DocProj dataset is a synthetic document image dataset initially proposed by the researchers of the DocProj method. This dataset contains 2,700 synthetic images, each with a resolution of 2400×1800. The images in the DocProj dataset have undergone geometric distortion and illumination distortion processing to simulate the distortion of real document images under different illumination conditions. Its high resolution and illumination distortion characteristics make it an ideal choice for training document illumination restoration models. To better adapt to the illumination restoration task of the IlluRestDiff model, the authors further preprocessed the DocProj dataset, removed some geometric distortion parts in the background, and reduced the difference between distorted images and undistorted images, making it more suitable for the training of the illumination restoration task.
[0098] The DocUNet Benchmark dataset is a real-world dataset for geometric correction and illumination restoration of document images and is also a widely used standard benchmark in the current field of document restoration. The DocUNet dataset contains 130 document images taken in real-world scenarios, covering documents of different contents and formats. These images often have complex problems such as geometric distortion, uneven illumination, and shadows. Compared with the DocProj dataset, the main challenge of DocUNet is to handle geometric distortion, especially in cases where the document is severely distorted, and it is necessary to combine geometric correction and illumination restoration techniques. To focus on the illumination restoration task, the authors used the DocTr method to perform geometric correction on the DocUNet dataset, thereby reducing the impact of geometric distortion on the experimental results and enabling the model to focus on solving problems such as uneven illumination, shadows, and light spots. The diversity and complexity of the DocUNet dataset provide a stringent evaluation criterion for the IlluRestDiff model, helping to verify the model's performance in real-world scenarios.
[0099] To further verify the generalization ability of the IlluRestDiff model, the authors designed and introduced the DocIll test set. DocIll is a newly constructed high-resolution document image dataset containing 110 images, each with a resolution of up to 3468×4624. The images in DocIll come from different real-world scenarios and are taken by mobile devices, thus presenting complex illumination distortion phenomena, including shadows, light spots, and local overexposure or underexposure. Different from the DocProj and DocUNet datasets, the images in DocIll not only have higher resolutions but also have more diverse contents and layouts, covering document structures with a mixture of tables, text, and images. In addition, the document images in DocIll also contain rich color information, not limited to black-and-white documents, which further increases the difficulty of illumination restoration, especially in maintaining image color consistency and details. Some images also have geometric distortion problems, and the authors also used the DocTr method for geometric correction to evaluate the performance of IlluRestDiff in solving illumination problems.
[0100] The combination of the DocProj, DocUNet, and DocIll datasets provides a comprehensive test platform for the IlluRestDiff model. DocProj is used for model training, DocUNet serves as the standard for evaluating the model's performance in real-world scenarios, and DocIll further verifies the model's generalization ability under high-resolution and complex illumination conditions. The diversity and complexity of these datasets ensure that the IlluRestDiff model can handle illumination restoration tasks in various real-world scenarios and perform well in improving OCR performance.
[0101] The IlluRestDiff model (framework) consists of two main modules, namely the Conditional Diffusion Module (CMD) and the Illumination Restoration Module (IRM). These two modules work together to address the problem of illumination distortion in document images, especially complex shadows and uneven illumination, and provide high-quality illumination restoration for document images.
[0102] The Conditional Diffusion Module (CMD) is one of the core components of the IlluRestDiff model and is responsible for generating features of shadow-free document images through Denoising Diffusion Probabilistic Models (DDPM). This module uses the forward diffusion and reverse restoration processes to gradually restore the document images covered by shadows. Specifically, in the forward process, Gaussian noise is gradually added to the document image to make it gradually blurred, while in the reverse process, the clarity of the image is restored through the inverse diffusion process. The CMD module adopts the classic U-Net architecture, which consists of four encoders and four decoder blocks, and can capture multi-scale feature information and reconstruct image details. At the input, the CMD module takes the document image with shadows as the conditional input and guides the illumination restoration process by predicting the score function. This score function reflects the feature differences between the shadow area and the shadow-free area and provides key information for restoring shadows and illumination distortion.
[0103] The Illumination Restoration Module (IRM) is responsible for performing illumination repair on the document image according to the score function provided by the CMD module. The IRM module adopts a coarse-to-fine illumination restoration strategy and consists of two cascaded U-Net networks. First, the coarse U-Net receives the shadow image as input and predicts the degradation map of this image, that is, the image distortion information caused by uneven illumination. The output of the coarse U-Net is a preliminary degradation map, which helps to determine the approximate range and location of the illumination distortion. Next, the fine U-Net further optimizes the degradation map generated by the coarse U-Net, combines the score function of the CMD module and the shadow image for refinement processing, and generates a more accurate degradation map. This fine degradation map can more accurately eliminate shadows and maintain the details and color information in the image. Compared with the block stitching operation of traditional methods, the IlluRestDiff model avoids complex block stitching steps by directly predicting the global degradation map, ensuring the consistency and integrity of the image after illumination restoration.
[0104] In addition, the two modules in the IlluRestDiff model are jointly trained. The conditional diffusion module and the illumination restoration module are co-optimized through different loss functions during training. The loss function of the CMD module is used to learn the score function, ensuring that the diffusion model can accurately recover the feature information of the shadowless image. The rough U-Net and the fine U-Net in the illumination restoration module each have their own supervised losses. The rough U-Net focuses on generating a preliminary degradation map, while the fine U-Net further optimizes the image quality through additional VGG losses and structural similarity (SSIM) losses, ensuring that the restored image is not only visually closer to the distortion-free original image but also structurally consistent.
[0105] Overall, these two modules of the IlluRestDiff model are tightly integrated to jointly achieve an efficient illumination restoration process. The CMD module learns the latent features of shadows and illumination distortions through the diffusion model, while the IRM module gradually repairs the illumination problems in the document image through coarse-to-fine degradation map prediction, especially showing excellent performance when dealing with high-resolution and complex illumination scenarios.
[0106] The training process of the IlluRestDiff model is a process of jointly optimizing the Conditional Diffusion Module (CMD) and the Illumination Restoration Module (IRM). The entire training process aims to achieve efficient illumination restoration of document images through the synergy of the two modules, especially to eliminate shadows and other illumination distortions in the images.
[0107] At the beginning of training, the Conditional Diffusion Module (CMD) is trained using the framework of the Denoising Diffusion Probability Model (DDPM). This module simulates the degradation and restoration of images through forward and reverse diffusion processes. In the forward diffusion process, the CMD module gradually adds Gaussian noise to the original shadowless image, making it gradually blur at multiple time steps. At each time step t, the CMD module receives a shadow image (as a conditional input) and the noisy image, and predicts the noise score function at the current step. This score function reflects the restoration direction and features of the image at this step, providing key information for illumination restoration. The goal of the CMD module is to learn to recover the shadowless original image features from the noisy image, which is achieved by minimizing the error of the noise prediction.
[0108] Meanwhile, the Illumination Restoration Module (IRM) is jointly trained with the Conditional Diffusion Module (CMD). The IRM adopts a coarse-to-fine strategy and uses two cascaded U-net networks for illumination restoration. First, the coarse U-net receives the shadow image as input and predicts its corresponding degradation map, which represents the distortion information caused by uneven illumination. The training objective of the coarse U-net is to minimize the error between the predicted degradation map and the ground-truth degradation map. Since the ground-truth degradation map is not directly available in the training data, the difference between the shadow image and the shadow-free image is used to generate a pseudo-ground-truth degradation map as the supervision signal.
[0109] After obtaining the coarse degradation map, the fine U-net further refines it. The fine U-net not only receives the coarse degradation map but also combines the score function output by the CMD module and the shadow image as input to generate a more accurate degradation map. The training objective of the fine U-net is to optimize the illumination restoration result through higher image quality and detail preservation. To ensure that the generated image is structurally consistent with the ground-truth shadow-free image, the fine U-net introduces the VGG perceptual loss and the Structural Similarity Index (SSIM) loss during training. These loss functions help the fine U-net maintain the text and content information of the document image while generating high-quality images.
[0110] The entire training process is driven by jointly optimizing the loss functions of the CMD module and the IRM module. The CMD module optimizes the learning of the score function by minimizing the noise prediction error, while the IRM module optimizes the illumination restoration result by minimizing the degradation map error, the VGG loss, and the SSIM loss. The synergistic effect of the two modules enables the IlluRestDiff model to effectively remove shadows and illumination distortions while preserving the details and content of the document image.
[0111] During the training process, the model is trained and validated on multiple document image datasets, including the synthetic DocProj dataset and the real-world DocUNet and DocIll datasets. By training on these diverse datasets, the IlluRestDiff model can learn the features under different illumination conditions and shadow types, demonstrating strong generalization ability, especially in the illumination restoration of high-resolution document images.
[0112] During the testing process of the IlluRestDiff model, after the model has been fully trained on the training set, multiple real-world scenarios and synthetic datasets are used to evaluate its performance in the task of document image illumination restoration. The main purpose of the test is to verify the model's performance in eliminating shadows, restoring illumination uniformity, and improving the accuracy of optical character recognition (OCR). During the test, the model demonstrates its generalization ability and practical application effects by processing high-resolution images and document images with complex illumination conditions.
[0113] The test dataset consists of two main parts: the DocUNet Benchmark and the DocIll test set. DocUNet is a widely used real-world dataset that contains document images in various contents and formats, and common illumination distortions such as shadows, local overexposure, or underexposure frequently occur in this dataset. The DocIll test set is specifically designed to evaluate the illumination restoration ability of high-resolution document images. The document images not only have high resolution but also contain rich color information and diverse illumination distortion phenomena.
[0114] During the test, first, the document images in the test dataset are input into the IlluRestDiff model. The conditional diffusion module (CMD) of the model utilizes the features learned during the training process to restore the shadow-free document features through forward and reverse diffusion processes. In the inference stage, the CMD module gradually removes noise according to the input shadowed image through the conditional diffusion mechanism, generating the image features after shadow elimination. These features are passed to the illumination restoration module for further processing. Next, the illumination restoration module (IRM) performs a coarse-to-fine illumination restoration process based on the score function information provided by the CMD module. First, the rough U-Net predicts the preliminary degradation map of the image, identifying the illumination non-uniform regions in the image. The fine U-Net further optimizes this map by refining the shadow edges and illumination transition regions, generating a more accurate degradation map. Finally, the illumination restoration module eliminates the shadows and illumination distortions in the image according to the fine degradation map, outputting a document image with uniform illumination and rich details.
[0115] To evaluate the performance of the IlluRestDiff model, various quantitative metrics and qualitative analyses were used during the testing process. Commonly used evaluation metrics include multi-scale structural similarity (MS-SSIM), character error rate (CER), and edit distance (ED). MS-SSIM is used to measure the structural similarity between the restored image and the original shadow-free image. A higher MS-SSIM value indicates better image restoration quality. CER and ED are used to evaluate the impact of illumination restoration on the OCR task. By comparing the OCR recognition accuracy before and after illumination restoration, it is verified whether the model can improve the accuracy of character recognition. In addition to quantitative evaluation, a large number of qualitative analyses were also carried out during the testing process to compare the effects of the IlluRestDiff model and existing methods when processing document images under different illumination conditions. Through the comparison of aspects such as shadow elimination, illumination consistency, color retention, and detail restoration, the test results show that the IlluRestDiff model can significantly improve the visual quality of document images while maintaining the integrity of the image content.
[0116] During the testing process, especially when processing high-resolution images (such as the DocIll dataset), IlluRestDiff demonstrated obvious advantages. Traditional stitching methods are prone to producing discontinuous illumination transitions on the image, while the IlluRestDiff model avoids such stitching artifacts through the prediction of global degradation mapping, making the output image have more uniform illumination and a more natural visual effect.
[0117] The model not only performs well in shadow processing but also can retain the key details and color information in the document image, which is particularly important when processing documents with complex color information.
[0118] Example 2: As Figure 5 shown, a document image illumination restoration system based on a diffusion model, which uses the document image illumination restoration method based on a diffusion model. The document image illumination restoration system based on a diffusion model includes:
[0119] A conditional diffusion module 11, which is used to predict a conditional score function with the shadow image as the condition.
[0120] The conditional diffusion module includes at least one encoder block and at least one decoder block. The encoder block is used to extract features, and the decoder block combines the shadow image and the extracted features to predict the score function.
[0121] An illumination restoration module 12, which includes a coarse network and a fine network. The coarse network is used to predict a preliminary degradation map, and the fine network combines the shadow image, the score function, and the preliminary degradation map to predict a fine degradation map.
[0122] Both the coarse network and the fine network are of U-Net structure. The coarse network is configured with a smaller number of convolutional layers and channels to reduce the computational complexity, while the fine network is configured with more convolutional layers and channels to refine the prediction results.
[0123] The restoration unit 13 is used to apply the fine degradation map to the shadow image to achieve the illumination restoration of the document image.
[0124] The above-mentioned document image illumination restoration system based on the diffusion model further includes a loss calculation unit, which is used to calculate the conditional diffusion module loss, the coarse network loss and the fine network loss, and jointly optimize the parameters of the conditional diffusion module and the illumination restoration module.
[0125] Construct a high-resolution and complex illumination restoration test set with rich layouts to evaluate the generalization ability of the model.
[0126] Compared with the current state-of-the-art illumination restoration methods, the IlluRestDiff algorithm has achieved a significant improvement of 4.55% in the ED metric and achieved the highest MS-SSIM metric, which indicates that the IlluRestDiff algorithm can better preserve the structure and details of the image, making the restored image more similar to the original distortion-free image. In addition, IlluRestDiff achieves the highest CER metric, which indicates that the image restored by IlluRestDiff has the most accurate OCR performance. Moreover, the quantitative results show that the MS-SSIM metric of the IlluRestDiff illumination restoration result is improved on the basis of the geometric correction by the DocTr-GEO algorithm, while other illumination restoration methods lead to a decrease in this metric.
[0127] In summary, the above embodiments of the present invention provide a method and system for document image illumination restoration based on the diffusion model, which realizes the effective restoration of the illumination of the document image, improves the efficiency of processing high-resolution images at the same time, and solves the problems of poor restoration effect, low processing efficiency and inconsistent generation results existing in the prior art when restoring the illumination of document images. Compared with the existing related technologies, the efficiency of the present invention in processing high-resolution images is increased by more than 13%.
Claims
1. A document image illumination restoration method based on a diffusion model, characterized in that: It includes the following steps: S100: construct a conditional diffusion module, which takes the shadow image as a conditional input and inputs the noise image at the same time, and predicts the conditional scoring function; wherein the noise image is generated by adding Gaussian noise to the shadow-free image, and the conditional diffusion module uses a U-Net model to predict the noise level of the noise image, and the U-Net model includes four encoder blocks and four decoder blocks, and features are extracted by the encoder blocks, and the decoder blocks predict the scoring function by combining the shadow image and the extracted features, and the conditional scoring function contains feature information between the shadow and shadow-free images, and is used to assist the following illumination restoration module to complete the image restoration task; S200: Design a document image illumination restoration module, adopt a coarse-to-fine processing strategy, first predict a preliminary degradation map through a coarse network, then combine the shadow image, the scoring function and the preliminary degradation map, and predict a more accurate degradation map through a fine network; the above conditional diffusion module and the illumination restoration module constitute the IlluRestDiff model; S300: applying the precise degradation mapping to the shadow image to achieve illumination restoration of the document image, and the final output image is the document image after restoration and removal of shadows; S400: Build a high-resolution, complex lighting restoration test set with rich layouts, that is, build a high-resolution dataset to evaluate the generalization ability of the model.
2. The document image illumination restoration method based on the diffusion model according to claim 1, characterized in that: In step S100, the conditional diffusion module consists of four encoder blocks and four decoder blocks, the number of output channels of each encoder block is 64, 128, 256 and 256 respectively, and each block contains normalization, rotation function and ResNet block.
3. The document image illumination restoration method based on the diffusion model according to claim 1, characterized in that: In step S200, the coarse network and the fine network are both U-Net structures, wherein the coarse network is used to preliminarily predict the degradation map, and the fine network is used to further refine the prediction result.
4. The document image illumination restoration method based on the diffusion model according to claim 1, characterized in that: It also includes constructing a loss function, which includes a conditional diffusion module loss function, a coarse illumination restoration loss function, and a fine illumination restoration loss function, which are used to jointly train the conditional diffusion module and the illumination restoration module. The calculation formula of the total loss function is: Among them, the total loss is the conditional diffusion model loss Coarse light restoration loss and fine illumination restoration loss The weighted sum of .
5. The document image illumination restoration method based on the diffusion model according to claim 4 is characterized in that: The calculation formula of the conditional diffusion module loss function is: Among them, G θ (I s ,x t ,t) represents the output of the conditional diffusion model CMD. It accepts the shadow image I s , time step t and the noisy image x t As input, the output model predicts the noise value. θ is the Gaussian noise actually added to the shadow-free image, which comes from the standard normal distribution ∈~N(0,I). t is the noisy image after t-step diffusion, and the calculation formula is Among them I f is a shadow-free image, t represents the diffusion time step, and ||·||1 represents the L1 norm, which is used to measure the difference between the noise predicted by the model and the actual noise.
6. The document image illumination restoration method based on the diffusion model according to claim 4, characterized in that: The calculation formula of the rough illumination restoration loss function is: Among them, C θ (I s ) represents the output of the rough U-net, i.e., the roughly predicted degradation map, m gt Represents the true value of the degradation map, i.e., the shadow image I s and the shadow-free image I f The difference between gt =I s -I f .
7. The document image illumination restoration method based on diffusion model according to claim 4, characterized in that: The calculation formula of the fine illumination restoration loss function is: Among them, R θ (I s ,m p , score) represents the output of the refined U-net, i.e., the refined predicted degradation map m, m gt is a true degenerate mapping. V(I f ,I r ) is the VGG perceptual loss, which means that after extracting features using the pre-trained VGG-19 network, the real shadow-free image I f And the corrected image I r The L1 loss of the VGG network is used to measure the similarity of high-level semantic features of images. SSIM (I f ,I r ) is the structural similarity metric (SSIM), which is used to measure the true shadow-free image I f And the corrected image I r The larger the SSIM value, the more similar the structures of the two images are. γ is the weight of VGG perceptual loss, which controls the trade-off between different loss terms.
8. The text image illumination restoration method based on the conditional diffusion model according to claim 1, characterized in that: In step S400, the data sources of the high-resolution data set include: a self-made high-resolution document image data set, a synthesized document image data set, and a real scene data set used for document image geometry correction and illumination restoration.
9. A document image illumination restoration system based on a diffusion model, which uses the document image illumination restoration method based on a diffusion model according to any one of claims 1 to 9, characterized in that: The document image illumination restoration system based on the diffusion model comprises: The conditional diffusion module is used to predict the conditional score function based on the shadow image; The illumination restoration module includes a coarse network and a fine network. The coarse network is used to predict a preliminary degradation map. The fine network combines the shadow image, the score function and the preliminary degradation map to predict a fine degradation map. The restoration unit is used to apply the refined degradation map to the shadow image to achieve illumination restoration of the document image.
10. The document image illumination restoration system based on diffusion model according to claim 9, characterized in that: The conditional diffusion module includes at least one encoder block and at least one decoder block, the encoder block extracts features, and the decoder block combines the shadow image and the extracted features to predict the score function.
Citation Information
Patent Citations
Light field projection method used for scene illumination recovery
CN104156916A
Cited By
Printing-shooting process image degradation simulation method, device and equipment based on image-to-image diffusion model
CN120411293A
Image degradation simulation method, device and equipment for printing-shooting process based on image-to-image diffusion model
CN120411293B