Lightweight interpretable image shadow removal model based on image reconstruction theory
Through a three-stage progressive network combined with image reconstruction theory, gradient descent and proximal mapping modules are designed, which solves the generalization ability and transparency of shadow removal technology in practical applications, and achieves an efficient and lightweight shadow removal effect.
Patent Information
- Application Number
- CN202411765355.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-04
- Publication Date
- 2025-07-08
AI Technical Summary
The existing shadow removal technology has insufficient generalization capabilities in actual applications, and the mechanism model is difficult to adapt to changeable scenarios. The data-driven model consumes a lot of computing resources and lacks transparency, making it difficult to deploy in edge devices.
A three-stage progressive network is adopted, combined with image reconstruction theory, a gradient descent module, a proximity mapping module and a spatial detail retention module are designed, and a lightweight interpretable image deshading model is constructed through the integration of image reconstruction theory and data-driven method.
Improves the shadow removal effect, ensures the interpretability of the model, and at the same time adapts to more practical application scenarios and reduces the demand for computing resources.
Smart Images

Figure BDA0005168790640000031 
Figure BDA0005168790640000032 
Figure BDA0005168790640000045
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and particularly relates to a lightweight interpretable image de-shadowing model based on image reconstruction theory. Background Art
[0002] A shadow is an area formed by an object blocking a light source, and appears as a darker part because light cannot reach directly. As a visual phenomenon widely existing in the natural environment, shadows bring rich three-dimensional sense and depth information to human visual perception. However, in computer vision tasks, the existence of shadows often brings many adverse effects. For example, in tasks such as object detection, image segmentation, visual measurement, and scene understanding, shadows can lead to misjudgment or block the true color and edges of objects, thereby affecting the recognition accuracy of the model and its generalization ability in various scenarios. These problems limit the wide application of computer vision technology in practical scenarios such as autonomous driving, drone monitoring, and intelligent security.
[0003] The formation mechanism of shadows is relatively complex and is jointly affected by multiple factors. First, the geometric shape, surface material, and orientation of an object will affect the shape, size, and edge characteristics of the shadow, making the shadow show extremely high morphological diversity. Second, the lighting conditions in the environment, including the number, intensity, direction, and color of light sources, etc., will also have a significant impact on the characteristics of the shadow. With the change of lighting conditions, the shadow may shift, stretch, deform, or even disappear or fade. These complex changes pose great challenges to shadow removal algorithms. Traditional rule-based methods often fail to cover various complex scenarios, while data-driven methods such as deep learning can achieve certain de-shadowing effects with the support of large-scale data, but their models are often difficult to interpret, have high computational costs, and have insufficient generalization performance, making it difficult to meet the requirements of real-time and resource-constrained scenarios.
[0004] An effective shadow removal method has become a key preprocessing step to improve the performance of computer vision systems in complex environments. By removing or weakening the influence of shadows before visual tasks, the computer vision system can more accurately identify object features, thereby improving the accuracy and robustness of downstream tasks. For example, in the autonomous driving scenario, shadow removal can help the system more clearly identify the outlines of roads and obstacles; in industrial inspection, removing shadows can ensure the accuracy of defect detection; in medical image analysis, shadow removal can reduce the influence of light on the boundaries of tissues and lesion areas and improve the accuracy of diagnosis.
[0005] Shadow removal techniques can be classified into two categories: mechanism-based models and data-driven models. Mechanism-based image de-shadowing methods accurately model shadows by comprehensively analyzing lighting conditions, surface properties, and camera characteristics based on the physical understanding of the lighting and shadow formation mechanisms. Common mechanism-based image de-shadowing methods can be divided into methods based on lighting models, methods based on color and brightness restoration, and methods based on edge and texture analysis. Among them: Methods based on lighting models attempt to reconstruct the direction and intensity of the light source by analyzing the lighting to understand how it affects objects in the scene, and then infer the shadow areas and perform de-shadowing processing. Methods based on color and brightness restoration focus on restoring the true color and brightness of the shadow areas to reduce or eliminate the impact of shadows. The most classic one is the color constancy theory. Methods based on edge and texture analysis distinguish shadows by analyzing the characteristics of edges, and at the same time reveal the surface details under the shadow through texture analysis to assist in removing the shadow impact. Data-driven methods learn how to recover the features of the shadow areas from shadowed images by training deep learning models using a large amount of data. Currently, the mainstream deep learning methods include: convolutional neural networks, generative adversarial networks, autoencoders, attention mechanisms, etc. Among them: Methods based on convolutional neural networks are the most commonly used. They use multi-layer convolutional neural networks to directly learn shadow removal from paired shadowed and shadow-free images, and can effectively capture local and global features, suitable for processing complex image content. Methods based on generative adversarial networks use the adversarial process between the generator network and the discriminator network to generate shadow-free images. The generator tries to create shadow-free images, while the discriminator tries to distinguish the generated images from the real shadow-free images, and finally generates high-quality images that are visually satisfactory. Methods based on attention mechanisms are currently highly popular. They are usually used in combination with other methods. By introducing attention mechanisms, the network can focus on the shadow areas in the image, thus removing shadows more effectively and improving the sensitivity and processing accuracy of the model to the shadow areas.
[0006] Although the shadow removal methods based on mechanism-based models and data-driven models have achieved certain results, the generalization ability of mechanism-based model methods is poor and it is difficult to be widely applied to diverse actual scenarios. Data-driven model methods represented by deep learning are often "black box" models, with a lack of transparency in their internal mechanisms and poor interpretability. In addition, these models consume a large amount of computing resources during training and deployment, which limits their application in edge devices and Internet of Things environments with limited computing power. Summary of the Invention
[0007] In order to overcome the deficiencies of existing shadow removal techniques, especially the bottlenecks in mechanism-based models and data-driven models, the present invention provides a lightweight and interpretable image de-shadowing model based on image reconstruction theory.
[0008] A lightweight interpretable image de-shadowing model based on image reconstruction theory of the present invention adopts a three-stage progressive network. Each stage is an iteration in the calculation process. The first two stages include a gradient descent module and a proximal mapping module, and the last stage includes a gradient descent module and a spatial detail preservation module.
[0009] In the first stage, the model receives the original image with shadows and the corresponding shadow mask image as inputs, learns the degradation process of the shadow area through the gradient descent module, and sends the obtained reconstruction result to the proximal mapping module. The proximal mapping module receives the reconstruction result of the gradient descent module and completes the shadow removal work in the first stage under the supervision of the shadow mask image to obtain the shadow-free image in the first stage; The second stage has the same structure and process as the first stage. The difference is that the gradient descent module in the second stage no longer receives the original image with shadows as input, but is replaced by the shadow-free image in the first stage; The third stage receives the shadow-free image and the shadow mask image from the second stage, sends them to the spatial detail preservation module after passing through the gradient descent module, and the spatial detail preservation module introduces the original image with shadows and the corresponding shadow mask image to adjust the image details to obtain the final de-shadowed image.
[0010] Model principle:
[0011] Image reconstruction is a process of finding the degradation mapping relationship between a low-quality image and a high-quality image, and its degradation process is defined as:
[0012] I lq = D⊙I hq + n (1)
[0013] Where: I lq is the low-quality image I hq is the high-quality image, D is the degradation matrix from the high-quality image to the low-quality image, and n is the additive noise.
[0014] In the case of an image shadow scene, the shadow usually only exists in a local image. Therefore, the shadow image is defined by the following formula:
[0015] I s = I m ⊙I s +(1 - I m )⊙I s (2)
[0016] Where: I s is the shadow image area, and I m is the shadow mask.
[0017] Integrating formula (2) into formula (1), expressing the formation of the shadow image as a degradation process of a shadow-free image, the optimized formula is as follows:
[0018] I s = (D⊙I m + 1 - I m )⊙I ns + n (3)
[0019] where: I ns is the non - shaded image.
[0020] Regarding formula (3) as a Bayesian problem, after representing it using the framework of unified maximum a posteriori and converting it into an energy function, the problem is expressed as:
[0021]
[0022] where: J is the regularization function, and λ is the hyperparameter used to balance the regularization term.
[0023] For the solution of formula (4), using PGD, it is represented as an iterative degradation process, and finally converted into two sub - problems: gradient descent and proximal mapping:
[0024]
[0025] where: ρ k is the step - size weighting value for the k - th iteration, is the feature after degradation processing in the k - th iteration, is the non - shaded image generated in the k - th iteration, and prox λ , J is the proximal operator.
[0026] Furthermore, the gradient - descent module is designed based on formula (5), using a data - driven method to predict the shadow - degradation process. It uses two deformable residual - block structures to respectively simulate the degradation matrix D and D T , and the rest remains consistent with formula (5). During the module training, by learning a large number of paired shadow and non - shadow image data features, the parameters in the residual structure are continuously adjusted to obtain the degradation matrix of the image shadow, specifically:
[0027] First, receive the non - shaded image generated in the previous iteration Multiply separately with the shadow mask I m , the non - shadow mask I nsmMultiply to obtain the shaded region and the non - shaded region; then, use the deformable residual block to degrade the shaded region and add it to the non - shaded region to obtain a new shadow - free image; after that, subtract the newly obtained shadow - free image from the shaded image, and multiply them with the shadow mask and the non - shadow mask respectively to obtain the corresponding subtracted shaded region and shadow - free region; subsequently, use the deformable residual block to degrade the shaded region and add it to the non - shaded region; finally, multiply by the step size of this iteration and subtract the shadow - free image generated in the previous iteration to obtain the final shadow degradation map of this iteration
[0028] The deformable residual block consists of a deformable convolutional layer DC and a ReLU activation function; the deformable convolutional layer enhances the model's adaptability to irregular shadows and local structural changes by dynamically adjusting the sampling positions of the convolutional kernels.
[0029] Furthermore, the design of the proximal mapping module is based on formula (6), uses the encoder - decoder core structure, and on this basis, uses a shadow attention adjuster and a shadow interaction module to enhance the block's perception and processing ability of the shaded region. At the same time, depth - wise separable convolution is used to lightweight the module, specifically as follows:
[0030] The module receives the output from the gradient descent module Extract features through a 3 * 3 depth - wise separable convolutional layer DSC and then send them to the channel attention module CAB to adjust the weights of each channel, enhance the response to information - rich channels, and suppress channels with less information, thereby obtaining a weighted feature map; next, the feature map is weight - adjusted under the action of the shadow attention adjuster SAA, and the adjusted feature map is sent to the encoder for feature extraction at three different scales. Each scale includes a residual block RB, an internal stage feature fusion sub - module ISSF, and a down - sampling layer Down; the features extracted by the encoder at each scale are passed to the corresponding scale decoder through a 1 * 1 convolution operation to assist in image reconstruction; the features extracted by the encoder are combined with the correlation between the shaded region and the non - shaded region through the shadow interaction module SIM to adjust the attention weights again, and then sent to the decoder for the decoding and reconstruction process of the shadow - free image. The decoded features are finally output as the shadow - free image generated in this iteration through the supervised attention module SAM The decoder also contains three sub - modules composed of an up - sampling layer UP and a residual block.
[0031] Furthermore, the spatial detail retention module, as the last module in the third stage, its output will be used as the output of the overall model. This module first receives the output from the gradient descent module in the third stage Feature extraction and enhancement are performed after passing through a 3×3 depthwise separable convolutional layer (DSC) and a channel attention module (CAB). Then, it is sent to a shadow attention adjuster (SAA) for feature adjustment and then to a depthwise separable convolutional layer. Finally, it is added to the shadow map to obtain the final output of the entire model, which is the shadow-free image.
[0032] Furthermore, during the model training process, a loss function L is designed. TSPLNet , L TSPLNet is composed of the superposition of losses in three stages, and each stage includes an edge loss L edge and a perceptual loss L perceptual . The calculation formula is as follows:
[0033]
[0034] where: k is the k-th stage, and λ1, λ2 are hyperparameters, and λ1 = 1.05 and λ2 = 1 are set.
[0035] The Laplacian operator is used to perform edge extraction on the original image and the image reconstructed in the k-th stage respectively, and the difference between the two extracted images is calculated. The formula is as follows:
[0036]
[0037] where: Δ is the Laplacian operator, I sf is the original shadow-free image, and ε is a stability term.
[0038] A pre-trained VGG19 network with a structure of 5 convolutional blocks is used to perform feature extraction on the image reconstructed in the k-th stage and the original shadow-free image, and 5 feature maps are extracted from the middle layers of the 5 convolutional blocks for the calculation of the perceptual loss. The formula is as follows:
[0039]
[0040] where: C l , H l and W l are the number of channels, height, and width of layer l respectively.
[0041] The beneficial technical effects of the present invention are as follows:
[0042] 1. The present invention integrates and improves image reconstruction and shadow removal from the theoretical and model structure levels, ensuring better interpretability of the model while improving the shadow removal effect, and performing lightweight design on the model to adapt to more practical application scenarios.
[0043] 2. The present invention introduces shadow mask information into the original mechanism model of image reconstruction, constructs a new formula for shadow removal and the corresponding solution method, and proposes a brand-new gradient descent module and proximal mapping module according to the solution method and the characteristics of shadows.
[0044] 3. The present invention designs a shadow attention adjuster to enhance the shadow perception and processing ability of the proximal mapping module. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 It is the overall structure diagram of the model of the present invention.
[0046] Figure 2 It is the schematic structural diagram of the gradient descent module of the present invention.
[0047] Figure 3 It is the schematic structural diagram of the proximal mapping module of the present invention.
[0048] Figure 4 It is the schematic structural diagram of the spatial detail retention module of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0049] The present invention will be further described in detail below in conjunction with the drawings and specific implementation methods.
[0050] The present invention combines the mechanism model and the data-driven model. For the shadow removal method based on the mechanism model, it mainly overcomes the problems that it lacks sufficient adaptability to the diverse shadow features in complex scenes, resulting in poor generalization ability and difficulty in coping with various environmental changes in practical applications. For the data-driven model, the present invention focuses on solving the problems such as the lack of transparency in practical applications due to its "black box" nature, making it difficult to effectively monitor and adjust the shadow removal process, as well as large computational resource consumption and difficulty in deployment in resource-limited embedded systems or edge devices.
[0051] A lightweight interpretable image de-shadowing model based on image reconstruction theory of the present invention adopts a three-stage progressive network. Each stage is an iteration in the calculation process. The model structure is as Figure 1 shown. The first two stages include a gradient descent module and a proximal mapping module, and the last stage includes a gradient descent module and a spatial detail retention module.
[0052] In the first stage, the model receives the original image with shadows and the corresponding shadow mask image as inputs, learns the degradation process of the shadow area through the gradient descent module, and sends the obtained reconstruction result to the proximal mapping module. The proximal mapping module receives the reconstruction result of the gradient descent module and completes the shadow removal work in the first stage under the supervision of the shadow mask image, obtaining the shadow-free image of the first stage; the second stage has the same structure and process as the first stage. The difference is that the gradient descent module in the second stage no longer receives the original image with shadows as input, but is replaced by the shadow-free image of the first stage; the third stage receives the shadow-free image and the shadow mask image from the second stage, passes them through the gradient descent module and then sends them to the spatial detail preservation module. The spatial detail preservation module introduces the original image with shadows and the corresponding shadow mask image to adjust the image details, obtaining the final shadow-removed image. Figure 1 In, the dotted connection lines between stages represent the features obtained through the proximal mapping module in the previous stage. The features of the previous stage are incorporated into the features of the next stage through the Shadow Attention Adjuster (SAA) to refine the feature representation; the line connection represents the inter-stage information transfer through the inter-stage feature fusion submodule (ISSF) in the proximal mapping module.
[0053] Model principle:
[0054] Image reconstruction is the process of finding the degradation mapping relationship between low-quality images and high-quality images. Its degradation process is defined as:
[0055] I lq = D ⊙ I hq + n (1)
[0056] Where: I lq is the low-quality image I hq is the high-quality image D is the degradation matrix from the high-quality image to the low-quality image n is the additive noise. It can be seen from formula (1) that the degradation matrix D directly acts on the high-quality image I hq , and its degradation process is to degrade the entire image. This method is more suitable for scenarios such as image denoising and image de-raining.
[0057] In the case of image shadow scenes, shadows usually only exist in local images. Therefore, the shadow image is defined as the following formula:
[0058] I s = I m ⊙ I s +(1 - I m ) ⊙ I s (2)
[0059] Where: I s is the shadow image region, and I m is the shadow mask (the image shadow mask can help distinguish the shadow part and the non - shadow part in the image, and is important auxiliary information in image shadow removal).
[0060] Therefore, in the image de - shadowing task, we hope to reconstruct the non - shadow image from the shadow image. We only need to reconstruct the shadow region of the image, and do not need to reconstruct the unaffected non - shadow region. In order to more accurately represent the degradation process of image shadow removal, formula (2) is incorporated into formula (1), and the formation of the shadow image is expressed as a degradation process of a non - shadow image. The optimized formula is as follows:
[0061] I s =(D⊙I m +1 - I m )⊙I ns +n (3)
[0062] Where: I ns is the non - shadow image.
[0063] For the shadow removal task, regarding formula (3) as a Bayesian problem, after representing it using the framework of unified maximum a posteriori and converting it into an energy function, this problem is expressed as:
[0064]
[0065] Where: J is a regularization function, whose purpose is to limit the optimization process and prevent overfitting or obtaining a recovery image that does not meet expectations. λ is a hyperparameter used to balance the regularization term.
[0066] For the solution of formula (4), using Proximal Gradient Descent (PGD), it is expressed as an iterative degradation process, and finally converted into two sub - problems: gradient descent (formula 5) and proximal mapping (formula 6):
[0067]
[0068] Where: ρ k is the step - size weighting value for the k - th iteration, is the feature after degradation processing in the k - th iteration, is the non - shadow image generated in the k - th iteration, and prox λ,J is the proximal operator.
[0069] Furthermore, the gradient descent module is designed based on formula (5). Since in the task of image shadow removal, the degradation process of the image is uncertain and it is difficult to directly solve formula (5), the module uses a data-driven method to predict the shadow degradation process. Two deformable residual block structures are used to respectively simulate the degradation matrix D and D T , and the rest remains the same as formula (5). During the module training, by learning the data features of a large number of paired shadow and non-shadow image data, the parameters in the residual structure are continuously adjusted to obtain the degradation matrix of the image shadow. The structure of the gradient descent module is as Figure 2 shown, specifically:
[0070] First, it receives the shadowless image generated by the previous iteration Multiply separately with the shadow mask I , and the non-shadow mask I nsm to obtain the shadow area and the non-shadow area; then, use the deformable residual block to degrade the shadow area and add it to the non-shadow area to obtain a new shadowless image; after that, subtract the newly obtained shadowless image from the shadow image, and multiply it separately with the shadow mask and the non-shadow mask to obtain the corresponding subtracted shadow area and shadowless area; subsequently, use the deformable residual block to degrade the shadow area and add it to the non-shadow area; finally, multiply it by the step size of this iteration and subtract the shadowless image generated by the previous iteration to obtain the shadow degradation map of this iteration
[0071] The deformable residual block consists of a deformable convolutional layer DC and a ReLU activation function; the deformable convolutional layer enhances the model's adaptability to irregular shadows and local structure changes by dynamically adjusting the sampling positions of the convolutional kernels. This is because the task of image shadow removal is different from image reconstruction tasks such as denoising and de-raining. The shadow area is usually affected by various factors such as the lighting angle and the object placement position, presenting an irregular shape. Deformable convolution can effectively capture these complex geometric transformations, thereby improving the effect of shadow removal. In contrast, traditional convolutional operations use a fixed receptive field and are difficult to adapt to such irregular shadow areas. In addition, traditional convolution has the same weight at each position, which makes it less sensitive to local structure changes in the image and difficult to handle complex local details.
[0072] Furthermore, the design of the proximal mapping module is based on formula (6) as the theoretical basis. It uses an encoder-decoder core structure, and on this basis, uses a shadow attention adjuster and a shadow interaction module to enhance the block's perception and processing ability of the shadow area. At the same time, depthwise separable convolution is used to lightweight the module. The structure of the proximal mapping module is as Figure 3 shown, specifically:
[0073] The module receives the output from the gradient descent module After extracting features through a 3×3 depthwise separable convolution layer (Depthwise Separable Convolution, DSC), the features are sent to the channel attention block (Channel Attention Block, CAB) to adjust the weights of each channel, enhance the response to information-rich channels, and suppress channels with less information, thereby obtaining a weighted feature map. Next, the feature map is weight-adjusted under the action of the shadow attention adjuster SAA, and the adjusted feature map is sent to the encoder for feature extraction at three different scales. Each scale includes a residual block (Residual Block, RB), an internal stage feature fusion sub-module ISSF, and a downsampling layer (Downsampling, Down). The features extracted by the encoder at each scale are passed to the corresponding scale decoder through a 1×1 convolution (Conv) operation to assist in image reconstruction. The features extracted by the encoder are adjusted for attention weights again through the shadow interaction module (Shadow Interaction Module, SIM) by combining the correlation between the shadow area and the non-shadow area, and then sent to the decoder for the decoding and reconstruction process of the shadow-free image. The decoded features are finally output as the shadow-free image generated in this iteration through the supervised attention module (Supervised Attention Module, SAM) The decoder also contains three sub-modules composed of an upsampling layer (Upsampling, UP) and residual blocks
[0074] Furthermore, the spatial detail retention module, as the last module in the third stage, its output will be used as the output of the overall model. We hope to generate richer high-resolution features and an image closer to the original image. Therefore, the proximal mapping module is simplified by canceling the encoder-decoder structure with downsampling and integrating the original image into the final input result. As Figure 4 shown, this module first receives the output from the gradient descent module in the third stage After feature extraction and enhancement through a 3×3 depthwise separable convolution layer DSC and the channel attention module CAB; then, it is sent to the shadow attention adjuster SAA for feature adjustment and then sent to the depthwise separable convolution layer; finally, it is added to the shadow map to obtain the shadow-free image finally output by the entire model
[0075] Furthermore, a loss function L is designed during the model training process TSPLNet , L TSPLNet is composed of the losses of three stages superimposed, and each stage includes an edge loss L edge and a perceptual loss Lperceptual , the calculation formula is as follows:
[0076]
[0077] Where: k is the k-th stage, λ1 and λ2 are hyperparameters, and λ1 = 1.05 and λ2 = 1 are set.
[0078] The main goal of the edge loss is to maintain the clarity and structural integrity of the image edges when the model performs shadow removal, and to keep the shadow edge details highly consistent with the original image. Therefore, the Laplace operator is used to extract the edges of the original image and the image reconstructed at the k-th stage respectively, and the difference between the two extracted images is calculated. The formula is as follows:
[0079]
[0080] Where: Δ is the Laplace operator, I sf is the original shadow-free image, and ε is the stability term.
[0081] The main goal of the perceptual loss is to measure the effect of shadow removal under higher-level features to ensure that the generated shadow-free image is closer to the target image. Therefore, a pre-trained VGG19 network with a structure of 5 convolutional blocks is used to extract features from the image reconstructed at the k-th stage and the original shadow-free image, and 5 feature maps are extracted from the middle layers of the 5 convolutional blocks for the calculation of the perceptual loss. The formula is as follows:
[0082]
[0083] Where: C l , H l and W l are the number of channels, height, and width of layer l respectively.
[0084] The present invention fuses and improves image reconstruction and shadow removal from the theoretical and model structure levels, ensures better interpretability of the model while improving the shadow removal effect, and performs lightweight design on the model to adapt to more practical application scenarios.
Claims
1. A lightweight interpretable image de-shadowing model based on image reconstruction theory, characterized in that A three-stage progressive network is adopted. Each stage is an iteration in the calculation process. The first two stages contain a gradient descent module and a proximal mapping module, and the last stage contains a gradient descent module and a spatial detail preservation module; In the first stage, the model receives the original image with shadows and the corresponding shadow mask image as inputs. The gradient descent module is used to learn the degradation process of the shadow area, and the obtained reconstruction result is sent to the proximal mapping module. The proximal mapping module receives the reconstruction result of the gradient descent module and completes the shadow removal work in the first stage under the supervision of the shadow mask image, obtaining the shadow-free image in the first stage. The second stage has the same structure and process as the first stage. The difference is that the gradient descent module in the second stage no longer receives the original image with shadows as input, but is replaced by the shadow-free image in the first stage. The third stage receives the shadow-free image and the shadow mask image from the second stage, and after passing through the gradient descent module, it is sent to the spatial detail preservation module. The spatial detail preservation module introduces the original image with shadows and the corresponding shadow mask image to adjust the image details, obtaining the final shadow-removed image; Model principle: Image reconstruction is a process of finding the degradation mapping relationship between low-quality images and high-quality images, and its degradation process is defined as: I lq = D ⊙ I hq + n(1) Where: I lq is the low-quality image I hq is the high-quality image, D is the degradation matrix from the high-quality image to the low-quality image, and n is the additive noise; In the case of image shadow scenarios, shadows usually only exist in local images. Therefore, the shadow image is defined by the following formula: I s = I m ⊙ I s + (1 - I m ) ⊙ I s (2) Where: I s is the shadow image area, and I m is the shadow mask; By integrating formula (2) into formula (1), the formation of the shadow image is expressed as a degradation process of a shadow-free image. The optimized formula is as follows: I s =(D⊙I m +1 - I m )⊙I ns +n (3) Wherein: I ns is a non-shaded figure; Regarding formula (3) as a Bayesian problem, after representing it using the framework of unified maximum a posteriori and converting it into an energy function, this problem is expressed as: where: J is a regularization function, and λ is a hyperparameter used to balance the regularization term; For the solution of formula (4), PGD is used to represent it as an iterative degradation process, and finally it is converted into two sub-problems of gradient descent and proximal mapping: Where: ρ k is the step size weighting value of the k-th iteration, is the feature after degradation processing in the k-th iteration, is the shadow-free image generated in the k-th iteration, prox λ,J is the proximal operator.
2. The lightweight interpretable image de-shadowing model based on image reconstruction theory according to claim 1, characterized in that, The gradient descent module is designed based on formula (5), using a data-driven method to predict the process of shadow degradation. Two deformable residual block structures are used to simulate the degradation matrix D and D T respectively, and the rest remains the same as formula (5). During the module training, by learning the feature data of a large number of paired shadow and non-shadow images, the parameters in the residual structure are continuously adjusted to obtain the degradation matrix of the image shadow, specifically: First, receive the shadowless image generated in the previous iteration Multiply separately with the shadow mask I m , the non-shadow mask I nsm to obtain the shadow area and the non-shadow area; then, use the deformable residual block to degrade the shadow area and add it to the non-shadow area to obtain a new shadowless image; then, subtract the newly obtained shadowless image from the shadow image, and multiply it separately with the shadow mask and the non-shadow mask to obtain the corresponding subtracted shadow area and shadowless area; subsequently, use the deformable residual block to degrade the shadow area and add it to the non-shadow area; finally, multiply it by the step length of this iteration and subtract the shadowless image generated in the previous iteration to obtain the final shadow degradation map of this iteration The deformable residual block consists of a deformable convolutional layer DC and a ReLU activation function; the deformable convolutional layer enhances the model's adaptability to irregular shadows and local structural changes by dynamically adjusting the sampling positions of the convolutional kernels.
3. The lightweight interpretable image de-shadowing model based on image reconstruction theory according to claim 2, characterized in that The design of the proximal mapping module is based on formula (6). It uses an encoder-decoder core structure. On this basis, a shadow attention adjuster and a shadow interaction module are used to enhance the block's perception and processing ability of the shadow area. At the same time, depthwise separable convolutions are used to lighten the weight of the module. Specifically: The module receives the output from the gradient descent module After extracting features through a 3×3 depthwise separable convolutional layer (DSC), the features are fed into the channel attention module (CAB) to adjust the weights of each channel, enhance the response to information-rich channels, and suppress channels with less information, thereby obtaining a weighted feature map. Next, the feature map is weighted by the shadow attention adjuster (SAA). The adjusted feature map is fed into the encoder to extract features at three different scales, each scale including a residual block (RB), an internal stage feature fusion sub-module (ISSF), and a downsampling layer (Down). The features extracted by the encoder at each scale are passed to the corresponding scale of the decoder through a 1×1 convolution operation to assist in image reconstruction. The features extracted by the encoder are adjusted for attention weights again through the shadow interaction module (SIM) by combining the correlation between the shadow area and the non-shadow area, and then fed into the decoder for the decoding and reconstruction process of the shadowless image. The decoded features are finally output as the shadowless image generated in this iteration through the supervised attention module (SAM). The decoder also contains three sub-modules consisting of an upsampling layer (UP) and a residual block.
4. A lightweight interpretable image de-shadowing model based on image reconstruction theory according to claim 1, characterized in that, The spatial detail preservation module, as the last module in the third stage, whose output will be the output of the overall model, first receives the output from the gradient descent module in the third stage. After passing through a 3×3 depthwise separable convolutional layer (DSC) and a channel attention module (CAB) for feature extraction and enhancement, it is then sent to a shadow attention adjuster (SAA) for feature adjustment and then to a depthwise separable convolutional layer. Finally, after adding it to the shadow map, the shadowless map that is the final output of the entire model is obtained.
5. A lightweight interpretable image de-shadowing model based on image reconstruction theory according to claim 1, characterized in that, Design the loss function L during model training TSPLNet , L TSPLNet is composed of the losses of three stages, and each stage contains an edge loss L edge and a perceptual loss L perceptual . The calculation formula is as follows: where: k is the k-th stage, λ1 and λ2 are hyperparameters, and λ1 = 1.05 and λ2 = 1 are set; The Laplacian operator is used to extract the edges of the original image and the image reconstructed in the k-th stage respectively, and the difference between the two extracted images is calculated. The formula is as follows: where: Δ is the Laplace operator, I sf is the original shadowless image, and ε is the stability term; A pre-trained VGG19 network with a structure of 5 convolutional blocks is used to extract features from the image reconstructed in the k-th stage and the original shadow-free image, and 5 feature maps are extracted from the middle layers of the 5 convolutional blocks for the calculation of perceptual loss. The formula is as follows: where: C l , H l and W l are the number of channels, height, and width of layer l, respectively.