Image shadow removal method based on double-branch network

Through the image shadow removal method based on the dual-branch network, shadow areas and non-shaded areas are processed respectively, shadow removal is performed using the light intensity and attenuation factor model, and the texture and color information are coordinated in the color repair module, the limitations of shadow removal in the existing technology in complex backgrounds and lighting situations are solved, and high-quality shadow removal effect is achieved.

CN120107119AActive Publication Date: 2025-06-06CHONGQING UNIV OF POSTS & TELECOMM

Patent Information

Application Number
CN202510267596.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-06
Estimated Expiration
2045-03-07

AI Technical Summary

Technical Problem

The prior art has limitations when dealing with shadow removal in complex backgrounds or complex lighting situations, making it difficult to understand the scene context semantics globally, resulting in poor shadow removal effect.

Method used

The image shadow removal method based on a dual-branch network is adopted, and shadow removal is performed using the relative light intensity and attenuation factor model to coordinate texture and color information in the color repair module.

Benefits of technology

It realizes high-quality shadow removal in a variety of natural light scenes, improves the image shadow removal effect, and can flexibly calculate pixel values ​​under different lighting conditions, significantly improving the accuracy and efficiency of shadow removal.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107119A_ABST
    Figure CN120107119A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of image processing, and particularly relates to an image shadow removal method based on a double-branch network. The method comprises the following steps: acquiring a shadow image to be processed and a corresponding mask image, and preprocessing the shadow image and the mask image to obtain a shadow region image and a non-shadow region image; respectively sending the shadow region image and the non-shadow region image into a shadow removal branch network and an identical mapping branch network for processing to obtain a shadow removal region image and a mapping image; performing pixel-by-pixel addition on the shadow-removed area image and the mapping image to obtain a preliminary shadow-removed image; the image with the shadow removed preliminarily is sent to a color restoration module for color restoration, and a clear image with the shadow removed is obtained; according to the method, shadow removal can be effectively realized, the color shift problem in a shadow removal result can be alleviated, and the method has a good application prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of image processing, and in particular relates to an image shadow removal method based on a double-branch network. Background Art

[0002] In the field of computer vision, shadows in images will reduce the quality of images to varying degrees, bringing considerable challenges to many downstream computer tasks. In the field of object detection, the presence of shadows will cause the boundaries of objects to be blurred, making it difficult to accurately identify the outline of the object. In the object detection task, accurate object boundaries are important information that can be used for target positioning and segmentation. Shadows will blur the boundaries of objects, increasing the difficulty of object detection. In the field of target tracking, the presence of shadows will blur the texture and details of the target surface. This may cause the target tracking algorithm based on texture features to be unable to accurately extract the feature information of the target, thereby reducing the performance of the tracking algorithm. In the field of face detection, shadows may cause part of the face area to be blocked or covered, making it impossible for the face detection algorithm to fully detect the face. The blurred boundary between the shadow area and the face area may cause the algorithm to produce incorrect detection results or miss the face. Therefore, image shadow removal is generally an upstream task for various computer vision tasks, which has a great impact on many other downstream tasks. Therefore, image shadow removal is a very challenging task with great reference value.

[0003] Shadow removal based on traditional methods: Traditional shadow removal methods generally focus on the underlying features of shadows, aiming to restore the material, brightness, gradient and other information of the shadow area. Some of them solve shadow-free images based on Gaussian mixture models and physical lighting models, some solve shadow-free images by calculating Poisson's equations through gradient fields, and some eliminate shadows by matching shadow areas with non-shadow areas. In general, these traditional shadow removal methods rely heavily on underlying features, which have inherent limitations in many cases. For example, these methods based on underlying features often have difficulty in understanding the semantics of the scene context from a global perspective, and therefore cannot solve shadow removal in complex backgrounds or complex lighting situations.

[0004] In recent years, with the release of shadow removal datasets and the development of deep learning research and computer hardware, shadow removal methods based on convolutional neural networks have made significant progress. Convolutional Neural Network is a deep neural network with a convolutional structure. Its main idea is based on the operation of the human visual system, that is, by using local receptive fields to extract local features of images, and combining these local features into global features at a higher level to achieve tasks such as image classification, recognition and segmentation. CNN can automatically learn features in images, and perform feature extraction and dimensionality reduction through operations such as convolution, pooling and activation. At the same time, it has the advantages of translation invariance and parameter sharing, which makes it widely used in the field of computer vision.

[0005] Using deep learning for shadow removal can simplify the process of physical modeling and manual feature extraction, and can achieve good removal effects. The application of convolutional neural networks in shadow removal can utilize its deep feature extraction capabilities to improve the accuracy and efficiency of shadow removal. Generally speaking, the deeper the number of neural network layers, the more complex the extracted features. However, many deep learning shadow removal methods have the problem of insufficient interpretability, making it difficult to effectively interpret and understand the results of shadow removal. Although some methods attempt to use illumination models for shadow removal and model the conversion between shadow pixels and shadow-free pixels as a linear relationship, their design ideas are relatively simple and fail to fully consider the nonlinear response characteristics of photoelectric sensors. In addition, current methods consider the same convolution operation for shadow and non-shadow areas, while ignoring the huge gap between the color mapping of shadow and non-shadow areas, resulting in poor quality of reconstructed images and heavy computational burden. Summary of the invention

[0006] In view of the shortcomings of the prior art, the present invention proposes an image shadow removal method based on a dual-branch network, the method comprising: obtaining a shadow image to be processed and its corresponding mask image and inputting them into a trained dual-branch network for processing to obtain a clear image after the shadow is removed;

[0007] The training process of the two-branch network includes:

[0008] S1: Obtain a shadow image to be processed and its corresponding mask image, preprocess the shadow image and the mask image to obtain a shadow area image and a non-shadow area image;

[0009] S2: Send the shadow area image and the non-shadow area image to the shadow removal branch network and the identity mapping branch network for processing respectively, to obtain the shadow removal area image and the mapping image;

[0010] S3: adding the de-shadowed area image and the mapped image pixel by pixel to obtain a preliminary de-shadowed image;

[0011] S4: sending the preliminary shadow-removed image to a color restoration module for color restoration to obtain a clear image after the shadows are removed;

[0012] S5: Calculate the total loss of the two-branch network and adjust the network parameters according to the total loss to obtain a trained two-branch network.

[0013] Preferably, the process of preprocessing the shadow image and the mask image includes:

[0014] The shadow image and the mask image are both scaled to 256×256; the shadow image and the mask image are multiplied pixel by pixel to obtain the shadow area image; the mask image is inverted and multiplied pixel by pixel with the shadow image to obtain the non-shadow area image.

[0015] Preferably, the process of sending the shadow area image and the non-shadow area image to the shadow removal branch network for processing respectively includes:

[0016] Input the shadow area image into the relative light intensity estimation network for processing to obtain the relative light intensity;

[0017] The non-shadow area image is input into the attenuation factor estimation network for processing to obtain the attenuation factor;

[0018] A response attenuation model is established according to the relationship between pixel value and illumination, and the response attenuation model is fitted according to the relative light intensity and attenuation factor to obtain a shadow-removed area image.

[0019] Furthermore, both the relative light intensity estimation network and the attenuation factor estimation network adopt a symmetrical encoding-decoding structure;

[0020] The encoder consists of 8 convolutional layers, with a convolution kernel of 4, a stride of 2, and a padding of 1; the number of input channels of the first convolutional layer is 4, and the number of output channels is 64; the number of input channels of the 2nd to 7th convolutional layers are 64, 128, 256, 512, 512, 512, and the number of output channels are 128, 256, 512, 512, 512, 512, respectively; LeakyRelu layers and BatchNorm layers are added before and after the 2nd to 7th convolutional layers respectively; the number of input channels of the 8th convolutional layer is 512, the number of output channels is 512, and a LeakyRelu layer is added before the 8th convolutional layer;

[0021] The decoder includes 8 deconvolution layers. The number of input channels of the first deconvolution layer is twice the number of output channels of the last convolution layer of the encoder. The parameters of each deconvolution layer are set to a convolution kernel size of 4, a stride of 2, and a padding of 1. The number of input channels of the 1st to 7th convolution layers are 1024, 2048, 2048, 2048, 1024, and 512, respectively, and the number of output channels are 1024, 1024, 1024, 1024, 512, respectively, and the number of output channels are 1024, 1024, 1024, 1024, 512, 256, and 128, respectively. Relu layers and BatchNorm layers are added before and after the 1st to 7th convolution layers, respectively. The number of input channels of the 8th convolution layer is 256, and the number of output channels is 3. At the same time, Relu layers and sigmoid layers are added before and after the 8th convolution layer, respectively.

[0022] Furthermore, the response attenuation model is expressed as:

[0023]

[0024] Among them, p represents the pixel value of the image after fitting the model, P s represents the shadow pixel value, k represents the incident light intensity ratio between the shadow-free area and the shadow area, a represents the attenuation factor corresponding to the object, and R represents the reflectivity.

[0025] Preferably, the identity mapping branch network consists of an encoder and a decoder;

[0026] The encoder is based on the SR-Net structure and consists of 6 convolutional layers. The 1st to 4th convolutional layers are cascaded with the instance normalization module and the LeakyReLU activation function, respectively, and the 5th convolutional layer is cascaded with the BatchNorm layer and the LeakyReLU activation function. The parameters of all convolutional layers are set to convolution kernel size 4×4, stride 2, padding 1, the number of input channels is 4, 64, 128, 256, 512 and 512, and the number of output channels is 64, 128, 256, 512, 512 and 512, respectively.

[0027] The decoder consists of 6 deconvolution layers. The first deconvolution layer is preceded and followed by a LeakyReLU activation layer and a BatchNorm layer, respectively. The second to fifth deconvolution layers are preceded and followed by a LeakyReLU activation layer and an instance normalization module, respectively. The sixth deconvolution layer is preceded and followed by a LeakyReLU activation layer and a tanh activation layer, respectively. The parameters of all deconvolution layers are set to a convolution kernel size of 4×4, a stride of 2, and a padding of 1. The number of input channels of the 6 deconvolution layers are 512, 1024, 1024, 512, 256, and 128, respectively, and the number of output channels are 512, 512, 256, 128, 64, and 3, respectively.

[0028] Preferably, the process of sending the preliminary shadow-removed image to the color restoration module for color restoration is expressed as:

[0029]

[0030] in, Represents the normalized pixel value, p h,w,c Indicates the initial pixel value of the image for preliminary shadow removal; represents the mean value of the pixels in the shadow area, represents the variance of the pixels in the shadow area, represents the variance of pixels in the non-shadow area, Represents the mean value of pixels in the non-shadow area.

[0031] Preferably, the formula for calculating the total loss of the dual-branch network is:

[0032] L=λ 1 L STF +λ 2 L FTF +λ 3 L coarse +λ 4 L fine

[0033] L STF =||I RES_S -I SF_S ||

[0034] L FTF =||I RES_SF -I SF_SF ||

[0035] L coarse =||I course -I SF ||

[0036] L fine =||I fine -I SF ||

[0037] Among them, L represents the total loss, L STF is the shadow removal branch loss, L FTF is the identity mapping branch loss, is the roughness loss, is the refinement loss, λ 1 ~λ 4 are the first to fourth hyperparameters, ||·|| 1 Indicates L 1 Distance, I RES_S Represents the image with shadow removed, I RES_SF To map the image, I courseTo preliminarily remove the shadow image, I fine To remove the shadows, I SF is a true shadow-free image, I SF_S and I SF_SF are the shadow area and non-shadow area of ​​the true shadow-free image, respectively.

[0038] The beneficial effects of the present invention are:

[0039] This paper proposes a dual-branch network image shadow removal method. This method uses different networks to process shadow areas and non-shadow areas respectively to solve the problem of task conflict and computational resource waste in previous shadow removal methods. Figure 2 The SRNet_STF branch of the SRNet proposes an image shadow removal method based on the excitation attenuation model to address the nonlinear response phenomenon of photoelectric sensors during color acquisition. The excitation attenuation model divides the image into two components, namely the relative light intensity component and the attenuation factor component. The relative light intensity is related to the reflected light intensity of the scene, and the attenuation factor is related to the material of the object. Since it is impossible to directly separate the relative light intensity and the attenuation factor from the image, a dual encoder structure is used to extract the features of different components in the image, and a dual decoder is used to predict the two components of the image. The components output by the decoder are constrained by the corresponding physical equations. Adaptively learn the complex illumination and color mapping relationship to fit the nonlinear transformation from shadow to non-shadow, thereby achieving accurate processing of the shadow area. In the non-shadow area such as Figure 2 The SRNet_FTF branch of the proposed method uses an identity mapping module and a lightweight convolution module to process in a low-cost computing manner, effectively saving computing resources. Considering that the non-shadow area contains richer background color information, in order to further strengthen the information flow from the non-shadow area to the shadow area and make the two more coordinated and unified, a new normalization strategy is proposed for color restoration. The implementation of this strategy will further optimize the overall performance of the dual-branch network in the shadow removal task and improve the image shadow removal effect. . Qualitative and quantitative experimental results show that the network of the present invention has good effectiveness in realizing shadow removal based on the illumination model and the response attenuation model, and shows excellent performance under a variety of natural light scenes and evaluation indicators, and can flexibly calculate pixel values ​​under different lighting conditions.

[0040] In view of the fact that the shadow and non-shadow areas in the shadow image should have similar texture and color distribution, and the non-shadow part contains important prior information, the present invention proposes a color restoration module. This module aims to transfer the color features of the non-shadow area to the shadow area to alleviate the color shift problem in the shadow removal result. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1It is a flow chart of the image shadow removal method based on the dual-branch network in the present invention;

[0042] Figure 2 It is a framework diagram of the double-branch network structure in the present invention;

[0043] Figure 3 This is a visual comparison effect diagram of the present invention on the ISTD+ test set;

[0044] Figure 4 This is a quantitative comparison result diagram of the present invention on the ISTD+ test set. DETAILED DESCRIPTION

[0045] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0046] The present invention proposes an image shadow removal method based on a dual-branch network, the method comprising the following steps: obtaining a shadow image to be processed and its corresponding mask image and inputting them into a trained dual-branch network for processing to obtain a clear image after the shadow is removed;

[0047] The dual-branch network of the present invention simplifies the Lambert illumination model in the shadow removal branch, and uses relative light intensity to describe the illumination relationship between shadow and non-shadow areas. According to the characteristics of the nonlinear response of the photoelectric sensor, an attenuation function in the form of an inverse is selected. It uses only one parameter (attenuation factor) to describe the excitation attenuation trend. The attenuation factor is only related to the material of the object and has nothing to do with the scene illumination. The pixel value is modeled as the integral of the attenuation factor on the reflected light intensity, and the response attenuation model can be derived. It divides the image into a material component (attenuation factor) and an illumination component (relative light intensity). A dual encoder-dual decoder structure is used to predict these two components. The two components are combined into a rough shadow removal image through corresponding physical constraints, and then the rough shadow removal image and the mapped image are channel-joined and sent to the color restoration module, and the mean and variance of the pixel values ​​in the non-shadow area are calculated to help color restoration in the shadow area.

[0048] like Figure 1 As shown, the training process of the dual-branch network includes:

[0049] S1: Obtain a shadow image to be processed and its corresponding mask image, preprocess the shadow image and the mask image to obtain a shadow area image and a non-shadow area image.

[0050] The shadow image and the mask image are both scaled to 256×256. Specifically, the shadow image and the mask image are scaled to 286×286 and then randomly cropped to obtain a cropped image of size 256×256. The cropped image is rotated and flipped with a probability of 50% to obtain a preliminarily processed shadow image and mask image. After that, the shadow image and the mask image are multiplied pixel by pixel to obtain a shadow area image. The mask image is inverted and multiplied pixel by pixel with the shadow image to obtain a non-shadow area image.

[0051] S2: Send the shadow area image and the non-shadow area image to the shadow removal branch network and the identity mapping branch network for processing respectively to obtain the removed shadow area image and the mapping image.

[0052] The shadow area image and the non-shadow area image are sent to the shadow removal branch network for processing to obtain the shadow area removed image. Figure 2 As shown in Figure 1, the shadow removal branch network, namely the SRNet_STF branch, includes a relative light intensity estimation network and an attenuation factor estimation network. The shadow area image is input into the relative light intensity estimation network to obtain the relative light intensity component, which is related to the reflected light intensity of the scene; the shadow area image is sent into the attenuation factor estimation network to obtain the attenuation factor component, which is related to the material of the object. The two components of the image are constrained by the corresponding physical equations and then reconstructed, as shown in Formula 17; the identity mapping branch network, namely the SRNet_FTF branch, reconstructs the image through a lighter convolutional SRNet.

[0053] In the present invention, both the relative light intensity estimation network and the attenuation factor estimation network adopt a symmetrical encoding-decoding structure, and their task objectives and gradient update paths are different, and the relative light intensity and attenuation factor are generated respectively; wherein:

[0054] The encoders are all 8-layer convolutional neural networks with 4 convolution kernels, a stride of 2, and a padding of 1, so that the width and height of the image are halved after each round of convolution. The first convolutional layer has 4 input channels and 64 output channels; the second to seventh convolutional layers have 64, 128, 256, 512, 512, 512 input channels and 128, 256, 512, 512, 512 output channels, respectively. In addition, LeakyRelu layers and BatchNorm layers are added before and after the second to seventh convolutional layers, respectively; the eighth convolutional layer has 512 input channels and 512 output channels, and a LeakyRelu layer is added before the eighth convolutional layer.

[0055] The decoder consists of 8 deconvolution layers. The number of input channels of the first deconvolution layer is twice the number of output channels of the last convolution layer of the encoder, so that the decoder can decode the features encoded by two encoders. The parameters of each deconvolution layer are set to a convolution kernel size of 4, a stride of 2, and a padding of 1, so that the width and height of the image are doubled after each round of convolution. The number of input channels of the 1st to 7th convolution layers are: 1024, 2048, 2048, 2048, 2048, 1024, 512, and the number of output channels are: 1024, 1024,

[0056] 1024, 1024, 512, 256, 128. In addition, Relu layers and BatchNorm layers are added before and after the 1st to 7th convolutional layers respectively. The number of input channels of the 8th convolutional layer is 256, and the number of output channels is 3. At the same time, Relu layers and sigmoid layers are added before and after the 8th convolutional layer respectively.

[0057] The encoder input size is 256×256. During the downsampling process, the intermediate feature size is halved after each convolution. After eight convolutions, the size changes from 256×256 to 1×1. During the upsampling process, the intermediate feature size doubles after each deconvolution. Through skip connections, the feature maps in the downsampling and upsampling processes are spliced, so that the underlying features and high-level features are integrated with each other, which helps to better consider the global and local features of the image.

[0058] Afterwards, a response attenuation model is established based on the relationship between pixel value and illumination, and the response attenuation model is fitted according to the relative light intensity and attenuation factor to obtain a shadow-free area image. The process of establishing the response attenuation model includes:

[0059] In the real world, most electronic cameras exhibit nonlinear response characteristics in the photoelectric conversion process. Specifically, as the light intensity increases, the change in pixel value is not proportional to the change in light intensity, but rather exhibits a nonlinear relationship.

[0060] Generally speaking, the reflectivity of an object is less affected by light intensity and can be approximately considered linear. For example, formula (1):

[0061] I R =I L ×R (1)

[0062] In the formula, I R is the reflected light intensity; I L is the incident light intensity; R is the reflectivity, which has nothing to do with the light intensity.

[0063] The intensity of the reflected light is proportional to the incident light, but the pixel value is not proportional to the reflected light, so the change in the pixel value is smaller than the change in the incident light.

[0064] The relationship between pixel value and reflected light intensity is modeled as an integral relationship. As shown in formula (2):

[0065] P = ∫ 0 x f(x)dx (2)

[0066] Where P represents the pixel value, x represents the current reflected light intensity, dx represents the small change in the reflected light intensity, and f(x) is the response function, which represents the pixel value increment caused by the increase in unit reflected light intensity when the reflected light intensity is x.

[0067] Due to the nonlinear response characteristics of the photoelectric sensor, the response value will gradually decrease as the reflected light intensity increases and eventually tend to zero. The present invention defines this phenomenon as response attenuation. Assuming that the light intensity is a continuous variable, the relationship between the reflected light intensity and the response value can be represented by a response curve.

[0068] Assume that in complete darkness, the response value is 1. This means that when there is no attenuation, 1 unit of light intensity causes the pixel value to increase by 1, as shown in formula (3):

[0069] f(0)=1 (3)

[0070] Since objects of different materials have different reflection characteristics to light, their response curves are usually different. When the same incident light shines on different objects, the spectral characteristics of the reflected light will be different. At the same time, the increase in light intensity will not lead to a decrease in pixel value, so the lower limit of the response curve is 0.

[0071] In order to meet the initial value defined by equation (3) and the lower bound of the response curve, the present invention uses a reciprocal function to model the attenuation of the response value, as shown in (4):

[0072]

[0073] Where f(x) is the response value corresponding to the reflected light intensity of x; a is the attenuation factor, which is related to the material of the object.

[0074] The relationship between the attenuation factor and the attenuation curve is that the response value decreases as the light intensity increases, and has a lower limit of 0. When the light intensity is the same, the larger the attenuation factor, the smaller the response value.

[0075] From equations (2) and (4), we can see that the relationship between the reflected light intensity and the pixel value is as shown in equation (5):

[0076]

[0077] Where: p is the pixel value (brightness); x is the reflected light intensity; a is the attenuation factor.

[0078] The present invention proposes a simplified illumination model, which regards the incident light intensity in the shadow-free area as the superposition of the ambient light intensity and the direct light intensity, so that the reflected light intensity can be expressed as a linear combination of the ambient light intensity and the direct light intensity, as shown in formula (6):

[0079] I r =R a *I a +R d *I d (6)

[0080] In the formula, I r Indicates the reflected light intensity; R a , R d Represent the reflectivity of ambient light and direct light respectively; I a ,I d Represents the light intensity of ambient light and direct light respectively.

[0081] Assume that the reflectivity of the object remains unchanged, and the reflectivity of direct light is equal to the reflectivity of diffuse reflection. It can be expressed as formula (7):

[0082] I r =R(I a +I d ) (7)

[0083] Assuming that the shadow area is illuminated only by ambient light, and the unshadowed area is illuminated by both ambient light and direct light, the pixel value of the shadow area can be obtained according to formula (5) as shown in formula (8):

[0084]

[0085] Where P s Represents the pixel value of the shadow area; I a Represents the intensity of ambient light; a is the attenuation factor, which is related to the material of the object; R is the reflectivity of the object.

[0086] The pixel value of the non-shadow area is shown in formula (9):

[0087]

[0088] Where P sf Indicates the pixel value of the non-shadow area, I a Indicates the ambient light intensity, I d It represents the intensity of direct light, a is the attenuation factor, and R is the reflectivity of the object.

[0089] The incident light intensity in the unshadowed area can be expressed as the ambient light intensity plus the direct light intensity, as shown in formula (10):

[0090] I sf =Ia +I d =kI a ,k>1 (10)

[0091] Where k represents the ratio of the incident light intensity between the shadow-free area and the shadow area.

[0092] According to formula (10), the conversion formula from shadow to non-shadow pixel value is derived as shown in formulas (11)-(12):

[0093]

[0094] By converting both sides of the equation (11) into exponential form, we can get equation (12):

[0095]

[0096] From equations (8) and (9), we can get:

[0097]

[0098] Substituting formula (13) into formula (12), we can obtain:

[0099]

[0100] Therefore, the transformation from shadow to non-shadow pixel value is as shown in equation (16):

[0101]

[0102] Where P sf is the value of the pixel without shadow; P s is the shadow pixel value; k is the incident light intensity ratio (relative light intensity) between the shadow-free area and the shadow area; a is the attenuation factor corresponding to the object.

[0103] From formula (16), we can see that the conversion of shadow to shadow-free pixel value requires two parameters, namely the attenuation factor a of the object and the relative light intensity k. The attenuation factor a of the object has nothing to do with the illumination and depends on the differential absorption of light of different wavelengths by the object. The relative light intensity k is related to the illumination conditions of the scene. The larger k is, the greater the proportion of direct light intensity to the total illumination intensity.

[0104] According to equations (8)-(9), the pixel values ​​of any area (shadow area and non-shadow area) can be summarized, that is, the response attenuation model is shown in equation (17):

[0105]

[0106] In the formula, p represents the pixel value of the image after fitting the model, a is the attenuation factor, which is related to the material; k is the relative light intensity, which is related to the lighting conditions. In the shadow-free area, k is the ratio of the total light intensity to the ambient light intensity, in the shadow area, k = 1, and in the penumbra area, k changes evenly; R is the reflectivity, which is related to the material.

[0107] The relative light intensity estimation network and the attenuation factor estimation network use the same symmetric encoding-decoding structure, and both the encoder and decoder are 8-layer convolutional neural networks with 4 convolution kernels and a step size of 2. The number of input channels is 4 (3 channels of the shadow image + 1 channel of the shadow mask), and the number of output channels is 3, which means that the excitation attenuation model considers the R, G, and B channels respectively.

[0108] The shadow area image and the non-shadow area image are sent to the identity mapping branch network for processing to obtain the mapping image. The identity mapping branch network consists of an encoder and a decoder, where:

[0109] The encoder is based on the SR-Net structure and consists of 6 convolutional layers. The 1st to 4th convolutional layers are cascaded with the instance normalization module and the LeakyReLU activation function, respectively, and the 5th convolutional layer is cascaded with the BatchNorm layer and the LeakyReLU activation function. The parameters of all convolutional layers are set to the convolution kernel size of 4×4, the stride of 2, and the padding of 1. The number of input channels is 4, 64, 128, 256, 512 and 512 respectively, and the number of output channels is 64, 128, 256, 512, 512 and 512 respectively.

[0110] The decoder consists of 6 deconvolution layers. The first deconvolution layer is preceded and followed by a LeakyReLU activation layer and a BatchNorm layer, the second to fifth deconvolution layers are preceded and followed by a LeakyReLU activation layer and an instance normalization module, and the sixth deconvolution layer is preceded and followed by a LeakyReLU activation layer and a tanh activation layer. The parameters of all deconvolution layers are set to a convolution kernel size of 4×4, a stride of 2, and a padding of 1. The number of input channels of the 6 deconvolution layers are 512, 1024, 1024, 512, 256, and 128, respectively, and the number of output channels are 512, 512, 256, 128, 64, and 3, respectively.

[0111] S3: Add the de-shadowed area image and the mapped image pixel by pixel to obtain a preliminary de-shadowed image.

[0112] S4: Send the preliminary shadow removal image to the color restoration module. Figure 2 The RefineNet module in is used to perform color restoration to obtain a clear image after removing shadows.

[0113] In the color restoration module, the mean and variance of the pixels in the non-shadow area and the shadow area are calculated respectively. Specifically, as shown in equations (18)-(20):

[0114] n total =n s +n n (18)

[0115]

[0116] Where n total Indicates that an image contains total pixels, n s Represents the shadow area pixels, n n represents the non-shadow area pixels, μ s and μ n Represent the mean values ​​of pixels in the shadow area and non-shadow area, respectively. and Represents the pixel variance of the non-shadow area and the pixel variance of the shadow area respectively. total represents the pixel mean of the entire image, Represents the variance of the pixels in the entire image.

[0117] Since the values ​​of the RGB channels in the shadow area are usually much lower than those in the unshadowed area, we can observe from the above formula that μ s , μ n compared to, and μ total Therefore, the mean and variance of the shadow area should be brought closer to the non-shadow area to make the whole image more harmonious.

[0118] The pixel mean and variance of the shadow area can be adjusted by the pixel mean and method normalization mapping of the non-shadow area. Specifically, the normalized value of the shadow pixel p located in (h, w, c) can be calculated by the following equations (21)-(23):

[0119]

[0120] Where p h,w,c and They are the initial pixel value and the normalized pixel value of the pixel with height h, width w and channel c (c∈R, G, B). ∈ is a minimum value to prevent the divisor from being zero. Represents the mean value of pixels in the shadow area of ​​the c channel, represents the variance of pixels in the shadow area of ​​the c channel, represents the variance of pixels in the non-shadow area of ​​the c channel, Indicates the mean value of pixels in the non-shadow area of ​​the c channel, NumRegion Represents the sum of the pixel values ​​of all pixels in the Region area. The Region area can be a shadow area or a non-shadow area.

[0121] S5: Calculate the total loss of the two-branch network and adjust the network parameters according to the total loss to obtain a trained two-branch network.

[0122] The formula for calculating the total loss of a two-branch network is:

[0123] L=λ 1 L STF +λ 2 L FTF +λ 3 L coarse +λ 4 L fine

[0124] L STF =||I RES_S -I SF_S ||

[0125] L FTF =||I RES_SF -I SF_SF ||

[0126] L coarse =||I course -I SF ||

[0127] L fine =||I fine -I SF ||

[0128] Among them, L represents the total loss, L STF is the shadow removal branch loss, L FTF is the identity mapping branch loss, is the roughness loss, is the refinement loss, λ 1 ~λ 4 are the first to fourth hyperparameters, ||·|| 1 Indicates L 1 Distance, I RES_S Represents the image with shadow removed, I RES_SF To map the image, I course To preliminarily remove the shadow image, I fine To remove the shadows, I SF is a true shadow-free image, I SF_S and I SF_SF are the shadow area and non-shadow area of ​​the true shadow-free image, respectively.

[0129] The network parameters are adjusted according to the loss function. When the loss function converges or reaches the maximum preset number of iterations, the training is stopped, the network parameters are maintained, and a trained dual-branch network is obtained. The shadow image to be processed is obtained together with the image mask and input into the trained dual-branch network for processing, and a clear image after removing the shadow can be obtained.

[0130] Evaluation of the present invention:

[0131] In order to quantify and compare the effectiveness of the present invention, the present invention is compared with algorithms such as DSC, AUTO-Exposure, SG-ShadowNet, BMNet, etc. on two public real-world data sets (SRD, ISTD). And three objective evaluation indicators are selected: root mean square error (RMSE) in LAB color space, peak signal to noise ratio (PSNR), and structural similarity index (SSIM). Given a shadow removal result and the corresponding shadow-free image, RMSE measures their average pixel error, PSNR measures their average pixel similarity, and SSIM measures their structural similarity. The lower the RMSE value, the better the generated result, and the higher the PSNR and SSIM values, the better the generated result.

[0132] Figure 3 The visual comparison of the processing results of the present invention and five shadow removal methods on real-world shadow images is shown. From left to right, the first column is the shadow image to be processed, followed by different shadow removal methods: Fu et al., DC-ShadowNet, Zhu et al., BMNet, ShadowDiffusion and the method proposed by the present invention, as well as real shadow-free images. The removal results of other methods have obvious artifacts at the edges of the shadows, or restore incorrect color and brightness information, while the dual-branch network of the present invention can reconstruct higher quality shadow-free images.

[0133] Figure 4 The quantitative comparison results of the dual-branch network image shadow removal method (DSR-Net) of the present invention and 9 shadow removal methods on the real shadow removal dataset SRD are shown, including the non-deep learning method Guo et al. and the deep learning-based methods DHAN, DC-ShadowNet, Fu et al., Zhu et al., BMNet, ShadowDiffusion, and ShadowFormer. Each removed image is divided into shadow area (S), non-shadow area (NS), and the entire image area (ALL), and the above three quantitative evaluation indicators are used to measure the shadow removal effects of different methods. Figure 4 The results show that the present invention has achieved comprehensive and excellent shadow removal results in terms of PSNR, SSIM and RMSE.

[0134] The above embodiments further illustrate the purpose, technical solutions and advantages of the present invention in detail. It should be understood that the above embodiments are only preferred implementation modes of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made to the present invention within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for removing image shadows based on a dual-branch network, characterized in that: include: Obtain the shadow image to be processed and its corresponding mask image and input them into the trained dual-branch network for processing to obtain a clear image after removing the shadow; The training process of the two-branch network includes: S1: Obtain a shadow image to be processed and its corresponding mask image, preprocess the shadow image and the mask image to obtain a shadow area image and a non-shadow area image; S2: Send the shadow area image and the non-shadow area image to the shadow removal branch network and the identity mapping branch network for processing respectively, to obtain the shadow removal area image and the mapping image; S3: adding the de-shadowed area image and the mapped image pixel by pixel to obtain a preliminary de-shadowed image; S4: sending the preliminary shadow-removed image to a color restoration module for color restoration to obtain a clear image after the shadows are removed; S5: Calculate the total loss of the two-branch network and adjust the network parameters according to the total loss to obtain a trained two-branch network.

2. According to the method for removing image shadows based on a dual-branch network in claim 1, it is characterized in that: The process of preprocessing the shadow image and the mask image includes: The shadow image and the mask image are both scaled to 256×256; the shadow image and the mask image are multiplied pixel by pixel to obtain the shadow area image; the mask image is inverted and multiplied pixel by pixel with the shadow image to obtain the non-shadow area image.

3. The image shadow removal method based on a dual-branch network according to claim 1 is characterized in that: The process of sending the shadow area image and the non-shadow area image to the shadow removal branch network for processing includes: Input the shadow area image into the relative light intensity estimation network for processing to obtain the relative light intensity; The non-shadow area image is input into the attenuation factor estimation network for processing to obtain the attenuation factor; A response attenuation model is established according to the relationship between pixel value and illumination, and the response attenuation model is fitted according to the relative light intensity and attenuation factor to obtain a shadow-removed area image.

4. The image shadow removal method based on a dual-branch network according to claim 3 is characterized in that: Both the relative light intensity estimation network and the attenuation factor estimation network adopt a symmetrical encoding-decoding structure; The encoder consists of 8 convolutional layers, with a convolution kernel of 4, a stride of 2, and a padding of 1; the number of input channels of the first convolutional layer is 4, and the number of output channels is 64; the number of input channels of the 2nd to 7th convolutional layers are 64, 128, 256, 512, 512, 512, and the number of output channels are 128, 256, 512, 512, 512, 512, respectively; LeakyRelu layers and BatchNorm layers are added before and after the 2nd to 7th convolutional layers respectively; the number of input channels of the 8th convolutional layer is 512, the number of output channels is 512, and a LeakyRelu layer is added before the 8th convolutional layer; The decoder includes 8 deconvolution layers. The number of input channels of the first deconvolution layer is twice the number of output channels of the last convolution layer of the encoder. The parameters of each deconvolution layer are set to a convolution kernel size of 4, a stride of 2, and a padding of 1. The number of input channels of the 1st to 7th convolution layers are 1024, 2048, 2048, 2048, 1024, and 512, respectively, and the number of output channels are 1024, 1024, 1024, 1024, 512, respectively, and the number of output channels are 1024, 1024, 1024, 1024, 512, 256, and 128, respectively. Relu layers and BatchNorm layers are added before and after the 1st to 7th convolution layers, respectively. The number of input channels of the 8th convolution layer is 256, and the number of output channels is 3. At the same time, Relu layers and sigmoid layers are added before and after the 8th convolution layer, respectively.

5. The image shadow removal method based on a dual-branch network according to claim 3 is characterized in that: The response attenuation model is expressed as: Among them, P represents the pixel value of the image after fitting the model, p s represents the shadow pixel value, k represents the incident light intensity ratio between the shadow-free area and the shadow area, a represents the attenuation factor corresponding to the object, and R represents the reflectivity.

6. The image shadow removal method based on a dual-branch network according to claim 1 is characterized in that: The identity mapping branch network consists of an encoder and a decoder; The encoder is based on the SR-Net structure and consists of 6 convolutional layers. The 1st to 4th convolutional layers are cascaded with the instance normalization module and the LeakyReLU activation function, respectively, and the 5th convolutional layer is cascaded with the BatchNorm layer and the LeakyReLU activation function. The parameters of all convolutional layers are set to convolution kernel size 4×4, stride 2, padding 1, the number of input channels is 4, 64, 128, 256, 512 and 512, and the number of output channels is 64, 128, 256, 512, 512 and 512, respectively. The decoder consists of 6 deconvolution layers. The first deconvolution layer is preceded and followed by a LeakyReLU activation layer and a BatchNorm layer, respectively. The second to fifth deconvolution layers are preceded and followed by a LeakyReLU activation layer and an instance normalization module, respectively. The sixth deconvolution layer is preceded and followed by a LeakyReLU activation layer and a tanh activation layer, respectively. The parameters of all deconvolution layers are set to a convolution kernel size of 4×4, a stride of 2, and a padding of 1. The number of input channels of the 6 deconvolution layers are 512, 1024, 1024, 512, 256, and 128, respectively, and the number of output channels are 512, 512, 256, 128, 64, and 3, respectively.

7. The image shadow removal method based on a dual-branch network according to claim 1 is characterized in that: The process of sending the initial shadow-removed image to the color restoration module for color restoration is expressed as: in, Represents the normalized pixel value, p h,w,c Indicates the initial pixel value of the image for preliminary shadow removal; represents the mean value of the pixels in the shadow area, represents the variance of the pixels in the shadow area, represents the variance of pixels in the non-shadow area, Represents the mean value of pixels in the non-shadow area.

8. The image shadow removal method based on a dual-branch network according to claim 1 is characterized in that: The formula for calculating the total loss of a two-branch network is: L=λ1L STF +λ2L FTF +λ3L coarse +λ4L fine L STF =||I RES_S -I SF_S || L FTF =||I RES_SF -I SF_SF || L coarse =||I course -I SF || L fine =||I fine -I SF || Among them, L represents the total loss, L STF is the shadow removal branch loss, L FTF is the identity mapping branch loss, is the roughness loss, is the refinement loss, λ1~λ4 are the first to fourth hyperparameters, ||·||1 represents the L1 distance, I RES_S Represents the image with shadow removed, I RES_SF To map the image, I course To preliminarily remove the shadow image, I fine To remove the shadows, I SF is a true shadow-free image, I SF_S and I SF_SF are the shadow area and non-shadow area of ​​the true shadow-free image, respectively.

Citation Information

Patent Citations

  • Non-paired image shadow removal method

    CN115146763A

  • Shadow removal method based on dynamic alignment and illumination perception convolution

    CN115937030A

  • Image shadow removing method

    CN118037593A

  • Face shadow removing method based on deep learning

    CN118071666A

  • Image shadow removing method and device

    CN118134823A

Cited By

  • Image data enhancement processing method for industrial visual inspection

    CN121120457A