Image Shadow Removal Method Based on Shadow Complexity-Aware Neural Network
By introducing shadow complexity perception mechanism and color style diversity enhancement methods in the shadow removal neural network, the problem of single color style and poor shadow complexity processing in the training data set in the prior art is solved, achieving a more efficient and robust shadow removal effect.
Patent Information
- Application Number
- CN202311039192.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-17
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2043-08-17
AI Technical Summary
The existing technology has two main problems in shadow removal: one is that the color style of the training data set is single and unbalanced, resulting in poor generalization performance of the model; the other is that the complexity of shadows is different. The existing methods usually use the entire neural network to process, resulting in poor performance and computing efficiency.
A neural network based on shadow complexity perception is proposed. By judging the complexity of shadows, it can reasonably allocate the computational amount and combine the color style diversity enhancement method to enhance the color diversity of the training data set. The network consists of grayscale structure information recovery branches and color information recovery branches, and dynamically adjusts the processing depth of the network according to the shadow complexity.
Improves shadow removal performance and model robustness, achieving a balance between shadow removal effects and computing resources, suitable for images of different color styles and shadow complexity.
Smart Images

Figure CN117058034B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to computer low-level vision, and in particular to an image shadow removal method based on a shadow complexity perception neural network. Background Art
[0002] A shadow is a dark area caused by an opaque object blocking light. In some cases, the shadow may occupy the entire three-dimensional space of the object, resulting in a decrease in the quality of the captured image. Shadows on the image may prevent these images from being used in a series of downstream computer vision tasks, including object detection and semantic segmentation. Therefore, the shadow removal task (including shadow detection and removal) has attracted extensive attention in the academic and industrial fields.
[0003] Before the advent of deep learning, image shadow removal methods relied heavily on the physical properties of shadows.For example, some methods use the shadow image gradient as prior knowledge (Finlayson G D, Drew M S, Lu C. Entropy minimization for shadow removal[J]. International Journal of Computer Vision, 2009, 85(1): 35-57. Gryka M, Terry M, Brostow G J. Learning to remove soft shadows[J]. ACM Transactions on Graphics(TOG), 2015, 34(5): 1-15.), some methods use the shadow color as prior information (Finlayson G D, Hordley S D, Lu C, et al. On the removal of shadows from images[J]. IEEE transactions on pattern analysis and machine intelligence, 2005, 28(1): 59-68. Zhang L, Zhang Q, Xiao C. Shadow remover: Image shadow removal based on illumination recovering optimization[J]. IEEE Transactions on Image Processing, 2015, 24(11): 4623-4636.), and some other methods use the relationship between the shadow region and the non-shadow region of the image as prior (Guo R, Dai Q, Hoiem D. Paired regions for shadow detection and removal[J]. IEEE transactions on pattern analysis and machine intelligence, 2012, 35(12): 2956-2967. Vicente TF Y, Hoai M, Samaras D. Leave-one-out kernel optimization for shadow detection and removal[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2017, 40(3): 682-695.).In addition, Gong et al. also proposed a method for user-interactive shadow removal (Gong H, Cosker D. Interactive removal and groundtruth for difficult shadow scenes[J]. JOSAA, 2016, 33(9): 1798-1811.).
[0004] With the development of deep learning and the emergence of large-scale shadow datasets, such as ISTD (Wang J, Li X, Yang J. Stacked conditional generative adversarial networks for jointly learning shadow detection and shadow removal[C] / / Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2018:1788-1797.), SRD (Qu L, Tian J, He S, et al. Deshadownet: A multi-context embedding deep network for shadow removal[C] / / Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2017:4067-4075.), shadow removal methods based on deep neural networks have become the mainstream methods. DeShadowNet (Qu L, Tian J, He S, et al. Deshadownet: A multi-context embedding deep network for shadow removal[C] / / Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2017:4067-4075.) constructs an end-to-end deep neural network with multiple contexts to solve the shadow removal problem. DSC (Hu X, Fu C W, Zhu L, et al. Direction-aware spatial context features for shadow detection and removal[J]. IEEE transactions on pattern analysis and machine intelligence, 2019, 42(11):2795-2808.) uses a direction-aware spatial attention module to obtain the context information of shadow images.CA-Net (Chen Z, Long C, Zhang L, et al. Canet: A context-aware network for shadow removal[C] / / Proceedings of the IEEE / CVF International Conference on ComputerVision. 2021:4743-4752.) constructs a two-layer context feature conduction model to conduct the features of non-shadow regions to shadow regions. Auto-Exposure (Fu L, Zhou C, Guo Q, et al. Auto-exposure fusion forsingle-image shadow removal[C] / / Proceedings ofthe IEEE / CVF conference oncomputer vision and pattern recognition. 2021:10571-10580.) regards the shadow removal problem as a problem of automatic fusion of multi-exposure images. SP+M-Net (Le H, Samaras D. Shadow removal viashadow image decomposition[C] / / Proceedings of the IEEE / CVF InternationalConference on ComputerVision. 2019:8578-8587.) and SP+M+I-Net (Le H, SamarasD. Physics-based shadow image decomposition for shadow removal[J].IEEETransactions on PatternAnalysis and Machine Intelligence, 2021, 44(12):9088-9101.) propose an illumination model to decompose shadows and design neural networks specifically for this purpose.EMDN (Zhu Y, Xiao Z, Fang Y, et al. Efficient Model-Driven Network for Shadow Removal [C] / / Proceedings of the AAAI Conference on Artificial Intelligence. 2022, 36(3): 3635-3643.) uses an iterative algorithm based on an interpretable illumination model to improve the performance of shadow removal using neural networks, while BM-Net (Zhu Y, Huang J, Fu X, et al. Bijective mapping network for shadow removal [C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2022: 5627-5636.) employs a reversible neural network to perform the shadow removal task, greatly reducing the number of required parameters and the amount of computation.
[0005] The emergence of generative adversarial networks has inspired a series of new image generation methods, and many shadow removal methods have also been studied along the lines of Style-GAN (Karras T, Laine S, Aila T. A style-based generator architecture for generative adversarial networks[C] / / Proceedings of the IEEE / CVF conference on computer vision and pattern recognition. 2019:4401-4410.) and Conditional-GAN (Isola P, Zhu J Y, Zhou T, et al. Image-to-image translation with conditional adversarial networks[C] / / Proceedings of the IEEE conference on computer vision and pattern recognition. 2017:1125-1134.). DHAN (Cun X, Pun C M, Shi C. Towards ghost-free shadow removal via dual hierarchical aggregation network and shadow matting GAN[C] / / Proceedings of the AAAI Conference on Artificial Intelligence. 2020, 34(07):10680-10687.) constructs a dual aggregation network and uses adversarial learning to generate more realistic shadow-free images.ST-CGAN (Wang J, Li X, Yang J. Stacked conditional generative adversarial networks for jointly learning shadow detection and shadow removal [C] / / Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2018:1788-1797.) and ARGAN (Ding B, Long C, Zhang L, et al. Argan: Attentive recurrent generative adversarial network for shadow detection and removal [C] / / Proceedings of the IEEE / CVF international conference on computer vision. 2019:10213-10222.) implement the joint task of shadow detection and removal using conditional generative adversarial networks. SG-ShadowNet (Wan J, Yin H, Wu Z, et al. Style-Guided Shadow Removal [C] / / Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part XIX. Cham: Springer Nature Switzerland, 2022:361-378.) regards shadow removal as an image generation task and uses Style-GAN to transfer the style from non-shadow regions to shadow regions.
[0006] The present invention discovers two prominent problems related to the training dataset and neural network design for shadow removal:
[0007] The first problem is that deep neural networks require diverse training data to improve generalization performance. However, existing shadow removal datasets are limited in color style, and there is a serious imbalance in the number of training samples with different color styles, which is exacerbated by shadows with different color styles widely existing in the human world environment. Therefore, the trained model may overmatch some specific color styles among them, resulting in poor shadow removal performance.
[0008] The second problem is that due to the different complexities of shadows, removing weak shadows requires less computational effort than strong shadows. Current methods usually use the entire neural network to process shadow images, which may not be optimal in terms of performance and computational efficiency. Summary of the Invention
[0009] The object of the present invention aims to address the above two problems existing in the prior art and provide a more efficient and better-performing method for removing image shadows. The present invention proposes a novel neural network based on shadow complexity perception, which can reasonably allocate the computational effort required to remove them by judging the complexity of shadows. In addition, the present invention also integrates a new step to enhance color style diversity to enhance the color diversity of the input training samples and solve the problem of the single and unbalanced color styles of the training set.
[0010] The present invention includes the following steps:
[0011] 1) Use the color style diversity enhancement method to perform color style transformation on the training set samples participating in training;
[0012] 2) Feed the samples after color transformation into the gray-scale structure information restoration branch of the neural network based on shadow complexity perception to restore the gray-scale structure information of the image;
[0013] 3) Judge the complexity of shadows according to the difference between the output result of the gray-scale structure information restoration branch and the gray-scale map of the input image;
[0014] 4) Feed the samples after color transformation and the output result of the gray-scale structure information restoration branch into the color information restoration branch of the neural network based on shadow complexity perception to restore the color information of the image; Images with low shadow complexity will exit the neural network in advance, while images with high shadow complexity will be processed through more parameters in the neural network;
[0015] 5) Calculate the L1 loss and gradient loss between the output result of the gray-scale structure information restoration branch and the gray-scale map of the shadowless image, and calculate the L1 loss, perceptual loss, and multi-exit distillation loss between the output result of the color information restoration branch and the shadowless image;
[0016] 6) Add up the various losses in different proportions as the loss of the entire network and perform backpropagation to train the neural network to obtain an excellent-performing shadow removal network.
[0017] In step 1), the color style diversity enhancement method refers to pre - defining ten different color style transformation matrices (all with dimensions of 3×3, corresponding to C0 - C9 in the following formulas respectively), namely chromaticity enhancement and chromaticity weakening matrices, red enhancement and red weakening matrices, green enhancement and green weakening matrices, blue enhancement and blue weakening matrices, primary color interference enhancement and primary color interference weakening matrices. The first two matrices can mainly increase the diversity of image illumination brightness, the following six matrices are used to increase or weaken specific colors, and the last two matrices are used to increase or decrease the color richness of a single picture.
[0018]
[0019]
[0020]
[0021]
[0022]
[0023]
[0024] Among them, α, β, and γ are all non - negative hyperparameters, and their corresponding values will be given in the implementation details.
[0025] The color transformation of the training samples participating in the training means that the same color style transformation is performed on both the shadow image - shadowless image (groundtruth) pairs.
[0026] The color transformation of the training samples participating in the training means that each time an iteration is performed, one of the above - mentioned ten color style transformations is randomly selected and multiplied by the picture sample matrix.
[0027] In step 2), the grayscale structure information restoration branch mainly includes a specially designed small U - Net module called StructureUBlock. This module takes the shadow image and the shadow mask as inputs and outputs the grayscale image after removing the shadow.
[0028] In step 3), the specific details of the method for judging shadow complexity are as follows:
[0029]
[0030] Among them, I s is the training sample, that is, the shadow image. G1, G2, and G3 respectively represent the three difficulty sets of simple, medium, and complex shadow complexity, mean() represents calculating the mean, The grayscale image without shadow representing the output of StructureUBlock, represents the input grayscale image with shadow. t1 and t2 are the thresholds for judging shadow complexity, and their corresponding values will be given in the implementation details.
[0031] In step 4), the color information recovery branch consists of three ColorUBlocks, and each ColorUBlock has its own exit, corresponding to simple shadow, medium shadow, and complex shadow respectively.
[0032] In step 5), calculate the L1 loss between the output result of the grayscale structure information recovery branch and the grayscale image of the shadowless image. The specific details are as follows:
[0033]
[0034] where, represents the grayscale image of the ground truth.
[0035] Calculate the gradient loss between the output result of the grayscale structure information recovery branch and the grayscale image of the shadowless image. The specific details are as follows:
[0036]
[0037] where, represents the gradient of the calculated image.
[0038] Calculate the L1 loss between the output result of the color information recovery branch and the shadowless image. The specific details are as follows:
[0039]
[0040]
[0041] Calculate the L1 loss of the shadow area and the non - shadow area separately. Among them, M represents the shadow mask, and I f represents the shadowless image output by the third ColorUBlock, represents the ground truth.
[0042] Calculate the perceptual loss between the output result of the color information recovery branch and the shadowless image. The specific details are as follows:
[0043]
[0044] where, VGG() represents the features of different layers extracted using the VGG16 network.
[0045] The different exit distillation losses between the output result of the calculated color information recovery branch and the shadowless image are as follows in detail:
[0046]
[0047] Among them, I f-1 / 2 refers to the shadowless image output by the first ColorUBlock or the second ColorUBlock, and StopGrad() represents the operation of stopping gradient calculation.
[0048] In step 6), the summation of each loss according to different ratios is as follows in detail:
[0049]
[0050]
[0051]
[0052] Among them, λ represents the weighting coefficient of different loss terms, and their corresponding values will be given in the implementation details.
[0053] Compared with the prior art, the present invention has the following outstanding advantages:
[0054] 1. Because the diversity of image color styles in the real world is considered, the present invention has done the work of enhancing the color style diversity of the training dataset, which makes the model trained with the data after data enhancement have extremely strong robustness and is used to solve the image shadows in various different scenarios in the real world;
[0055] 2. Because the complexity of shadows in the real world is considered to be non-unique, the present invention first evaluates the complexity of shadows for this phenomenon, and then reasonably allocates the number of parameters and the amount of calculation for the test samples according to the evaluated shadow complexity, achieving a balance between the shadow removal effect and the number of model parameters and the amount of calculation.
[0056] 3. The network structure and data enhancement method proposed by the present invention belong to a plug-and-play method and can be easily combined with future shadow removal methods and similar tasks to improve their performance. Brief Description of the Drawings
[0057] Figure 1 It is a schematic diagram of the method framework of the embodiment of the present invention.
[0058] Figure 2 It is a schematic diagram of the structural composition of the ColorUBlock of the embodiment of the present invention. Detailed Embodiment
[0059] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the following embodiments will further illustrate the present invention in conjunction with the accompanying drawings.
[0060] The method framework diagram of the embodiment of the present invention is as Figure 1 shown.
[0061] 1. Network structure description
[0062] The neural network designed by the present invention consists of two parts, namely the grayscale structure information restoration branch ( Figure 1 the upper half) and the color information restoration branch ( Figure 1 the lower half).
[0063] The function of the grayscale structure information restoration branch is to initially remove shadows and generate a shadow-free grayscale version of the image. This branch mainly consists of a StructureUBlock. The first half of the StructureUBlock is stacked with convolution operations and pooling operations, and its main function is to extract the features of the image and compress the resolution of the image. The second half of the StructureUBlock consists of an attention module and a PixelShuffle module, and its main function is to aggregate features at different levels and restore the resolution.
[0064] The function of the color information restoration branch is to generate the final shadow-free image. This branch is composed of three ColorUBlocks connected in series. Inside the ColorUBlock, it is mainly composed of CodConv ( Figure 2 ), which uses a parallel structure of dilated convolutions with different dilation rates to effectively extract the corresponding image features under different receptive fields. This branch is a three-output branch. For images with simple shadows, a good restoration result can be obtained only by passing through the first ColorUBlock; for images with difficult shadows, a good restoration result can be obtained only by passing through three ColorUBlocks.
[0065] 2. Motivation analysis
[0066] The present invention realizes two prominent problems related to the training dataset for shadow removal and the neural network design.
[0067] First of all, deep neural networks require diverse training data to improve generalization performance. However, the existing shadow removal datasets are limited in color style, and there is a serious imbalance in the number of training samples with different color styles. This problem is exacerbated by the shadows with different color styles widely existing in the human world environment. Therefore, the trained model may overmatch some specific color styles, resulting in poor shadow removal performance.
[0068] The second problem is that due to the different complexities of shadows, removing weak shadows requires less computational effort than strong shadows. Current methods usually use the entire neural network to process shadow images, which may not be optimal in terms of performance and computational efficiency.
[0069] 3. Training Instructions
[0070] The embodiments of the present invention include the following steps:
[0071] 1) Use the color style diversity enhancement method to perform color style transformation on the training set samples participating in the training;
[0072] 2) Send the samples after color transformation into the gray-scale structure information recovery branch of the neural network based on shadow complexity perception to recover the gray-scale structure information of the image;
[0073] 3) Judge the complexity of the shadow according to the difference between the output result of the gray-scale structure information recovery branch and the gray-scale image of the input image;
[0074] 4) Send the samples after color transformation and the output result of the gray-scale structure information recovery branch into the color information recovery branch of the neural network based on shadow complexity perception to recover the color information of the image. Images with low shadow complexity will exit the neural network in advance, while images with high shadow complexity will be processed through more parameters in the neural network;
[0075] 5) Calculate the L1 loss and gradient loss between the output result of the gray-scale structure information recovery branch and the gray-scale image of the shadowless image, and calculate the L1 loss, perceptual loss, and multi-exit distillation loss between the output result of the color information recovery branch and the shadowless image.
[0076] 6) Add up the various losses according to different proportions as the loss of the entire network and perform backpropagation to train the neural network to obtain an excellent performance shadow removal network.
[0077] In step 1), the color style diversity enhancement method means that the present invention pre-defines ten different color style transformation matrices (all with dimensions of 3×3, corresponding to C0-C9 of the following formulas respectively), namely chromaticity enhancement and chromaticity weakening matrices, red enhancement and red weakening matrices, green enhancement and green weakening matrices, blue enhancement and blue weakening matrices, primary color interference enhancement and primary color interference weakening matrices; the first two matrices can mainly increase the diversity of image illumination brightness, the following six matrices are used to increase or weaken specific colors, and the last two matrices are used to increase or weaken the color richness of a single picture.
[0078]
[0079]
[0080]
[0081]
[0082]
[0083]
[0084] Among them, α, β, and γ are all non-negative hyperparameters. In the present invention, α and β are set to 0.2, and γ is set to 0.1. It should be noted that these three hyperparameters can be fine-tuned according to different specific tasks, but it is necessary not to exceed a reasonable range, otherwise it will seriously affect the performance of the trained model.
[0085] The color transformation of the training samples participating in the training refers to performing the same color style transformation on the shadow image - shadowless image (groundtruth) pair.
[0086] The color transformation of the training samples participating in the training means randomly selecting one of the above ten color style transformations for each iteration and performing a matrix multiplication operation with the picture samples.
[0087] In step 2), the grayscale structure information restoration branch mainly includes a specially designed small U-Net module called StructureUBlock. This module takes the shadow image and the shadow mask as inputs and outputs the grayscale image after removing the shadow.
[0088] In step 3), the specific details of the method for judging shadow complexity are as follows:
[0089]
[0090] Among them, I s is the training sample, that is, the shadow image; G1, G2, and G3 respectively represent the three difficulty sets of simple, medium, and complex shadow complexity, mean() represents calculating the mean value, represents the shadowless grayscale image output by StructureUBlock, Represents the input grayscale shadow image. t1 and t2 are the thresholds for judging shadow complexity respectively. Based on statistical and observational results, in the present invention, t1 is set to 0.2 and t2 is set to 0.3. It should be noted that these two hyperparameters can also be adjusted according to the actual situation. If their values are made smaller, the overall computational amount of the model may increase, and correspondingly, the shadow removal effect will also improve; conversely, if their values are made larger, more samples will be judged as samples with relatively simple shadow complexity, so the computational amount of the model may decrease, but the shadow removal effect may be slightly affected. This set of hyperparameters set in the present invention balances the shadow removal effect and the computational amount, and will enable the model to reach an optimal state.
[0091] In step 4), the color information restoration branch consists of three ColorUBlocks, and each ColorUBlock has its own exit, corresponding to simple shadows, medium shadows, and complex shadows respectively.
[0092] In step 5), calculate the L1 loss between the output result of the grayscale structure information restoration branch and the grayscale image of the shadowless image. The specific details are as follows:
[0093]
[0094] Among them, Represents the grayscale image of the ground truth.
[0095] Calculate the gradient loss between the output result of the grayscale structure information restoration branch and the grayscale image of the shadowless image. The specific details are as follows:
[0096]
[0097] Among them, Represents the gradient of the calculated image.
[0098] Calculate the L1 loss between the output result of the color information restoration branch and the shadowless image. The specific details are as follows:
[0099]
[0100]
[0101] In the present invention, the L1 losses of the shadow area and the non - shadow area are calculated separately. Among them, M represents the shadow mask, and I f Represents the shadowless image output by the third ColorUBlock, Represents the ground truth.
[0102] Calculate the perceptual loss between the output result of the color information recovery branch and the shadowless image. The specific details are as follows:
[0103]
[0104] Among them, VGG() represents the features of different layers extracted using the VGG16 network.
[0105] Calculate the different exit distillation loss between the output result of the color information recovery branch and the shadowless image. The specific details are as follows:
[0106]
[0107] Among them, I f-1 / 2 refers to the shadowless image output by the first ColorUBlock or the second ColorUBlock. StopGrad() represents the operation of stopping gradient calculation. According to the actual situation during training, if the shadow in the training sample belongs to simple complexity, the output of the first ColorUBlock will be used as the final result of the model, and the multi-exit distillation loss will also be calculated using it; if the shadow in the training sample belongs to medium complexity, the output of the second ColorUBlock will be used as the final result of the model, and the multi-exit distillation loss will also be calculated using it; if the shadow in the training sample belongs to difficult complexity, this distillation loss is not required.
[0108] In step 6), add up each loss according to different proportions. The specific details are as follows:
[0109]
[0110]
[0111]
[0112] Among them, λ represents the weighting coefficient of different loss terms. In the present invention, λ grad is set to 0.1, λ s is set to 100, λ ns is set to 20, λ kd is set to 50, λ structure is set to 100. The present invention implements an end-to-end neural network and jointly trains the gray-scale structure information branch and the color branch in a joint training manner.
[0113] 4. Implementation details
[0114] The image shadow removal method based on the shadow complexity-aware neural network proposed by the present invention uses three datasets, namely ISTD, SRD, and ISTD+, for effect evaluation. The present invention is implemented using Pytorch and trained and tested using a single NVIDIA 3090. In the training phase, the input image is resized to 400×400, the batch size is set to 2, and a total of 600 rounds of training are carried out. The present invention uses the Adam optimizer, with the initial learning rate set to 0.0002 and gradually decreased during the training process. In the testing phase, the resolution is set to 640×480 (ISTD dataset and ISTD+ dataset) or 840×640 (SRD dataset).
[0115] 5. Application Field
[0116] The present invention can be applied to the field of computer low-level vision to remove shadows in single images. Table 1 shows the performance comparison of the present invention with other shadow removal methods on the ISTD dataset. The present invention uses three metrics to measure performance. The larger the PSNR, the better the effect; the larger the SSIM, the better the effect; the smaller the MAE, the better the effect. Among them, bold indicates the best result and underlined indicates the second-best result.
[0117] Table 1
[0118]
[0119] As can be seen from Table 1, the method of the present invention significantly leads other methods in the recovery effect of the shadow area. The PSNR increases from 36.95 to 38.81 (an increase of 1.86) compared with EMDN, the SSIM increases from 0.9880 to 0.9905 (an increase of 0.0025) compared with BM-Net, and the MAE decreases from 6.65 to 6.28 (a decrease of 0.37) compared with ARGAN, and it has fewer parameters and less computational complexity.
[0120] Table 2 shows the performance comparison of the present invention with other shadow removal methods on the SRD dataset.
[0121] Table 2
[0122]
[0123] As can be seen from Table 2, the method of the present invention significantly leads other methods in the recovery effect of the shadow area. The PSNR increases from 35.05 to 36.44 (an increase of 1.39) compared with BM-Net, the SSIM increases from 0.9810 to 0.9826 (an increase of 0.0016) compared with BM-Net, and the MAE decreases from 6.35 to 6.20 (a decrease of 0.15) compared with ARGAN.
[0124] Table 3 shows the performance comparison of the present invention with other shadow removal methods on the ISTD+ dataset.
[0125] Table 3
[0126]
[0127] As can be seen from Table 3, the present invention also achieves relatively good performance.
[0128] Therefore, in summary, on most of the currently public shadow datasets, the present invention has obtained the best results, which proves the effectiveness of the present invention.
[0129] The above embodiments are only preferred embodiments of the present invention and should not be considered as limiting the scope of implementation of the present invention. All equivalent changes and improvements made within the scope of the application of the present invention shall still fall within the scope covered by the patent of the present invention.
Claims
1. An image shadow removal method based on a shadow complexity-aware neural network, characterized in that Including the following steps: 1) Use the color style diversity enhancement method to perform color style transformation on the training set samples participating in the training; The specific steps of the color style diversity enhancement method are as follows: Pre-define ten different color style transformation matrices, namely the chromaticity enhancement matrix and the chromaticity weakening matrix, the red enhancement matrix and the red weakening matrix, the green enhancement matrix and the green weakening matrix, the blue enhancement matrix and the blue weakening matrix, the primary color interference enhancement matrix and the primary color interference weakening matrix, all with a dimension of 3×3, corresponding to C0 - C9 of the following formulas respectively; The first two matrices are used to increase the diversity of image illumination brightness, the following six matrices are used to increase or weaken specific colors, and the last two matrices are used to increase or weaken the color richness of a single picture; Among them, α, β, and γ are all non-negative hyperparameters; 2) Send the samples after color transformation into the gray-scale structure information recovery branch based on the shadow complexity perception neural network to recover the gray-scale structure information of the image; The gray-scale structure information recovery branch includes a specially designed small U-Net module called StructureUBlock, which takes the shadow image and the shadow mask as inputs and outputs the gray-scale image after removing the shadow; 3) Judge the shadow complexity according to the difference between the output result of the gray-scale structure information recovery branch and the gray-scale image of the input image; The method for judging the shadow complexity is as follows: Among them, G1, G2, and G3 respectively represent three difficulty sets of simple, medium, and complex shadow complexities, mean() represents calculating the mean value, represents the shadowless grayscale image output by StructureUBlock, represents the input shadow grayscale image; t1 and t2 are respectively the thresholds for judging shadow complexity; 4) Send the samples after color transformation and the output result of the gray-scale structure information recovery branch into the color information recovery branch based on the shadow complexity perception neural network to recover the color information of the image; Images with low shadow complexity will exit the neural network in advance, while images with high shadow complexity will be processed through more parameters in the neural network; The color information recovery branch consists of three ColorUBlocks, and each ColorUBlock has its own exit, corresponding to simple shadow, medium shadow, and complex shadow respectively; 5) Calculate the L1 loss and gradient loss between the output result of the gray-scale structure information recovery branch and the gray-scale image of the shadowless image, and calculate the L1 loss, perceptual loss, and multi-exit distillation loss between the output result of the color information recovery branch and the shadowless image; 6) Add up the various losses according to different proportions as the loss of the entire network and perform backpropagation to train the neural network to obtain an excellent shadow removal network.
2. The method for removing image shadows based on a shadow complexity-aware neural network according to claim 1, characterized in that In step 1), when performing color style transformation on the training set samples participating in the training, the same color style transformation is performed on the shadow image-shadowless image pair.
3. The image shadow removal method based on a shadow complexity-aware neural network according to claim 1, characterized in that In step 1), when performing color style transformation on the training set samples participating in the training, randomly select one from the ten color style transformations each time and perform the operation of matrix multiplication with the picture samples.
4. The image shadow removal method based on a shadow complexity-aware neural network according to claim 1, characterized in that In step 5), the calculation of the L1 loss between the output result of the gray-scale structure information recovery branch and the gray-scale image of the shadowless image is as follows: Among them, The grayscale image representing the ground truth; Calculate the gradient loss between the output result of the gray-scale structure information recovery branch and the gray-scale image of the shadowless image as follows: Among them, represents the gradient of the computed image; The L1 loss between the output result of the calculated color information recovery branch and the shadowless image is as follows: Calculate the L1 loss of the shaded region and the non - shaded region separately; where M represents the shadow mask, and I f represents the shadow - free image output by the third ColorUBlock, represents the ground truth; The perceptual loss between the output result of the calculated color information recovery branch and the shadowless image is as follows: where VGG() represents the features of different layers extracted using the VGG16 network; The multi-exit distillation loss, that is, the different exit distillation losses between the output result of the calculated color information recovery branch and the shadowless image, is as follows: Among them, I f-1 / 2 refers to the shadowless image output by the first ColorUBlock or the second ColorUBlock, and StopGrad() represents the operation of stopping gradient calculation.
5. The method for removing image shadows based on a shadow complexity-aware neural network according to claim 4, characterized in that In step 6), the summation of each loss according to different ratios is as follows: where λ represents the weighting coefficient of different loss terms.
Citation Information
Patent Citations
Color image enhancement method based on bright channel filtering
CN103578084A
Blocking perception Hash tracking method with shadow removing
CN105989611A