Image rain and snow removal method based on exposure imaging model and modular deep network

By using a physical imaging-based nonlinear model and modular deep network, the problems of unrealistic rain and snow models and insufficient robustness in existing rain and snow image restoration methods are solved, achieving faster and more accurate rain and snow removal results.

CN115471414BActive Publication Date: 2026-03-27TAIYUAN UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-24
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing deep learning-based rain and snow image restoration methods have limitations in rain and snow removal performance, especially the linear superposition rain model, which leads to background pixel saturation and inaccurate data, affecting network robustness.

Method used

A nonlinear rain and snow exposure imaging model based on physical imaging is adopted. Local and global U-shaped rain and snow detection subnetworks, shrinking exposure time estimation subnetworks and stacked multi-scale residual rain and snow removal subnetworks are designed. Rain and snow images are estimated through multi-branch networks and occlusions are removed step by step.

Benefits of technology

It achieves more accurate description of rain and snow imaging characteristics and more realistic image restoration, improves the robustness and speed of rain and snow removal algorithms, and can effectively handle diverse rain and snow images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115471414B_ABST
    Figure CN115471414B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of image restoration, and discloses an image rain and snow removing method based on an exposure imaging model and a modular deep network. In view of the fact that a simplified linear superposition rain model in an existing single image rain removing method based on deep learning can synthesize unreal training and test rain image data sets, and then affect the network architecture design based on the model and the rain removing effect, a nonlinear rain and snow exposure imaging model is provided, the model considers the exposure factor of rain and snow imaging, and the model can more accurately describe the imaging characteristics of rain and snow and the rain and snow shielding from semi-transparent to opaque in an image. A diversified rain and snow data set with different exposure times is synthesized by using the model, and a novel rain and snow removing network is designed. Since the network fuses the estimation of the exposure time, compared with the existing rain and snow removing network, the network can more effectively remove the rain and snow with different shielding.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of image restoration, and particularly relates to an image rain and snow removal method based on an exposure imaging model and a modular deep network. BACKGROUND

[0002] An image is the basis of vision and the source of data for a computer vision system to complete autonomous perception and understanding of the environment. In an outdoor unstructured visual perception system, adverse weather (rain, snow, fog) can form occlusions and affect the imaging characteristics of background objects. This causes a certain degradation in vision and information in images taken in adverse weather, which seriously affects the robustness and environmental adaptability of many high-level vision tasks of outdoor robots, such as semantic segmentation, target recognition and tracking, unmanned machine vision perception and navigation, etc., making the reliability of the vision system difficult to meet actual requirements. Therefore, image restoration in adverse weather conditions has important theoretical significance and application value for improving the robustness and environmental adaptability of machine vision. How to restore clear image data from a large amount of visual observation data in adverse weather to improve the autonomous perception and adaptability of robots to the environment has become a research hotspot in the fields of robots, pattern recognition and computer vision.

[0003] Rain and snow image technology is usually to restore clear images only according to rain and snow degraded image data, which is a pathological problem in mathematical modeling and solving, i.e. belongs to image blind restoration. Due to the lack of sufficient information to uniquely determine the real clear image, it is usually necessary to qualitatively and quantitatively analyze the rain and snow to form prior constraints, and the image restoration is to obtain the best estimate value of the real image under certain constraint criteria on the premise of meeting these constraints. Image restoration in adverse weather has become a key problem to be broken through in current machine vision and its applications, and its research progress has important theoretical significance and application value for improving the robustness and environmental adaptability of machine vision methods and systems. At present, this rain and snow restoration technology mainly includes two types: optimization type restoration method based on prior information and restoration method based on deep learning. The image restoration technology based on prior optimization mainly obtains the best estimate of the clear image under the constraint of prior information. This kind of method has a common limitation: their effectiveness depends largely on the accuracy of the prior information, and they can usually only process some rain and snow degraded images that meet specific prior conditions. The image restoration technology based on deep learning is to use a deep network to establish a mapping relationship from degraded rain and snow images to clear images. At present, these deep learning-based methods are usually superior to prior optimization type methods in terms of restoration effect.

[0004] The existing rain and snow image restoration methods based on deep learning mainly include: a completely data-driven deep rain and snow removal network and a model and data double-driven deep rain removal network. The completely data-driven deep rain and snow removal network can obtain clear rain and snow-free images, but since most of these methods introduce new network structures to achieve better rain / snow removal performance, the existing rain and snow removal network structures become more and more complex, but the rain and snow removal effect does not obviously improve. The introduction of new network structures has appeared a bottleneck in theory and actual effect for the rain and snow removal method based on deep network. In another kind of rain and snow model and data double-driven deep rain removal network, the model of the rain streak image is a linear superposition model. The model separates the clean background image and the linear superposition of the rain streak image. The linear superposition rain model synthesizes the rain streak image, and the direct superposition of pixels will saturate the grayish white pixels of the background, which makes the rain streak image data not real, and directly affects the robustness of the data-driven rain removal network. SUMMARY

[0005] In the existing rain and snow model and data double-driven rain deep network method, most of them are based on the simplified linear superposition rain model to design the rain removal network, and the simplified linear superposition rain model will synthesize the unrealistic training and test rain image dataset, and then affect the network architecture design based on the model and the rain removal effect. The present application proposes a nonlinear rain and snow exposure imaging model based on physical imaging, which can more accurately describe the rain and snow imaging characteristics and the rain and snow occlusion from semi-transparent to opaque in the image. Based on this model, a diversified rain and snow dataset with different exposure times is synthesized, and a novel rain and snow removal network is designed. The network includes a local and global U-shaped rain / snow detection sub-network, a shrinkage exposure time estimation sub-network and a stacked multi-scale residual rain and snow removal sub-network. The local and global U-shaped rain / snow detection sub-network not only uses the residual dense connection U-shaped auto-encoding structure to fully utilize all levels of features, but also introduces a local-global attention block between the encoder and the decoder to improve the pixel-level estimation accuracy. The stacked multi-scale residual rain and snow removal sub-network fuses multi-scale features through inter-layer residual, and gradually recovers the rain and snow image through the stacked structure. The method has good robustness, strong generalization ability and fast processing speed, and can adapt to multiple public rain and snow image datasets.

[0006] In order to achieve the above purpose, the application adopts the following technical scheme:

[0007] An image rain and snow removal method based on an exposure imaging model and a modular deep network, comprising the following steps:

[0008] Step 1, a nonlinear rain and snow exposure imaging model is established, which can be expressed as:

[0009]

[0010] where I rs represents the rain / snow image intensity, I c is the clean background image, K represents the average irradiance of rain / snow, and the rain / snow image intensity I rs is the clean background image I c , the dynamic rain / snow image M, and the nonlinear combination of the exposure time T;

[0011] Step 2, a deep rain / snow removal network is established, the network includes a local and global U-shaped rain / snow detection sub-network, a shrinkage type exposure time estimation sub-network, and a stacked multi-scale residual rain / snow removal sub-network; the network uses a multi-branch network to simultaneously estimate the rain / snow image and the exposure time, after applying the exposure imaging model to obtain a preliminary rain / snow-free image, a reinforcement sub-network, i.e., the stacked multi-scale residual rain / snow removal sub-network, is used to obtain a fine rain / snow-free image;

[0012] Step 3, the rain / snow image is input into the local and global U-shaped rain / snow detection sub-network, the global and local features of the rain / snow image are extracted, and a rain / snow streak image is estimated;

[0013] Step 4, the rain / snow image is input into the shrinkage type exposure time estimation sub-network, the global information is sampled into a single variable value to obtain an exposure time estimation value, and then the estimation value is up-sampled to the same size as the input image for loss estimation;

[0014] Step 5, the rain / snow streak image and the exposure time obtained in steps 3 and 4 are used to obtain a preliminary recovered rain / snow image through the nonlinear rain / snow exposure imaging model in step 1.

[0015] Step 6, the preliminary recovered rain / snow image obtained in step 4 is input into the stacked multi-scale residual rain / snow removal sub-network to further remove rain / snow occlusion of different degrees, and a final fine rain / snow-free image is obtained;

[0016] Step 7, the loss functions of the local and global U-shaped rain / snow detection sub-network, the shrinkage type exposure time estimation sub-network, and the stacked multi-scale residual rain / snow removal sub-network are set, and under the constraint of the loss functions, the network model parameters are trained and learned through stepwise iteration of each sub-network on the rain / snow dataset until the network converges;

[0017] Step 8, the effectiveness of the algorithm in removing rain / snow is quantitatively and qualitatively verified on synthetic and real rain / snow images, and compared with typical algorithms.

[0018] Further, the specific process of establishing the nonlinear rain / snow exposure imaging model in step 1 is as follows:

[0019] During the exposure time T, the intensity Irs is the intensity of static rain streaks or snowflakes R s with the background R b is a linear combination of:

[0020]

[0021] where τ represents the time that rain streaks or snowflakes fall through a pixel, since the background is almost static in the exposure time T, R b can be approximated as a constant, thus the above equation is rewritten as:

[0022]

[0023] In the equation, the time τ and the speed of rain / snow fall are related, and the uncertainty of which directly leads to the diversity of rain / snow imaging results; therefore, the integral of static rain / snow with respect to time τ can be regarded as a dynamic rain / snow map M, which enriches the diversity of rain / snow imaging transparency and introduces motion blur; TR b is the clean background image I c ; thus the above equation can be rewritten as:

[0024]

[0025] where represents the average intensity of rain / snow, represents the dynamic rain / snow map M; in order to make each term of the equation have certain physical meaning, the equation is further rewritten in the following form:

[0026]

[0027] where * represents element-wise multiplication, because represents the dynamic rain / snow map M, TR b is the clean background image I c , so the above equation is written as:

[0028]

[0029] Since the average irradiance of rain / snow is determined by the characteristics of rain / snow itself, it can be approximated as a constant K; let k = KT, the equation is further written as:

[0030]

[0031] where, I rs represents the rain / snow image intensity, I c is the clean background image, and K represents the average irradiance of rain / snow, in this rain / snow model, the rain / snow image intensity I rs is the clean background image I c, a nonlinear combination of dynamic rain / snow map M and exposure time T. This model takes into account the exposure factors of imaging, and can more accurately describe the imaging characteristics of rain and snow and the rain and snow occlusion from translucent to opaque in the image compared with the linear superposition model and the simplified nonlinear model.

[0032] Further, the local and global U-shaped rain / snow detection sub-network in step 2 is a residual dense U-shaped auto-encoder, which includes an encoder composed of four-stage residual dense blocks and down-sampling layers, and a decoder composed of global-local attention modules, four-stage residual dense blocks and up-sampling layers; the shrinkage type exposure time estimation sub-network is composed of eight-stage down-sampling operations and adaptive pooling layers; the stacked multi-scale residual rain and snow removal sub-network is two parallel residual multi-scale blocks stacked in a residual manner.

[0033] Further, the encoder is used to capture context information such as density and direction; the decoder is used to capture low-level information such as position, transparency and edge; the global-local attention module is used to extract global and local semantic features of the encoder, further improving the feature representation capability of the self-network; the residual dense block as a feature extraction unit is used to fully utilize the features of all convolutional layers, including hierarchical feature fusion (HFF) and residual learning (RL), wherein the HFF operation can be represented as:

[0034] F HFF =f HFF ([F s ,x1,x2...,x l ]),

[0035] where [F s ,x1,...,x l-1 ] represents the concatenation of the feature maps extracted by the front-end layers, F HFF represents the fused feature map, f HFF represents a convolution mapping function;

[0036] The output of the residual dense block can be represented as:

[0037] F RDB =F s +F HFF ,

[0038] where F s represents the input feature map of the residual dense block;

[0039] The output of the l-th layer of the residual dense block can be represented as:

[0040]

[0041] where f ds(·) represents a composite function of four consecutive operations: batch normalization operation, linear activation function, convolution layer, and leaky layer.

[0042] Further, the parallel residual multi-scale block is composed of a front-end convolution block, a multi-scale block, and a back-end convolution block; the front-end convolution block is used to increase the number of channels, and the output thereof is input into the multi-scale module to extract multi-scale features; then, the multi-scale features and the output of the front-end convolution block are concatenated and input into a 1*1 convolution layer to realize dimension reduction and information fusion; finally, the output of the front-end convolution block and the output of the 1*1 convolution layer are fused in a residual manner and input into a 3*3 convolution layer to obtain a rain / snow removed image; for the multi-scale block, hierarchical features are repeatedly used in a parallel residual manner to realize multi-scale representation of the features, which can be described as:

[0043]

[0044] wherein represents the output of the i-th path of the multi-scale block, F1 represents the input of the multi-scale block, f Conv3 represents a 3*3 convolution operation.

[0045] Further, the local feature extraction path in step 3 is composed of three layers of convolution layers in series, which can be represented as:

[0046] F l = δW3(δW2(δ(W1F)))

[0047] wherein W1, W2, and W3 represent convolution kernel weights, δ represents an activation function ReLU, and F represents an input feature map;

[0048] The global feature extraction path is composed of an average pooling layer, a maximum pooling layer, and a 1*1 convolution layer; the average pooling layer and the maximum pooling layer extract global common features and unique features of the input image respectively, and then all the features are concatenated and input into the 1*1 convolution layer to reduce the dimension and fuse the information together; the global feature extraction path can be represented as:

[0049] F g = δ(W4[F a ,F m ]) )

[0050] wherein F a and F m represent average pooling features and maximum pooling features respectively, and W4 is a weight;

[0051] Finally, the global features and the local features are multiplied to form a local-global feature map; the final output of the local-global block is represented as:

[0052] F o = Fl *F g .

[0053] Further, the specific operation process of step 4 is as follows:

[0054]

[0055] wherein f down,i , g ad and respectively represent down-sampling operation, adaptive pooling and up-sampling operation.

[0056] Further, the loss function of the local and global U-shaped rain / snow map detection sub-network in step 7 is as follows:

[0057]

[0058] wherein and M respectively represent the predicted rain / snow map and the corresponding true value, N is the total number of training data, and are the weights for balancing the two losses, is the absolute error loss, and SSIM loss represents the structural similarity loss.

[0059] The loss function of the shrinkage exposure time estimation sub-network is as follows:

[0060]

[0061] wherein and k respectively represent the estimated exposure time and the corresponding true value;

[0062] The loss function of the stacked multi-scale residual rain / snow removal sub-network is as follows:

[0063]

[0064] wherein and I c respectively represent the estimated rain / snow image and the corresponding true value, and are the weights for balancing the importance between the absolute error loss and the SSIM loss.

[0065] Further, the method of iterative training in step 7 is as follows: firstly, the local and global U-shaped rain / snow detection sub-network and the shrinkage exposure time estimation sub-network are optimized separately by using the rain / snow image, then the stacked multi-scale residual rain / snow removal sub-network is optimized, and finally the whole rain / snow removal network is optimized to obtain the final optimization result.

[0066] Compared with the prior art, the present application has the following advantages:

[0067] 1. The method of the present application establishes a nonlinear rain and snow exposure imaging model, which considers the exposure factors of imaging, and compared with the linear superposition model and the simplified nonlinear model, it can more accurately describe the imaging characteristics of rain and snow and the rain and snow shielding in the image from translucent to opaque.

[0068] 2. The present application establishes a rain and snow removal network including a local and global U-shaped rain and snow detection sub-network, a shrinkage exposure time estimation sub-network and a stacked multi-scale residual rain and snow removal sub-network, wherein the local and global U-shaped rain and snow detection sub-network uses a global-local attention mechanism to realize the perception of global and local features of rain and snow, the shrinkage exposure time estimation sub-network can sample the global information to a single variable value to obtain the exposure time estimation value through downsampling operation and adaptive pooling layer, and the stacked multi-scale residual rain and snow removal sub-network can further remove rain and snow from the preliminary restored rain and snow image through a parallel residual multi-scale module to extract multi-scale features, so as to gradually remove rain and snow strips of different shielding degrees and shielding scales. The parallel residual multi-scale module adopts a residual fusion mode inside and outside the multi-branch. The multi-branch inside realizes the increase of different branch receptive fields through a parallel residual mode, so that the features are reused and the calculation redundancy caused by independent learning of each branch is reduced. The residual learning mode is also used for feature fusion outside, so that the multi-scale module simultaneously fuses low-level and multi-scale high-level features.

[0069] 3. The method of the present application establishes a global-local attention mechanism, a parallel residual multi-scale module and a double-driven learning mode. Through the embedded rain and snow exposure imaging model, the imaging model is effectively used to guide the image restoration process of the deep network, and the explainability and robustness of the deep network rain and snow removal algorithm are improved. In addition, in order to avoid the difference between large errors and small errors caused by the square penalty in the square loss, so as to cause the restored image to be excessively smooth, the network uses the average absolute error L al as the loss function, and in order to make the restored normal illumination image more consistent with the human visual effect, the SSIM loss is introduced to constrain the learning of network parameters. The weighted balanced rain and snow detection loss function, exposure time loss function and rain and snow removal loss function proposed by the method can better constrain the training of the network, and obtain the final clear rain and snow-free image. Compared with other typical rain and snow removal methods, the method of the present application has good performance and robustness. BRIEF DESCRIPTION OF DRAWINGS

[0070] Figure 1 is the deep network framework of the present application based on the exposure imaging model;

[0071] Figure 2 is the residual dense module used in the present application;

[0072] Figure 3is the global-local attention module used in the present application;

[0073] Figure 4 is the parallel residual multiscale module used in the present application;

[0074] Figure 5 is the comparison of the restoration results of the synthetic rain image by the present application and other typical methods;

[0075] Figure 6 is the comparison of the restoration results of the synthetic snow image by the present application and other typical methods;

[0076] Figure 7 is the comparison of the restoration results of the real rain image by the present application and other typical methods;

[0077] Figure 8 is the comparison of the restoration results of the real snow image by the present application and other typical methods;

[0078] Figure 9 is the comparison of the estimated rain and snow images, exposure time and real values by the present application;

[0079] Figure 10 is the quantitative comparison of the present application and other typical methods on different exposure rain and snow images; DETAILED DESCRIPTION

[0080] In order to make the purpose, technical scheme and advantages of the present application more clear, the present application will be further described in detail below in combination with the drawings. The description here, if involving specific examples, is only used to explain the present application and does not limit the present application.

[0081] Example 1

[0082] As shown in the following figure, the rain and snow image restoration method based on the exposure imaging model and the modular deep network in the present embodiment includes the following steps: Figure 1 Step 1, a nonlinear rain and snow exposure imaging model is established. During the exposure time T, the intensity I rs of the rain and snow image is a linear combination of the intensity R s of the static rain streaks or snowflakes and the intensity R b of the background:

[0083]

[0084]

[0085] where τ represents the time for the rain streaks or snowflakes to fall through a pixel. Since the background is almost static during the exposure time T, R b can be approximated as a constant. Therefore, the above formula is rewritten as:

[0086]

[0087] In the equation, the time τ is related to the speed of rain / snow falling, and its uncertainty directly leads to the diversity of rain / snow imaging results. Therefore, the integral of static rain / snow with respect to time τ can be regarded as a dynamic rain / snow map M, which enriches the diversity of rain / snow imaging transparency and introduces motion blur. TR b is the clean background image I c . In this way, the above equation can be rewritten as:

[0088]

[0089] where represents the average intensity of rain / snow, represents the dynamic rain / snow map M. In order to make each term of the equation have certain physical meaning, the equation is further rewritten as follows:

[0090]

[0091] where * represents element-wise multiplication. Because represents the dynamic rain / snow map M, TR b is the clean background image I c , so the above equation is written as:

[0092]

[0093] Since the average irradiance of rain / snow is determined by the characteristics of rain / snow itself, it can be approximated as a constant K. Let k = KT, and the equation is further written as:

[0094]

[0095] where, I rs represents the rain / snow image intensity, I c is the clean background image, K represents the average irradiance of rain / snow, and in this rain / snow model, the rain / snow image intensity I rs is a nonlinear combination of the clean background image I c , the dynamic rain / snow map M, and the exposure time T. It not only can accurately describe the imaging characteristics of rain / snow, but also can generate rain / snow images that are more visually realistic and diverse.

[0096] Step 2, establish a deep rain / snow removal network, including a local and global U-shaped rain / snow detection sub-network, a shrinkage type exposure time estimation sub-network, and a stacked multi-scale residual rain / snow removal sub-network.

[0097] The local and global U-shaped rain and snow detection sub-network is a residual dense U-shaped auto-encoder, which includes an encoder composed of four-stage residual dense blocks and down-sampling layers, and a decoder composed of a global-local attention module, four-stage residual dense blocks and up-sampling layers. Both low-level information (such as position, transparency and edge) and context information (such as density and direction) are important. In order to capture low-level information and context information at the same time, a U-shaped (U-Net) structure is adopted to capture context information through a contraction path (encoder) and low-level information through a symmetric expansion path (decoder). At the same time, residual dense blocks (RDB) are used as feature extraction units in each stage of the U-Net to fully utilize the features of all convolutional layers. In addition, unlike the traditional U-Net which directly connects the encoder and the decoder, the application introduces a local-global attention mechanism between the encoder and the decoder to extract global and local semantic features of the contraction path at the same time, further improving the feature representation capability of the U-Net. Figure 1 The branches on the top show the detailed structure of the local-global U-shaped rain and snow detection sub-network (LGUN).

[0098] The residual dense block contains hierarchical feature fusion and residual learning, where the hierarchical feature fusion operation can be represented as:

[0099] F HFF = f HFF ([F s ,x1,x2...,x l ]) (7)

[0100] where [F s ,x1,...,x l-1 ] represents the concatenation of the feature maps extracted by the front-end layers, F HFF represents the fused feature map, f HFF represents a convolutional mapping function;

[0101] The output of the residual dense block can be represented as:

[0102] F RDB = F s +F HFF (8)

[0103] where F s represents the input feature map of the residual dense block;

[0104] The output of the l-th layer of the residual dense block can be represented as:

[0105]

[0106] where f ds (·) represents a composite function of four consecutive operations of batch normalization operation, linear activation function, convolutional layer and leaky layer.

[0107] The shrinkage exposure time estimation subnetwork is composed of eight-stage down-sampling operations and adaptive pooling layers;

[0108] The stacked multi-scale residual rain-snow removal subnetwork is two parallel residual multi-scale blocks stacked in a residual manner.

[0109] Step 3, input the rain / snow image into the local and global U-shaped rain-snow detection subnetwork, extract the global and local features of the rain / snow image, and estimate the rain-snow streak image.

[0110] Figure 4 The detailed internal structure of the local-global attention mechanism is shown. The local feature extraction path is composed of three layers of ordinary convolution layers in series, which gradually extract local information. The local feature extraction can be expressed as:

[0111] F l = δW3(δW2(δ(W1F))) (10)

[0112] Where W1, W2 and W3 represent the convolution kernel weights, δ represents the activation function ReLU, and F represents the input feature map; the global feature extraction path is composed of average pooling layers, maximum pooling layers and 1*1 convolution layers. The average and maximum pooling layers extract the global common features and unique features of the input image respectively, and then all the features are concatenated into a 1*1 convolution layer to reduce the dimension and fuse the information together. The global feature extraction path can be expressed as:

[0113] F g = δ(W4[F a ,F m ])) (11)

[0114] Where F a and F m represent the average pooling features and the maximum pooling features respectively, and W4 is the weight. Finally, the global features are multiplied by the local features to form the local-global feature map. The final output of the local-global block is expressed as:

[0115] F o = F l *F g (12)

[0116] Step 4, input the rain / snow image into the shrinkage exposure time estimation subnetwork, since the exposure time is a global quantity, it needs to be estimated by integrating the global features of the rain / snow image, so first the eight-stage down-sampling operation and the adaptive pooling layer are used to sample the global information to a single variable value to obtain the exposure time estimation value. Then, this estimation value is up-sampled to the same size as the input image for loss estimation. This continuous operation can be described as:

[0117]

[0118] where f down,i , g ad and denote down-sampling operation, adaptive pooling and up-sampling operation respectively. The detailed internal structure of the shrinkage exposure time estimation subnetwork is shown in the middle branch, which mainly uses general convolution operation. Figure 1

[0119] Step 5, the rain / snow streaks map and exposure time estimated in step 3 and 4 are input into the preliminary restoration image obtained by the rain / snow exposure imaging model in step 1;

[0120] Step 6, the preliminary restoration image is input into the stacked multi-scale residual rain / snow removal subnetwork, and the preliminary restoration image is further enhanced and restored. Since the occlusion level and size of rain streaks / snowflakes in the rain / snow image are different, a stacked multi-scale residual network is used to gradually remove different degrees of occlusion. Generally, multi-scale features can provide different receptive fields to reconstruct the image. The rain / snow removal subnetwork realizes the extraction of multi-scale features through parallel residual, and embeds the multi-scale module into the residual block to form a hybrid residual multi-scale block. It consists of three parts: a front-end convolution block, a multi-scale block and a back-end convolution block. The front-end convolution block is used to increase the number of channels, and its output is input into the multi-scale module to extract multi-scale features; then, the multi-scale features are concatenated with the output of the front-end convolution block and input into a 1x1 convolution layer to realize dimension reduction and information fusion. Finally, the output of the front-end convolution block and the output of the 1x1 convolution layer are fused in a residual manner to obtain a rain / snow removal image through a 3x3 convolution layer. For the multi-scale block, it realizes multi-scale representation of features by repeatedly using hierarchical features in a parallel residual manner, which can be described as:

[0121]

[0122] where denotes the output of the i-th path of the multi-scale block, F1denotes the input of the multi-scale block, f Conv3 denotes a 3*3 convolution operation. The multi-scale module uses a residual fusion method inside and outside the multi-branch. The multi-branch inside realizes the increase of the receptive field of different branches through a parallel residual manner, so that the features are reused and the computational redundancy caused by independent learning of each branch is reduced. The outside also uses a residual learning method to fuse features, so that the multi-scale module simultaneously fuses low-level and multi-scale high-level features.

[0123] ​Step 7, setting the loss functions of the local and global U-shaped rain and snow map detection sub-network, the shrinkage type exposure time estimation sub-network and the stacked multi-scale residual rain and snow removal sub-network, and iteratively training each sub-network by learning the network model parameters on the dataset under the constraint of the loss function until the network converges.

[0124] In order to avoid the square penalty in the square loss to amplify the difference between large errors and small errors, resulting in the over-smoothing of the restored image. The network uses the mean absolute error L al As a loss function, in order to make the restored normal illumination image more consistent with human visual effects, the SSIM loss is introduced to constrain the learning of network parameters.

[0125] The loss function of the local and global U-shaped rain and snow map detection sub-network is:

[0126]

[0127] Wherein and M represent the predicted rain / snow map and the corresponding true value respectively. N is the total number of training data, and is the weight of balancing the two losses, is the absolute error loss, indicates the structural similarity loss;

[0128] The loss function of the shrinkage type exposure time estimation sub-network is:

[0129]

[0130] Wherein and k represent the estimated exposure time and the corresponding true value respectively.

[0131] The loss function of the stacked multi-scale residual rain and snow removal sub-network can be written as:

[0132]

[0133] Wherein and I c represent the estimated rain and snow image and the corresponding true value respectively, and is the weight of balancing the importance between the absolute error loss and the SSIM loss.

[0134] The total loss function of the deep network model can be represented as:

[0135] L all = L m + L k + L c (18)

[0136] The method for iterative training is:

[0137] First, the local and global U-shaped rain and snow detection sub-networks and the shrinkage exposure time estimation sub-network are optimized separately using rain and snow images, then the stacked multi-scale residual rain and snow removal sub-network is optimized, and finally the whole rain and snow removal network is optimized to obtain the final optimization result. The training process is as follows:

[0138]

[0139]

[0140] To verify the effectiveness of the algorithm, we compared the algorithm with mainstream algorithms on synthetic rain and snow datasets and real rain and snow datasets. The public synthetic datasets used include Rain1200, Rain100H, Rain100L and Snow100K. Since it is difficult to capture real rain and snow image pairs of different exposure times, a multi-exposure rain and snow database RSQD is synthesized. First, background images under 15 exposure times (from 1 / 10s to 1 / 250s) are shot; then rain / snow streak images are generated using various features (transparency, scale and density) of rain streaks and snowflakes, and a nonlinear exposure imaging rain and snow model is used to synthesize the rain and snow image dataset. The real rain dataset used is Real1000.

[0141] To quantitatively evaluate and compare the performance of each algorithm on the labeled dataset, for the evaluation of rain / snow streak images and rain / snow-free images, the signal-to-noise ratio PSNR and the structural similarity SSIM are used as evaluation indicators. For exposure time, the average absolute error (AAE) and the average relative error (ARE) between the estimated value and the true value are used to measure the accuracy of the estimated exposure time. In the comparison of quantitative experimental results, the method of the present application is compared with five mainstream methods on seven labeled databases. The five methods include: (a) DDN, (b) DID-MDN, (c) SPANet, (d) MSPFN and (e) Syn2Real. The comparison results are shown in Tables 1 and 2. As can be seen from the tables, the method of the present application is superior to other methods in terms of PSNR and SSIM values. The PSNR value of the method of the present application exceeds the second best method by 2.00dB, 3.22dB, 1.43dB, 0.12dB, 2.94dB, 3.38dB and 3.98dB respectively on the seven datasets. The SSIM value is also higher than the second best method, and the SSIM value on the seven datasets is higher than the second best method by 0.7%, 2.5%, 5.2%, 0.4%, 1.6%, 3.0% and 2.1% respectively. These quantitative results all verify the effectiveness of the algorithm of the present application.

[0142] Table 1 shows the mean PSNR / SSIM results of the state-of-the-art method on rain / snow datasets.

[0143]

[0144] Table 2 shows the mean PSNR / SSIM values ​​for evaluating state-of-the-art methods on rain / snow datasets.

[0145]

[0146] Furthermore, this invention also provides a visual, qualitative comparative evaluation of the performance of various rain / snow removal methods. An intuitive visual comparison of rain removal results on a synthetic dataset is shown below. Figure 5 As shown in the figure, the first and second rows reveal that other methods generally do not handle opaque rain streaks well and tend to produce rain streak artifacts in rainless images. For semi-transparent rain streaks (e.g., Figure 5 (The third and fourth rows) While these methods achieve better overall visual results, they also tend to leave slight rain streaks in rainless images. Compared to these methods, the method of the present invention leaves fewer rain streaks and has higher PSNR and SSIM. Snow removal results on synthetic snow images are shown below. Figure 6 As shown in the figure, DDN and DID-MDN typically only remove a small number of semi-transparent snowflakes, with almost no effect on large, opaque snowflakes. SPANet and MSPFN methods tend to leave some large snowflakes in snow-free images. The Syn2Real algorithm can remove most snowflakes, achieving considerable visual results, but it also leaves some large snowflakes and is lower than the method of this invention in terms of quantitative metrics PSNR and SSIM. In contrast, the method of this invention can more thoroughly remove snowflakes of various features in the image while preserving background details to some extent.

[0147] Deraining results from the Real1000 real rain image database are as follows: Figure 7 As shown in the figure, DDN and SPANet can only handle small rain streaks and cannot process long rain streaks. DID-MDN and MSPFN can generally only handle a few semi-transparent rain streaks and have almost no effect on opaque rain streaks. The Syn2Real method achieves better visual results, but still leaves some long and opaque rain streaks. Compared with these methods, the method of this invention can remove real rain streaks better. The snow removal result on the real snow image is shown below. Figure 8As shown in the figure, DDN, DID-MDN, SPANet and MSPFN can only remove a small amount of semi-transparent snowflakes, and there are still a large number of residual snowflakes in both sparse and dense snowflake images. Syn2Real can remove most of the sparse snowflakes, but some snowflakes with unobvious features are still left. In contrast, the method of the present application can process both sparse and dense snowflakes. The better rain / snow removal visual effects on synthetic and real rain / snow images verify the effectiveness of the method of the present application.

[0148] In addition, the rain / snow image and exposure time predicted by the method of the present application are also quantitatively and qualitatively evaluated. The quantitative results on the database RQD and SQD are listed in Table 3. As can be seen from the table, the PSNR values on the rain / snow image dataset are all greater than 30, and the SSIM values are all greater than 0.9. This shows that the detected rain / snow image of the present application is close to the true value. For the exposure time, the average absolute error (AAE) is less than 0.01 s, and the average relative error (ARE) is less than 0.2. The visual effect comparison between the estimated value and the true value is as shown in Figure 9 As can be seen from the figure, the rain / snow image and exposure time estimated by the method of the present application are very close to their true values. For the exposure time, the present application uses a color map to represent the length of time, with red representing the longest exposure time value (1 / 10 s) and blue representing the shortest exposure time value (1 / 250 s). The quantitative and qualitative results show that the method of the present application can not only effectively remove the rain / snow in the image, but also accurately estimate the rain / snow image and exposure time.

[0149] Table 3 Quantitative evaluation of rain / snow image and exposure time on RQD and SQD

[0150]

[0151] In order to analyze the rain / snow removal effect of the image under different exposure times, further quantitative comparison is made with some mainstream rain / snow removal methods under ten exposure times (1 / 15 s-1 / 125 s), and the results are as shown in Figure 10 As can be seen from the figure, the following conclusions can be drawn. First, the longer the exposure time lasts, the lower the PSNR and SSIM are, and the more serious the occlusion caused by rain / snow in the rain / snow image is (see the black line). Second, as the exposure time shortens, the PSNR and SSIM values on the rain / snow removal image of all methods except DID-MDN increase, that is, semi-transparent rain / snow is easier to remove than opaque rain / snow. Third, the method of the present application is superior to other methods at each exposure time, which shows that estimating the exposure time is necessary for effective rain / snow removal.

Claims

1. An image rain and snow removal method based on an exposure imaging model and a modular deep network, characterized in that, The method comprises the following steps: Step 1, establishing a nonlinear rain / snow exposure imaging model as follows: ; wherein represents the rain / snow image intensity, is the clean background image, represents the average irradiance of rain / snow, the rain / snow image intensity is the clean background image , the dynamic rain / snow image and the exposure time in a non-linear combination; Step 2, establishing a deep rain / snow removal network, which comprises a local and global U-shaped rain / snow detection sub-network, a shrinkage exposure time estimation sub-network and a stacked multi-scale residual rain / snow removal sub-network; Step 3, inputting the rain / snow image into the local and global U-shaped rain / snow detection sub-network to extract the global and local features of the rain / snow image and estimate the rain / snow stripe image; Step 4, inputting the rain / snow image into the shrinkage exposure time estimation sub-network, sampling the global information to a single variable value to obtain the exposure time estimation value, and then up-sampling the estimation value to the same size as the input image for loss estimation; Step 5, obtaining the preliminary restored rain / snow image through the nonlinear rain / snow exposure imaging model of step 1 by using the rain / snow stripe image and the exposure time obtained in steps 3 and 4; Step 6, inputting the preliminary restored rain / snow image obtained in step 4 into the stacked multi-scale residual rain / snow removal sub-network to further remove rain / snow occlusion of different degrees and obtain the final refined rain / snow-free image; Step 7, setting the loss functions of the local and global U-shaped rain / snow detection sub-network, the shrinkage exposure time estimation sub-network and the stacked multi-scale residual rain / snow removal sub-network, and training the network model parameters through step-by-step iterative training of each sub-network on the rain / snow dataset under the constraint of the loss function until the network converges; Step 8, quantitatively and qualitatively verifying the effectiveness of the algorithm in removing rain / snow on the synthetic and real rain / snow images and comparing with typical algorithms; The local and global U-shaped rain / snow detection sub-network in step 2 is a residual dense U-shaped auto-encoder, which comprises an encoder composed of four-stage residual dense blocks and down-sampling layers and a decoder composed of a global and local attention module, four-stage residual dense blocks and up-sampling layers; the shrinkage exposure time estimation sub-network in step 2 is composed of eight-stage down-sampling operations and adaptive pooling layers; and the stacked multi-scale residual rain / snow removal sub-network in step 2 is two parallel residual multi-scale blocks stacked in a residual manner; The loss function of the local and global U-shaped rain / snow detection sub-network in step 7 is: ; wherein and denote the predicted rain / snow map and the corresponding ground truth, respectively, is the total number of training data, and are the weights balancing the two losses, is the absolute error loss, denotes the structural similarity loss; The loss function of the shrinkage exposure time estimation sub-network is: ; wherein and denote the estimated exposure time and the corresponding true value, respectively; The loss function of the stacked multi-scale residual rain / snow removal sub-network is: ; wherein and respectively denote the estimated rain / snow image and the corresponding ground truth, and is a weight balancing the importance between the absolute error loss and the SSIM loss.

2. The method of claim 1, wherein: The specific process of the nonlinear rain and snow exposure imaging model in step 1 is that during the exposure time , the intensity of the rain and snow image is a linear combination of the intensity of the static rain stripe or snowflake and the intensity of the background: ; where represents the time it takes for a rain streak or snowflake to fall through a pixel, since the background is almost stationary during the exposure time can be approximated as a constant, so rewrite the above equation as:​ ; In the equation, time and the speed of rain / snow fall, whose uncertainty directly leads to the diversity of rain / snow imaging results; therefore, the integral of static rain / snow over time can be regarded as a dynamic rain / snow map , which enriches the diversity of rain / snow imaging transparency and introduces motion blur; is a clean background image ; thus the above equation can be rewritten as: ; wherein represents the average intensity of rain / snow, represents a dynamic rain / snow map ; in order to give each term of the equation a certain physical meaning, the equation is further rewritten in the following form: ; where represents element-wise multiplication, because represents the dynamic rain / snow map , is the clean background image , so the above equation is written as: ; Since the average irradiance of rain / snow is determined by the properties of the rain / snow itself, it can be approximated as a constant ; let , the equation further writes as: ; wherein represents the rain / snow image intensity, is the clean background image, represents the average irradiance of rain / snow, in this rain / snow model, the rain / snow image intensity is the clean background image , the dynamic rain / snow image and the exposure time in a non-linear combination.

3. The method of claim 1, wherein: The encoder is used to capture context information such as density and direction; the decoder is used to capture low-level information such as position, transparency and edge; the global and local attention module is used to extract the global and local semantic features of the encoder to further improve the feature representation capability of the network; the residual dense block is used as a feature extraction unit to fully utilize the features of all convolutional layers and contains hierarchical feature fusion and residual learning, wherein the hierarchical feature fusion operation can be represented as: ; wherein, denotes concatenation of the feature maps extracted by the front-end layers, denotes the fused feature maps, denotes a convolution mapping function; The output of the residual dense block can be represented as: ; wherein, denotes the input feature map of the residual dense block; Residual dense block The output of the layer can be represented as: ; wherein represents a composite function of four successive operations of batch normalization operation, linear activation function, convolution layer and leaky layer.

4. The method of claim 1, wherein: The parallel residual multiscale block realizes the multiscale representation of features in a parallel residual manner, and is composed of a front-end convolution block, a multiscale block and a back-end convolution block; the front-end convolution block is used to increase the number of channels, and the output thereof is sent to the multiscale block to extract multiscale features; then, the multiscale features and the output of the front-end convolution block are connected together and input to a 1x1 convolution layer to realize dimension reduction and information fusion; finally, the output of the front-end convolution block and the output of the 1x1 convolution layer are fused together in a residual manner to obtain a rain / snow removed image through a convolution layer; the multiscale representation can be described as: ​ ; wherein denotes the output of the i-th path of the multi-scale block, denotes the input of the multi-scale block, denotes a 3*3 convolution operation.

5. The method of claim 1, wherein: The path for extracting local features in step 3 is composed of three layers of convolutional layers connected in series, and the local feature extraction can be represented as: ; wherein 、 and denote convolution kernel weights, denotes an activation function ReLU, and F denotes an input feature map; The global feature extraction path is composed of an average pooling layer, a maximum pooling layer and a 1*1 convolution layer; the average pooling layer and the maximum pooling layer respectively extract common features and unique features of the global input image, and then all the features are concatenated together to input the 1*1 convolution layer to reduce the dimension and fuse the information together; the global feature extraction path can be represented as: ; wherein and respectively represent the average-pooled feature and the max-pooled feature, is a weight; Finally, the global feature is multiplied with the local feature to form a local-global feature map; the final output of the local-global block is represented as: 。 6. The method of claim 1, wherein, The specific operation process of the step 4 is as follows: ; wherein , and denote down-sampling operation, adaptive pooling and up-sampling operation, respectively.

7. The method of claim 1, wherein, The iterative training method in the step 7 is as follows: firstly, the rain and snow image is used to separately optimize the local and global U-shaped rain and snow detection sub-network and the shrinkage type exposure time estimation sub-network, then the stacked multi-scale residual rain and snow removal sub-network is optimized, finally the whole deep rain and snow removal network is optimized to obtain the final optimization result.