Image raindrop removal method based on reverse deraining process

By adopting deep learning methods in image raindrop removal, combining the forward raindrop removal network with multiple attention mechanisms and U-shaped encoder decoder structure, and combining the reverse raindrop removal network, the problem of insufficient robustness of raindrop removal in the existing technology is solved, and better image quality recovery is achieved.

CN116402720BActive Publication Date: 2025-05-13FUZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310396874.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-14
Publication Date
2025-05-13
Estimated Expiration
2043-04-14

AI Technical Summary

Technical Problem

The prior art is not robust enough when removing raindrops in outdoor images, making it difficult to deal with raindrops of different degrees of blur, brightness and size, resulting in a degradation of image quality.

Method used

Using a deep learning-based method, a forward rain removal network with multiple attention mechanisms and U-shaped encoder decoder structure is designed, and combined with the reverse rain removal network, the forward rain network is guided to improve the feature representation learning ability through supervision of losses, thereby removing raindrops.

Benefits of technology

Through this method, the effect of image raindrop removal is significantly improved, the ability to extract raindrop features is enhanced, and the image quality is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116402720B_ABST
    Figure CN116402720B_ABST
Patent Text Reader

Abstract

The present invention relates to an image raindrop removal method based on the guidance of the reverse rain removal process. The method comprises the following steps: step S1, performing data preprocessing, including data pairing, data random cropping, and data enhancement processing, to obtain a training data set; step S2, designing a multiple attention module, using pixel attention, channel attention, and spatial attention to extract image features; step S3, designing a U-shaped encoder-decoder network based on the multiple attention module, which is used to extract features of the image at different scales; step S4, designing an image feature fusion module, which fuses the image features and the original image to obtain a fused restored image; step S5, designing a forward raindrop removal network and a reverse raindrop removal network, and using the reverse raindrop removal network to guide the training of the forward raindrop removal network; step S6, inputting the attached raindrop image into the trained forward raindrop removal network, and outputting a clean image with raindrops removed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the fields of image processing and computer vision, and in particular to an image raindrop removal method guided by a reverse rain removal process. Background Art

[0002] With the continuous development of unmanned driving, smart cities, intelligent transportation and other fields, how to use various image acquisition devices to effectively collect outdoor images and improve the quality of the acquired outdoor images are problems that need to be solved in these fields. High-quality outdoor images can facilitate the subsequent use of image recognition systems for automatic navigation, vehicle tracking, pedestrian detection and other tasks. However, when collecting outdoor images, it is particularly susceptible to weather changes. Rain is a common weather phenomenon. When raindrops adhere to glass windows or camera lenses, they hinder the visibility of the background scene, reduce the quality of the image, and cause serious image degradation, such as reduced contrast, reduced saturation, and blurred image background. This causes the loss of detailed information on the image, which greatly reduces the use value of the image. Therefore, removing raindrops from the image to restore the clean background is of great research significance for improving the stability and applicability of outdoor computer vision systems.

[0003] At present, image raindrop removal methods are generally divided into two categories, namely traditional machine learning-based methods and deep learning-based methods. Traditional machine learning-based methods usually use a filter to decompose it into high-frequency and low-frequency parts, and then distinguish the rain component and non-rain component in the high-frequency part through dictionary learning. Most traditional machine learning-based methods rely on artificially pre-set parameters to extract image features and optimize performance. They can only extract shallow information in the image and lack deep reasoning about the content, making them not robust to input changes, such as raindrops with different blur levels, brightness and sizes. The deep learning-based method is a data-driven method that uses a large amount of data to train a neural network. The powerful feature learning representation ability of the neural network can better extract image features and achieve better raindrop removal effects.

[0004] Based on the deep learning method, this paper proposes an image raindrop removal method guided by the reverse rain removal process. This method uses a multiple attention mechanism and a U-shaped encoder-decoder structure to build a forward rain removal network to extract the raindrop features of the image, and designs a reverse rain removal network to receive the features of the forward rain removal network for reconstructing the rain map. By applying a supervised loss, the feature representation learning ability of the forward rain removal network is promoted, thereby obtaining a better rain removal effect. Summary of the invention

[0005] The object of the present invention is to provide an image raindrop removal method based on the guidance of a reverse rain removal process, which can remove raindrops attached to an image.

[0006] To achieve the above object, the technical solution of the present invention is: a method for removing raindrops from an image based on a reverse rain removal process, comprising the following steps:

[0007] Step S1, perform data preprocessing, including data pairing, data random cropping, and data enhancement processing, to obtain a training data set;

[0008] Step S2: design a multiple attention module and use pixel attention, channel attention and spatial attention to extract image features;

[0009] Step S3, designing a U-shaped encoder-decoder network based on multiple attention modules to extract features at different scales of the image;

[0010] Step S4, designing an image feature fusion module, which fuses the image features and the original image to obtain a fused restored image;

[0011] Step S5, designing a forward raindrop removal network and a reverse raindrop removal network, and using the reverse raindrop removal network to guide the training of the forward raindrop removal network;

[0012] Step S6: input the attached raindrop image into a trained forward raindrop removal network guided by a reverse rain removal process, and output a clean image with the raindrops removed.

[0013] In one embodiment of the present invention, step S1 is specifically implemented as follows:

[0014] Step S11, pairing the attached raindrop image and the corresponding clean image;

[0015] Step S12, randomly cropping each attached raindrop image of size H×W×3 into an image of size p×p×3, and using the same random cropping method for the corresponding clean image, where H and W are the height and width of the attached raindrop image and the clean image, and p is the height and width of the cropped image;

[0016] Step S13, randomly adopt one of the following eight enhancement methods to perform data enhancement on the training paired images: keep the original image, vertically flip, rotate 90 degrees, rotate 90 degrees and then vertically flip, rotate 180 degrees, rotate 180 degrees and then vertically flip, rotate 270 degrees, and rotate 270 degrees and then vertically flip; after data enhancement, the training data set is obtained, and the attached raindrop image is recorded as I rain , and its paired clean image is denoted as I clean .

[0017] In one embodiment of the present invention, step S2 is specifically implemented as follows:

[0018] Step S21: Design pixel attention module and input image features Where H′ and W′ represent the height and width of the feature, respectively. After 3×3 convolution, PReLU activation function and 3×3 convolution, the features are obtained. After 1×1 convolution, PReLU activation function, 1×1 convolution and Sigmoid activation function, the pixel attention weight Att is obtained. p , With Att p Multiply and Add up to get pixel attention weighted features The specific formula is as follows:

[0019]

[0020]

[0021]

[0022] Among them, Conv 3×3 Represents a 3×3 convolutional layer, Conv 1×1 represents a 1×1 convolutional layer, PReLU(·) and Sigmoid(·) represent activation functions;

[0023] Step S22: Design a channel attention module and input image features After 3×3 convolution, PReLU activation function and 3×3 convolution, the features are obtained. The channel attention weight Att is obtained after channel average pooling, 1×1 convolution, PReLU activation function, 1×1 convolution and Sigmoid activation function. c , With Att c Multiply and Add up to get pixel attention weighted features The specific formula is as follows:

[0024]

[0025]

[0026]

[0027] Among them, Conv 3×3Represents a 3×3 convolutional layer, Conv 1×1 represents a 1×1 convolutional layer, PReLU(·) and Sigmoid(·) represent activation functions, and CAP(·) represents channel average pooling;

[0028] Step S23: Design a spatial attention module and input image features After 3×3 convolution, PReLU activation function and 3×3 convolution, the features are obtained. After spatial average pooling and spatial maximum pooling, feature F is obtained. s1 and F s2 , Splice along the channel F s1 and F s2 , and then obtain the spatial attention weight Att through 1×1 convolution and Sigmoid activation function s , With Att s Multiply and Add up to get pixel attention weighted features The specific formula is as follows:

[0029]

[0030]

[0031]

[0032] Att s =Sigmoid(Conv 1×1 (Concat(F s1 , F s2 )))

[0033]

[0034] Among them, Conv 3×3 Represents a 3×3 convolutional layer, Conv 1×1 represents a 1×1 convolutional layer, PReLU(·) and Sigmoid(·) represent activation functions, SAP(·) represents spatial average pooling, SMP(·) represents spatial maximum pooling, and Concat represents concatenation along the channel;

[0035] Step S24: Design a multi-attention module and input image feature F in , F in After the pixel attention module of step S21, the channel attention module of step S22 and the spatial attention module of step S23, the output feature F is obtained. out , The specific formula is as follows:

[0036] F out =SAB(CAB(PAB(F in )))

[0037] Among them, PAB is the pixel attention module, CAB is the channel attention module, and SAB is the spatial attention module.

[0038] In one embodiment of the present invention, step S3 is specifically implemented as follows:

[0039] Step S31: Design a shallow feature extraction module, with an image I as input. I undergoes 3×3 convolution, PRe LU activation function and 3×3 convolution to obtain the shallow feature F shallow , The specific formula is as follows:

[0040] F shallow =Conv 3×3 (PReLU(Conv 3×3 (I rain )))

[0041] Among them, Conv 3×3 represents a 3×3 convolutional layer, and PReLU(·) represents the activation function;

[0042] Step S32: Design an encoder in a U-shaped encoder-decoder network based on multiple attention modules, including three multiple attention modules; the input of the encoder is the shallow feature F in step S31 shallow , F shallow After multiple attention modules, the encoding feature E1 is obtained. E1 is downsampled and multi-attention modules are used to obtain the encoding feature E2. E2 is downsampled and multi-attention module to obtain the encoding feature E3. The specific formula is as follows:

[0043] E1=MA(F shallow )

[0044] E2=MA(DownSample(E1))

[0045] E3=MA(DownSample(E2))

[0046] Among them, MA represents multiple attention modules, DownSample represents image downsampling;

[0047] Step S33, design a decoder in a U-shaped encoder-decoder network based on multiple attention modules, including three multiple attention modules; the input of the decoder is the encoded feature E3 in step S32, and E3 is subjected to the multiple attention modules to obtain the decoded feature D3, D3 is upsampled and added to the encoded feature E2, and then the decoded feature D2 is obtained through the multiple attention modules. D2 is upsampled and added to the encoded feature E1, and then the decoded feature D1 is obtained through the multiple attention modules. The specific formula is as follows:

[0048] D3=MA(E3)

[0049] D2=MA(UpSample(D3)+E2)

[0050] D1=MA(UpSample(D2)+E1)

[0051] Among them, MA stands for multiple attention modules and UpSample stands for image upsampling.

[0052] In one embodiment of the present invention, step S4 is specifically implemented as follows:

[0053] Design an image feature fusion module, the input of which is the decoded feature D and image I. The decoded feature D is convolved by 3×3 and then added to the image I to obtain the intermediate result I mid , I mid After 3×3 convolution and Sigmoid activation function, we get the attention map AttMap. D is convolved by 1×1 and then multiplied by AttMap and then added to D to obtain the intermediate feature F mid , F mid After 3×3 convolution and intermediate result I mid Add to get the output image I out , The specific formula is as follows:

[0054] I mid =I+Conv 3×3 (D)

[0055] AttMap=Sigmoid(Conv 3×3 (I mid ))

[0056] F mid =D+AttMap×Conv 3×3 (D)

[0057] I out =I mid +Conv 3×3 (F mid )

[0058] Among them, Conv 3×3 represents a 3×3 convolutional layer, and Sigmoid(·) represents the activation function.

[0059] In one embodiment of the present invention, step S5 is specifically implemented as follows:

[0060] Step S51, design a forward raindrop removal network, which consists of a shallow feature extraction module, an encoder and decoder in a U-shaped encoder-decoder network based on a multiple attention module, and an image feature fusion module; the input of the network is the attached raindrop image I rain , I rain After the shallow feature extraction module in step S31, The encoder and decoder in the U-shaped encoder-decoder network based on the multiple attention module in step S32 and step S33 obtain the encoded features and decode features and I rain After the image feature fusion module, the derained image I is obtained. derain , the specific formula is as follows:

[0061]

[0062]

[0063]

[0064] Among them, ShallowBlock represents the shallow feature extraction module, Encoder represents the encoder, Decoder represents the decoder, and FusionBlock represents the image feature fusion module;

[0065] Step S52, design a reverse raindrop removal network, which consists of a shallow feature extraction module, an encoder and a decoder in a U-shaped encoder-decoder network based on a multiple attention module, and an image feature fusion module. The encoder and decoder of the network receive the encoded features in step S51. and decode features As the input feature of this part of the network; and the attached raindrop image I in step S51 rain The corresponding clean image I clean As the input image, I clean After the shallow feature extraction module in step S31, Plus Then the encoding features are obtained through multiple attention modules After downsampling The encoding features are obtained through multiple attention modules After downsampling The encoding features are obtained through multiple attention modules Plus Decoding features are obtained through multiple attention modules After upsampling and encoding features Add, and then get the decoding features through multiple attention modules After upsampling and encoding features Add, and then get the decoding features through multiple attention modules and I clean After the image feature fusion module, the restored attached raindrop image I is obtained re-rain , the specific formula is as follows:

[0066]

[0067]

[0068]

[0069]

[0070]

[0071]

[0072]

[0073]

[0074] Among them, ShallowBlock represents the shallow feature extraction module, MA represents the multiple attention module, DownSample represents the image downsampling, UpSample represents the image upsampling, and FusionBlock represents the image feature fusion module;

[0075] Step S53, design a loss function, and use the reverse deraining process to guide the forward deraining network training; randomly extract an attached raindrop image I from the training data set in step S1 rain And its corresponding clean image I clean , I rain After the forward deraining network, we get deraining Iderain, I clean After the reverse deraining network, the rained image I is obtained re-rain , calculate the following loss, and use the Adam optimizer to update the network parameters until the loss converges; the specific loss function formula is as follows:

[0076] L=L forward +L reverse

[0077]

[0078]

[0079] Where L forward is the forward deraining network loss, L reverse is the inverse deraining network loss, PSNR is the peak signal-to-noise ratio metric, and SSIM is the structural similarity metric.

[0080] In one embodiment of the present invention, the specific implementation of step S6 is: inputting the image of attached raindrops to be tested into a trained forward rain removal network guided by a reverse rain removal process, and outputting a clean image with raindrops removed.

[0081] Compared with the prior art, the present invention has the following beneficial effects: The present invention proposes an image raindrop removal method based on the guidance of the directional rain removal process. The present invention designs a multiple attention mechanism and a U-shaped encoder-decoder structure to build a forward rain removal network for extracting raindrop features of the image, and designs a reverse rain removal network to receive the features of the forward rain removal network for reconstructing the rain map. By applying a supervised loss, the feature representation learning ability of the forward rain removal network is promoted to obtain a better rain removal effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0082] Figure 1 The present invention is an implementation flow chart of the method according to the embodiment of the present invention.

[0083] Figure 2 is a structural diagram of a multiple attention module in an embodiment of the present invention.

[0084] Figure 3 4 is a structural diagram of a forward rain removal network in an embodiment of the present invention.

[0085] Figure 4 It is a structural diagram of the image feature fusion module in an embodiment of the present invention.

[0086] Figure 5 It is a structural diagram of a forward rain removal network guided by a reverse rain removal network in an embodiment of the present invention. DETAILED DESCRIPTION

[0087] The technical solution of the present invention is described in detail below in conjunction with the accompanying drawings.

[0088] The object of the present invention is to provide an image raindrop removal method based on the reverse rain removal process. Figure 1 As shown, the following steps are included:

[0089] Step S1, perform data preprocessing, including data pairing, data random cropping, and data enhancement processing, to obtain a training data set;

[0090] Step S2: design a multiple attention module and use pixel attention, channel attention and spatial attention to extract image features;

[0091] Step S3, designing a U-shaped encoder-decoder network based on multiple attention modules to extract features at different scales of the image;

[0092] Step S4, designing an image feature fusion module, which fuses the image features and the original image to obtain a fused restored image;

[0093] Step S5, designing a forward raindrop removal network and a reverse raindrop removal network, and using the reverse raindrop removal network to guide the training of the forward raindrop removal network;

[0094] Step S6: input the attached raindrop image into the trained forward raindrop removal network, and output a clean image with the raindrops removed.

[0095] Furthermore, step S1 includes the following steps:

[0096] Step S11, pairing the attached raindrop image and the corresponding clean image;

[0097] Step S12, randomly cropping each attached raindrop image of size H×W×3 into an image of size p×p×3, and using the same random cropping method for the corresponding clean image, where H and W are the height and width of the attached raindrop image and the clean image, and p is the height and width of the cropped image;

[0098] Step S13, randomly use one of the following eight enhancement methods to perform data enhancement on the training paired images: keep the original image, vertically flip, rotate 90 degrees, rotate 90 degrees and then vertically flip, rotate 180 degrees, rotate 180 degrees and then vertically flip, rotate 270 degrees, and rotate 270 degrees and then vertically flip. After data enhancement, the training data set is obtained, and the attached raindrop image is recorded as I rain , and its paired clean image is denoted as I clean .

[0099] Furthermore, if Figure 2 As shown, step S2 includes the following steps:

[0100] Step S21: Design pixel attention module and input image features Where H′ and W′ represent the height and width of the feature, respectively. After 3×3 convolution, PReLU activation function and 3×3 convolution, the features are obtained. After 1×1 convolution, PReLU activation function, 1×1 convolution and Sigmoid activation function, the pixel attention weight Att is obtained. p , With Att p Multiply and Add up to get pixel attention weighted features The specific formula is as follows:

[0101]

[0102]

[0103]

[0104] Among them, Conv 3×3 Represents a 3×3 convolutional layer, Conv 1×1 represents a 1×1 convolutional layer, and PReLU(·) and Sigmoid(·) represent activation functions.

[0105] Step S22: Design a channel attention module and input image features After 3×3 convolution, PReLU activation function and 3×3 convolution, the features are obtained. The channel attention weight Att is obtained after channel average pooling, 1×1 convolution, PReLU activation function, 1×1 convolution and Sigmoid activation function. c , With Att c Multiply and Add up to get pixel attention weighted features The specific formula is as follows:

[0106]

[0107]

[0108]

[0109] Among them, Conv 3×3 Represents a 3×3 convolutional layer, Conv 1×1represents a 1×1 convolutional layer, PReLU(·) and Sigmoid(·) represent activation functions, and CAP(·) represents channel average pooling.

[0110] Step S23: Design a spatial attention module and input image features After 3×3 convolution, PReLU activation function and 3×3 convolution, the features are obtained. After spatial average pooling and spatial maximum pooling, feature F is obtained. s1 and F s2 , Splice along the channel F s1 and F s2 , and then obtain the spatial attention weight Att through 1×1 convolution and Sigmoid activation function s , With Att s Multiply and Add up to get pixel attention weighted features The specific formula is as follows:

[0111]

[0112]

[0113]

[0114] Att s =Sigmoid(Conv 1×1 (Concat(F s1 , F s2 )))

[0115]

[0116] Among them, Conv 3×3 Represents a 3×3 convolutional layer, Conv 1×1 represents a 1×1 convolutional layer, PReLU(·) and Sigmoid(·) represent activation functions, SAP(·) represents spatial average pooling, SMP(·) represents spatial maximum pooling, and Concat represents concatenation along the channel.

[0117] Step S24: Design a multi-attention module and input image feature F in , F in After the pixel attention module in step S21, the channel attention module in step S22, and the spatial attention module in step S23, the output feature F is obtained. out , The specific formula is as follows:

[0118] F out =SAB(CAB(PAB(F in )))

[0119] Among them, PAB is the pixel attention module, CAB is the channel attention module, and SAB is the spatial attention module.

[0120] Furthermore, if Figure 3 As shown, step S3 includes the following steps:

[0121] Step S31: Design a shallow feature extraction module, with an image I as input. I undergoes 3×3 convolution, PRe LU activation function and 3×3 convolution to obtain the shallow feature F shallow , The specific formula is as follows:

[0122] F shallow =Conv 3×3 (PReLU(Conv 3×3 (I rain )))

[0123] Among them, Conv 3×3 represents a 3×3 convolutional layer, and PReLU(·) represents the activation function.

[0124] Step S32: Design an encoder in a U-shaped encoder-decoder network based on multiple attention modules, including three multiple attention modules in step S24. The input of the encoder is the shallow feature F in step S31. shallow , F shallow The multiple attention modules obtain the encoding feature E1, E1 is downsampled and multi-attention modules are used to obtain the encoding feature E2. E2 is downsampled and multi-attention module to obtain the encoding feature E3. The specific formula is as follows:

[0125] E1=MA(F shallow )

[0126] E2=MA(DownSample(E1))

[0127] E3=MA(DownSample(E2))

[0128] Among them, MA represents multiple attention modules and DownSample represents image downsampling.

[0129] Step S33, design a decoder in a U-shaped encoder-decoder network based on multiple attention modules, including three multiple attention modules in step S24. The input of the decoder is the encoded feature E3 in step S32, and E3 is decoded by the multiple attention modules to obtain the decoded feature D3. D3 is upsampled and added to the encoded feature E2, and then the decoded feature D2 is obtained through the multiple attention modules. D2 is upsampled and added to the encoded feature E1, and then the decoded feature D1 is obtained through the multiple attention modules. The specific formula is as follows:

[0130] D3=MA(E3)

[0131] D2=MA(UpSample(D3)+E2)

[0132] D1=MA(UpSample(D2)+E1)

[0133] Among them, MA stands for multiple attention modules and UpSample stands for image upsampling.

[0134] Furthermore, if Figure 4 As shown, step S4 includes the following steps:

[0135] Step S41: Design an image feature fusion module, the input of which is the decoded feature D and image I. Feature D is convolved by 3×3 and then added to image I to obtain the intermediate result I mid , I mid After 3×3 convolution and Sigmoid activation function, we get the attention map AttMap. D is convolved by 1×1 and then multiplied by AttMap and added to D to obtain the intermediate feature F mid , F mid After 3×3 convolution and intermediate result I mid Add to get the output image I out , The specific formula is as follows:

[0136] I mid =I+Conv 3×3 (D)

[0137] AttMap=Sigmoid(Conv 3×3 (I mid ))

[0138] F mid =D+AttMap×Conv3×3 (D)

[0139] I out =I mid +Conv 3×3 (F mid )

[0140] Among them, Conv 3×3 represents a 3×3 convolutional layer, and Sigmoid(·) represents the activation function.

[0141] Furthermore, if Figure 5 As shown, step S5 includes the following steps:

[0142] Step S51, design a forward raindrop removal network, which consists of a shallow feature extraction module, a U-shaped encoder and decoder network based on a multiple attention module, and an image feature fusion module. The input of the network is the attached raindrop image I rain , I rain After the shallow feature extraction module in step S31, The encoding features are obtained through the U-shaped encoder and decoder based on the multiple attention modules in step S32 and step S33. and decode features and I rain After the image feature fusion module in step S41, the derained image I is obtained. derain , the specific formula is as follows:

[0143]

[0144]

[0145]

[0146] Among them, ShallowBlock represents the shallow feature extraction module, Encoder represents the encoder, Decoder represents the decoder, and FusionBlock represents the image feature fusion module.

[0147] Step S52, design a reverse raindrop removal network, which consists of a shallow feature extraction module, a U-shaped encoder and decoder network based on a multiple attention module, and an image feature fusion module. The encoder and decoder of the network receive the encoded features in step S51. and decode features As the input feature of this part of the network. rain The corresponding clean image I clean As the input image, I cleanAfter the shallow feature extraction module in step S31, Plus Then the encoding features are obtained through the multiple attention modules in step S24 After downsampling The encoding features are obtained through multiple attention modules After downsampling The encoding features are obtained through multiple attention modules Plus Decoding features are obtained through multiple attention modules After upsampling and encoding features Add, and then get the decoding features through multiple attention modules After upsampling and encoding features Add, and then get the decoding features through multiple attention modules and I clean The restored attached raindrop image I is obtained by the image feature fusion module in step S41. re-rain , the specific formula is as follows:

[0148]

[0149]

[0150]

[0151]

[0152]

[0153]

[0154]

[0155]

[0156] Among them, ShallowBlock represents the shallow feature extraction module, MA represents the multiple attention module, DownSample represents the image downsampling, UpSample represents the image upsampling, and FusionBlock represents the image feature fusion module.

[0157] Step S53: Design a loss function and use the reverse deraining process to guide the forward deraining network training. Randomly extract an attached raindrop image I from the enhanced training set in step S1 rainAnd its corresponding clean image I clean , I rain After the forward deraining network, we get deraining I derain , I clean After the reverse deraining network, the rained image I is obtained re-rain , calculate the following loss, and use the Adam optimizer to update the network parameters until the loss converges. The specific loss function formula is as follows:

[0158] L=L forward +L reverse

[0159]

[0160]

[0161] Where L forward is the forward deraining network loss, L reverse is the inverse deraining network loss, PSNR is the peak signal-to-noise ratio metric, and SSIM is the structural similarity metric.

[0162] Furthermore, step S6 is implemented as follows:

[0163] The image with attached raindrops to be tested is input into the trained forward deraining network guided by the reverse deraining process, and a clean image with raindrops removed is output.

[0164] The above are preferred embodiments of the present invention. Any changes made according to the technical solution of the present invention, as long as the resulting functions do not exceed the scope of the technical solution of the present invention, belong to the protection scope of the present invention.

Claims

1. A method for removing raindrops from an image based on a reverse deraining process, characterized in that: The steps include: Step S1, perform data preprocessing, including data pairing, data random cropping, and data enhancement processing, to obtain a training data set; Step S2: design a multiple attention module and use pixel attention, channel attention and spatial attention to extract image features; Step S3, designing a U-shaped encoder-decoder network based on multiple attention modules to extract features at different scales of the image; Step S4, designing an image feature fusion module, which fuses the image features and the original image to obtain a fused restored image; Step S5, designing a forward raindrop removal network and a reverse raindrop removal network, and using the reverse raindrop removal network to guide the training of the forward raindrop removal network; Step S6, inputting the attached raindrop image into the trained forward raindrop removal network guided by the reverse rain removal process, and outputting a clean image with the raindrops removed; The step S3 is specifically implemented as follows: Step S31: Design a shallow feature extraction module, with an image I as input. I obtains the shallow feature F through 3×3 convolution, PReLU activation function and 3×3 convolution. shallow , The specific formula is as follows: F shallow =Conv 3×3 (PReLU(Conv. 3×3 (I rain ))) Among them, Conv 3×3 represents a 3×3 convolutional layer, and PReLU(·) represents the activation function; Step S32: Design an encoder in a U-shaped encoder-decoder network based on multiple attention modules, including three multiple attention modules; the input of the encoder is the shallow feature F in step S31 shallow , F shallow After multiple attention modules, the encoding feature E1 is obtained. E1 is downsampled and multi-attention modules are used to obtain the encoding feature E2. E2 is downsampled and multi-attention module to obtain the encoding feature E3. The specific formula is as follows: E1=MA(F shallow ) E2=MA(DownSample(E1)) E3=MA(DownSample(E2)) Among them, MA represents multiple attention modules, DownSample represents image downsampling; Step S33, design a decoder in a U-shaped encoder-decoder network based on multiple attention modules, including three multiple attention modules; the input of the decoder is the encoded feature E3 in step S32, and E3 is subjected to the multiple attention modules to obtain the decoded feature D3, D3 is upsampled and added to the encoded feature E2, and then the decoded feature D2 is obtained through the multiple attention modules. D2 is upsampled and added to the encoded feature E1, and then the decoded feature D1 is obtained through the multiple attention modules. The specific formula is as follows: D3=MA(E3) D2=MA(UpSample(D3)+E2) D1=MA(UpSample(D2)+E1) Among them, MA represents multiple attention modules, and UpSample represents image upsampling; The step S5 is specifically implemented as follows: Step S51, design a forward raindrop removal network, which consists of a shallow feature extraction module, an encoder and decoder in a U-shaped encoder-decoder network based on a multiple attention module, and an image feature fusion module; the input of the network is the attached raindrop image I rain , I rain After the shallow feature extraction module in step S31, The encoder and decoder in the U-shaped encoder-decoder network based on the multiple attention module in step S32 and step S33 obtain the encoded features and decode features and I rain After the image feature fusion module, the derained image I is obtained. derain , the specific formula is as follows: Among them, ShallowBlock represents the shallow feature extraction module, Encoder represents the encoder, Decoder represents the decoder, and FusionBlock represents the image feature fusion module; Step S52, design a reverse raindrop removal network, which consists of a shallow feature extraction module, an encoder and a decoder in a U-shaped encoder-decoder network based on a multiple attention module, and an image feature fusion module. The encoder and decoder of the network receive the encoded features in step S51. and decode features As the input feature of the network; and the attached raindrop image I in step S51 rain The corresponding clean image I clean As the input image, I clean After the shallow feature extraction module in step S31, Plus Then the encoding features are obtained through multiple attention modules After downsampling The encoding features are obtained through multiple attention modules After downsampling The encoding features are obtained through multiple attention modules Plus Decoding features are obtained through multiple attention modules After upsampling and encoding features Add, and then get the decoding features through multiple attention modules After upsampling and encoding features Add, and then get the decoding features through multiple attention modules and I clean After the image feature fusion module, the restored attached raindrop image I is obtained re-rain , the specific formula is as follows: Among them, ShallowBlock represents the shallow feature extraction module, MA represents the multiple attention module, DownSample represents the image downsampling, UpSample represents the image upsampling, and FusionBlock represents the image feature fusion module; Step S53, design a loss function, and use the reverse deraining process to guide the forward deraining network training; randomly extract an attached raindrop image I from the training data set in step S1 rain And its corresponding clean image I clean , I rain After the forward deraining network, we get deraining I derain , I clean After the reverse deraining network, the rained image I is obtained re-rain , calculate the following loss, and use the Adam optimizer to update the network parameters until the loss converges; the specific loss function formula is as follows: L=L forward +L reverse Where L forward is the forward deraining network loss, L reverse is the inverse deraining network loss, PSNR is the peak signal-to-noise ratio metric, and SSIM is the structural similarity metric.

2. The image raindrop removal method based on reverse rain removal process guidance according to claim 1 is characterized in that: The step S1 is specifically implemented as follows: Step S11, pairing the attached raindrop image and the corresponding clean image; Step S12, randomly cropping each attached raindrop image of size H×W×3 into an image of size p×p×3, and using the same random cropping method for the corresponding clean image, where H and W are the height and width of the attached raindrop image and the clean image, and p is the height and width of the cropped image; Step S13, randomly adopt one of the following eight enhancement methods to perform data enhancement on the training paired images: keep the original image, vertically flip, rotate 90 degrees, rotate 90 degrees and then vertically flip, rotate 180 degrees, rotate 180 degrees and then vertically flip, rotate 270 degrees, and rotate 270 degrees and then vertically flip; after data enhancement, the training data set is obtained, and the attached raindrop image is recorded as I rain , and its paired clean image is denoted as I clean .

3. The image raindrop removal method based on reverse rain removal process guidance according to claim 1 is characterized in that: The step S2 is specifically implemented as follows: Step S21: Design pixel attention module and input image features Where H' and W' represent the height and width of the feature, respectively. After 3×3 convolution, PReLU activation function and 3×3 convolution, the features are obtained. After 1×1 convolution, PReLU activation function, 1×1 convolution and Sigmoid activation function, the pixel attention weight Att is obtained. p , With Att p Multiply and Add up to get pixel attention weighted features The specific formula is as follows: Among them, Conv 3×3 Represents a 3×3 convolutional layer, Conv 1×1 represents a 1×1 convolutional layer, PReLU(·) and Sigmoid(·) represent activation functions; Step S22: Design a channel attention module and input image features After 3×3 convolution, PReLU activation function and 3×3 convolution, the features are obtained. The channel attention weight Att is obtained after channel average pooling, 1×1 convolution, PReLU activation function, 1×1 convolution and Sigmoid activation function. c , With Att c Multiply and Add up to get pixel attention weighted features The specific formula is as follows: Among them, Conv 3×3 Represents a 3×3 convolutional layer, Conv 1×1 represents a 1×1 convolutional layer, PReLU(·) and Sigmoid(·) represent activation functions, and CAP(·) represents channel average pooling; Step S23: Design a spatial attention module and input image features After 3×3 convolution, PReLU activation function and 3×3 convolution, the features are obtained. After spatial average pooling and spatial maximum pooling, feature F is obtained. s1 and F s2 , Splice along the channel F s1 and F s2 , and then obtain the spatial attention weight Att through 1×1 convolution and Sigmoid activation function s , With Att s Multiply and Add up to get pixel attention weighted features The specific formula is as follows: Att s =Sigmoid(Conv 1×1 (Concat(F s1 ,F s2 ))) Among them, Conv 3×3 Represents a 3×3 convolutional layer, Conv 1×1 represents a 1×1 convolutional layer, PReLU(·) and Sigmoid(·) represent activation functions, SAP(·) represents spatial average pooling, SMP(·) represents spatial maximum pooling, and Concat represents concatenation along the channel; Step S24: Design a multi-attention module and input image feature F in , F in After the pixel attention module of step S21, the channel attention module of step S22 and the spatial attention module of step S23, the output feature F is obtained. out , The specific formula is as follows: F out =SAB(CAB(PAB(F in ))) Among them, PAB is the pixel attention module, CAB is the channel attention module, and SAB is the spatial attention module.

4. The image raindrop removal method based on reverse rain removal process guidance according to claim 1 is characterized in that: The step S4 is specifically implemented as follows: Design an image feature fusion module, the input of which is the decoded feature D and image I. The decoded feature D is convolved by 3×3 and then added to the image I to obtain the intermediate result I mid , I mid After 3×3 convolution and Sigmoid activation function, we get the attention map AttMap. D is convolved by 1×1 and then multiplied by AttMap and added to D to obtain the intermediate feature F mid , F mid After 3×3 convolution and intermediate result I mid Add to get the output image I out , The specific formula is as follows: I mid =I+Conv 3×3 (D) AttMap=Sigmoid(Conv 3×3 (I mid )) F mid =D+AttMap×Conv 3×3 (D) I out =I mid +Conv 3×3 (F mid ) Among them, Conv 3×3 represents a 3×3 convolutional layer, and Sigmoid(·) represents the activation function.

5. The image raindrop removal method based on reverse rain removal process guidance according to claim 1 is characterized in that: The specific implementation of step S6 is as follows: inputting the image of attached raindrops to be tested into a trained forward rain removal network guided by a reverse rain removal process, and outputting a clean image with raindrops removed.

Citation Information

Patent Citations

  • Single-image rain removing method based on convolutional neural network double-branch attention generation

    CN112070690A

  • Image rain imprint removal method and system and storage medium

    CN115775213A