Single image deraining method based on neural structure search and rain density guidance

Through the method based on neural structure search, the rain density inference network and single image rain removal network are automatically designed, and the rain density map is used to guide the rain removal process, solving the problem of time-consuming and low efficiency of manual design network structure in the existing technology, and achieving efficient and automatic single image rain removal effect.

CN115272097BActive Publication Date: 2025-05-06SOUTH CHINA UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210665296.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-14
Publication Date
2025-05-06
Estimated Expiration
2042-06-14

AI Technical Summary

Technical Problem

The existing single-image rain removal method requires manual design of network structure, which is time-consuming and inefficient, making it difficult to effectively deal with complex scenarios.

Method used

The structure of rain density inference network and single image rain removal network is automatically searched by using neural structure search method, and the rain density map is used as a guide to improve rain removal effect.

Benefits of technology

Without manual design of network structure, the network structure suitable for the current task is automatically searched, which improves the efficiency and quality of image rain removal, and reduces the cumbersomeness of neural network structure design.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115272097B_ABST
    Figure CN115272097B_ABST
Patent Text Reader

Abstract

The invention discloses a single image deraining method based on neural structure search and rain density guidance, comprising the following steps: S1. determining the structural search space of a rain density inference network and a single image deraining network, and constructing the rain density inference network and the single image deraining network; S2. inputting a rainy image into the rain density inference network, and the rain density inference network automatically searches for a network structure; S3. freezing the structural parameters of the rain density inference network, inputting a rainy image into the rain density inference network, and training the network weight of the rain density inference network; S4. inputting a rainy image into the rain density inference network, and the rain density inference network outputs a rain density map, inputting the rainy image and the rain density map into the single image deraining network, and the single image deraining network automatically searches for a network structure; S5. freezing the structural parameters of the single image deraining network, inputting the rainy image and the rain density map into the single image deraining network, training the network weight of the single image deraining network, freezing its network weight, and obtaining a deraining model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer vision and image restoration, and in particular to a single image rain removal method based on neural structure search and rain density guidance. Background Art

[0002] In recent years, with the continuous development of artificial intelligence and digitalization, photographic equipment (such as mobile phones, surveillance cameras, professional cameras, aerial photography equipment, etc.) has begun to spread throughout all aspects of people's lives, and with it comes the demand for image processing technology. Among them, rainy days, as a common weather climate, often cause distortion, blurring and other degradation phenomena in the images taken by photographic equipment, making subsequent image applications impossible. For this reason, the demand for image deraining has arisen, especially single image deraining technology, which is more difficult than video deraining because it cannot utilize the rich information between video frames.

[0003] Among the single image deraining methods, they can be roughly divided into two categories. One is the traditional method based on sparse representation, Gaussian mixture model, and guided filtering. This type of method relies on physical models and mathematical derivation to solve, uses priors to model rain streaks and background images, and solves by converting the entire process into an iterative process. However, simple modeling will make it difficult for the model to cope with complex scenes, and overly complex modeling will cause the difficulty of solving the problem to increase geometrically and be inefficient. These limitations hinder the application of traditional methods in practical scenarios. The other type is based on deep learning, which tends to directly build a network framework and learn the mapping relationship from rainy images to rain-free background images in a data-driven way. It usually imposes other constraints to solve the problem of model accuracy in the deraining process. It has more advantages than traditional methods in dealing with complex scenes and efficiency.

[0004] However, in practice, the network structures of almost all existing deep learning-based single image deraining methods rely on manual design and are subject to many limitations: 1. Finding an effective network structure requires a lot of manpower, and image restoration performance is very sensitive to the network structure. In particular, multi-scale networks are usually composed of multiple sub-networks, which further increases the difficulty of manual design; 2. Recently, many researchers have proposed using rain density maps as a guide for single image deraining. However, these methods require not only manual design of single image deraining networks, but also manual design of rain density map networks, which further increases the workload; 3. Multi-scale networks need to consider scale fusion strategies. Fusing information of different scales too early or too late will have a great impact on network performance. Summary of the invention

[0005] The purpose of the present invention is to overcome the cumbersome and time-consuming neural network structure design required in the prior art, and to provide a single image deraining method based on neural structure search and rain density guidance. The neural structure search can adaptively search for a network structure suitable for the current deraining task, and at the same time introduce a rain density inference network as a guiding condition to guide the single image deraining network to perform image deraining and improve the quality of restored images.

[0006] In order to achieve the above object, the technical solution adopted by the present invention is as follows:

[0007] A single image deraining method based on neural structure search and rain density guidance comprises the following steps:

[0008] S1. Determine the structural search space of the rain density inference network and the single image deraining network, and construct the rain density inference network and the single image deraining network;

[0009] S2. Inputting the rain image into the rain density inference network and automatically searching the network structure of the rain density inference network;

[0010] S3. Freeze the structural parameters of the rain density inference network, input the rainy image into the rain density inference network, and train the network weights of the rain density inference network;

[0011] S4. Freeze the network weights of the rain density inference network, input the rainy image into the rain density inference network, the rain density inference network outputs a rain density map, input the rainy image and the rain density map into the single image deraining network, and automatically search for the network structure of the single image deraining network;

[0012] S5. Freeze the structural parameters of the single image deraining network, input the rainy image and the rain density map into the single image deraining network, train the network weights of the single image deraining network, freeze the network weights of the single image deraining network, and obtain the deraining model.

[0013] Furthermore, both the rain density inference network and the single image deraining network use the neural structure search method to automatically search for the network structure.

[0014] Furthermore, the rain density inference network includes a parallel module P, a fusion module F and a transformation module T, and the single image deraining network includes a parallel module P, a fusion module F, a transformation module T and an attention module A.

[0015] Furthermore, the structural search space of the rain density inference network is a multi-scale cell unit C composed of a parallel module P and a fusion module F, and the structural search space of the single image deraining network is composed of a multi-scale cell unit C composed of a parallel module P and a fusion module F and an attention module A.

[0016] Furthermore, the parallel module P and the fusion module F are arranged in parallel in the multi-scale cell unit C. During training, the output of the multi-scale cell unit C is the weighted sum of the parallel module P and the fusion module F. The output calculation formula of the multi-scale cell unit C is as follows:

[0017]

[0018] in, represents a set of features output by the transformation module T, ω p ,ω f Respectively represent the weights of the parallel module P and the fusion module F, ω p ,ω f Initialize to a random number between 0 and 1.

[0019] Furthermore, the attention module A includes a plurality of attention cells A cell , attention cell unit A cell There are m different attention operations in parallel, and the attention cell unit A during training cell The output in is the weighted sum of the outputs of each attention operation, and the attention cell unit A cell The output calculation formula is as follows:

[0020]

[0021] Where m represents the number of different attention operations, and f i It means that the input feature f is evenly divided into s parts in the channel dimension and the i-th segmentation feature is taken, T k represents the kth attention operation, α k is the weight parameter of the kth attention operation, y i is the ith output feature.

[0022] Furthermore, the parallel module P performs the same convolution operation on features of different resolutions, and the number of input features is consistent with the number of output features. The parallel module P is expressed by the following formula:

[0023]

[0024] in, Indicates that the feature length and width scale is the input rainy image The channel dimension is 2n times that of the rainy image; MLP(·) represents multi-layer perceptron.

[0025] Furthermore, the fusion module F fuses features of different resolutions with each other. If a feature with a large resolution is fused to a feature with a small resolution, the feature with a large resolution is downsampled using a convolution operation with a step size of 2 and a convolution kernel size of 3x3; for downsampling with a scale difference of 2t times, where t>1, t convolutions are used for downsampling, and the number of feature channels is kept unchanged during the first t-1 convolutions, and the last convolution sets the number of output feature channels to the number of channels of the feature with a small resolution; if a feature with a small resolution is fused to a feature with a large resolution, the number of feature channels of the feature with a small resolution is first adjusted using a convolution operation with a convolution kernel size of 1x1, and then upsampled using the nearest neighbor interpolation method.

[0026] Furthermore, the final output of each scale is obtained by adding the features of all other scales that have undergone scale changes to the features of this scale, and then activating it through the ReLU function. The fusion module F is expressed by the following formula:

[0027]

[0028] Among them, t up (·) represents the upsampling operation performed by the nearest neighbor interpolation method, t down (·) represents the downsampling operation performed by the convolution operation with a stride of 2 and a kernel size of 3x3, and add(·) represents the pixel addition operation.

[0029] Furthermore, the conversion module T scales the input features to their original length and width through convolution operations. Doubling the channel dimension, the calculation process of the conversion module T is expressed as follows:

[0030]

[0031] in, Indicates that the feature is the original feature in terms of length and width It is 2 times the original feature in the channel dimension n times; F c (·) indicates a convolution operation with a stride of 2; outputs the newly added features By the input The features are obtained through convolution, and each time through the conversion module T only adds a feature of different resolution, that is, the conversion module T only adds generate The remaining features are directly obtained from the input features.

[0032] Compared with the prior art, the present invention has the following beneficial effects:

[0033] 1. The present invention adopts the idea of ​​neural structure search in the network structure design of both the rain density inference network and the single image deraining network, without the need to manually design the network structure. Compared with the traditional network structure based on experimental and intuitive design, the network structure generated based on data-driven generation is more suitable for current data and tasks, and can alleviate the tedious and time-consuming neural network structure design to a certain extent.

[0034] 2. In the training process, the present invention makes full use of the rain streak density information jointly generated by the rainy images and the rainless images in the database, and uses the neural structure search idea to train a rain density inference network that can effectively extract the rain density information in the rainy images. The rain streak density information contained in the rain density map generated by the rain density inference network can guide the single image deraining network to perform the rainy image deraining task more accurately.

[0035] 3. When designing the structural search space of the single image deraining network and the rain density inference network, the present invention does not use the attention module A in the rain density inference network with a simpler task objective according to the difference in tasks, thereby reducing the complexity of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 It is a flowchart of the single image deraining method based on neural structure search and rain density guidance of the present invention.

[0037] Figure 2 Schematic diagram of the structure of the multi-scale cell unit C in an embodiment of the present invention.

[0038] Figure 3 is the attention cell unit A in the embodiment of the present invention cell The structural block diagram of .

[0039] Figure 4 Schematic diagram of the attention module A in an embodiment of the present invention.

[0040] Figure 5 This is the overall network block diagram of the single image deraining method based on neural structure search and rain density guidance of the present invention. DETAILED DESCRIPTION

[0041] The single image rain removal method based on neural structure search and rain density guidance of the present invention is further described below in conjunction with the accompanying drawings and specific embodiments.

[0042] See also Figure 1The present invention discloses a single image deraining method based on neural structure search and rain density guidance, comprising the following steps: S1. Determine the structural search space of the rain density inference network and the single image deraining network, and construct the rain density inference network and the single image deraining network; S2. Input the rainy image into the rain density inference network, and automatically search the network structure of the rain density inference network; S3. Freeze the structural parameters of the rain density inference network, input the rainy image into the rain density inference network, and train the network weight of the rain density inference network; S4. Freeze the network weight of the rain density inference network, input the rainy image into the rain density inference network, the rain density inference network outputs a rain density map, input the rainy image and the rain density map into the single image deraining network, and automatically search the network structure of the single image deraining network; S5. Freeze the structural parameters of the single image deraining network, input the rainy image and the rain density map into the single image deraining network, train the network weight of the single image deraining network, freeze the network weight of the single image deraining network, and obtain the deraining model. When the rainy image is input into the final deraining model, the single image deraining effect can be achieved.

[0043] The present invention introduces the idea of ​​neural structure search in the network structure design of the rain density inference network and the single image deraining network, so that the network can automatically learn different network structures and effectively cope with image processing tasks with different requirements. In order to more accurately process rain streaks in rainy images, a rain density inference network based on neural structure search is introduced, and a single image deraining network is guided by a binary rain density map to achieve a more accurate deraining effect.

[0044] Specifically, obtain the real rain density map required for rain density inference network training. The method of obtaining the real rain density map is as follows:

[0045] Existing databases such as Rain800 and Rian1400 provide rainy-no-rainy image pairs for training. i,j ) n×n and the rain-free background image B=(b i,j ) n×n The relationship between the real rain density map M = (m i,j ) n×n , real rain density map M=(m i,j ) n×n It is expressed as:

[0046]

[0047] Among them, m i,j represents the pixel at coordinate (i, j) in the rain density map M, o i,j represents the pixel at coordinate (i, j) in the rainy image O, b i,jRepresents the pixel at coordinate (i, j) in the rain-free background image B. The length and width of the image are both n pixels. The rain density map M represents the difference between the rainy image and the rain-free image, indicating the location of the rain streaks in the image. According to the number of rain streaks in different areas of the image, the rain streak density information of the area can be obtained to guide the subsequent single image deraining network to perform more accurate deraining.

[0048] (1) Determine the structural search space of the rain density inference network

[0049] In the rain density inference network, a multi-scale network structure is used to retain more image information, and the idea of ​​neural structure search is used to allow the network to determine the timing of multi-scale information fusion according to the current task, thereby achieving higher image restoration performance. Specifically, in the rain density inference network, the method is as follows:

[0050] First, a simple MLP (Multi-Layer Perceptron) is used to extract image features. For each RGB input rain image x, the method for extracting image features is:

[0051] f in =MLP(x)

[0052] Among them, f in It represents the features extracted after the MLP network. Its size is consistent with x, and the dimension is increased from RGB 3D to 32D.

[0053] Then, the obtained feature f in Input to a transformation module T to obtain multi-scale features. Figure 2 As shown, the conversion module T scales the input features to the original length and width through convolution operations. Doubling the channel dimension adds a feature of different resolution. The calculation process of the transformation module T can be expressed as:

[0054]

[0055] in, Indicates that the feature is the original feature in terms of length and width It is 2 times the original feature in the channel dimension n times; F c (·) indicates a convolution operation with a stride of 2; outputs the newly added features By the input The features are obtained through convolution, and each time through the conversion module T only adds a feature of different resolution, that is, the conversion module T only adds generate The remaining features are directly obtained from the input features.

[0056] After obtaining the multi-scale features, they are input into the multi-scale cell unit C composed of the parallel module P and the fusion module F. Figure 2 As shown in the figure, the parallel module P and the fusion module F are distributed in parallel in the multi-scale cell unit C. During the training phase, different weights ω are assigned to each parallel module P and fusion module F. p ,ω f , the output of each step in the multi-scale cell unit C is the weighted sum of the parallel module P and the fusion module F. The output calculation formula of the multi-scale cell unit C is as follows:

[0057]

[0058] in, represents the characteristic output of the conversion module T, ω p and ω f Respectively represent the weights of the parallel module P and the fusion module F, ω p and ω f is initialized to a random number between 0 and 1. After the training is completed, according to the weight ω of the parallel module P p and the weight ω of the fusion module F f The size of the network is determined by selecting the network channel with the largest weight (parallel module P or fusion module F) as part of the overall network.

[0059] Specifically, for each parallel module P and fusion module F, the multi-scale features keep the scale and channel dimension unchanged after input and output. The difference is that the features of different resolutions in the parallel module P do not exchange information, and the output features of the current scale only depend on the input features of the same scale. The input is directly output after the MLP (multi-layer perceptron) operation. The parallel module P can be expressed as:

[0060]

[0061] in, Indicates that the feature length and width scale is the input rainy image The channel dimension is 2n times that of the rainy image; MLP(·) represents multi-layer perceptron.

[0062] The output feature of the current scale of the fusion module F is the fusion of the input features of all scales. If a feature with a large resolution is fused to a feature with a small resolution, the feature with a large resolution is downsampled using a convolution operation with a step size of 2 and a convolution kernel size of 3x3. For downsampling with a scale difference of 2t times, where t>1, t convolutions are used for downsampling. The number of feature channels is kept unchanged during the first t-1 convolutions, and the last convolution sets the number of output feature channels to the number of channels of the feature with a small resolution. If a feature with a small resolution is fused to a feature with a large resolution, the number of feature channels is first adjusted using a convolution operation with a convolution kernel size of 1x1 for the feature with a small resolution, and then upsampled using the nearest neighbor interpolation method.

[0063] The final output of each scale is obtained by adding the features of all other scales that have undergone scale changes to the features of this scale, and then activating it through the ReLU function. The fusion module F can be expressed as follows:

[0064]

[0065] Among them, t up (·) represents the upsampling operation performed by the nearest neighbor interpolation method, t down (·) represents the downsampling operation performed by the convolution operation with a stride of 2 and a kernel size of 3x3, and add(·) represents the pixel addition operation.

[0066] like Figure 2 As shown, each multi-scale cell unit C contains four parallel modules P and four fusion modules F arranged in parallel. Figure 5 As shown in the figure, the rain density inference network contains multiple multi-scale cell units C, and the multi-scale cell units C are connected using the conversion module T. Therefore, as the number of multi-scale cell units C increases, the number of corresponding features of different scales also increases. After the first input single-scale feature passes through the conversion module T, the channel dimension of its feature is expanded to [32,64], after the second input multi-scale feature passes through the conversion module T, the channel dimension of the feature is expanded to [32,64,128], and after the third input multi-scale feature passes through the conversion module T, the channel dimension of the feature is expanded to [32,64,128,256]. After the three conversion modules T, the feature scales are 1,

[0067] After passing the last multi-scale cell unit C, the multi-scale features generated by the network are input into the feature fusion module F, and the feature with 32 channels in the output multi-scale features is taken as the output. Then, the output is passed through an MLP (multi-layer perceptron) network to obtain the rain density map M. out The loss L of the rain density inference network during training is:

[0068] L = MSE (M out ,M)

[0069] Among them, M out represents the rain density map inferred by the rain density inference network, M represents the real rain density map, and MSE represents the mean square loss.

[0070] The network was trained using back propagation and gradient descent methods, with the learning rate set to 0.001. The training was stopped when the loss L of the rain density inference network was observed to be stable, and the trained model was obtained.

[0071] After the rain density inference network structure search training is completed, for each multi-scale cell unit C, the parallel module P and the fusion module F are compared and binarized to 0 or 1:

[0072]

[0073] According to the network structure weight {ω p ,ω f}, and the final structural parameters {γ p ,γ f} and freeze it, thus generating the searched rain density inference network, which is expressed as follows:

[0074]

[0075] When γ p When γ is 1, f is 0, indicating that the rain density inference network selects the parallel module P in the path selection here; when γ p When γ is 0, f is 1, indicating that the rain density inference network selects the fusion module F in the path selection here.

[0076] At this point, the network structure search of the rain density inference network is completed. The rainy image is input again to train the rain density inference network after the frozen structure. The MSE loss and gradient descent method are used to train the weights in the network to obtain the final rain density inference network model.

[0077] (2) Determine the structural search space of the single image deraining network

[0078] The single image rain removal network is similar to the rain density inference network, both of which use the neural structure search method to search the network structure. However, since the output of the network is an RGB image, it is undoubtedly more complicated than a simple one-dimensional binary rain density map. Figure 5As shown in the figure, an attention module A is introduced behind each multi-scale cell unit C in the rain density inference network to obtain a single image deraining network. The features of each scale in a set of multi-scale features output by the multi-scale cell unit C will be independently enhanced by the attention module A.

[0079] like Figure 4 As shown, the attention module A takes the input feature F in Split the channel equally into s parts, and get a set of segmentation features {f1,f2,...,f s}. Each segmentation feature f i Through an attention cell unit A cell . Attention Cell Unit A cell The structure is as Figure 3 As shown, attention cell unit A cell There are m different attention operations in parallel, and the attention cell unit A during training cell The output feature of is the weighted sum of the output features of all attention operations, and the attention cell unit A cell The output is expressed as follows:

[0080]

[0081] Where m represents the number of different attention operations, and f i represents the i-th segmentation feature, T k represents the kth attention operation, α k is the weight parameter of the kth attention operation, y i is the ith output feature.

[0082] During training, the softmax function is used to set the weight parameter α k Normalized to between 0 and 1, the formula is as follows:

[0083]

[0084] Among them, α k represents the weight parameter of the k-th attention operation, β k Represents the normalized weight parameter α k Get the value.

[0085] After the training is completed, except for the weight parameter with the largest value which is assigned a value of 1, the other weight parameters are assigned a value of 0, thereby realizing the selection of the attention operation, which is expressed by the following formula:

[0086]

[0087] The overall process of attention module A is as follows Figure 4As shown, for a set of segmentation features {f1,f2,...,f s}, first obtain a corresponding set of intermediate features {U1,U2,...,U s}.

[0088] For the intermediate feature U1, it is input to the attention cell unit A by the segmentation feature f1 cell For the intermediate feature U2, in addition to the segmentation feature f2, its calculation also depends on the intermediate feature U1 calculated previously, that is, the segmentation feature f2 and the intermediate feature U1 are input into the attention cell unit A respectively. cell In the above example, the output results of the two are added together to obtain the intermediate feature U2. Similarly, the calculation of the intermediate feature U3 depends not only on the segmentation feature f3, but also on the intermediate features U1 and U2 calculated previously. That is, the segmentation feature f3, the intermediate feature U1 and the intermediate feature U2 are input into the attention cell unit A respectively. cell In the example above, the output results of the three are added together to obtain the intermediate feature U3. Similarly, a set of intermediate features {U1, U2, ..., U s The process of} is expressed by the formula as follows:

[0089]

[0090] Get the intermediate features {U1,U2,...,U s}, the segmentation features {f1,f2,...,f s} and intermediate features {U1,U2,...,U s}. For a set of segmentation features {f1,f2,...,f s}, split features f1,f2,...,f s Add, and then input the added result into SENet to get the output feature f sum . It can be expressed as follows:

[0091]

[0092] For a set of intermediate features {U1,U2,...,U s}, convert U1, U2,…,U s and f sum Merge along the channel dimension to form a new feature U sum , feature U sum Through the convolution operation, the number of channels is restored to the same as the input feature F in The two are consistent, and the result is input into SENet to obtain the final output feature F out , which can be expressed as follows:

[0093] Fout =SENet(Conv(Cat(U1,U2,...,U s ,f sum )))

[0094] The single image deraining network uses MSE (mean square error) and perceptual loss as the total loss of the network:

[0095]

[0096] Among them, y′ represents the output result of the single image deraining network, B represents the real rain-free background image, and L p Represents the perceptual loss. Similar to the rain density inference network, after passing through the last multi-scale cell unit C, the multi-scale features generated by the network are input into the feature fusion module F, and the features with 32 channels are taken as output, which are input into the MLP (multi-layer perceptron) network to obtain the final derained output image y′.

[0097] (3) Training strategy

[0098] In order to make the network training tend to be stable, in the network structure search stage, the rain density inference network and the single image deraining network do not add the network structure weight parameter ω to the back propagation for update in the first 10 epochs, but keep the initial random initialization value. After 10 epochs of training, the network tends to be more stable. At this time, the network structure weight parameter ω is added to the back propagation for update to avoid large fluctuations in the network during training.

[0099] In summary, the present invention has the following beneficial effects:

[0100] 1. The present invention adopts the idea of ​​neural structure search in the network structure design of both the rain density inference network and the single image deraining network, without the need to manually design the network structure. Compared with the traditional network structure based on experimental and intuitive design, the network structure generated based on data-driven generation is more suitable for current data and tasks, and can alleviate the tedious and time-consuming neural network structure design to a certain extent.

[0101] 2. In the training process, the present invention makes full use of the rain streak density information generated by the combination of rainy images and rainless images in the database, and uses the neural structure search idea to train a rain density inference network that can effectively extract rain density information from rainy images. The rain streak density information contained in the rain density map generated by the rain density inference network can guide the single image deraining network to perform the rainy image deraining task more accurately.

[0102] 3. When designing the structural search space of the single image deraining network and the rain density inference network, the present invention does not use the attention module A in the rain density inference network with a simpler task objective according to the difference in tasks, thereby reducing the complexity of the model.

[0103] The above description is a detailed description of the preferred feasible embodiments of the present invention, but the embodiments are not intended to limit the scope of the patent application of the present invention. All equivalent changes or modified changes completed under the technical spirit disclosed by the present invention should fall within the patent scope covered by the present invention.

Claims

1. A single image rain removal method based on neural structure search and rain density guidance, characterized in that: The following steps are involved: S1. Determine the structural search space of the rain density inference network and the single image deraining network, and construct the rain density inference network and the single image deraining network; S2. Input the rain image into the rain density inference network, and the rain density inference network automatically searches for the network structure; S3. Freeze the structural parameters of the rain density inference network, input the rainy image into the rain density inference network, and train the network weights of the rain density inference network; S4. Freeze the network weights of the rain density inference network, input the rainy image into the rain density inference network, the rain density inference network outputs a rain density map, input the rainy image and the rain density map into the single image deraining network, and the single image deraining network automatically searches for a network structure; S5. Freeze the structural parameters of the single image deraining network, input the rainy image and the rain density map into the single image deraining network, train the network weights of the single image deraining network, freeze the network weights of the single image deraining network, and obtain a deraining model; The rain density inference network includes a parallel module P, a fusion module F, and a conversion module T. The single image deraining network includes a parallel module P, a fusion module F, a conversion module T, and an attention module A. The structural search space of the rain density inference network is a multi-scale cell unit C composed of a parallel module P and a fusion module F. The structural search space of the single image deraining network is composed of a multi-scale cell unit C composed of a parallel module P and a fusion module F and an attention module A. The parallel module P and the fusion module F are arranged in parallel in the multi-scale cell unit C. During training, the output of the multi-scale cell unit C is the weighted sum of the parallel module P and the fusion module F. The output calculation formula of the multi-scale cell unit C is as follows: in, represents a set of features output by the transformation module T, ω p and ω f Respectively represent the weights of the parallel module P and the fusion module F, ω p ,ω f Initialize to a random number between 0 and 1; The attention module A includes multiple attention cells A cell , attention cell unit A cell There are m different attention operations in parallel, and the attention cell unit A during training cell The output in is the weighted sum of the outputs of each attention operation, and the attention cell unit A cell The output calculation formula is as follows: Where m represents the number of different attention operations, and f i It means that the input feature f is evenly divided into s parts in the channel dimension and the i-th segmentation feature is taken, T k represents the kth attention operation, α k is the weight parameter of the kth attention operation, y i is the i-th output feature; The parallel module P performs the same convolution operation on features of different resolutions. The number of input features is consistent with the number of output features. The parallel module P is expressed as follows: in, Indicates that the feature length and width scale is the input rainy image The channel dimension is 2n times that of the rainy image; MLP(·) represents multi-layer perceptron; The fusion module F fuses features of different resolutions. If a feature with a large resolution is fused to a feature with a small resolution, the feature with a large resolution is downsampled using a convolution operation with a step size of 2 and a convolution kernel size of 3x3. For downsampling with a scale difference of 2t times, where t>1, t convolutions are used for downsampling. The number of feature channels is kept unchanged during the first t-1 convolutions, and the last convolution sets the number of output feature channels to the number of channels of the feature with a small resolution. If a feature with a small resolution is fused to a feature with a large resolution, the number of feature channels is first adjusted using a convolution operation with a convolution kernel size of 1x1 for the feature with a small resolution, and then upsampled using the nearest neighbor interpolation method. The final output of each scale is obtained by adding the features of all other scales that have undergone scale changes to the features of this scale, and then activating it through the ReLU function. The fusion module F is expressed by the following formula: Among them, t up (·) represents the upsampling operation performed by the nearest neighbor interpolation method, t down (·) represents the downsampling operation performed by the convolution operation with a stride of 2 and a kernel size of 3x3, and add(·) represents the pixel addition operation; The conversion module T scales the input features to their original length and width through convolution operations. Doubling the channel dimension, the calculation process of the conversion module T is expressed as follows: in, Indicates that the feature is the original feature in terms of length and width It is 2 times the original feature in the channel dimension n times; F c (·) indicates a convolution operation with a stride of 2; outputs the newly added features By the input The features are obtained through convolution, and each time through the conversion module T only adds a feature of different resolution, that is, the conversion module T only adds generate The remaining features are directly obtained from the input features.

2. The single image rain removal method based on neural structure search and rain density guidance as claimed in claim 1, characterized in that: Both the rain density inference network and the single image deraining network use the neural structure search method to automatically search for the network structure.

Citation Information

Patent Citations

  • Single image rain removal method based on composite residual network and deep supervision

    CN111062892A

  • Method of De-Raining Based on Validity-Aware for Single Image Rain Removal

    KR102402749B1