Improved KPN network image deraining method, system and storage medium
By improving the KPN network, the introduction of deep separable convolution and spatial attention mechanisms has been solved, and the problems of poor image rain removal treatment and low computational efficiency in the prior art have been solved, achieving more efficient image rain removal enhancement effect and lower computational complexity.
Patent Information
- Application Number
- CN202210843808.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-18
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2042-07-18
AI Technical Summary
The existing image rain removal treatment method based on KPN network has poor image enhancement effect, excessive network parameters, large hardware loss, and low computing efficiency.
By improving the KPN network, a deep separable convolution and spatial attention mechanism was introduced. The feature extraction network layer includes a depth separable convolution layer, a mean pooling layer and a bicubital interpolation upsampling layer. The ordinary convolution network layer introduces a spatial attention mechanism. The convolution network layer at different scales and the depth separable convolution network layer are used for rain removal enhancement.
It effectively improves the effect of image enhancement, reduces the parameter quantity and computational complexity of the convolutional neural network, improves the computing efficiency, and is suitable for terminal hardware deployment.
Smart Images

Figure CN115311155B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing, and in particular to a method, system and storage medium for removing rain from an image based on an improved KPN network. Background Art
[0002] Digital images are everywhere in our production and life. Image degradation is easily caused by noise corruption, camera shake, object motion, resolution limitation, haze, rain streaks or their combination during imaging. Image degradation is generally irreversible and can only be processed with the help of image enhancement algorithms. Among them, image degradation caused by raindrops falling on the camera has attracted more and more attention.
[0003] The main ideas of image enhancement for raindrop removal include removing raindrop noise from the original input image through filtering or other operations, and generating raindrop-free images through adversarial generative networks. The fundamental idea of image rain removal and enhancement is to obtain the mapping relationship from noisy images to noise-free images. There are traditional denoising and enhancement methods and deep learning-based rain removal algorithms. Traditional denoising and enhancement methods mainly include spatial domain filtering and frequency domain filtering. They can remove some noise, but are difficult to remove complex noise. In addition, the calculation is complex, blind denoising cannot be performed, and the time consumption is long. They are not real-time when deployed and used in terminals. With the rise of deep learning, more and more neural network models are used in image enhancement. The main popular ones are CNN, Resnet, u-net, Hinet and GAN series. These network models all have innovative breakthroughs in network architecture. As the network evolves more and more complex, the number of network parameters is also increasing, which restricts the real-time performance of the algorithm and places high requirements on terminal hardware. Many of the deep learning-based rain removal methods are based on rain removal pattern assumptions or prior knowledge. The mathematical model of raindrops is expressed as:
[0004] I=(1-M)⊙B+R Formula (1)
[0005] Among them, I represents the degraded image affected by raindrops, M represents the binary mask (a mask of 1 indicates that the pixels in the area are affected by raindrops, otherwise it indicates that the pixels in the area are not affected by raindrops), B represents the image background, and R represents the effect of raindrops, which represents a complex mixture of background information and light reflected by the environment and penetrating raindrops. At this time, the rain removal network requires a lot of fine-tuning optimization processes, which is very time-consuming and cannot cover various situations in real rainfall scenes. Image enhancement algorithms that are independent of rain models can solve this problem well. Among them, the KPN algorithm absorbs the ideas of excellent denoising algorithms. The algorithm does not make modeling assumptions about how rain is generated in the image. Its network structure integrates u-net and kernel prediction. A pixel-by-pixel filter kernel is predicted through the u-net network. This filter kernel is used to convolve with the input noise image for denoising. It has great advantages in image deraining efficiency. While obtaining similar effects to the deraining SOTA algorithm, the deraining of a single image can be kept within 10ms, which can be easily deployed on terminal hardware.
[0006] Usually, the evaluation indicators of image enhancement accuracy are PSNR (Peak Signal to Noise Ratio) and SSIM (structural similarity index). PSNR is an objective standard for evaluating images, and the calculation formula is:
[0007]
[0008] SSIM is a measure of the similarity between two images. One of the two images used is an uncompressed, distortion-free image, and the other is a distorted image, where the value ranges from 0 to 1, 1 is completely consistent, and 0 is completely inconsistent. Assume that two given M*N images are X and Y, where the mean, standard deviation of X, and the covariance of X and Y are represented by ux, σx, and σxy, respectively, and the comparison functions of brightness, contrast, and structure are defined as follows:
[0009]
[0010]
[0011]
[0012] SSIM(X,Y)=[l(X,Y)] α [c(X,Y)] β [s(X,Y)] γ Formula (6)
[0013] Qing Guo et al. proposed to use the RainMix data augmentation method to expand the real scene data, and then train the KPN network for rain removal. The speed and effectiveness of the image rain removal task (involving raindrop removal, mainly rain line removal) are good. Bin Zhang et al. proposed the AME-KPN (attention mechanism enhanced kernel prediction networks) image denoising method, which uses an almost cost-free attention module to refine the feature map and further utilize the inter-frame and intra-frame redundancy of the noisy image. Among them, the prediction kernel roughly restores the clean pixels at its corresponding position through adaptive convolution operations, and the weighted sum is calculated to compensate the prediction kernel. Talmaj et al. proposed a multi-KPN image denoising method. The multi-KPN prediction kernel has more than one size, but kernels of different sizes. Kernels of different sizes help to extract different information from the image, thereby achieving better reconstruction, and kernel fusion ensures that the extracted information is retained while maintaining computational efficiency. Summary of the invention
[0014] The main purpose of the present invention is to provide an image deraining method, device and readable storage medium based on an improved KPN network, aiming to solve the technical problems of poor image enhancement effect, excessive network parameters, large hardware loss and low computational efficiency in the existing image deraining processing method based on the KPN network.
[0015] In the first aspect, a method for image deraining based on an improved KPN network, wherein the improved KPN network includes a feature extraction network layer, a common convolutional network layer, a convolutional network layer of different scales, and a depth-separable convolutional network layer;
[0016] The method comprises the following steps:
[0017] Input the preprocessed image into the feature extraction network layer to obtain the output image after feature extraction;
[0018] The output graph after feature extraction is input into the common convolutional network layer that introduces the spatial attention mechanism to obtain the prediction kernel;
[0019] The obtained prediction kernel is input into convolutional network layers of different sizes to obtain prediction maps of different receptive field features;
[0020] The features of each predicted image are weighted and summed, and input into the depthwise separable convolutional layer to obtain the deraining enhanced image.
[0021] In some embodiments, the step of inputting the preprocessed image into the feature extraction network layer of the improved KPN network model to obtain the output graph of the feature extraction network layer includes:
[0022] The feature extraction network layer includes a first network sublayer, a second network sublayer, a third network sublayer and a fourth network sublayer, the first network sublayer includes a depthwise separable convolution layer, the second network sublayer includes a mean pooling layer and a depthwise separable convolution layer, the third network sublayer includes a bicubic interpolation upsampling layer, a depthwise separable convolution layer and a rectified linear unit activation layer, and the fourth network sublayer includes a bicubic interpolation upsampling layer;
[0023] Perform convolution processing on the input image through the depth-separable convolution layer of the first network sublayer to obtain a feature map after convolution processing;
[0024] Performing pooling processing on the mean of the feature map after convolution processing through the mean pooling layer of the second network sublayer, and performing convolution processing through the depthwise separable convolution layer to obtain the feature map after pooling and convolution processing;
[0025] The feature map after pooling and convolution processing is interpolated and up-sampled through the bicubic interpolation upsampling layer of the third network sublayer, the depth-separable convolution layer performs convolution processing, and the rectified linear unit activation layer performs correction activation processing, and a residual connection is performed with the feature map after pooling and convolution processing obtained by the second network sublayer to obtain the feature extraction layer output map.
[0026] In some embodiments, the step of performing convolution processing on the input image through a depthwise separable convolution layer to obtain a feature map after the convolution processing includes:
[0027] Performing convolution processing on the input image through a depthwise separable convolution layer, wherein the depthwise separable convolution layer includes a depthwise convolution layer and a pointwise convolution layer;
[0028] Perform a single channel convolution operation on the input image through the deep convolution layer to obtain the feature map after the convolution operation;
[0029] The feature map after the convolution operation is integrated through point-by-point convolution to complete the feature extraction and obtain the feature map after convolution processing.
[0030] In some embodiments, the number of the second network sublayer and the third network sublayer is multiple, and the step of performing interpolation upsampling processing on the feature map after pooling and convolution processing through the bicubic interpolation upsampling layer of the third network sublayer, performing convolution processing on the depthwise separable convolution layer, and performing correction activation processing on the rectified linear unit activation layer, and performing residual connection with the feature map after pooling and convolution processing obtained by the second network sublayer, and obtaining the output map of the feature extraction layer includes:
[0031] Performing interpolation upsampling processing on the feature map after pooling and convolution processing through the bicubic interpolation upsampling layer of the third network sublayer, performing convolution processing through the depthwise separable convolution layer, and performing correction activation processing through the rectified linear unit activation layer, to obtain a corrected and activated feature map;
[0032] The corrected activated feature maps output by each third network sub-layer and the convolution-processed feature maps output by each second network sub-layer are residually connected to obtain the feature maps after residual connection as the feature maps after pooling and convolution processing.
[0033] In some embodiments, before the step of inputting the preprocessed image into the feature extraction network layer of the improved KPN network model to obtain the output image after feature extraction, the following steps are also included:
[0034] Adjust the size of the input image and obtain the resized image;
[0035] The preprocessed image is obtained by performing data augmentation processing on the resized image.
[0036] In some embodiments, the data enhancement process includes but is not limited to the following operations:
[0037] Random rotation, horizontal flipping and normalization.
[0038] In some embodiments, the step of inputting the output graph after feature extraction into a common convolutional network layer that introduces a spatial attention mechanism to obtain a prediction kernel specifically includes the following steps:
[0039] The output graph after feature extraction is input into the common convolutional network layer that introduces the spatial attention mechanism, and the weight matrix is learned for all channels on the two-dimensional plane.
[0040] The weight matrix is attached to the output graph after feature extraction to obtain the prediction kernel.
[0041] In a second aspect, the present application provides an image deraining system based on an improved KPN network, wherein the improved KPN network includes a feature extraction network layer, a common convolutional network layer, a convolutional network layer of different scales, and a depth-separable convolutional network layer;
[0042] The improved KPN network image deraining system comprises:
[0043] The feature extraction module is used to input the preprocessed image into the feature extraction layer of the improved KPN network model to obtain the output image after feature extraction;
[0044] A common convolution module, which is connected to the feature extraction module, is used to input the output image after feature extraction into the common convolution network layer that introduces the spatial attention mechanism to obtain the prediction kernel; a different-scale convolution module, which is connected to the common convolution module, is used to input the obtained prediction kernel into the convolution network layers of different sizes to obtain prediction images of different receptive field features;
[0045] The depthwise separable convolution module is connected to the different-scale convolution modules for performing weighted summation of the features of the acquired prediction images and inputting the weighted summation into the depthwise separable convolution layer to obtain a derained enhanced image.
[0046] In a third aspect, the present application provides an improved KPN network image deraining device, the improved KPN network image deraining device comprising a processor, a memory, and an improved KPN network image deraining program stored in the memory and executable by the processor, wherein when the improved KPN network image deraining program is executed by the processor, the steps of the improved KPN network image deraining method as described above are implemented.
[0047] In a fourth aspect, the present application provides a readable storage medium, on which is stored an improved KPN network image deraining program, wherein when the improved KPN network image deraining program is executed by a processor, the steps of the improved KPN network image deraining method as described above are implemented.
[0048] The present invention realizes image deraining processing based on an improved KPN network image deraining method, introduces deep separable convolution, replaces conventional convolution with deep separable convolution, greatly reduces the number of parameters and computational complexity of convolutional neural network; introduces spatial attention matrix, improves the feature expression of key areas of feature map, and effectively improves the effect of image enhancement. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 A schematic diagram of a KPN network structure for image rain removal in one of the prior art solutions in the prior art;
[0050] Figure 2 It is a schematic diagram of another image deraining KPN network structure in the prior art solution 1 in the prior art;
[0051] Figure 3 A schematic diagram of the KPN network structure for image rain removal of the second prior art solution in the prior art;
[0052] Figure 4 A schematic diagram of the KPN network structure for image rain removal of the prior art solution 3 in the prior art;
[0053] Figure 5 A schematic diagram of the KPN network structure for image rain removal according to the fourth prior art solution in the prior art;
[0054] Figure 6 A schematic diagram of the KPN network structure for image rain removal according to the fifth prior art solution in the prior art;
[0055] Figure 7 Schematic diagram of the KPN network structure based on the improved KPN network image deraining method provided in an embodiment of the present application;
[0056] Figure 8 It is a method flow chart based on the improved KPN network image deraining method provided by an embodiment of the present application;
[0057] Fig. 9 It is a functional module block diagram of the image deraining system based on the improved KPN network provided in an embodiment of the present application.
[0058] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0059] It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.
[0060] In the prior art, such as Figure 1 As shown in the figure, the implementation of scheme 1 is as follows: the KPN network inputs N+1 consecutive frames of images, where N frames are continuous multi-noise images and one frame of image has estimated noise. N frames of noisy images are filtered and denoised by the pixel-by-pixel filter output by the network to generate N frames of clean denoised images, and after alignment and averaging, a final noise-free image Y can be generated. Among them, the 2nd to 5th layers of KPN are all 3 layers of convolution plus a mean pooling layer, and the 6th to 8th layers are all Bilinear interpolation upsampling first, then convolution, and skip connection is required to output pixel kernels. The pixel kernels are convolved with the corresponding frames of the original input respectively, and finally weighted average is performed to obtain the output image. The network structure is simple and clear, with few model parameters, and some rain lines can be removed, but the PSNR and SSIM are worse than SOTA, and large raindrops can hardly be removed.
[0061] Another solution adopts the same network structure as Solution 1, the structure is as follows Figure 2 shown.
[0062] like Figure 3 As shown, the second prior art solution is implemented as follows:
[0063] Some improvements have been made to the KPN network. The prediction kernel of multi-KPN has multiple sizes instead of just one size, and these kernels are fused into one kernel based on each pixel. Kernels of different sizes help extract different information from the image, resulting in better reconstruction, and kernel fusion ensures that the extracted information is retained while maintaining computational efficiency.
[0064] like Figure 4 As shown, the existing technical solution three is implemented as follows:
[0065] The deformable kernel prediction network improves the deficiency of the KPN algorithm that it cannot be used for color image denoising, and introduces a deformable convolution kernel to improve the denoising performance of the network. This method replaces ordinary convolution with deformable convolution, that is, an offset variable is added to the position of each sampling point in the convolution kernel. Through these variables, the convolution kernel can sample arbitrarily near the current position. This method increases the complexity of the algorithm, increases hardware consumption, and does not effectively improve the use effect.
[0066] like Figure 5 As shown, the prior art solution 4 is implemented by a dual-channel deformable KPN as follows:
[0067] This method proposes a two-channel deformable kernel prediction network structure. It can be seen that N frames of Bayer images and their corresponding RGB images are sent to two channels of the network respectively. The network has two identical u-net-like channels. The end of each channel is divided into two channels, namely the prediction kernel feature map and the offset feature map. The feature maps are multiplied in pairs to obtain the final prediction kernel and offset feature map. The information of the two channels is exchanged and fused with each other, making the output feature map more robust. At the end of the network, we use the offset feature map to offset the multi-frame Bayer images, and then use the prediction kernel to perform the final denoising operation on the offset Bayer images to obtain the denoised multi-frame Bayer images. Finally, the single-frame Bayer image can be restored after mean processing.
[0068] like Figure 6 As shown, the prior art solution five is implemented as follows:
[0069] This method proposes a KPN network structure enhanced by the attention mechanism for image enhancement. The almost cost-free attention module is used to extract feature maps, effectively utilizing the redundancy between and within frames of noisy images. The AME-KPN network includes feature extraction, output branch, and reconstruction modules. When extracting features, the KPN network introduces the Channel Attention and Spatital Attention structures.
[0070] AME-KPNs outputs pixel-by-pixel spatial adaptive kernels, residual maps, and corresponding weight maps, where the prediction kernel roughly restores the clean pixels at its corresponding position through adaptive convolution operations, and then the residuals are weighted and summed to compensate for the limited receptive field of the prediction kernel. Simulations and actual experiments verify the robustness of the proposed AME-KPN.
[0071] The above-mentioned KPN network-based image deraining schemes have absorbed the ideas of excellent denoising algorithms. The algorithm predicts a pixel-by-pixel filter kernel through the u-net network, and uses this filter kernel to convolve with the input noisy image for denoising. It has great advantages in image deraining efficiency, but there is still some gap between PSNR and SSIM and the deraining SOTA algorithm.
[0072] In view of this, the present application provides an image deraining method based on an improved KPN network, which effectively solves the technical problems of the existing image deraining processing method based on the KPN network, such as poor image enhancement effect, too many network parameters, large hardware loss, and low computational efficiency.
[0073] First, please refer to Figure 7 and Figure 8 , an embodiment of the present invention provides an image deraining method based on an improved KPN network, wherein the improved KPN network includes a feature extraction network layer, a common convolutional network layer, a convolutional network layer of different scales, and a depth-separable convolutional network layer;
[0074] The method comprises the following steps:
[0075] Step S1, input the preprocessed image into the feature extraction network layer to obtain the output image after feature extraction;
[0076] Step S2: Input the output graph after feature extraction into the common convolutional network layer that introduces the spatial attention mechanism to obtain the prediction kernel;
[0077] Step S3: input the obtained prediction kernel into convolutional network layers of different sizes to obtain prediction maps of different receptive field features;
[0078] Step S4: perform weighted summation of the features of each prediction image obtained, and input the summation into a depthwise separable convolutional layer to obtain a rain-removing enhanced image.
[0079] The present invention realizes image deraining processing based on an improved KPN network image deraining method, introduces deep separable convolution, replaces conventional convolution with deep separable convolution, greatly reduces the number of parameters and computational complexity of convolutional neural network; introduces spatial attention matrix, improves the feature expression of key areas of feature map, and effectively improves the effect of image enhancement.
[0080] In one embodiment, the step of inputting the preprocessed image into the feature extraction network layer of the improved KPN network model to obtain the output graph of the feature extraction network layer includes:
[0081] The feature extraction network layer includes a first network sublayer, a second network sublayer, a third network sublayer and a fourth network sublayer, the first network sublayer includes a depthwise separable convolution layer, the second network sublayer includes a mean pooling layer and a depthwise separable convolution layer, the third network sublayer includes a bicubic interpolation upsampling layer, a depthwise separable convolution layer and a rectified linear unit activation layer, and the fourth network sublayer includes a bicubic interpolation upsampling layer;
[0082] Perform convolution processing on the input image through the depth-separable convolution layer of the first network sublayer to obtain a feature map after convolution processing;
[0083] Performing pooling processing on the mean of the feature map after convolution processing through the mean pooling layer of the second network sublayer, and performing convolution processing through the depthwise separable convolution layer to obtain the feature map after pooling and convolution processing;
[0084] The feature map after pooling and convolution processing is interpolated and up-sampled through the bicubic interpolation upsampling layer of the third network sublayer, the depth-separable convolution layer performs convolution processing, and the rectified linear unit activation layer performs correction activation processing, and a residual connection is performed with the feature map after pooling and convolution processing obtained by the second network sublayer to obtain the feature extraction layer output map.
[0085] In a more specific embodiment, the first network sublayer is Figure 7 The leftmost three-layer depth-separable convolutional layer of the box feature extraction network layer in , and the second network sublayer is Figure 7 Each of the 2-5 layers in the box includes a mean pooling layer and three depth-separable convolutional layers from left to right. The second network sublayers are Figure 7 Each of the 6-8 layers in the box includes a Bicubic interpolation upsampling layer (bicubic interpolation upsampling layer), two depth-separable convolutional layers and a ReLU activation layer (rectified linear unit activation layer). The fourth network sublayer is Figure 7 The rightmost Bicubic interpolation upsampling layer in the box (bicubic interpolation upsampling layer).
[0086] In one embodiment, the step of performing convolution processing on the input image through a depthwise separable convolution layer to obtain a feature image after the convolution processing includes:
[0087] Performing convolution processing on the input image through a depthwise separable convolution layer, wherein the depthwise separable convolution layer includes a depthwise convolution layer and a pointwise convolution layer;
[0088] Perform a single channel convolution operation on the input image through the deep convolution layer to obtain the feature map after the convolution operation;
[0089] The feature map after the convolution operation is integrated through point-by-point convolution to complete the feature extraction and obtain the feature map after convolution processing.
[0090] This application introduces a depthwise separable convolution layer, which replaces conventional convolution with depthwise separable convolution. Depthwise separable convolution consists of depthwise convolution and pointwise convolution. A convolution kernel of depthwise convolution is responsible for one channel, and a channel is convolved by only one convolution kernel. The number of feature maps after completion is the same as the number of channels of the input layer, and the feature map cannot be expanded. Therefore, Pointwise Convolution is needed to combine these feature maps to generate a new feature map, which has a relatively low number of parameters and computational cost. By splitting the correlation between the spatial dimension and the channel dimension or the depth dimension, the number of parameters required for convolution calculation is reduced. First, the single channel is convolved using depth convolution, and then the point convolution in the form of 1×1 convolution is used to integrate all channel information at a single feature point position to complete feature extraction. Because the original intention of the design of the 1×1 convolution kernel is to reduce the number of parameters and calculations, the number of parameters of the depthwise separable convolution is reduced by 8-9 times compared to the standard convolution, which greatly reduces the number of parameters and computational complexity of the convolutional neural network (CNN). At the same time, although the number of parameters has dropped significantly, the accuracy has not been affected.
[0091] In the process of interpolation upsampling, bilinear is not used, but Bicubic interpolation upsampling is used. Compared with bilinear, bicubic has a larger amount of calculation, but the latter has a better effect of interpolation restoration features.
[0092] In one embodiment, the number of the second network sublayer and the third network sublayer is multiple, and the step of performing interpolation upsampling processing on the feature map after pooling and convolution processing by the bicubic interpolation upsampling layer of the third network sublayer, performing convolution processing by the depthwise separable convolution layer, and performing correction activation processing by the rectified linear unit activation layer, and performing residual connection with the feature map after pooling and convolution processing obtained by the second network sublayer, and obtaining the output map of the feature extraction layer includes:
[0093] Performing interpolation upsampling processing on the feature map after pooling and convolution processing through the bicubic interpolation upsampling layer of the third network sublayer, performing convolution processing through the depthwise separable convolution layer, and performing correction activation processing through the rectified linear unit activation layer, to obtain a corrected and activated feature map;
[0094] The corrected activated feature maps output by each third network sub-layer and the convolution-processed feature maps output by each second network sub-layer are residually connected to obtain the feature maps after residual connection as the feature maps after pooling and convolution processing.
[0095] In a more specific embodiment, the 6th layer of the feature extraction network layer performs Bicubic interpolation upsampling on the feature map output by the 5th layer, passes through the depthwise separable convolution and ReLU activation layer, and performs skip connection with the 4th layer. The 7th layer performs Bicubic interpolation upsampling on the feature map output by the 6th layer, passes through the depthwise separable convolution and ReLU activation layer, and performs skip connection (residual connection) with the 2nd layer. The 8th layer performs Bicubic interpolation upsampling on the feature map output by the 7th layer, passes through the depthwise separable convolution and ReLU activation layer, and performs skip connection with the 2nd layer.
[0096] In one embodiment, before the step of inputting the preprocessed image into the feature extraction network layer of the improved KPN network model to obtain the output image after feature extraction, the following steps are also included:
[0097] Adjust the size of the input image and obtain the resized image;
[0098] The preprocessed image is obtained by performing data augmentation processing on the resized image.
[0099] In one embodiment, the data enhancement process includes but is not limited to the following operations:
[0100] Random rotation, horizontal flipping and normalization.
[0101] In one embodiment, the input image is first resized to a size of 480×480, and then data enhancement processing is performed, including random rotation and horizontal flipping, and normalization processing.
[0102] In one embodiment, the step of inputting the output graph after feature extraction into a common convolutional network layer that introduces a spatial attention mechanism to obtain a prediction kernel specifically includes the following steps:
[0103] The output graph after feature extraction is input into the common convolutional network layer that introduces the spatial attention mechanism, and the weight matrix is learned for all channels on the two-dimensional plane.
[0104] The weight matrix is attached to the output graph after feature extraction to obtain the prediction kernel.
[0105] Not all regions in an image are equally important in contributing to a task. Only regions relevant to the task are of concern. For convolutional neural networks, each layer of CNN outputs a feature map, and the spatial attention mechanism is to learn a weight matrix for the feature map on a two-dimensional plane for all channels, and a weight is learned for each pixel. These weights represent the importance of a certain spatial position information, forming a spatial attention matrix, which is attached to the original feature map to increase useful features and weaken useless features, thereby achieving the effect of feature screening and enhancement.
[0106] In one embodiment, the prediction kernel is convolved with the input image in different sizes to obtain prediction images with different receptive field features, the prediction image features are weighted and summed, and then passed through a 5×1 depthwise separable convolution layer to finally output a rain-removed and enhanced image.
[0107] The improved KPN network model provided in this application is tested using image pairs under different conditions, and each group of image pairs consists of a clean, noise-free true image and a noisy, degraded image. The test data set covers various scenes such as indoor, outdoor, underground, parking lots and non-parking lots in multiple dimensions such as different weather, different time, and different lighting conditions. The PSNR and SSIM between the model's rain-removed image and the true image are calculated using formula (2) and formula (3) to formula (6). The results show that the image enhancement effect is improved, and the SSIM and PSNR are improved by 0.12 and 3.54 compared with scheme 1, and by 0.06 and 3.58 respectively compared with scheme 5.
[0108] Second, please refer to Fig. 9 , the present application provides an image deraining system based on an improved KPN network, wherein the improved KPN network includes a feature extraction network layer, a common convolutional network layer, a convolutional network layer of different scales, and a depth-separable convolutional network layer;
[0109] The improved KPN network-based image deraining system includes a feature extraction module 100, a common convolution module 200, a different-scale convolution module 300 and a depth-separable convolution module 400. The feature extraction module 100 is used to input the preprocessed image into the feature extraction layer of the improved KPN network model to obtain an output image after feature extraction; the common convolution module 200 is communicated with the feature extraction module 100 and is used to input the output image after feature extraction into the common convolution network layer that introduces the spatial attention mechanism to obtain a prediction kernel; the different-scale convolution module 300 is communicated with the common convolution module 200 and is used to input the obtained prediction kernel into convolution network layers of different sizes to obtain prediction images of different receptive field features; the depth-separable convolution module 400 is communicated with the different-scale convolution module 300 and is used to weighted sum the features of each obtained prediction image and input it into the depth-separable convolution layer to obtain a deraining enhanced image.
[0110] In a third aspect, the present application provides an improved KPN network image deraining device, the improved KPN network image deraining device comprising a processor, a memory, and an improved KPN network image deraining program stored in the memory and executable by the processor, wherein when the improved KPN network image deraining program is executed by the processor, the steps of the improved KPN network image deraining method as described above are implemented.
[0111] In a fourth aspect, the present application provides a readable storage medium, on which is stored an improved KPN network image deraining program, wherein when the improved KPN network image deraining program is executed by a processor, the steps of the improved KPN network image deraining method as described above are implemented.
[0112] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or system. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or system including the element.
[0113] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0114] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes a number of instructions for a terminal device to execute the methods described in each embodiment of the present invention.
[0115] The above are only preferred embodiments of the present invention, and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A method for image deraining based on an improved KPN network, characterized in that: The improved KPN network includes a feature extraction network layer, a common convolutional network layer, a convolutional network layer of different scales and a depth-separable convolutional network layer; The improved KPN network image deraining method comprises the following steps: Input the preprocessed image into the feature extraction network layer to obtain the output image after feature extraction; The output graph after feature extraction is input into the common convolutional network layer that introduces the spatial attention mechanism to obtain the prediction kernel; The obtained prediction kernel is input into convolutional network layers of different sizes to obtain prediction maps of different receptive field features; The features of each predicted image are weighted and summed, and input into the depthwise separable convolutional layer to obtain the deraining enhanced image; The step of inputting the preprocessed image into the feature extraction network layer of the improved KPN network model to obtain the output graph of the feature extraction network layer includes: The feature extraction network layer includes a first network sublayer, a second network sublayer, a third network sublayer and a fourth network sublayer, the first network sublayer includes a depthwise separable convolution layer, the second network sublayer includes a mean pooling layer and a depthwise separable convolution layer, the third network sublayer includes a bicubic interpolation upsampling layer, a depthwise separable convolution layer and a rectified linear unit activation layer, and the fourth network sublayer includes a bicubic interpolation upsampling layer; Perform convolution processing on the input image through the depth-separable convolution layer of the first network sublayer to obtain a feature map after convolution processing; Performing pooling processing on the mean of the feature map after convolution processing through the mean pooling layer of the second network sublayer, and performing convolution processing through the depthwise separable convolution layer to obtain the feature map after pooling and convolution processing; The feature map after pooling and convolution processing is interpolated and up-sampled through the bicubic interpolation upsampling layer of the third network sublayer, the depth-separable convolution layer performs convolution processing, and the rectified linear unit activation layer performs correction activation processing, and a residual connection is performed with the feature map after pooling and convolution processing obtained by the second network sublayer to obtain the feature extraction layer output map.
2. The improved KPN network image deraining method according to claim 1, characterized in that: The step of performing convolution processing on the input image through the depthwise separable convolution layer to obtain the feature image after the convolution processing comprises: Performing convolution processing on the input image through a depthwise separable convolution layer, wherein the depthwise separable convolution layer includes a depthwise convolution layer and a pointwise convolution layer; Perform a single channel convolution operation on the input image through the deep convolution layer to obtain the feature map after the convolution operation; The feature map after the convolution operation is integrated through point-by-point convolution to complete the feature extraction and obtain the feature map after convolution processing.
3. The image deraining method based on the improved KPN network as claimed in claim 1, characterized in that: The number of the second network sublayer and the third network sublayer is multiple, and the step of performing interpolation upsampling processing on the feature map after pooling and convolution processing through the bicubic interpolation upsampling layer of the third network sublayer, performing convolution processing on the depthwise separable convolution layer, and performing correction activation processing on the rectified linear unit activation layer, and performing residual connection with the feature map after pooling and convolution processing obtained by the second network sublayer, and obtaining the output map of the feature extraction layer includes: Performing interpolation upsampling processing on the feature map after pooling and convolution processing through the bicubic interpolation upsampling layer of the third network sublayer, performing convolution processing through the depthwise separable convolution layer, and performing correction activation processing through the rectified linear unit activation layer, to obtain a corrected and activated feature map; The corrected activated feature maps output by each third network sub-layer and the convolution-processed feature maps output by each second network sub-layer are residually connected to obtain the feature maps after residual connection as the feature maps after pooling and convolution processing.
4. The improved KPN network image deraining method according to claim 1, characterized in that: Before the step of inputting the preprocessed image into the feature extraction network layer of the improved KPN network model to obtain the output image after feature extraction, the following steps are also included: Adjust the size of the input image and obtain the resized image; The preprocessed image is obtained by performing data augmentation processing on the resized image.
5. The improved KPN network image deraining method as claimed in claim 4, characterized in that: The data enhancement process includes but is not limited to the following operations: Random rotation, horizontal flipping and normalization.
6. The improved KPN network image deraining method according to claim 1, characterized in that: The step of inputting the output graph after feature extraction into a common convolutional network layer that introduces a spatial attention mechanism to obtain a prediction kernel specifically includes the following steps: The output graph after feature extraction is input into the common convolutional network layer that introduces the spatial attention mechanism, and the weight matrix is learned for all channels on the two-dimensional plane. The weight matrix is attached to the output graph after feature extraction to obtain the prediction kernel.
7. An image deraining system based on an improved KPN network, characterized in that: The improved KPN network includes a feature extraction network layer, a common convolutional network layer, a convolutional network layer of different scales and a depth-separable convolutional network layer; The improved KPN network image deraining system comprises: The feature extraction module is used to input the preprocessed image into the feature extraction layer of the improved KPN network model to obtain the output image after feature extraction; A common convolution module, which is connected to the feature extraction module, is used to input the output image after feature extraction into the common convolution network layer that introduces the spatial attention mechanism to obtain the prediction kernel; a different-scale convolution module, which is connected to the common convolution module, is used to input the obtained prediction kernel into the convolution network layers of different sizes to obtain prediction images of different receptive field features; A depthwise separable convolution module is connected to the different-scale convolution modules for performing weighted summation of features of the acquired prediction images and inputting the weighted summation into the depthwise separable convolution layer to obtain a derained enhanced image. The feature extraction module is also used for the feature extraction network layer including a first network sublayer, a second network sublayer, a third network sublayer and a fourth network sublayer, the first network sublayer includes a depth-separable convolution layer, the second network sublayer includes a mean pooling layer and a depth-separable convolution layer, the third network sublayer includes a bicubic interpolation upsampling layer, a depth-separable convolution layer and a rectified linear unit activation layer, and the fourth network sublayer includes a bicubic interpolation upsampling layer; the input image is convolved through the depth-separable convolution layer of the first network sublayer to obtain the convolution processing The method further comprises the following steps: performing pooling processing on the mean of the feature map after convolution processing by the mean pooling layer of the second network sublayer, and performing convolution processing on the depthwise separable convolution layer to obtain the feature map after pooling and convolution processing; performing interpolation upsampling processing on the feature map after pooling and convolution processing by the bicubic interpolation upsampling layer of the third network sublayer, performing convolution processing on the depthwise separable convolution layer, and performing correction activation processing on the rectified linear unit activation layer, and performing residual connection with the feature map after pooling and convolution processing obtained by the second network sublayer to obtain the output map of the feature extraction layer.
8. An image deraining device based on an improved KPN network, characterized in that: The improved KPN network image deraining device comprises a processor, a memory, and an improved KPN network image deraining program stored in the memory and executable by the processor, wherein when the improved KPN network image deraining program is executed by the processor, the steps of the improved KPN network image deraining method as described in any one of claims 1 to 6 are implemented.
9. A readable storage medium, characterized in that: The readable storage medium stores an improved KPN network image deraining program, wherein when the improved KPN network image deraining program is executed by a processor, the steps of the improved KPN network image deraining method according to any one of claims 1 to 6 are implemented.