Image rain removal algorithm capable of increasing receptive field and realizing multi-scale fusion
By using algorithms that enlarge the receptive field and multi-scale fusion in image rain removal technology, combining feature extraction network, hollow convolution and FPN fusion, the problem of poor results in existing rain removal technology when dealing with complex rain modes is solved, and efficient and high-quality rain removal effect is achieved.
Patent Information
- Application Number
- CN202510029913.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-05-06
AI Technical Summary
The existing rain removal technology is poor in handling complex and variable rain modes, has limited performance, and is difficult to fully reflect reality. Especially in the fields of autonomous driving and security monitoring, it is necessary to deal with rainfall problems under various weather and road conditions.
The image rain removal algorithm with enlarged receptive fields and multi-scale fusion is adopted. The feature extraction network combines SEBlock to capture global feature information, combines hollow convolution to expand the receptive fields, capture wider context information, and effectively integrate multi-scale features through FPN fusion, and ultimately improves rain removal performance through average fusion strategy.
It realizes efficient and high-quality rain removal technology, removes blurred rain marks, and integrates multiple scales to enhance the sense of reality, accurately treats raindrops, significantly improving the efficiency and quality of rain removal.
Smart Images

Figure CN119941580A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision, and in particular to an image deraining algorithm for increasing receptive field and multi-scale fusion. Background Art
[0002] Rain removal technology can improve the clarity and visual effects of images, and is crucial for applications in video surveillance, autonomous driving, intelligent transportation, and weather forecasting. Filter-based methods are one of the earliest proposed rain removal methods. They remove raindrop noise by filtering the image. This type of method is simple and easy to implement, with fast calculation speed, but the rain removal effect is limited and may cause image distortion. Sparse coding-based methods convert image processing problems into sparse coding problems. Although this type of method can handle complex raindrop noise, the calculation complexity is high. Deep learning methods have made significant progress in rain removal effects, but this method usually requires a large amount of training data and computing resources, and may face the problem of insufficient generalization ability in practical applications.
[0003] Although rain removal technology has made some progress, it still faces many unsolved challenges. Existing rain removal methods mostly rely on physics, mathematics, image processing or sparse coding, but they are not effective when dealing with complex and changeable rain patterns and their performance is limited. Existing methods are limited when dealing with real rain patterns because they are mostly based on assumptions and cannot fully reflect reality. The effect of rain removal algorithms is affected by multiple factors such as raindrop size, shape, and speed, and needs to be selected based on actual application scenarios and needs. In the fields of autonomous driving, security monitoring, etc., rain removal technology needs to deal with rainfall problems in various weather, road conditions, and different scenarios and shooting angles.
[0004] Based on the above problems, the present invention proposes a rain removal method that is not only flexible and efficient in adapting to the size of rain shapes, but also does not require iterative optimization. Through innovative feature extraction and multi-scale prediction, it not only maintains the overall rain removal of the image without distortion, but also pays attention to detail processing. It demonstrates its advantages on synthetic and real rainy day datasets, achieving high-performance and efficient rain removal. Summary of the invention
[0005] The purpose of the present invention is to provide an image deraining algorithm with enlarged receptive field and multi-scale fusion, efficient feature extraction and hole convolution, FPN fusion to improve deraining efficiency and quality, remove rain mark blur, multi-scale restoration fusion to enhance realism, and accurately process raindrops.
[0006] To achieve the above object, the present invention provides an image deraining algorithm with increased receptive field and multi-scale fusion, which restores the original state of the image through image filtering technology, including the following steps:
[0007] S1. Feature extraction network: Design feature extraction components in the form of traditional convolutional neural networks;
[0008] S2. Obtain feature maps: After the feature extraction network, four feature maps of different sizes are obtained;
[0009] S3, dilated convolution fusion: Convolution fusion is performed on the acquired feature maps. The kernel is expanded by inserting spaces between the elements of the convolution kernel, thereby expanding the receptive field and capturing a wider range of contextual information.
[0010] S4, FPN fusion: FPN fusion is performed on the acquired feature maps. The feature pyramid is fused layer by layer through upsampling and 1*1 convolution. The length and width of each layer are halved to achieve effective integration of multi-scale features.
[0011] S5, average fusion: average fusion is performed on the feature maps after dilated convolution fusion and FPN fusion, and a new fused feature image is generated by weighted averaging the pixel values of multiple feature images.
[0012] Preferably, the image filtering technology formula is:
[0013]
[0014] in, For the derained image, I r is a rainy image, and K is a filtering operation.
[0015] Preferably, the feature extraction component in S1 is: after using three convolutional layers, the introduced SEBlock adaptively recalibrates the response intensity of each channel in the feature map by modeling the interdependence between channels, thereby effectively capturing and utilizing the global feature information of the image.
[0016] Preferably, the description formula of the global feature information is:
[0017]
[0018] Among them, z c The global feature description of the cth channel, H is the number of pixels in the vertical direction, W is the number of pixels in the horizontal direction, F i,j,c is the pixel value at the i-th row, j-th column, and c-th channel;
[0019] The formula for generating channel weights is:
[0020] s=σ(W 2 δ(W 1 z));
[0021] Among them, W 1 and W 2 is the weight matrix, σ is the Sigmoid activation function, δ is the ReLU activation function, and z is the global feature description vector of all channels;
[0022] The pixel formula for assigning weights to each channel is:
[0023]
[0024] Among them, s c is the weight for each channel.
[0025] Preferably, the sizes of the feature maps in S2 are: 128*128*128, 64*64*256, 32*32*512 and 16*16*512 respectively.
[0026] Preferably, the formula for outputting the feature map after convolution fusion in S3 is:
[0027]
[0028] Among them, y(m,n) is the element of the output feature map, r is the expansion rate, and w(i,j) is the weight of the convolution kernel.
[0029] Preferably, the FPN fusion formula in S4 is:
[0030] FPN = bottom layer upsampling + previous layer convolution.
[0031] Preferably, the average fusion formula in S5 is:
[0032] F(i,j)=a·A(i,j)+b·B(i,j);
[0033] Among them, F(i,j) is the pixel value of the fused image at position (i,j), A(i,j) and B(i,j) are the pixel values of the two original images at position (i,j), and a and b are weights.
[0034] Therefore, the present invention adopts the above-mentioned image deraining algorithm with increased receptive field and multi-scale fusion, and the technical effects are as follows:
[0035] 1. Achieved efficient and high-quality rain removal technology: By designing a unique feature extraction component, the feature extraction process becomes faster and more efficient. Combining the dilated convolution and FPN fusion technology, the rain removal effect is more natural and harmonious, the details are more delicate, and the blur problem in the image restoration process is effectively solved. It also speeds up the training process, thereby significantly improving the rain removal efficiency;
[0036] 2. Improve the quality and realism of rain removal. During the rain removal process, we focus on eliminating the blur caused by rain marks, thereby optimizing the restoration results. At the same time, we use the feature maps obtained through training to perform multi-scale restoration and fusion processing, which can more accurately respond to different raindrops and further improve the quality of rain removal. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 It is a flow chart of the technical route of the present invention;
[0038] Figure 2 It is the feature extraction network flow chart of the present invention;
[0039] Figure 3 This is a flow chart of the CSEBLOCK module of the present invention;
[0040] Figure 4 This is the FPN fusion network flow chart of the present invention;
[0041] Figure 5 This is the flow chart of the dilated convolutional fusion network of the present invention;
[0042] Figure 6 It is a comparison diagram of the actual effects of the present invention. DETAILED DESCRIPTION
[0043] The technical solution of the present invention is further described below through the accompanying drawings and embodiments.
[0044] Unless otherwise defined, the technical terms or scientific terms used in the present invention should be understood by people with ordinary skills in the field to which the present invention belongs. The words "first", "second" and similar words used in the present invention do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprise" and similar words mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0045] Embodiment 1
[0046] like Figure 1-Figure 6 As shown, the present invention provides an image deraining algorithm with increased receptive field and multi-scale fusion, aiming to restore the original state of the image through image filtering technology and remove effects such as occlusion, fog and motion blur caused by rainy days.
[0047]
[0048] in, For the derained image, I r is a rainy image, and K is a filtering operation.
[0049] First, a feature extraction component is designed in the form of a traditional CNN through a feature extraction network. After using three convolutional layers, the introduced SEBlock adaptively recalibrates the response strength of each channel in the feature map by modeling the interdependence between channels, effectively capturing and utilizing the global feature information of the image.
[0050] For the feature map of the cth channel, its global feature description is obtained by calculating the sum of all pixel values of the channel and then dividing it by the total number of pixels in the feature map (i.e. H×W). This step can be expressed by the following formula:
[0051]
[0052] Among them, z c The global feature description of the cth channel, H is the number of pixels in the vertical direction, W is the number of pixels in the horizontal direction, F i,j,c is the pixel value at the i-th row, j-th column, and c-th channel;
[0053] Next, a convolutional layer is used to generate a weight vector s, where each element s c The weight of the cth channel of the channel is processed by the Sigmoid activation function and the ReLU activation function. This step can be expressed by the following formula:
[0054] s=σ(W 2 δ(W 1 z));
[0055] Among them, W 1 and W 2 is the weight matrix, σ is the Sigmoid activation function, δ is the ReLU activation function, and z is the global feature description vector of all channels;
[0056] The weight s of each channel is obtained c , assigning these weights to each pixel in the corresponding channel. This step can be regarded as a channel-level weighted operation on the feature map, which can be expressed by the following formula:
[0057]
[0058] Among them, s c is the weight for each channel.
[0059] After the feature extraction network, four feature maps of different sizes are obtained, namely 128*128*128, 64*64*256, 32*32*512 and 16*16*512. Dilated convolution fusion and FPN fusion are used for the four feature maps of different sizes respectively.
[0060] Atrous convolution fusion dilates the kernel by inserting spaces between the elements of the convolution kernel, thereby expanding the receptive field and capturing a wider range of contextual information. The dilation rate parameter L indicates that the range of the kernel is to be expanded, that is, L-1 spaces are inserted between the kernel elements. When L = 1, no spaces are inserted between the kernel elements, and it becomes a standard convolution. The formula for the output feature map after convolution fusion is:
[0061]
[0062] Among them, y(m,n) is the element of the output feature map, r is the expansion rate, and w(i,j) is the weight of the convolution kernel.
[0063] Through convolution and upsampling operations, the size of the four feature maps is adjusted to 256*256, the number of channels is superimposed using the connection operation, and the features are fused using convolution and dilated convolution to obtain a 256*256*3 feature map.
[0064] FPN fusion is to fuse the feature pyramid layer by layer through upsampling and 1*1 convolution, and the length and width of each layer are halved to achieve effective integration of multi-scale features. The FPN fusion formula is:
[0065] FPN = bottom layer upsampling + previous layer convolution;
[0066] This will result in two 256*256*3 feature maps, which are then averaged and fused.
[0067] The feature maps after dilated convolution fusion and FPN fusion are averaged and a new fused feature image is generated by weighted averaging the pixel values of multiple feature images. The average fusion formula is:
[0068] F(i,j)=a·A(i,j)+b·B(i,j);
[0069] Among them, F(i,j) is the pixel value of the fused image at position (i,j), A(i,j) and B(i,j) are the pixel values of the two original images at position (i,j), and a and b are weights.
[0070] As shown in Tables 1 and 2, Table 1 shows the performance index parameters of the model in rain100L, and Table 2 shows the performance index parameters of the model of the present invention in rain100H. By comparison, it can be seen that the model of the present invention shows relatively stable performance on both data sets, especially in retaining the overall structure and details of the image.
[0071] Table 1 Model performance index parameters in rain100L
[0072] Model PSNR SSIM SFNet 38.21 0.974 HINet 37.28 0.97 MPRNet 36.40 0.965 MSPFN 32.40 0.933 OURS 36.2 0.977
[0073] Table 2 Performance index parameters of the model in rain100H
[0074]
[0075]
[0076] Therefore, the present invention adopts the above-mentioned image deraining algorithm with increased receptive field and multi-scale fusion, and effectively captures global feature information through feature extraction network combined with SEBlock to improve the deraining effect. The atrous convolution is used to expand the receptive field, capture a wider range of contextual information, and enhance the feature fusion capability. FPN fusion realizes the effective integration of multi-scale features and improves the quality of image detail restoration. At the same time, the algorithm adopts an average fusion strategy, combining the advantages of atrous convolution fusion and FPN fusion, to further improve the image deraining performance.
[0077] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that they can still modify or replace the technical solution of the present invention with equivalents, and these modifications or equivalent replacements cannot cause the modified technical solution to deviate from the spirit and scope of the technical solution of the present invention.
Claims
1. An image deraining algorithm with increased receptive field and multi-scale fusion, characterized in that: Restoring the original state of the image through image filtering technology includes the following steps: S1. Feature extraction network: Design feature extraction components in the form of traditional convolutional neural networks; S2. Obtain feature maps: After the feature extraction network, four feature maps of different sizes are obtained; S3, dilated convolution fusion: Convolution fusion is performed on the acquired feature maps. The kernel is expanded by inserting spaces between the elements of the convolution kernel, thereby expanding the receptive field and capturing a wider range of contextual information. S4, FPN fusion: FPN fusion is performed on the acquired feature maps. The feature pyramid is fused layer by layer through upsampling and 1*1 convolution. The length and width of each layer are halved to achieve effective integration of multi-scale features. S5, average fusion: average fusion is performed on the feature maps after dilated convolution fusion and FPN fusion, and a new fused feature image is generated by weighted averaging the pixel values of multiple feature images.
2. The image deraining algorithm for increasing the receptive field and integrating multi-scale fusion according to claim 1, characterized in that: The image filtering technique formula is: in, For the derained image, I r is a rainy image, and K is a filtering operation.
3. The image deraining algorithm for increasing the receptive field and integrating multi-scale fusion according to claim 1, characterized in that: The feature extraction component in S1 is as follows: After using three convolutional layers, the introduced SEBlock adaptively recalibrates the response intensity of each channel in the feature map by modeling the interdependence between channels, thereby effectively capturing and utilizing the global feature information of the image.
4. The image deraining algorithm for increasing the receptive field and integrating multi-scale fusion according to claim 3, characterized in that: The description formula of the global feature information is: Among them, z c The global feature description of the cth channel, H is the number of pixels in the vertical direction, W is the number of pixels in the horizontal direction, F i,j,c is the pixel value at the i-th row, j-th column, and c-th channel; The formula for generating channel weights is: s = σ(W2δ(W1z)); Among them, W1 and W2 are weight matrices, σ is the Sigmoid activation function, δ is the ReLU activation function, and z is the global feature description vector of all channels; The pixel formula for assigning weights to each channel is: Among them, s c is the weight for each channel.
5. The image deraining algorithm for increasing the receptive field and multi-scale fusion according to claim 1, characterized in that: The sizes of the feature maps in S2 are: 128*128*128, 64*64*256, 32*32*512 and 16*16*512 respectively.
6. The image deraining algorithm for increasing the receptive field and integrating multi-scale fusion according to claim 1, characterized in that: The formula for outputting the feature map after convolution fusion in S3 is: Among them, y(m,n) is the element of the output feature map, r is the expansion rate, and w(i,j) is the weight of the convolution kernel.
7. The image deraining algorithm for increasing the receptive field and integrating multi-scale fusion according to claim 1, characterized in that: The FPN fusion formula in S4 is: FPN = bottom layer upsampling + previous layer convolution.
8. The image deraining algorithm for increasing the receptive field and integrating multi-scale fusion according to claim 1, characterized in that: The average fusion formula in S5 is: F(i,j)=a·A(i,j)+b·B(i,j); Among them, F(i,j) is the pixel value of the fused image at position (i,j), A(i,j) and B(i,j) are the pixel values of the two original images at position (i,j), and a and b are weights.