Image rain removal method and device, equipment and storage medium

By constructing a rain pattern direction attention network and multi-scale feature map, combined with adaptive attention fusion and hierarchical supervision mechanism, the problems of insufficient generalization ability and high computing cost of single-image rain removal technology are solved, and efficient rain pattern removal and calculation optimization are achieved.

CN120147191APending Publication Date: 2025-06-13WUXI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510216710.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing single-image rain removal technology faces the problems of insufficient generalization capabilities and high computing costs in actual applications, resulting in poor rain removal results on unseen scenarios and resource-constrained devices.

Method used

A method for rain removal is proposed. By constructing a rain pattern direction attention network, using multi-scale feature maps and cascaded codec block structures, combining adaptive attention fusion, multi-grained direction attention and hierarchical supervision mechanisms, the model structure is optimized and the parameter quantity is reduced.

Benefits of technology

It realizes accurate capture and removal of rain pattern features, improves rain removal effect, optimizes computing efficiency, reduces storage requirements, and shows better PSNR and SSIM on multiple benchmark data sets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147191A_ABST
    Figure CN120147191A_ABST
Patent Text Reader

Abstract

The invention discloses an image rain removal method and device, equipment and a storage medium, and relates to the field of image processing, a plurality of groups of image data pairs are acquired, and a model training set and a test set are constructed; constructing a multi-scale feature map based on each group of image data, and constructing a rain stripe direction attention network according to the scale of the multi-scale feature map; respectively calculating mean square error loss, content loss and edge loss of each level according to the cascade structure of the coding and decoding blocks, and constructing a hybrid multi-scale loss function by combining the predicted stage rain removal images of each level; and updating model parameters of the rain stripe direction attention network based on the loss function value, and simulating and outputting a rain-removed image according to the rain image after iteration model training is completed. According to the scheme, the direction and distribution characteristics of rainwater can be accurately concerned, a hierarchical supervision mechanism is combined, a staged rain removal image is generated at each hierarchy, loss is calculated earlier, model learning is guided in a refined manner, the calculation efficiency is optimized while the rain stripe removal effect is improved, and the storage requirement is reduced while efficient operation is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of image processing, and in particular to an image deraining method, apparatus, device and storage medium. Background Art

[0002] As an important branch of artificial intelligence, computer vision is widely used in image recognition, medical image analysis and other fields, significantly improving the performance of intelligent systems. However, the negative impact of bad weather, especially rainy days, on image quality has seriously hindered the application of outdoor vision systems such as autonomous driving and intelligent monitoring. Single Image Deraining (SID) is an important underlying visual task in the field of computer vision. Its core goal is to reconstruct a clear rain-free background image from a degraded image disturbed by rain streaks to facilitate subsequent visual tasks. The deep learning-based deraining method can effectively cope with the diversity and complexity of rainy day images through implicit feature learning, providing technical support for the stable application of intelligent vision systems in multiple scenarios.

[0003] In recent years, researchers have been constantly exploring new network architectures and optimization strategies, making significant progress in single image deraining technology in both theoretical research and practical applications. The classic supervised learning baseline method Derain Net uses a simple convolutional neural network (CNN) to achieve a preliminary deraining effect by learning the nonlinear mapping between paired clear images and rainy images. Subsequently, researchers proposed a variety of network structures such as multi-scale fusion and multi-stage progressive. These improvements significantly enhanced the ability of end-to-end learning and laid the foundation for improving the performance of deraining tasks. In recent years, thanks to the powerful ability of the Transformer model in modeling non-local information, a series of Transformer-based methods have emerged. Among them, the latest research work such as Transformer (IDT) and Sparse Transformer (DRSformer) has shown significant performance advantages in single image deraining tasks.

[0004] However, the current single image deraining technology still faces two key challenges in practical applications: On the one hand, due to the significant differences in rain streak morphology, density and transparency in different scenes, the generalization ability of existing methods is insufficient. When the model is applied to unseen scenes, the deraining effect is prone to performance degradation, affecting the reliability and applicability of the visual system. On the other hand, although the Transformer-based method performs well in the deraining task due to its non-local information modeling ability, its high computational cost and large number of parameters limit its application in resource-constrained devices. Therefore, how to improve the deraining effect while optimizing the model structure and reducing the number of parameters is still an important problem that needs to be solved. Summary of the invention

[0005] An embodiment of the present application provides an image de-raining method, device and storage medium, which can improve the de-raining effect while optimizing the model structure and reducing the number of parameters.

[0006] On the one hand, the present application provides an image de-raining method, and the method includes:

[0007] Obtain several groups of image data pairs, and construct a model training set and a test set; each group of image data pairs includes two rainy images and rain-free images with the same picture content;

[0008] Construct multi-scale feature maps based on each group of image data pairs, and construct a rain streak direction attention network according to the scales of the multi-scale feature maps;

[0009] Calculate the mean square error loss, content loss and edge loss at each level respectively according to the cascade structure of the encoding and decoding blocks, and construct a hybrid multi-scale loss function by combining the stage de-rained images predicted at each level and the real rainy images in the validation set;

[0010] Update the model parameters of the rain streak direction attention network based on the loss function value, and simulate and output a de-rained image according to the rainy image after the iterative model training is completed.

[0011] Specifically, the rain streak direction attention network sequentially includes cascaded encoding blocks, a bottleneck layer and cascaded decoding blocks according to the data flow direction; among them, the cascaded encoding blocks and decoding blocks correspond one by one according to the levels, and the number of cascades is the same as the number of multi-scale feature maps;

[0012] Each encoding block includes a multi-granularity direction attention (MGDA) module, an encoding adaptive attention integration (AAIB) module, a downsampling module and an encoding bidirectional dynamic interaction aggregation (BDIA) module;

[0013] The MGDA module generates a rain streak direction attention map according to the constructed multi-scale feature maps and outputs it to the encoding AAIB module; the encoding AAIB module performs feature fusion according to the rain streak direction attention map and the feature map output by the encoding to generate an encoding fusion feature map; the encoding BDIA module constructs a multi-scale supervision structure according to the encoding fusion feature map and the output of the HSAM module in the corresponding decoding block; the downsampling module performs feature extraction and encoding output on the encoding fusion feature map;

[0014] Each decoding block includes a decoding AAIB module, a hierarchical supervision attention (HSAM) module, an upsampling module and a decoding BDIA module;

[0015] The upsampling module receives the input from the previous stage decoding and outputs after upsampling; the HSAM module outputs the stage de-rained image and the stage feature map based on the upsampled input and the rain streak direction attention map of the corresponding coding block; the stage feature map is concatenated with the intermediate feature map output by the corresponding hierarchical coding BDIA and then input into one input of the decoding AAIB module, and the rain streak direction attention map of the corresponding hierarchy is input into the other input of the decoding AAIB module, and the decoding fusion feature map is output through the decoding AAIB module; the decoding BDIA module located at the subsequent stage generates the intermediate feature map based on the input decoding fusion feature map and the output of the previous stage decoding BDIA module.

[0016] Specifically, the coding AAIB module and the decoding AAIB module have the same structure, including a cascaded front layer network and a back layer network;

[0017] The front layer network sequentially includes a layer normalization unit, a point convolution layer, a dilated convolution layer, a simple gating unit, an adaptive attention integration AAIU unit, and a point convolution layer;

[0018] The back layer network sequentially includes a layer normalization unit, a first split path and a second split path in parallel, and a point convolution layer;

[0019] The AAIU unit generates an attention map according to the set learning rate parameter, and the back layer network generates a fusion feature map based on the attention map and the rain streak direction attention map;

[0020] The expression of the feature map flow in the coding AAIB module is as follows:

[0021] F mid =F in +PConv(AAIU(SG(DConv(PConv(LN(F in ))))))×A rain

[0022] F mid1 ,F mid2 =Split(LN(F mid ))

[0023] F out =F mid +PConv(Dconv(PConv(F mid1 ))×SG(Dconv(PConv(F mid2 ))))×A rain

[0024] Among them, F midis the output of the front - layer network; PConv is point convolution; DConv is dilated convolution; AAIU is the Adaptive Attention Integration Unit; SG is the Simple Gating Operation; LN is the Layer Normalization Operation; Split is channel splitting, F mid1 ,F mid2 represents the feature maps of two paths after splitting; F out represents the fused feature map.

[0025] Specifically, the AAIU unit includes:

[0026] The first branch, sequentially including an average pooling layer, a point convolution layer, a simple gating unit, a point convolution layer, and an activation layer; the average pooling layer of the first branch extracts global information, and through the point convolution and the simple gating unit, it extracts the channel - level feature distribution, and outputs the channel attention map through the activation layer;

[0027] The second branch, sequentially including a point convolution layer, a normalization layer, a point convolution layer, a simple gating unit, a point convolution layer, and an activation layer; the second branch extracts spatial attention based on the normalization layer and the simple gating unit, and generates the first spatial attention map through the activation layer;

[0028] The feature map flow expression in the AAIU unit is as follows:

[0029]

[0030] X AAIU =(α×A cha )×(β×A spa )×X f

[0031] where, A cha is the channel attention map; A spa is the first spatial attention map; X AAIU is the output feature map of the AAIU; P is the average pooling layer operation; is the Sigmoid activation function; is the Softmax activation function; α and β are learning parameters.

[0032] Specifically, the HSAM module includes a first branch and a second branch, and commonly inputs the up - sampled feature map:

[0033] The first branch sequentially includes an AAIU unit, a first point convolution layer, a second point convolution layer, and a Sigmoid activation layer; the upsampled feature map generates a residual image through the AAIU unit and the first point convolution layer; a rainy image corresponding to the hierarchical scale is added after the first point convolution layer to generate a stagewise rain-removing image at the corresponding scale, and the stagewise rain-removing image is subtracted from the corresponding rainless image to generate various sub-loss functions at the corresponding scale; all the sub-loss functions at all scales form a mixed multi-scale loss function; the stagewise rain-removing image generates a second spatial attention map through the second point convolution layer and the sigmoid activation layer;

[0034] The second branch sequentially includes a third point convolution layer and a RELU activation layer; the feature map generated after the upsampled feature map passes through the second path is multiplied pixel by pixel with the second spatial attention map and the rain streak direction attention map, and the residual between the upsampled input feature map is introduced to output a predicted feature map.

[0035] Specifically, the BDIA module includes dual-path feature inputs, one is the encoded block feature and the other is the decoded block feature;

[0036] The BDIA module sequentially includes a first point convolution layer, a connection layer, a simple gating unit, a pooling layer, and a second point convolution layer;

[0037] The two inputs of the encoded block feature and the decoded block feature are respectively input into the first point convolution layer and concatenated through the connection layer; after concatenation, channel splitting is performed through the simple gating unit and then multiplied, and then max pooling operation is performed through the pooling layer; the feature after pooling and point convolution is multiplied element by element with the feature before pooling, and the residual between the encoded block feature is introduced to output an intermediate feature map;

[0038] The feature map flow expression in the BDIA module is as follows:

[0039]

[0040] where k is the level of the encoding and decoding blocks, and represents the k-level encoded feature and decoded feature; MaxPool represents the max pooling operation; F temp is the output after the simple gating logic, and F fused is the intermediate feature map.

[0041] Specifically, the construction of the mixed multi-scale loss function includes:

[0042] The mean square error loss, content loss, and edge loss corresponding to each level of the encoding and decoding blocks are respectively constructed as follows:

[0043]

[0044] where d represents the corresponding level or scale; represents the mean square error loss function; represents the content loss function; represents the edge loss function; A d is the rain streak direction attention map predicted by the MGDA module at the corresponding scale; M d is the binary rain pattern map generated by the threshold method; X d is the stage rain removal image at the corresponding scale, Y d is the rain-free image at the corresponding scale; ∈ is the smoothing parameter; Δ represents the Laplacian operator;

[0045] A hybrid multi-scale loss function L is constructed based on the mean square error loss, content loss, and edge loss total as follows:

[0046]

[0047] where ω d represents the weights of the content loss and edge loss at scale d; λ is used to control the relative importance of the two loss functions; η d represents the weight of the mean square error loss at scale d.

[0048] On the other hand, the present application provides an image rain removal device, and the device includes:

[0049] A data construction module, configured to obtain several groups of image data pairs, and construct a model training set and a test set; each group of image data pairs includes two rainy images and rain-free images with the same picture content;

[0050] A model construction module, configured to construct multi-scale feature maps based on each group of image data pairs, and construct a rain streak direction attention network according to the scales of the multi-scale feature maps;

[0051] A loss function calculation module, configured to calculate the mean square error loss, content loss, and edge loss at each level respectively according to the cascade structure of the encoding and decoding blocks, and construct a hybrid multi-scale loss function by combining the stage rain removal images predicted at each level with the real rainy images in the verification set;

[0052] An optimization output module, configured to update the model parameters of the rain streak direction attention network based on the loss function values, and simulate and output a rain removal image according to the rainy image after the iterative model training is completed.

[0053] In another aspect, the present application provides a computer device, which includes a processor and a memory. At least one instruction, at least one program, a code set or an instruction set is stored in the memory, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the image de-raining method described in the above aspect.

[0054] In another aspect, the present application provides a computer-readable storage medium, in which at least one instruction, at least one program, a code set or an instruction set is stored, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement the image de-raining method described in the above aspect.

[0055] The beneficial effects brought by the technical solutions provided in the embodiments of the present application at least include:

[0056] Combined with adaptive attention fusion, multi-granularity directional attention, bidirectional dynamic interaction aggregation and hierarchical supervision mechanism, accurate capture and effective removal of rain streak features are achieved. Under the end-to-end training framework, this method uses a hybrid multi-scale loss function to optimize the model, ensuring the stability and convergence speed of multi-scale feature learning.

[0057] Through the multi-level position attention mechanism, this method can guide the model to accurately focus on the direction and distribution characteristics of rainwater. At the same time, combined with the hierarchical supervision mechanism, stage de-rained images are generated at each level, calculating the loss earlier to finely guide the model learning, thereby accelerating the convergence of the model. This method significantly improves the phenomenon of insufficient rainwater feature extraction and loss of restored image details in traditional single-image de-raining techniques, providing an efficient solution for the de-raining task. Compared with traditional de-raining methods, the present application improves the rain streak removal effect while optimizing the computing efficiency, reducing the storage requirements while ensuring efficient operation. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 is a flowchart of the image de-raining method provided by the embodiments of the present application;

[0059] Figure 2 shows a schematic diagram of the network structure of the RaDA-Net network model provided by the embodiments of the present application;

[0060] Figure 3 is a schematic diagram of the network structure of the AAIB module provided by the present application;

[0061] Figure 4 shows a schematic diagram of the network structure of the MGDA module;

[0062] Figure 5 shows a schematic diagram of the effect of the single-image de-raining method of the RaDA-Net network structure;

[0063] Figure 6 Shows a schematic diagram of the network structure of the HSAM module;

[0064] Figure 7 Shows a schematic diagram of the network structure of the BDIA module;

[0065] Figure 8 Is a comparison result graph of objective evaluation indicators of different methods on a common benchmark dataset;

[0066] Figure 9 Shows a structural block diagram of an image de-raining device provided by an embodiment of the present application;

[0067] Figure 10 Shows a structural block diagram of a computer device provided by an exemplary embodiment of the present application. Detailed implementation manners

[0068] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.

[0069] As used herein, "a plurality of" means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.

[0070] Figure 1 Is a flowchart of an image de-raining method provided by an embodiment of the present application, including the following steps:

[0071] S1. Obtain several groups of image data pairs, and construct a model training set and a test set; each group of image data pairs includes two rainy images and non-rainy images with the same picture content;

[0072] In some possible implementation manners, samples with paired rainy images and non-rainy images can be obtained from a publicly available synthetic rain image dataset (such as Rain200H). To match the input requirements of the model and improve the generalization ability of the model under different scenarios and rain pattern distributions, preprocessing operations are performed on the original data.

[0073] S2. Construct multi-scale feature maps based on each group of image data pairs, and construct a rain pattern direction attention network according to the scales of the multi-scale feature maps;

[0074] The preprocessing process is to first randomly crop the dataset into image patches of 255*255, then perform random flipping and rotation, and finally divide the dataset into a training set and a validation set in a ratio of 8:2. After that, a multi-scale feature map can be constructed for encoding inputs of different scales.

[0075] As Figure 2 shown, this example designs a Rain Directional Attention Network (RaDA-Net) based on the mainstream encoder-decoder idea. The network consists of four core modules: Adaptive Attention Integration Block (AAIB), Multi-Granularity Directional Attention Module (MGDA), Bidirectional Dynamic Interaction Aggregator (BDIA), and Hierarchical Supervision Attention Module (HSAM).

[0076] S3. Calculate the mean square error loss, content loss, and edge loss at each level respectively according to the cascaded structure of the encoder-decoder blocks, and construct a hybrid multi-scale loss function by combining the stagewise de-rained images predicted at each level and the real rainy images in the validation set;

[0077] S4. Update the model parameters of the rain direction attention network based on the loss function value, and simulate and output de-rained images according to the rainy images after the iterative model training is completed.

[0078] Figure 2 The RaDA-Net network structure in contains cascaded encoder blocks, a bottleneck layer, and cascaded decoder blocks in sequence according to the input order. The levels of the cascaded encoder blocks and decoder blocks are set according to the scales of the multi-scale feature maps. Each encoder block is set to contain an MGDA module, an encoding AAIB module, a downsampling module DOWN, and an encoding BDIA module. Each decoder block is set to contain a decoding AAIB module, an HSAM module, an upsampling module UP, and a decoding BDIA module.

[0079] The rainy image of the original size is input through the first convolutional layer to extract a shallow feature map. Where H represents the width of the image, W represents the height of the image, and C represents the number of channels of the image. These features are then fed into an encoder-decoder framework. The following text application takes a three-layer framework as an example for illustration. The encoding AAIB module in the first-level encoding block inputs the original rainy image through ordinary convolution. The downsampling module in the third-level encoding block outputs and connects to the bottleneck layer. The outputs of the bottleneck layer are respectively connected to the upsampling module and the decoding BDIA module in the third-level decoding block. The three-level decoding BDIA modules are cascaded. The output of the first-level decoding BDIA module is convolved through ordinary convolution and then superimposed on the rainy image to output the de-rained image.

[0080] Each level of the encoding block is provided with two feature inputs for inputting feature maps of the same scale. The sizes of the feature maps input and output by the encoding block decrease gradually level by level, and the sizes of the rainy images input and output by the decoding block increase gradually level by level. In some embodiments, the downsampling of the encoding block is implemented through a 3×3 convolution and a pixel unshuffle function.

[0081] The AAIB module adaptively integrates information of different scales in the spatial and channel dimensions, as Figure 2 shown. The AAIB module of each encoding block and decoding block includes two inputs. The first input is the output of the cascaded module, and the second input is the output of the MGDA module. Specifically, in the encoding block, the first feature map of the first input of the encoding AAIB module is the encoding fusion feature map. The first feature map is obtained by gradually encoding and extracting the rainy feature map of the original size. The MGDA module inputs the second feature map. The second feature map is input after size transformation of the original rainy feature map, and is used to extract the rain streak direction attention map and output it to the encoding AAIB module. The output of the encoding AAIB module is respectively connected to the inputs of the encoding BDIA module and the DOWN module. The DOWN module therein is used to gradually reduce the size of the feature map, and the encoding BDIA module is used to connect to the corresponding decoding block to construct a gradually supervised network structure.

[0082] Specifically, the MGDA module generates rain streak direction attention maps of different scales according to the second feature map, and the encoding AAIB module performs feature attention fusion according to the rain streak direction attention map and the first feature map input from the previous stage respectively, and outputs the encoding fusion feature map. In this process, the encoding AAIB module performs an element-wise dot product operation on the rain streak direction attention map and the first feature map, giving higher weights to the rain stripe regions, which prompts the network to consciously focus on the rainfall regions throughout the de-raining stage. In addition, the encoding AAIB modules at different levels can adaptively adjust the fusion degree of spatial attention and channel attention according to the number of channels of the feature map and the set learning weights (α and β values, see Figure 3 ) specifically.

[0083] The bottleneck layer is located between the encoding block and the decoding block. In this application, the AAIB module is also used as the bottleneck layer structure, and the feature map after fusion in the encoding block part is sent into the decoding block part. In some embodiments, the number of channels of the feature map in each layer of the encoding and decoding block is set to 32, 64, 128, and 256, and the number of AAIBs corresponding to each layer is 2, 4, 4, and 6. That is Figure 2 In the first encoding block, feature maps with sizes of H×W×3 and H×W×32 are respectively input. The size of the attention map extracted in each layer remains unchanged, and the channels are normalized. After feature extraction and downsampling, they respectively form and feature maps, which then flow into the subsequent stage.

[0084] The feature map F is downsampled layer by layer, and deeper representation information is extracted through multiple convolutions. At the same time, the spatial dimension is compressed, enabling the model to focus on the global structure of the image. As the number of layers increases, the spatial resolution of the feature map decreases, but the semantic information becomes richer. When the feature flows to the bottleneck layer, the decoding block part starts to gradually restore the image information.

[0085] In the decoding block part, the adaptive attention fusion module (AAIB), the hierarchical supervised attention module (HSAM), and the bidirectional dynamic interaction aggregation module (BDIA) work together. The subsequent UP module receives the decoded fusion feature map output by the previous decoding AAIB module, and after upsampling, it is sent into the HSAM module for image prediction. The decoding AAIB module receives the fused and spliced input of the corresponding hierarchical encoding BDIA module and HSAM module, aiming to use the features in the encoding block for guidance and reuse. The HSAM module receives the output of upsampling and the rain streak direction attention map extracted from the corresponding encoding block, and simulates and outputs the stage de-rained image and the stage feature map. The subsequent decoding BDIA module receives the output of the current decoding AAIB module and the output of the previous decoding BDIA module, and finally simulates the de-rained image. The decoding BDIA module takes the high-level features of the bottleneck layer as input, combines with the output of the previous decoding BDIA module, and can transfer and obtain the fused features of multiple scales and depths to the greatest extent through consecutive BDIA modules.

[0086] The UP module corresponds to the DOWN module. In the decoding part, the feature size needs to be restored level by level. Therefore, it is necessary to first use the UP module to increase the dimension, and then send the upsampled rainy feature map into the HSAM module. The UP module is jointly completed by a convolutional layer and the pixelshuffle function.

[0087] The HSAM module introduced in the decoding part mainly learns the residual de-raining features through the convolution and attention fusion method. Then, based on the rainy image feature map after dimension elevation and the second feature map input by the corresponding encoding block, a stage de-raining image is generated. It should be noted that each decoding block part will generate a stage de-raining image at its respective scale. The stage de-raining image is used to calculate the stage loss with the real rain-free image (Ground Truth). Taking the three-stage encoding and decoding structure as an example, a total of three stage de-raining images and a final output target rain-free image are generated. Loss functions (i.e., Loss4, Loss3, Loss2, and Loss1) are constructed between these four predicted rain-free images and the corresponding real rain-free images respectively. Finally, a hybrid multi-scale loss function L is constructed based on all the loss functions total 。

[0088] Because it is multi-scale supervised learning, the HSAM module in the decoding block will also input the rain streak direction attention map output by the MGDA module in the corresponding encoding block, and combine it with the stage de-raining image to output a stage feature map. This step makes full use of the rain streak features in the rain streak direction attention map to further guide the network model to learn and utilize useful detail features. The stage feature map output by the HSAM module is concatenated with the feature output of the corresponding encoding BDIA module in channels and then sent to the decoding AAIB module in the decoding block

[0089] The decoding process upsamples the features to restore a higher resolution, and at the same time fuses the multi-scale information from the encoding block to retain the detail features. Through multiple layers of decoding, the model gradually reconstructs a clear image after de-raining. At the last stage of the decoding block, a 3×3 convolution is used to obtain a residual image Adding it to the input original image (rainy image) N can obtain the final clear de-raining image C, that is, C = N + R

[0090] Figure 3 is the schematic diagram of the network structure of the AAIB module provided by this application. Whether in the encoding or decoding block, the AAIB module includes two parts: the front layer and the back layer. The front layer part sequentially includes a layer normalization unit (layernorm), a point convolution layer, a dilated convolution layer, a simple gating unit, an adaptive attention integration AAIU unit, and a point convolution layer. The AAIB module first performs layer normalization on the input data, then performs channel mapping through the point convolution layer, expands the receptive field through the dilated convolution, and then increases the non-linear learning ability of the model through the channel-level splitting and multiplication operations of the simple gating unit

[0091] Next, the data enters the core adaptive attention integration unit (AAIU). Specifically, this part consists of two branches

[0092] The first branch: sequentially includes an average pooling layer, a point convolution layer, a simple gated unit, a point convolution layer, and an activation layer. The first branch uses average pooling to extract global information, and combines point convolution with the simple gated unit to learn the channel-level feature distribution, and finally generates a channel attention map through the Sigmoid activation function.

[0093] The second branch: sequentially includes a point convolution layer, a normalization layer, a point convolution layer, a simple gated unit, a point convolution layer, and an activation layer.

[0094] In the first half of the AAIB module structure, the point convolution layer combines Softmax (a normalization function) to calculate spatial attention to strengthen the information expression of the local feature region, and finally also uses the Sigmoid activation function to generate the first spatial attention map. Further, the first spatial attention map and the channel attention map can be fused with each other as feature weights, and multiplied by the original features fed into the AAIU unit. In some embodiments, during the fusion process, the channel attention and spatial attention can be multiplied by the learnable parameters α and β respectively, so that the network can adaptively adjust the importance of the two different attentions. Finally, the fused features are further combined with the rain streak direction attention map, and the information flow is enhanced through residual (the path between the layer normalization in the first half and the current feature fusion) connection.

[0095] The latter part of the AAIB module sequentially includes a layer normalization unit (layernorm), a first split path and a second split path in parallel, and a point convolution layer. The first split path sequentially includes point convolution and ordinary convolution, and the second split path sequentially includes a point convolution layer, an ordinary convolution layer, and a simple gated unit.

[0096] In the second half of the AAIB module structure, after layer normalization, the data is divided into two parts in the channel dimension and undergoes different convolution processes respectively. For the two divided feature maps, operations of 1×1 point convolution and 3×3 ordinary convolution are performed. One branch additionally introduces a simple gated unit, and then the two parts of the feature maps are multiplied to implement a further gated mechanism to filter useful features. Finally, the fused feature map is adjusted in the channel dimension through a layer of point convolution, and multiplied by the rain streak direction attention map again. At the same time, the information flow is optimized through residual (the path between the layer normalization in the second half and the current feature fusion) connection, so as to generate a fused feature map (the output of the encoding block is the encoded fused feature map, and the output of the decoding block is the decoded fused feature map).

[0097] Mathematically, assume the feature map input to the AAIU unit is The rain streak direction attention map is where H represents the width of the feature map, W represents the height of the feature map, and C represents the number of channels. The flow process of the feature map in the module can be expressed as:

[0098] F mid = F in + PConv(AAIU(SG(DConv(PConv(LN(F in )))))) × A rain

[0099] F mid1 ,F mid2 = Split(LN(F mid ))

[0100] F out = F mid + PConv(Dconv(PConv(F mid1 )) × SG(Dconv(PConv(F mid2 )))) × A rain

[0101] Among them, F mid is the output of the first half of the module; PConv is point convolution; DConv is dilated convolution; AAIU is the adaptive attention integration unit; SG is the simple gating operation; LN is the layer normalization operation; Split means dividing the feature map in the channel dimension, F mid1 , F mid2 represents the two-path feature maps after division; F out represents the integrated feature map.

[0102] The AAIU unit is the core of the AAIB module and can adaptively adjust the integration degree of spatial attention and channel attention according to different inputs. Mathematically, given the input feature map The flow process of this feature map in the module can be expressed as:

[0103]

[0104] X AAIU = (α × A cha ) × (β × A spa ) × X f

[0105] Among them, A cha is the channel attention map; A spa is the first spatial attention map; X AAIU is the output feature map of the AAIU; P is the average pooling layer operation (AvgPooling); is the Sigmoid activation function; is the Softmax activation function; α and β are learning parameters used to dynamically optimize the weights of spatial attention and spatial attention; SG is the simple gating operation.

[0106] Simple gating operations have been widely used in the field of image restoration in recent years. Its main function is to provide non-linear expression ability for the model, while avoiding complex activation functions in traditional models. Mathematically, given the input feature map as The flow process of this feature map in SG can be expressed as:

[0107] X f1 ,X f2 =Split(X f )

[0108] X sg =X f1 ×X f2

[0109] X f1 ,X f2 represent the two-channel feature maps after splitting inside the simple gating unit, and X sg represents the output of element-wise multiplication of channels.

[0110] Figure 4 The schematic diagram of the network structure of the MGDA module is shown. The MGDA module mainly processes the input features, generates a directional attention map using the spatial attention mechanism, and captures the coarse-grained and fine-grained directional features through a multi-level convolutional structure, so as to output a rain streak direction attention map that fuses rich context information.

[0111] This structure inputs the second feature map at the corresponding scale and introduces two levels of features for directional feature extraction, specifically including a coarse-grained feature extraction path and a fine-grained feature extraction path, and finally outputs a rain streak direction attention map.

[0112] The coarse-grained feature extraction path successively includes a common convolutional layer, a RELU activation layer, a point convolutional layer, a Sigmoid activation layer, a pooling layer, a common convolutional layer, and a Sigmoid activation layer. After the Sigmoid activation layer outputs the attention map, it is split into four channel feature maps (A up 、A tight 、A down 、A left ) according to channels, forming four branches. The function of this part is to provide a preliminary directional attention distribution for subsequent weighted calculations by the feature extraction module.

[0113] In the coarse-grained extraction path, at the RELU activation layer, a dilated convolutional layer with a dilation rate of 2 and four parallel branches is also introduced to capture the global rainwater distribution information with a larger receptive field and perform weighting according to the corresponding channel feature maps, enabling the module to initially identify the directionality of rain streaks. The outputs of the four dilated convolutional layers are respectively weighted with the four-channel features Figure 1A mapping multiplies the dilated convolution features in different directions with the four-channel feature maps element by element respectively, so as to form a feature map with explicit direction dependence relationship. After the four-way direction dependence relationship feature maps are concatenated, they are connected to a pooling layer (the concatenated output of the parallel average pooling operation AvgPool and max pooling MaxPool), which further refines the information, enhances the significant features, and reduces the influence of redundant information at the same time. After the pooling output, it is refined by a 3×3 ordinary convolution and activated by Sigmoid to obtain a rough attention map, which is multiplied by the feature map output by the RELU activation layer, so as to obtain a rough feature map after coarse-grained feature extraction.

[0114] The fine-grained extraction path successively includes a first RELU activation layer, a layer normalization unit, an ordinary convolution layer, a second RELU activation layer, a point convolution layer, and a Sigmoid activation layer. After the layer normalization unit, a four-way parallel dilated convolution layer with a dilation rate of 1 is also introduced to perform more accurate direction feature extraction. Similar to and corresponding to the coarse-grained stage, the fine-grained features are multiplied and weighted with the four-way channel feature maps corresponding to the coarse-grained stage in the up, down, left, and right directions respectively, so as to enhance the ability to model the local rain structure. Finally, the fine-grained features are concatenated in the channel dimension (to obtain an attention map with the same size as the original), and then the dimensionality is increased and decreased through a 3×3 convolution and a 1×1 convolution, and the final rain streak direction attention map integrating rich context information is output through the Sigmoid activation function. Mathematically, given the input feature map The flow process of the feature map in the module can be expressed as follows:

[0115]

[0116] Among them, A up , A right , A down , A left are the attention maps in the four spatial directions; are the four dilated convolutions with a dilation rate of i; is the four-way channel feature maps obtained from, where k = 1 represents the coarse-grained feature extraction stage and k = 2 represents the fine-grained feature extraction stage; Cat represents the concatenation operation in the channel dimension; A rough is the coarse-grained attention map; A rain is the rain streak direction attention map; Pool is the pooling operation.

[0117] It should be noted that although the division of the channel dimension itself does not directly determine the spatial direction, through the independent training of the MGDA module, the attention maps of different channels can adaptively learn and map to the feature weights in four directions. During the training process, the MGDA module not only serves as a part of the overall model, and its parameter optimization is not only affected by the global loss function.

[0118] Figure 5 The schematic diagram of the effect of the single-image deraining method based on the rain streak direction attention network (RaDA-Net) is shown. From left to right, the first two columns are the input image pairs, which are the rainy image and the rain-free image respectively, and the third column is the derained image output according to the RaDA-Net network model. The rightmost is the decoding stage, using the rain streak direction feature map output by the MGDA module (corresponding to the rain streak direction attention map).

[0119] Figure 6 The schematic diagram of the network structure of the HSAM module is shown. The HSAM module mainly supervises the deraining results at different scales, outputs the stagewise derained images and calculates the corresponding loss values to accelerate the model convergence with earlier error backpropagation. The HSAM module includes a first branch and a second branch, and the two branches use the feature maps at the corresponding levels (after upsampling) as inputs. The first branch sequentially includes an AAIU unit, a first point convolutional layer, a second point convolutional layer, and a Sigmoid activation layer. The second branch includes a third point convolutional layer and a RELU activation layer.

[0120] Feature map First, the residual image is generated through the AAIU unit and the first point convolutional layer in the first branch Subsequently, the residual image is added to the rainy image N at the corresponding level scale to obtain the stagewise derained image Meanwhile, the stagewise derained image is used to calculate the loss with the rain-free image at the same scale, so as to optimize the deraining effect of the model under multi-scale supervision. Among them, the stagewise derained image and the corresponding rain-free image at each scale are subtracted to generate various sub-loss functions at the corresponding scale; all the sub-loss functions at all scales form the mixed multi-scale loss function.

[0121] Furthermore, based on the stagewise derained image, the second spatial attention map is generated through the second point convolutional layer and the sigmoid activation layer Subsequently, this second spatial attention map and the rain streak direction attention map A rain jointly participate in feature modulation and are element-wise multiplied with the feature map processed by the third point convolutional layer and the RELU activation layer in the second branch. Finally, an upsampled residual connection (the path between the input feature map and the feature fusion) is introduced to ensure the integrity of the information flow and output the predicted feature map.

[0122] The HSAM module has two main advantages: First, it can supervise the de-raining learning effect of the model at different scales and finely guide the model learning by calculating the loss earlier, thus accelerating the convergence of the model; Second, by using the stage de-raining image to generate a spatial attention map, it can effectively suppress useless information and enhance the expression of important features.

[0123] Figure 7 The schematic diagram of the network structure of the BDIA module is shown. Whether it is the BDIA module in the encoding block or the decoding block, it receives feature maps from both the encoding and decoding directions. For example Figure 2 For the BDIA module in the encoding block, it comes from the output of the AAIB module in the encoding block and the output of the HSAM module in the decoding block respectively; for the BDIA module in the decoding block, it comes from the output of the AAIB module in the decoding block and the output of the previous-level BDIA module respectively. Tracing back, it ultimately comes from the encoding block direction and is obtained through passing through the bottleneck layer. For ease of description, the input from the encoding block direction is called the encoding block feature, and the input from the decoding block (output of the previous-level BDIA module) direction is called the decoding fusion feature map. This module sequentially includes a first point convolutional layer, a connection layer, a simple gating unit, a pooling layer, and a second point convolutional layer.

[0124] Specifically, the two-way inputs of the encoding block feature and the decoding block feature respectively pass through the first point convolution to align the channel dimensions and extract key information. Subsequently, the features of both are concatenated in the channel dimension through the connection layer and pass through a simple gating mechanism. This gating mechanism enhances the non-linear expression ability of the features through multiplication after channel splitting. Then the fused features enter a max pooling (MaxPool) branch to extract the most significant local information and perform further channel dimensionality reduction adjustment through point convolution. The features after pooling and point convolution are multiplied element-wise with the features before pooling to strengthen the key features at the channel level, and at the same time, the integrity of the information flow is maintained through the residual connection with the encoding block features. Such a design enables the final output to retain both the features of the current encoding block and make full use of the upsampled features in the corresponding decoding block.

[0125] Mathematically, given the encoding block feature map as The decoding block feature map is Then the overall process of the BDIA module is summarized as follows:

[0126]

[0127] where k is the level of the encoding and decoding blocks, and represent the k-level encoding feature and decoding feature; MaxPool represents the max pooling operation; F temp is the output after passing through the simple gating logic, Ffused It is the intermediate feature map of the output.

[0128] In some embodiments, regarding the process of constructing the loss function, this example uses a hybrid multi-scale loss function to optimize the parameters of RaDa-Net end-to-end. To ensure the multi-scale feature learning effect, this model adopts a composite supervision strategy composed of content loss (Charbonnier Loss), edge loss (Edge Loss), and mean square error loss (Mean Square Error).

[0129] As Figure 2 shown, the network calculates the content loss and edge loss at all four scales (d = 1, 2, 3, 4) of the decoding block, and additionally applies rain attention supervision at the first three scales. Finally, the losses at different scales are fused through a hierarchical weighting strategy, enabling the network to adaptively learn global to local information. Specifically, the expression of the total loss function is:

[0130]

[0131] where and represent the content loss and edge loss calculated at scale d, which are used to maintain the consistency of the image across scales and the stability of the image structure; represents the mean square error loss calculated at scale d, which is used to optimize the learning ability of the MGDA module in the local rain streak area. In fact, if the encoding and decoding levels are not limited, the above can also be expressed as:

[0132]

[0133] Here, s represents the maximum encoding and decoding levels.

[0134] In terms of weight allocation, the network uses different hierarchical weighting strategies to balance the loss contributions of different scales. The weights ω d of the content loss and edge loss are set at the four scales as:

[0135] ω d = [1.0, 0.5, 0.25, 0.125], d = 1, 2, 3, 4

[0136] That is, the loss weight of the low resolution is smaller, while the high resolution layer (d = 1) has the largest weight. In addition, the weight η d of the mean square error mse loss of the rain attention supervision is set as:

[0137] η d = [1.0, 0.5, 0.125], d = 1, 2, 3

[0138] The parameter λ is used to control the relative importance of the two loss functions, and its value is empirically set to 0.05. The expression of the Charbonnier loss function is:

[0139]

[0140] where X d is the temporary rain-removal result at the corresponding scale d, Y d is the rain-free ground truth image at the corresponding scale d, and the smoothing parameter ∈ = 1×10 -3 is a constant, which is used to retain the robustness of the L1 loss while enhancing the numerical stability. The expression of the edge loss function is:

[0141]

[0142] where Δ represents the Laplacian operator, and the smoothing parameter ∈ = 1×10 -3 is a constant. At the attention mechanism level, the MGDA module supervises the attention map learning through the mean square error loss, and its expression is:

[0143]

[0144] where A d is the rain streak direction attention map predicted by the MGDA module at the corresponding scale, where M d is the binary map of rain streaks generated by the threshold method, and its calculation process can be described as:

[0145]

[0146] where R d represents the rainy image, and τ is the preset threshold. The attention map A d is compared with M d pixel by pixel to obtain a binary map indicating the positions of the rain streaks. In the binary map, 1 represents the pixels covered by rain, and 0 represents the pixels not covered by rain.

[0147] During the iterative training process of the model, through forward propagation, backpropagation, and parameter optimization, a network model that meets the progress conditions is finally obtained. After obtaining the output through forward propagation, the difference between the rain-removal result and the rain-free ground truth is calculated based on the hybrid multi-scale loss function, and the error gradient is transmitted to the trainable parameters of each network layer using backpropagation. During the training process, the number of training epochs is set to 200 on all datasets. The optimization process uses the AdamW optimizer (β1 = 0.9, β2 = 0.999, weight decay = 0.0001), with an initial learning rate of 2×10 -4 and gradually reduced to 1×10 -6。The patch size is set to 256×256 pixels, and the batch size is set to 2. The operating system used for training is Ubuntu 20.04.6 LTS. The model is trained using Nvidia GeForce RTX 4070Ti Super, and Nvidia CUDA 12.1 is used to accelerate the training speed of the GPU. The implementation is based on the PyTorch framework.

[0148] After the model training is completed, convergence verification and effect evaluation are carried out. Quantitative analysis of the model output on the validation set is performed to evaluate the rain removal performance and convergence stability of the model. During the training iteration process, if the fluctuation amplitude of the PSNR on the validation set is less than the preset threshold (set to 0.01 in this experiment) within 30 consecutive epochs, it is determined that the model has converged and the training can be stopped; otherwise, the forward propagation and backward propagation cycles continue until the convergence condition is met.

[0149] After the model is trained, the peak signal-to-noise ratio (PSNR) and structural similarity index (SSIM), which are commonly used metrics in the field of image restoration, are used to evaluate the visual quality of the image and the consistency of the structural information on the corresponding test set.

[0150] Peak Signal-to-Noise Ratio (PSNR): It is one of the extremely important evaluation metrics in the field of digital signal processing, representing the ratio of the maximum possible signal power of a signal to the actual power of the destructive noise. Since the dynamic range of the signal values is extremely wide, it is usually expressed in logarithmic decibel (dB) units. The calculation formula is:

[0151]

[0152] Among them, Max represents the maximum value in the spatial domain where images X and Y are located. MSE is the mean square error, and the calculation formula is:

[0153]

[0154] Among them, X and Y represent the reconstructed image and the clear image respectively, and n and m are the maximum sizes of the image in the horizontal and vertical dimensions.

[0155] Structural Similarity (SSIM) is used to measure the similarity between two images. It comprehensively compares the images from three aspects: brightness, contrast, and structural information, and evaluates the image quality in a way that is more in line with human visual perception. The calculation formula is:

[0156]

[0157] where μ and σ represent the mean (luminance information) and standard deviation (contrast information) respectively, and C 1 , C 2 is a fixed constant to prevent the denominator from approaching zero, and σ XY is the covariance of images X and Y.

[0158] In practical applications, the model needs to handle complex and ever-changing real-world scenarios. The model is applied to the real-world rain removal task, processing real rainy-day images to remove rain and generating corresponding restored images to evaluate its generalization ability on non-synthetic data.

[0159] To verify the effectiveness of the network model, this solution conducts rain removal experiments on multiple publicly available benchmark datasets, including Rain200L / H, DDN-Data, and Rain13K to evaluate the performance. Rain200L and Rain200H contain 1800 synthetic rain images for training and 200 for testing respectively. DDN-Data consists of 12600 synthetic images with different rain directions and density levels, among which 1200 and 1400 rain images are used for testing. Rain13K is integrated from multiple existing datasets. The test set is divided into five parts, namely Rain100H, Rain100L, Test100, Test2800, and Test1200, containing 13712 pairs of images for training and 4298 test images. Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity (SSIM) are used as objective evaluation metrics.

[0160] This method is quantitatively compared with a variety of mainstream single-image rain removal methods, including the prior-based method (DSC), the convolutional neural network (CNN)-based methods (RESCAN, PreNet, MSPFN, MPRNet), and the recent Transformer-based methods (SwinIR, Uformer-B).

[0161] The comparison results are as Figure 8As shown, the optimal values of each dataset are marked in bold yellow. Compared with other CNN- and Transformer-based methods, the RaDA-Net proposed in the present invention has achieved significant performance improvements on multiple datasets. The average PSNR has increased by 0.49 dB compared with other state-of-the-art methods. Especially on the Rain100L and Test1200 datasets, the PSNR has increased by 1.69 dB and 0.99 dB respectively. RaDA-Net has reached the top level among CNN-based methods and even shows strong competitiveness compared with Transformer-based methods. In particular, it maintains a better balance between computational cost and memory overhead and has excellent cost performance.

[0162] Table 1 Comparison of Computational Costs of Different Methods

[0163]

[0164] For the single-image rain removal task, computational cost is also an important factor to consider. Table 1 shows the computational costs of various methods, covering the number of trainable model parameters and the number of floating-point operations (FLOPs) of each method when processing images of size 256×256. From the data in the table, in terms of the number of parameters, the RaDA-Net model proposed in the present invention has 13.07M parameters, slightly lower than 13.35M of MSPFN and 20.1M of MPRNet with the same CNN structure, and 74% lower than 50.9M of Uformer-B based on Transformer. As for the number of floating-point operations (FLOPs), RaDA-Net is only 77.17G, significantly lower than 595.5G of MSPFN, 752.1G of SwinIR, 572.9G of MPRNet, and even lower than 89.46G of Uformer-B. This shows that the proposed method has a low computational cost while ensuring performance and is more cost-effective in practical applications.

[0165] In summary, the single-image rain removal method and network model designed in this application combine adaptive attention fusion, multi-granularity directional attention, bidirectional dynamic interaction aggregation, and hierarchical supervision mechanism to achieve precise capture and effective removal of rain streak features. Under the end-to-end training framework, the method uses a hybrid multi-scale loss function to optimize the model, ensuring the stability and convergence speed of multi-scale feature learning.

[0166] Through a multi-level position attention mechanism, this method can guide the model to accurately focus on the direction and distribution characteristics of rainwater. At the same time, combined with a hierarchical supervision mechanism, it generates stagewise rain-removed images at each level, calculates the loss earlier to finely guide the model learning, and thus accelerates the convergence of the model. This method significantly improves the phenomenon of insufficient rainwater feature extraction and loss of restored image details in traditional single-image rain-removal techniques, providing an efficient solution for the rain-removal task.

[0167] A large number of experiments show that compared with traditional rain-removal methods, this application not only improves the rain streak removal effect but also optimizes the computational efficiency, reducing the storage requirements while ensuring efficient operation. In addition, on multiple benchmark test sets, this method shows better PSNR and SSIM, and the reconstructed image texture is clearer and the visual perception quality is higher.

[0168] Figure 9 An image rain-removal device provided by an embodiment of this application is shown. The device includes:

[0169] A data construction module 910, configured to obtain several groups of image data pairs and construct a model training set and a test set; each group of image data pairs includes two rainy images and a rain-free image with the same picture content;

[0170] A model construction module 920, configured to construct multi-scale feature maps based on each group of image data pairs and construct a rain streak direction attention network according to the scales of the multi-scale feature maps;

[0171] A loss function calculation module 930, configured to calculate the mean square error loss, content loss, and edge loss at each level respectively according to the cascade structure of the encoding and decoding blocks, and construct a hybrid multi-scale loss function by combining the stagewise rain-removed images predicted at each level with the real rainy images in the validation set;

[0172] An optimization output module 940, configured to update the model parameters of the rain streak direction attention network based on the loss function value, and simulate and output a rain-removed image according to the rainy image after the iterative model training is completed.

[0173] The image rain-removal device provided by the embodiment of this application can be applied to the image rain-removal method provided in the above embodiment. For relevant details, refer to the above method embodiment. The implementation principle and technical effect are similar and will not be elaborated here.

[0174] It should be noted that the image de-raining device provided in the embodiments of the present application is only illustrated by the above division of each functional module / functional unit. In actual applications, the above functions can be allocated to different functional modules / functional units according to needs, that is, the internal structure of the image de-raining device is divided into different functional modules / functional units to complete all or part of the functions described above. In addition, the implementation manner of the image de-raining method provided in the above method embodiment and the implementation manner of the image de-raining device provided in this embodiment belong to the same concept. For the specific implementation process of the image de-raining device provided in this embodiment, please refer to the above method embodiment and will not be elaborated here.

[0175] Figure 10 The block diagram of the structure of a computer device provided by an exemplary embodiment of the present application is shown. It can be a desktop computer, a laptop computer, a handheld computer, a cloud server, or other computer devices. The computer device may include, but is not limited to, a processor and a memory. Among them, the processor and the memory can be connected through a bus or other means. Among them, the processor can be a central processing unit (CPU). The processor can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, graphics processing units (GPUs), embedded neural network processors (NPUs), or other dedicated deep learning coprocessors, discrete gate or transistor logic devices, discrete hardware components, etc. chips, or combinations of the above various types of chips.

[0176] The processor may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor may also include a main processor and a coprocessor. The main processor is a processor used to process data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.

[0177] The memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as the program instructions / modules corresponding to the methods in the above embodiments of the present application. The processor executes various functional applications and data processing of the processor by running the non-transitory software programs, instructions, and modules stored in the memory, that is, to implement the methods in the above method embodiments. The memory may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created by the processor, etc. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely set relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0178] The peripheral device interface can be used to connect at least one peripheral device related to I / O (Input / Output) to the processor and the memory. In some embodiments, the processor, the memory, and the peripheral device interface are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor, the memory, and the peripheral device interface can be implemented on a separate chip or circuit board, and this embodiment does not limit this.

[0179] The embodiments of the present application also disclose a computer-readable storage medium. Specifically, the computer-readable storage medium is used to store a computer program, and when the computer program is executed by a processor, the methods in the above method embodiments are implemented. Those skilled in the art can understand that to implement all or part of the processes in the above method embodiments of the present application, it can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above method embodiments. Among them, the storage medium can be a magnetic disk, an optical disc, a read-only memory (ROM), a random access memory (RAM), a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD), etc.; the storage medium can also include a combination of the above types of memories.

[0180] This specific embodiment is only an interpretation of the present invention, and it is not a limitation of the present invention. Those skilled in the art can make modifications without creative contributions to this embodiment as needed after reading this specification, but as long as it is within the scope of the claims of the present invention, it is protected by the patent law.

Claims

1. An image deraining method, characterized in that: The method comprises: Obtain several sets of image data pairs to construct model training sets and test sets; each set of image data pairs contains two rainy images and non-rainy images with the same picture content; A multi-scale feature map is constructed based on each set of image data pairs, and a rain streak direction attention network is constructed according to the scale of the multi-scale feature map; According to the cascade structure of the codec blocks, the mean square error loss, content loss and edge loss at each level are calculated respectively, and a hybrid multi-scale loss function is constructed by combining the predicted staged derained images at each level with the real rainy images in the validation set. The model parameters of the rain streak direction attention network are updated based on the loss function value. After the iterative model training is completed, the rain-free image is simulated and output based on the rainy image.

2. The method according to claim 1, characterized in that The rain streak directional attention network includes cascaded encoding blocks, bottleneck layers, and cascaded decoding blocks in sequence according to the data flow direction; wherein the cascaded encoding blocks and decoding blocks correspond one to one according to the levels, and the number of cascades is consistent with the number of multi-scale feature maps; Each encoding block contains a multi-granularity directional attention MGDA module, an encoding adaptive attention fusion AAIB module, a downsampling module, and an encoding bidirectional dynamic interaction aggregation BDIA module; The MGDA module generates a rain streak directional attention map based on the constructed multi-scale feature map and outputs it to the encoding AAIB module; the encoding AAIB module performs feature fusion based on the rain streak directional attention map and the feature map of the encoding output to generate an encoding fusion feature map; the encoding BDIA module constructs a multi-scale supervision structure based on the encoding fusion feature map and the output of the HSAM module in the corresponding decoding block; the downsampling module performs feature extraction and encoding output on the encoding fusion feature map; Each decoding block contains a decoding AAIB module, a hierarchical supervised attention HSAM module, an upsampling module, and a decoding BDIA module; The upsampling module receives the previous decoding input and performs upsampling output; the HSAM module outputs a staged derained image and a stage feature map based on the upsampling input and the rain mark directional attention map of the corresponding coding block; the stage feature map is spliced ​​with the intermediate feature map output by the corresponding level encoding BDIA and input into one input of the decoding AAIB module, and the rain mark directional attention map of the corresponding level is input into the other input of the decoding AAIB module, and the decoding AAIB module outputs a decoding fusion feature map; the decoding BDIA module at the subsequent stage generates an intermediate feature map based on the input decoding fusion feature map and the output of the previous stage decoding BDIA module.

3. The method according to claim 2, characterized in that The encoding AAIB module has the same structure as the decoding AAIB module, including a cascaded front-layer network and a back-layer network; The front layer network sequentially comprises a layer normalization unit, a point convolution layer, an expansion convolution layer, a simple gating unit, an adaptive attention integration AAIU unit, and a point convolution layer; The back-layer network sequentially comprises a layer normalization unit, a first splitting path and a second splitting path connected in parallel, and a point convolution layer; The AAIU unit generates an attention map according to the set learning rate parameter, and the back-layer network generates a fusion feature map based on the attention map and the rain streak direction attention map; The characteristic graph flow expression in the encoding AAIB module is as follows: F mid =F in +PConv(AAIU(SG(DConv(PConv(LN(F in ))))))×A rain F mid1 ,F mid2 =Split(LN(F mid )) F out =F mid +PConv(Dconv(PConv(F mid1 ))×SG(Dconv(PConv(F mid2 ))))×A rain Among them, F mid is the output of the previous network layer; PConv is point convolution; DConv is dilated convolution; AAIU is adaptive attention fusion unit; SG is simple gating operation; LN is layer normalization operation; Split is channel division, F mid1 , F mid2 Represents the feature map of the two paths after division; F out Represents the fused feature map.

4. The method according to claim 3, characterized in that The AAIU unit includes: The first branch includes an average pooling layer, a point convolution layer, a simple gating unit, a point convolution layer, and an activation layer in sequence; the average pooling layer of the first branch extracts global information, extracts channel-level feature distribution through point convolution and a simple gating unit, and outputs a channel attention map through the activation layer; The second branch includes a point convolution layer, a normalization layer, a point convolution layer, a simple gate unit, a point convolution layer and an activation layer in sequence; the second branch extracts spatial attention based on the normalization layer and the simple gate unit, and generates a first spatial attention map through the activation layer; The feature map flow expression in the AAIU unit is as follows: X AAIU =(α×A cha )×(β×A spa )×X f Among them, A cha is the channel attention map; A spa is the first spatial attention map; X AAIU is the output feature map of AAIU; P is the average pooling layer operation; is the Sigmoid activation function; is the Softmax activation function; α and β are learning parameters.

5. The method according to claim 2, characterized in that: The HSAM module includes a first branch and a second branch, which jointly input the upsampled feature map: The first branch includes the AAIU unit, the first point convolution layer, the second point convolution layer, and the Sigmoid activation layer in sequence; the upsampled feature map generates a residual image through the AAIU unit and the first point convolution layer; a rainy image of the corresponding level scale is added after the first point convolution layer to generate a staged derained image of the corresponding scale, and the staged derained image is subtracted from the corresponding rain-free image to generate various sub-loss functions at the corresponding scale; all sub-loss functions at all scales constitute a mixed multi-scale loss function; the staged derained image passes through the second point convolution layer and the sigmoid activation layer to generate a second spatial attention map; The second branch includes the third point convolution layer and the RELU activation layer in sequence; The feature map generated by the upsampled feature map after passing through the second path is multiplied pixel by pixel with the second spatial attention map and the rain streak direction attention map, and the residual between the upsampled input feature map is introduced to output the predicted feature map.

6. The method according to claim 2, characterized in that The BDIA module includes two feature inputs, one for encoding block features and the other for decoding block features; The BDIA module sequentially comprises a first point convolution layer, a connection layer, a simple gating unit, a pooling layer, and a second point convolution layer; The two inputs of the encoding block feature and the decoding block feature are respectively input into the first point convolution layer and spliced ​​through the connection layer; after splicing, they are split into channels by a simple gating unit and then multiplied, and then passed through the pooling layer for maximum pooling operation; the features after pooling and point convolution are multiplied element by element with the features before pooling, and then the residual between them and the encoding block features is introduced to output the intermediate feature map; The characteristic graph flow expression in the BDIA module is as follows: Where k is the level of the codec block, and represents the k-level encoding features and decoding features; MaxPool represents the maximum pooling operation; F temp It is output through simple gating logic, F fused is the intermediate feature map.

7. The method according to claim 2, characterized in that: The constructing of a hybrid multi-scale loss function includes: The mean square error loss, content loss and edge loss corresponding to each level of encoding block and decoding block are constructed respectively as follows: Among them, d represents the corresponding level or scale; represents the mean square error loss function; represents the content loss function; represents the marginal loss function; A d is the rain streak direction attention map predicted by the MGDA module at the corresponding scale; M d is a rain streak binary image generated by the threshold method; X d is the staged derained image at the corresponding scale, Y d is a rain-free image at the corresponding scale; ∈ is a smoothing parameter; Δ represents the Laplace operator; A hybrid multi-scale loss function L is constructed based on mean square error loss, content loss and edge loss. total ,as follows: The ω d represents the weight of content loss and edge loss at scale d; λ is used to control the relative importance of the two loss functions; η d Represents the weight of the mean squared error loss at scale d.

8. An image deraining device, characterized in that: The device comprises: A data construction module is used to obtain several sets of image data pairs to construct a model training set and a test set; each set of image data pairs contains two rainy images and non-rainy images with the same picture content; A model building module is used to build a multi-scale feature map based on each set of image data pairs, and to build a rain streak direction attention network according to the scale of the multi-scale feature map; The loss function calculation module is used to calculate the mean square error loss, content loss and edge loss at each level according to the cascade structure of the codec block, and to construct a hybrid multi-scale loss function by combining the predicted staged derained images at each level with the real rainy images in the validation set; The optimized output module is used to update the model parameters of the rain streak direction attention network based on the loss function value, and to simulate and output the rain-free image based on the rainy image after the iterative model training is completed.

9. A computer device, characterized in that: The computer device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the image deraining method according to any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that: The readable storage medium stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement the image deraining method according to any one of claims 1 to 8.