Image rain removal method, device and equipment and computer readable storage medium
The image deraining method, which uses a multi-scale structure and dual-frequency feature fusion mechanism, solves the problem of detail loss or blurring during the deraining process caused by the similarity between rain marks and image details, and achieves more accurate raindrop recognition and background detail preservation.
Patent Information
- Application Number
- CN202510667410.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-10-03
AI Technical Summary
Existing image deraining technologies easily lead to detail loss or blurring during the deraining process when processing the similarity of high-frequency features and geometric features between rain marks and image details.
A multi-scale contextual attention module and a multi-layer adaptive gated attention residual module are adopted, combined with a dual-frequency feature fusion mechanism. The multi-scale contextual attention module captures the localized channel features and spatial features of the image, and the dual-frequency feature fusion module dynamically adjusts the weights of different frequency bands to optimize the feature image.
It improves the accuracy of raindrop recognition and positioning, enhances the sensitivity and ability to distinguish different rainfall patterns, effectively decouples and calibrates image frequency, retains background details, and solves the problem of detail loss or blurring during rain removal.
Smart Images

Figure CN120746879A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of image deraining, and in particular to an image deraining method, apparatus, device, and computer-readable storage medium. Background Art
[0002] Image deraining is an important topic in computer vision. It aims to remove raindrops and other artifacts from rain-affected images, restoring clarity and visual quality. Rain streaks, noise, and image distortion introduced by rain not only degrade image quality but also impact the performance of vision systems in fields such as autonomous driving, surveillance, and outdoor photography.
[0003] In recent years, deep learning, particularly convolutional neural networks (CNNs), has made significant progress in image deraining. These methods typically learn the mapping between rain-affected images and clear images, removing raindrops and restoring background details. SPANet proposed a spatial attention network to remove rain streaks from a local to global perspective. DARGNet proposed a dual attention network that combines spatial and channel attention mechanisms to process raindrops and background regions separately. Uformer, drawing on the architectures of U-Net and transformer, proposed self-attention with non-overlapping windows to capture more global dependencies. With the development of deep learning technology, deraining methods have begun to integrate frequency analysis and deep neural networks. By training convolutional neural networks to identify and process specific frequency components, they learn which features in the frequency domain are associated with raindrops. Methods based on frequency channel attention mechanisms focus on processing the frequency features of images, enhancing feature subgraphs using frequency components to improve deraining results.
[0004] Under complex climatic conditions, image deraining tasks face the challenge of accurately identifying and removing rain streaks of various orientations and densities. The similarity between rain streaks and image details in high-frequency and geometric features often leads to technical issues such as loss or blurring of details during the deraining process. Summary of the Invention
[0005] The present application provides an image deraining method, apparatus, device, and computer-readable storage medium, which solve the technical problem of loss or blurring of details during the deraining process due to the similarity between rain marks and image details in high-frequency features and geometric features.
[0006] In view of this, a first aspect of the present application provides an image deraining method, the method comprising:
[0007] S1, obtain the original image;
[0008] S2. Perform shallow feature extraction on the original image using a preset feature extractor to obtain a shallow feature map of the original image;
[0009] S3, using a multi-scale contextual attention module to capture the localized channel features and expanded spatial features in the shallow feature map, and combining it with a multi-layer adaptive gated attention residual module to fuse the channel features, spatial features, and shallow feature map to generate a preliminary fused feature map;
[0010] S4. After capturing the high-frequency information and low-frequency information in the preliminary fusion feature map through the dual-frequency feature fusion module, the weights of different frequency bands in the preliminary fusion feature map are dynamically adjusted, and the features of the preliminary fusion feature map are optimized according to the updated weights to obtain a deep fusion feature map.
[0011] S5. Compare the deep fusion feature map with the original image to generate a derained image of the original image.
[0012] Optionally, the multi-scale context attention module specifically includes:
[0013] A channel feature extraction submodule consisting of three groups of 1×1 convolution modules and a 3×3 convolution module, and a spatial feature extraction submodule consisting of a 7×7 depthwise separable convolution module;
[0014] Among them, the query Q, key K, and the merged result of the second value V of the shallow feature map generated by the three groups of 1×1 convolution modules and 3×3 convolution modules of the channel feature extraction submodule are combined with the third value V of the depth-separable convolution module to obtain a preliminary fused feature map.
[0015] Optionally, step S3 specifically includes:
[0016] S31. Input the shallow feature map into the residual block of the multi-layer adaptive gated attention to obtain the input vector X∈R H×W×C ;
[0017] S32, after mapping the input vector X to three times the channel dimension through three groups of 1×1 convolution modules of the channel feature extraction submodule, the channel-level information is encoded through three groups of 3×3 convolution modules of the channel feature extraction submodule to generate the query Q, key K, and first value V of the shallow feature map respectively;
[0018] S33, after normalizing the query Q and key K, generate the attention weight by calculating the matrix product;
[0019] S34. Perform a weighted summation of the attention weight and the first value V to obtain a second value V of the attention mechanism;
[0020] S35. Add the third value V output by the 7×7 depthwise separable convolution module of the spatial feature extraction submodule of the input vector X to the second value V of the attention mechanism, and obtain a preliminary fused feature map through 1×1 convolution.
[0021] Optionally, step S4 specifically includes:
[0022] S41, extracting low-frequency information in the preliminary fusion feature map through global average pooling;
[0023] S42, obtaining high-frequency information in the preliminary fusion feature map by calculating the difference between the shallow feature map and the low-frequency information;
[0024] S43. After adding the low-frequency information and the high-frequency information element by element, dynamically adjust the weights of different frequency bands in the preliminary fusion feature map, and optimize the features of the preliminary fusion feature map according to the updated weights to obtain a deep fusion feature map.
[0025] A second aspect of the present application provides an image deraining device, comprising:
[0026] An acquisition unit, used for acquiring an original image;
[0027] A shallow feature extraction unit is used to extract shallow features of the original image using a preset feature extractor to obtain a shallow feature map of the original image;
[0028] The contextual attention unit is used to capture the localized channel features and expanded spatial features in the shallow feature map through a multi-scale contextual attention module, and fuse the channel features, spatial features and shallow feature map with the residual module of multi-layer adaptive gated attention to generate a preliminary fused feature map;
[0029] The dual-frequency feature fusion unit is used to capture the high-frequency information and low-frequency information in the preliminary fusion feature map through the dual-frequency feature fusion module, dynamically adjust the weights of different frequency bands in the preliminary fusion feature map, and optimize the features of the preliminary fusion feature map according to the updated weights to obtain a deep fusion feature map;
[0030] The image deraining unit is used to compare the deep fusion feature map with the original image to generate a derained image of the original image.
[0031] Optionally, the context attention unit is specifically configured to:
[0032] Input the shallow feature map into the residual block of the multi-layer adaptive gated attention to obtain the input vector X∈R H ×W×C ;
[0033] After mapping the input vector X to three times the channel dimension through three sets of 1×1 convolution modules of the channel feature extraction submodule, the channel-level information is encoded through three sets of 3×3 convolution modules of the channel feature extraction submodule to generate the query Q, key K, and first value V of the shallow feature map respectively;
[0034] After normalizing the query Q and key K, the attention weight is generated by calculating the matrix product;
[0035] The attention weight is weighted and summed with the first value V to obtain the second value V of the attention mechanism;
[0036] The third value V output by the 7×7 depth-wise separable convolution module of the spatial feature extraction submodule of the input vector X is added to the second value V of the attention mechanism, and a preliminary fusion feature map is obtained through 1×1 convolution.
[0037] Optionally, the dual-frequency feature fusion unit is specifically configured to:
[0038] Extract low-frequency information from the preliminary fusion feature map through global average pooling;
[0039] By calculating the difference between the shallow feature map and the low-frequency information, the high-frequency information in the preliminary fusion feature map is obtained;
[0040] After adding the low-frequency information and the high-frequency information element by element, the weights of different frequency bands in the preliminary fusion feature map are dynamically adjusted, and the preliminary fusion feature map is optimized according to the updated weights to obtain a deep fusion feature map.
[0041] A third aspect of the present application provides an image deraining device, the device comprising a processor and a memory:
[0042] The memory is used to store program code and transmit the program code to the processor;
[0043] The processor is configured to execute the steps of the image deraining method as described in the first aspect according to the instructions in the program code.
[0044] A fourth aspect of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium is used to store program code, and the program code is used to execute the method described in the first aspect.
[0045] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages:
[0046] In this application, an image deraining method, apparatus, device, and computer-readable storage medium are provided. Based on a contextual attention mechanism, while processing the channel dimension information of the image, spatial dimension considerations are also introduced, which can more accurately identify and locate raindrops, enhance sensitivity and discrimination capabilities for different rainfall patterns, and further achieve effective decoupling and calibration of image frequencies through a dual-frequency feature fusion mechanism, thereby enhancing the retention of background details during the deraining process and solving the technical problem of detail loss or blurring during the deraining process due to the similarity between rain marks and image details in high-frequency and geometric features. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 This is a flow chart of the image deraining method in an embodiment of the present application;
[0048] Figure 2 This is a structural diagram of an image rain removal device in an embodiment of the present application;
[0049] Figure 3 This is a structural diagram of an image deraining device in an embodiment of the present application;
[0050] Figure 4 This is a schematic diagram of the framework of the image deraining model in the embodiment of the present application;
[0051] Figure 5 This is a structural diagram of the context attention module in an embodiment of the present application;
[0052] Figure 6 Schematic diagram of the structure of the dual-frequency feature fusion module in the embodiment of the present application. DETAILED DESCRIPTION
[0053] In order to help those skilled in the art better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this application.
[0054] This application designs an image deraining method, apparatus, device and computer-readable storage medium to solve the technical problem of loss or blurring of details during the deraining process due to the similarity between rain marks and image details in high-frequency features and geometric features.
[0055] For easier understanding, see Figure 1 , Figure 1 This is a flow chart of the image deraining method according to an embodiment of the present application. Figure 1 As shown, specifically:
[0056] S1, obtain the original image;
[0057] It should be noted that the original images can come from a variety of common image acquisition devices, such as digital cameras and mobile phone cameras. When these images are taken on rainy days, they exhibit rain streaks, noise, and image distortion due to the interference of rain, resulting in a decrease in image quality, affecting the visual effect and subsequent image analysis tasks.
[0058] For example, suppose we want to remove rain from an outdoor street scene photo taken on a rainy day. The photo is dark overall and has obvious rain marks. The outlines of buildings, roads, and pedestrians are blurred by the rain. This is the original image we want to process.
[0059] S2. Perform shallow feature extraction on the original image using a preset feature extractor to obtain a shallow feature map of the original image;
[0060] It should be noted that the preset feature extractor is usually composed of convolutional layers in a convolutional neural network (CNN). Its purpose is to extract shallow feature maps from the original image. These shallow feature maps contain some basic image features, such as edges and textures, and provide a foundation for subsequent more in-depth feature processing and analysis.
[0061] For example, consider feeding the aforementioned street view photo into a pre-defined feature extractor, such as a simple convolutional neural network consisting of convolutional layers. After the convolution operation, the resulting shallow feature map likely highlights basic features such as edges and lines. These features serve as a preliminary interpretation of the original image, outlining the general outlines and basic shapes of objects.
[0062] S3, using a multi-scale contextual attention module to capture the localized channel features and expanded spatial features in the shallow feature map, and combining it with a multi-layer adaptive gated attention residual module to fuse the channel features, spatial features, and shallow feature map to generate a preliminary fused feature map;
[0063] Furthermore, the multi-scale context attention module specifically includes:
[0064] A channel feature extraction submodule consisting of three groups of 1×1 convolution modules and a 3×3 convolution module, and a spatial feature extraction submodule consisting of a 7×7 depthwise separable convolution module;
[0065] Among them, the query Q, key K, and the merged result of the second value V of the shallow feature map generated by the three groups of 1×1 convolution modules and 3×3 convolution modules of the channel feature extraction submodule are combined with the third value V of the depth-separable convolution module to obtain a preliminary fused feature map.
[0066] It's important to note that the channel feature extraction submodule, consisting of three groups of 1×1 and 3×3 convolutional modules, captures localized channel features through operations on convolution kernels of different sizes. The 1×1 convolution can combine and filter channels, changing their dimensions and extracting inter-channel correlations; the 3×3 convolution can capture spatial features within a certain range. Together, they provide a better understanding of raindrop channel characteristics, such as the directionality and density distribution of rain streaks.
[0067] The spatial feature extraction submodule, consisting of a 7×7 depthwise separable convolution module, uses larger convolution kernels to expand the receptive field, capturing the spatial interactions between raindrops and background objects over a wider area. This allows for more comprehensive spatial features, helping to understand the distribution of raindrops throughout the image and their relative positions to background objects.
[0068] Furthermore, the step S3 specifically includes:
[0069] S31. Input the shallow feature map into the residual block of the multi-layer adaptive gated attention to obtain the input vector X∈R H×W×C ;
[0070] It's important to note that the residual block is a special neural network structure that introduces bypass connections, allowing the input to be passed directly to the output and added to the transformed features. This structure helps alleviate the vanishing gradient problem in deep network training, allowing the network to more effectively learn useful features. Here, by inputting shallow feature maps into it, we can perform preliminary feature integration and optimization, resulting in an input vector more suitable for subsequent processing.
[0071] S32, after mapping the input vector X to three times the channel dimension through three groups of 1×1 convolution modules of the channel feature extraction submodule, the channel-level information is encoded through three groups of 3×3 convolution modules of the channel feature extraction submodule to generate the query Q, key K, and first value V of the shallow feature map respectively;
[0072] It's important to note that mapping the input vector to a higher channel dimension through the 1×1 convolution module increases the network's ability to express features, providing more channel resources for subsequent feature extraction. The subsequent 3×3 convolution module performs spatial convolution operations on these high-dimensional channel features, encoding channel-level information and extracting more meaningful features for generating query Q, key K, and value V. This forms the basis of the attention mechanism, helping to better focus on raindrop features.
[0073] S33, after normalizing the query Q and key K, generate the attention weight by calculating the matrix product;
[0074] It's important to note that normalization is performed to keep the values of query Q and key K within an appropriate range, facilitating the subsequent matrix product calculation. The matrix product operation can be viewed as a similarity measure between query Q and key K, reflecting the degree of similarity between different positions in the feature space. The Softmax function converts these similarity scores into probabilistic weights, ensuring that the sum of the weights for each row is 1. This better represents the relative importance of different positions and provides a basis for the subsequent weighted summation operation.
[0075] S34. Perform a weighted summation of the attention weight and the first value V to obtain a second value V of the attention mechanism;
[0076] It's important to note that the weighted summation process actually combines features at different locations in the first value V according to attention weights, with features at locations with higher weights contributing more to the result. This is equivalent to redistributing and integrating the features of the input vector, allowing the network to focus more on key features related to raindrops, thereby enhancing the representation of raindrop characteristics and preparing for subsequent feature fusion.
[0077] S35. Add the third value V output by the 7×7 depthwise separable convolution module of the spatial feature extraction submodule of the input vector X to the second value V of the attention mechanism, and obtain a preliminary fused feature map through 1×1 convolution.
[0078] It's important to note that adding the third value V and the second value V simply fuses the spatial information obtained by the spatial feature extraction submodule with the channel-wise feature information processed by the attention mechanism. These two types of information complement each other. These features are then further integrated through 1×1 convolutions. Adjusting the number of channels, for example, makes the initial fused feature map more compact and effective, better representing the characteristics of raindrops in the image and their relationship to the background.
[0079] S4. After capturing the high-frequency information and low-frequency information in the preliminary fusion feature map through the dual-frequency feature fusion module, the weights of different frequency bands in the preliminary fusion feature map are dynamically adjusted, and the features of the preliminary fusion feature map are optimized according to the updated weights to obtain a deep fusion feature map.
[0080] It's important to note that the core of the dual-frequency feature fusion module lies in the separate processing and fusion of high-frequency and low-frequency features of an image. Low-frequency features primarily reflect the overall structure and background information of the image, while high-frequency features contain image details and edge information, such as the texture of raindrops. Global average pooling extracts low-frequency information to reveal the overall background structure of the image; calculating the difference between the high-frequency information highlights detailed features such as raindrops. Combining these two methods and dynamically adjusting the weights achieves a better balance between raindrop removal and background preservation, effectively removing raindrop features while preserving background details.
[0081] Furthermore, the step S4 specifically includes:
[0082] S41, extracting low-frequency information in the preliminary fusion feature map through global average pooling;
[0083] It's important to note that global average pooling is a simple feature extraction method that compresses the feature map of each channel into a single value by calculating the global average value across each channel. This helps extract low-frequency information in the image, namely the overall background information, while ignoring local details and variations, providing a reference for background structure for subsequent processing.
[0084] S42, obtaining high-frequency information in the preliminary fusion feature map by calculating the difference between the shallow feature map and the low-frequency information;
[0085] It's important to note that by calculating the difference between the shallow feature map and the low-frequency information, we can obtain the portion of the image that deviates from the overall background characteristics, namely the high-frequency information. This high-frequency information mainly contains details and changes in the image, such as object edges and textures. The characteristics of raindrops are also largely reflected in this high-frequency information.
[0086] S43. After adding the low-frequency information and the high-frequency information element by element, dynamically adjust the weights of different frequency bands in the preliminary fusion feature map, and optimize the features of the preliminary fusion feature map according to the updated weights to obtain a deep fusion feature map.
[0087] S5. Compare the deep fusion feature map with the original image to generate a derained image of the original image.
[0088] It's important to note that the deep fusion feature map contains optimized raindrop and background feature information. By comparing it with the original image, these features can be transformed into the final derained image using a specific image reconstruction algorithm or network structure. This process is equivalent to remapping the processed feature information back into the image space, generating a qualified derained image.
[0089] See also Figure 2, Figure 2 This is a structural diagram of an image rain removal device according to an embodiment of the present application. Figure 2 As shown, specifically:
[0090] An acquisition unit 201 is used to acquire an original image;
[0091] A shallow feature extraction unit 202 is configured to extract shallow features from the original image using a preset feature extractor to obtain a shallow feature map of the original image;
[0092] The context attention unit 203 is used to capture the localized channel features and expanded spatial features in the shallow feature map through a multi-scale context attention module, and fuse the channel features, spatial features and shallow feature map with the residual module of the multi-layer adaptive gated attention to generate a preliminary fused feature map;
[0093] The dual-frequency feature fusion unit 204 is configured to capture high-frequency information and low-frequency information in the preliminary fusion feature map through the dual-frequency feature fusion module, dynamically adjust the weights of different frequency bands in the preliminary fusion feature map, and optimize the preliminary fusion feature map according to the updated weights to obtain a deep fusion feature map.
[0094] The image deraining unit 205 is used to compare the depth fusion feature map with the original image to generate a derained image of the original image.
[0095] Furthermore, the context attention unit is specifically used to:
[0096] Input the shallow feature map into the residual block of the multi-layer adaptive gated attention to obtain the input vector X∈R H ×W×C ;
[0097] After mapping the input vector X to three times the channel dimension through three sets of 1×1 convolution modules of the channel feature extraction submodule, the channel-level information is encoded through three sets of 3×3 convolution modules of the channel feature extraction submodule to generate the query Q, key K, and first value V of the shallow feature map respectively;
[0098] After normalizing the query Q and key K, the attention weight is generated by calculating the matrix product;
[0099] The attention weight is weighted and summed with the first value V to obtain the second value V of the attention mechanism;
[0100] The third value V output by the 7×7 depth-wise separable convolution module of the spatial feature extraction submodule of the input vector X is added to the second value V of the attention mechanism, and a preliminary fusion feature map is obtained through 1×1 convolution.
[0101] Furthermore, the dual-frequency feature fusion unit is specifically used to:
[0102] Extract low-frequency information from the preliminary fusion feature map through global average pooling;
[0103] By calculating the difference between the shallow feature map and the low-frequency information, the high-frequency information in the preliminary fusion feature map is obtained;
[0104] After adding the low-frequency information and the high-frequency information element by element, the weights of different frequency bands in the preliminary fusion feature map are dynamically adjusted, and the preliminary fusion feature map is optimized according to the updated weights to obtain a deep fusion feature map.
[0105] The present application also provides another image deraining device, such as Figure 3 For ease of explanation, only the parts related to the embodiments of the present application are shown. For specific technical details not disclosed, please refer to the method part of the embodiments of the present application. The terminal can be any terminal device including a mobile phone, a tablet computer, a personal digital assistant (English full name: Personal Digital Assistant, English abbreviation: PDA), a sales terminal (English full name: Point of Sales, English abbreviation: POS), a car computer, etc., taking the mobile phone as an example:
[0106] Figure 3 The block diagram shows a partial structure of a mobile phone related to the terminal provided in the embodiment of the present application. Figure 3 The mobile phone includes components such as a radio frequency (RF) circuit 1010, a memory 1020, an input unit 1030, a display unit 1040, a sensor 1050, an audio circuit 1060, a wireless fidelity (WiFi) module 1070, a processor 1080, and a power supply 1090. Those skilled in the art will appreciate that Figure 3 The mobile phone structure shown in the figure does not constitute a limitation to the mobile phone, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.
[0107] The following combination Figure 3 A detailed introduction to the various components of a mobile phone:
[0108] RF circuitry 1010 can be used to receive and transmit signals during information transmission or calls. Specifically, it receives downlink information from the base station and transmits it to processor 1080 for processing. It also transmits uplink data to the base station. Typically, RF circuitry 1010 includes, but is not limited to, an antenna, at least one amplifier, a transceiver, a coupler, a low-noise amplifier (LNA), a duplexer, and more. RF circuitry 1010 can also communicate with the network and other devices via wireless communications. The above-mentioned wireless communications may use any communication standard or protocol, including but not limited to Global System of Mobile Communications (Global System of Mobile communication, English abbreviation: GSM), General Packet Radio Service (English full name: General Packet Radio Service, GPRS), Code Division Multiple Access (English full name: Code Division Multiple Access, English abbreviation: CDMA), Wideband Code Division Multiple Access (English full name: Wideband Code Division Multiple Access, English abbreviation: WCDMA), Long Term Evolution (English full name: Long Term Evolution, English abbreviation: LTE), email, Short Messaging Service (English full name: Short Messaging Service, SMS), etc.
[0109] The memory 1020 can be used to store software programs and modules. The processor 1080 executes the various functional applications and data processing of the mobile phone by running the software programs and modules stored in the memory 1020. The memory 1020 may mainly include a program storage area and a data storage area. The program storage area may store an operating system and at least one application required for a function (such as a sound playback function, an image playback function, etc.); the data storage area may store data created based on the use of the mobile phone (such as audio data, a phone book, etc.). In addition, the memory 1020 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state memory device.
[0110] The input unit 1030 can be used to receive input digital or character information, and to generate key signal input related to the user settings and function control of the mobile phone. Specifically, the input unit 1030 may include a touch panel 1031 and other input devices 1032. The touch panel 1031, also known as a touch screen, can collect user touch operations on or near it (such as operations performed by the user using any suitable object or accessory such as a finger, stylus, etc. on or near the touch panel 1031) and drive the corresponding connection device according to a pre-set program. Optionally, the touch panel 1031 may include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the user's touch direction and detects the signal caused by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device and converts it into touch point coordinates, which are then sent to the processor 1080. It can also receive commands sent by the processor 1080 and execute them. In addition, the touch panel 1031 can be implemented using various types such as resistive, capacitive, infrared and surface acoustic wave. In addition to the touch panel 1031, the input unit 1030 may further include other input devices 1032. Specifically, the other input devices 1032 may include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and a joystick.
[0111] The display unit 1040 can be used to display information input by the user or information provided to the user and various menus of the mobile phone. The display unit 1040 may include a display panel 1041. Optionally, the display panel 1041 may be configured in the form of a liquid crystal display (English full name: Liquid Crystal Display, English abbreviation: LCD), an organic light-emitting diode (English full name: Organic Light-Emitting Diode, English abbreviation: OLED), etc. Further, the touch panel 1031 may cover the display panel 1041. When the touch panel 1031 detects a touch operation on or near it, it is transmitted to the processor 1080 to determine the type of touch event. Subsequently, the processor 1080 provides corresponding visual output on the display panel 1041 according to the type of touch event. Although in Figure 3 In the embodiment, the touch panel 1031 and the display panel 1041 are used as two independent components to realize the input and output functions of the mobile phone, but in some embodiments, the touch panel 1031 and the display panel 1041 can be integrated to realize the input and output functions of the mobile phone.
[0112] The mobile phone may also include at least one sensor 1050, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor, wherein the ambient light sensor may adjust the brightness of the display panel 1041 according to the brightness of the ambient light, and the proximity sensor may turn off the display panel 1041 and / or the backlight when the mobile phone is moved to the ear. As a type of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in all directions (generally three axes), and can detect the magnitude and direction of gravity when stationary. It can be used for applications that identify the posture of the mobile phone (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc.; as for other sensors that the mobile phone can also be configured with, such as gyroscopes, barometers, hygrometers, thermometers, infrared sensors, etc., they will not be described here.
[0113] Audio circuit 1060, speaker 1061, and microphone 1062 provide an audio interface between the user and the phone. Audio circuit 1060 converts received audio data into electrical signals and transmits them to speaker 1061, which then converts them into sound signals for output. Microphone 1062, on the other hand, converts collected sound signals into electrical signals, which are then received by audio circuit 1060 and converted into audio data. The audio data is then processed by processor 1080 and transmitted to, for example, another phone via RF circuit 1010, or stored in memory 1020 for further processing.
[0114] WiFi is a short-range wireless transmission technology. The mobile phone can help users send and receive emails, browse the web and access streaming media through the WiFi module 1070. It provides users with wireless broadband Internet access. Figure 3 A WiFi module 1070 is shown, but it is understandable that it is not an essential component of the mobile phone and can be omitted as needed without changing the essence of the invention.
[0115] Processor 1080 is the control center of the mobile phone, connecting all parts of the mobile phone using various interfaces and circuits. By running or executing software programs and / or modules stored in memory 1020 and accessing data stored in memory 1020, it performs various functions of the mobile phone and processes data, thereby providing overall monitoring of the mobile phone. Optionally, processor 1080 may include one or more processing units; preferably, processor 1080 may integrate an application processor and a modem processor, wherein the application processor primarily handles the operating system, user interface, and application programs, while the modem processor primarily handles wireless communications. It is understood that the modem processor may not be integrated into processor 1080.
[0116] The mobile phone also includes a power supply 1090 (such as a battery) for supplying power to various components. Preferably, the power supply can be logically connected to the processor 1080 through a power management system, thereby managing charging, discharging, and power consumption through the power management system.
[0117] Although not shown, the mobile phone may also include a camera, a Bluetooth module, etc., which will not be described in detail here.
[0118] In the embodiment of the present application, the processor 1080 included in the terminal further has the following functions:
[0119] S1, obtain the original image;
[0120] S2. Perform shallow feature extraction on the original image using a preset feature extractor to obtain a shallow feature map of the original image;
[0121] S3, using a multi-scale contextual attention module to capture the localized channel features and expanded spatial features in the shallow feature map, and combining it with a multi-layer adaptive gated attention residual module to fuse the channel features, spatial features, and shallow feature map to generate a preliminary fused feature map;
[0122] S4. After capturing the high-frequency information and low-frequency information in the preliminary fusion feature map through the dual-frequency feature fusion module, the weights of different frequency bands in the preliminary fusion feature map are dynamically adjusted, and the features of the preliminary fusion feature map are optimized according to the updated weights to obtain a deep fusion feature map.
[0123] S5. Compare the deep fusion feature map with the original image to generate a derained image of the original image.
[0124] An embodiment of the present application further provides a computer-readable storage medium for storing program code, wherein the program code is used to execute any one of the implementations of the image deraining method described in the aforementioned embodiments.
[0125] In an embodiment of the present application, a method, apparatus, device, and computer-readable storage medium for image deraining are provided. Based on a contextual attention mechanism, while processing the channel dimension information of the image, spatial dimension considerations are also introduced, which can more accurately identify and locate raindrops, enhance sensitivity and differentiation capabilities for different rainfall patterns, and further achieve effective decoupling and calibration of image frequencies through a dual-frequency feature fusion mechanism, thereby enhancing the retention of background details during the deraining process and solving the technical problem of detail loss or blurring during the deraining process due to the similarity between rain marks and image details in high-frequency and geometric features.
[0126] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0127] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can, for example, be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0128] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0129] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0130] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0131] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0132] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (English full name: Read-Only Memory, English abbreviation: ROM), random access memory (English full name: Random Access Memory, English abbreviation: RAM), disk or optical disk and other media that can store program code.
[0133] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. An image deraining method, characterized in that: include: S1, obtain the original image; S2. Perform shallow feature extraction on the original image using a preset feature extractor to obtain a shallow feature map of the original image; S3, using a multi-scale contextual attention module to capture the localized channel features and expanded spatial features in the shallow feature map, and combining it with a multi-layer adaptive gated attention residual module to fuse the channel features, spatial features, and shallow feature map to generate a preliminary fused feature map; S4. After capturing the high-frequency information and low-frequency information in the preliminary fusion feature map through the dual-frequency feature fusion module, the weights of different frequency bands in the preliminary fusion feature map are dynamically adjusted, and the features of the preliminary fusion feature map are optimized according to the updated weights to obtain a deep fusion feature map. S5. Compare the deep fusion feature map with the original image to generate a derained image of the original image.
2. The image deraining method according to claim 1, characterized in that: The multi-scale context attention module specifically includes: A channel feature extraction submodule consisting of three groups of 1×1 convolution modules and a 3×3 convolution module, and a spatial feature extraction submodule consisting of a 7×7 depthwise separable convolution module; Among them, the query Q, key K, and the merged result of the second value V of the shallow feature map generated by the three groups of 1×1 convolution modules and 3×3 convolution modules of the channel feature extraction submodule are combined with the third value V of the depth-separable convolution module to obtain a preliminary fused feature map.
3. The image deraining method according to claim 2, characterized in that: The step S3 specifically includes: S31. Input the shallow feature map into the residual block of the multi-layer adaptive gated attention to obtain the input vector X∈R H ×W×C ; S32, after mapping the input vector X to three times the channel dimension through three groups of 1×1 convolution modules of the channel feature extraction submodule, the channel-level information is encoded through three groups of 3×3 convolution modules of the channel feature extraction submodule to generate the query Q, key K, and first value V of the shallow feature map respectively; S33, after normalizing the query Q and key K, generate the attention weight by calculating the matrix product; S34. Perform a weighted summation of the attention weight and the first value V to obtain a second value V of the attention mechanism; S35. Add the third value V output by the 7×7 depthwise separable convolution module of the spatial feature extraction submodule of the input vector X to the second value V of the attention mechanism, and obtain a preliminary fused feature map through 1×1 convolution.
4. The image deraining method according to claim 1, characterized in that: The step S4 specifically includes: S41, extracting low-frequency information in the preliminary fusion feature map through global average pooling; S42, obtaining high-frequency information in the preliminary fusion feature map by calculating the difference between the shallow feature map and the low-frequency information; S43. After adding the low-frequency information and the high-frequency information element by element, dynamically adjust the weights of different frequency bands in the preliminary fusion feature map, and optimize the features of the preliminary fusion feature map according to the updated weights to obtain a deep fusion feature map.
5. An image deraining device, characterized in that: include: An acquisition unit, used for acquiring an original image; A shallow feature extraction unit is used to extract shallow features of the original image using a preset feature extractor to obtain a shallow feature map of the original image; The contextual attention unit is used to capture the localized channel features and expanded spatial features in the shallow feature map through a multi-scale contextual attention module, and fuse the channel features, spatial features and shallow feature map with the residual module of multi-layer adaptive gated attention to generate a preliminary fused feature map; The dual-frequency feature fusion unit is used to capture the high-frequency information and low-frequency information in the preliminary fusion feature map through the dual-frequency feature fusion module, dynamically adjust the weights of different frequency bands in the preliminary fusion feature map, and optimize the features of the preliminary fusion feature map according to the updated weights to obtain a deep fusion feature map; The image deraining unit is used to compare the deep fusion feature map with the original image to generate a derained image of the original image.
6. The image rain removal device according to claim 6, characterized in that: The context attention unit is specifically used to: Input the shallow feature map into the residual block of the multi-layer adaptive gated attention to obtain the input vector X∈R H×W×C ; After mapping the input vector X to three times the channel dimension through three sets of 1×1 convolution modules of the channel feature extraction submodule, the channel-level information is encoded through three sets of 3×3 convolution modules of the channel feature extraction submodule to generate the query Q, key K, and first value V of the shallow feature map respectively; After normalizing the query Q and key K, the attention weight is generated by calculating the matrix product; The attention weight is weighted and summed with the first value V to obtain the second value V of the attention mechanism; The third value V output by the 7×7 depth-wise separable convolution module of the spatial feature extraction submodule of the input vector X is added to the second value V of the attention mechanism, and a preliminary fusion feature map is obtained through 1×1 convolution.
7. The image rain removal device according to claim 5, characterized in that: The dual-frequency feature fusion unit is specifically used for: Extract low-frequency information from the preliminary fusion feature map through global average pooling; By calculating the difference between the shallow feature map and the low-frequency information, the high-frequency information in the preliminary fusion feature map is obtained; After adding the low-frequency information and the high-frequency information element by element, the weights of different frequency bands in the preliminary fusion feature map are dynamically adjusted, and the preliminary fusion feature map is optimized according to the updated weights to obtain a deep fusion feature map.
8. An image deraining device, characterized in that: The device includes a processor and a memory: The memory is used to store program code and transmit the program code to the processor; The processor is configured to execute the image deraining method according to any one of claims 1 to 4 according to instructions in the program code.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium is used to store program code, and the program code is used to execute the image deraining method according to any one of claims 1 to 4.