A Single Image Rain Removal Method Based on Spatial-Frequency Co-modeling
Patent Information
- Application Number
- CN202610832305.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-10
- Publication Date
- 2026-09-01
AI Technical Summary
[0008]本发明的目的在于提供一种基于空间-频率协同建模的单张图像去雨方法,以解决上述背景技术中提出的现有单张图像去雨技术中背景提取网络空间局限大、计算代价高昂,导致背景高频结构纹理极易与线状雨痕产生混淆、误伤的问题,以及现有技术依赖于像素级加减进行解耦拆分,缺乏局部非线性退化动态感知,极易在重雨、边缘或高光区域引发严重的纹理缺失、伪影及色调扭曲的问题
本发明利用二维快速实数傅里叶变换开辟频域路径GFEM,在频域中直接拦截和捕获宏观的全局低频无雨背景轮廓,从物理机制上规避了空间域雨痕高频成分对背景恢复的干扰,完美保护了物体的边界与高频纹理;
Smart Images

Figure CN122675701A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision and digital image processing technology, specifically to a single image deraining method based on spatial-frequency co-modeling. Background Technology
[0002] In outdoor vision systems, single-image rain removal is a crucial digital image preprocessing technique. Severe weather, primarily characterized by rainfall, causes extensive degradation in images acquired by outdoor imaging devices. Rain streaks and rain curtains not only severely obscure object edges and details but also introduce complex nonlinear degradation, leading to a sharp decline in the accuracy of subsequent advanced computer vision tasks such as object detection and semantic segmentation. Therefore, efficiently and completely removing rain streak noise and restoring high-fidelity rain-free backgrounds has significant commercial value and practical application implications.
[0003] Traditional image deraining methods primarily rely on manually designed statistical priors, which are less robust to multi-scale, high-density, or complex heavy rain scenes, easily leading to incomplete deraining or over-smoothing of the background structure. In recent years, with the development of deep learning, image restoration techniques based on convolutional neural networks and self-attention mechanisms have made significant progress. Since rain streaks typically manifest as highly directional, high-density local high-frequency noise in the spatial domain, traditional spatial domain methods often attempt to construct deep spatial convolutional or self-attention matrices to spatially separate rain streaks from the background.
[0004] However, CNNs are limited by a finite receptive field, and the global matrix calculation of Transformers based on the self-attention mechanism leads to extremely high computational complexity and memory consumption. This makes it easy for existing technologies to face technical bottlenecks such as excessive computational overhead, loss of detail texture, and deformation of background structure when dealing with complex multi-scale rain streak degradation in the spatial domain.
[0005] One existing technology is a typical single-path direct mapping image deraining scheme based on traditional convolutional neural networks. It directly feeds the rainy image into the spatial path for continuous convolution and feature stacking, and then outputs the restored background image or predicted rain streak residuals. This scheme has extremely weak stripping ability under complex nonlinear degradation. When rain streaks are tightly intertwined with the background, the network tends to erase the object edges and key high-frequency textures of the background itself in order to remove the rain streaks, resulting in large areas of over-smoothing in the derained image. At the same time, it is designed entirely for rain streak residuals of a specific form and lacks generalization and robustness in multi-source degradation scenarios.
[0006] Existing technique two is a typical lightweight rain removal method that uses a purely spatial domain approach. It typically introduces a Laplacian pyramid to downsample the image multiple times in the spatial domain and uses a shallow convolutional network with very few parameters to independently denoise sub-band features at each scale. While this approach simplifies the network structure and the number of feature channels to achieve lightweight design, it lacks the ability to model complex, dense rain streaks in a non-local, global manner due to the use of shallow spatial convolutions. This results in incomplete rain removal and the easy retention of noticeable rain streak artifacts. Furthermore, during the upsampling and restoration process in the spatial domain pyramid, the denoising features are prone to geometric misalignment, leading to severe blurring or distortion of previously clear object edges and high-frequency textures in the recovered rain-free image.
[0007] Existing technology three is a multi-scale fusion and decomposition network for deraining and low-level visual image restoration of a single image. It constructs an asymmetric dual-path mutual representation network: the rain streak flow branch uses downsampling and cascaded spatial Transformer blocks in the spatial domain to focus on complex and varied non-local rain streak patterns; the background flow branch is completely confined to the traditional spatial domain, using ordinary CNN convolutions and channel attention blocks to encode local background content. This scheme separates the background by forcing pixel-level matrix point-to-point subtraction and addition operations between the two paths through coupled representation blocks. However, its background path is completely confined to the traditional spatial domain, resulting in high computational cost and inefficient acquisition of a global receptive field covering the entire image. More critically, rain streaks exhibit linear texture in the spatial domain. When there are naturally occurring high-frequency structures in the background with similar direction and scale, the mechanism easily confuses the high-frequency background with the high-frequency rain streaks, leading to accidental damage to the background. Furthermore, it relies on pixel moments. Direct addition and subtraction of arrays severely lacks awareness of local nonlinear spatial entanglement in the image. In areas of heavy rain, highlights, or image edges, this can easily lead to incomplete subtraction, resulting in severe image distortion, tone distortion, and edge blur artifacts. Summary of the Invention
[0008] The purpose of this invention is to provide a single-image deraining method based on spatial-frequency co-modeling, in order to solve the problems of existing single-image deraining techniques mentioned above, such as the large spatial limitation and high computational cost of the background extraction network, which makes it easy for high-frequency background textures to be confused with linear rain streaks and cause accidental damage; and the problems of existing techniques relying on pixel-level addition and subtraction for decoupling and decomposition, lacking dynamic perception of local nonlinear degradation, which easily causes serious texture loss, artifacts and tone distortion in heavy rain, edge or highlight areas.
[0009] Therefore, this invention provides a single-image rain removal method based on spatial-frequency co-modeling, comprising the following steps: S1: Feature Initialization and Multi-Scale Parallel Feature Flow Construction: After the original rainy image is input into the system, shallow representation extraction is first performed through an image patch embedding module. This module contains a standard convolutional layer with a kernel size of 3×3, which projects the image from the 3-channel RGB space to a high-dimensional feature channel space. Subsequently, the first set of shallow spatial detail encoding is performed through a built-in parallel two-branch depthwise separable convolution. Following this, a three-scale parallel feature processing network topology is constructed: the main scale recovery branch directly preserves the original feature map resolution. The channel capacity is set to a fixed base value; the 1 / 2 scale recovery branch uses bilinear interpolation to reduce the spatial height and width geometry of the feature map to 1 / 2 of the original image, and is followed by a 1×1 channel-level convolutional layer to expand the feature channel dimension to 2C; the 1 / 4 scale recovery branch uses bilinear interpolation to further downsample and shrink the spatial height and width geometry to 1 / 4 of the original image, and uses a 1×1 convolutional layer to expand the feature channel dimension to 4C.
[0010] S2: Parallel Spatiotemporal Deraining Recovery and Dual-Domain Background Feature Reconstruction Based on the FIRM Module: The frequency domain heuristic interaction and representation module FIRM is deployed in parallel on the constructed main scale branch, half-scale branch, and quarter-scale branch for distributed iterative recovery. In the single-layer FIRM backbone, the input feature map is simultaneously copied and fed independently into two highly asymmetric spatiotemporal collaborative deraining recovery paths for deep reconstruction. S21: Spatial Domain Background Recovery Branch, i.e., d1 branch: The feature map stream first inputs an efficient enhanced depthwise separable convolutional channel attention block (CAB). It utilizes two layers of depthwise separable convolutions combined with synchronously cascaded adaptive global average pooling to extract global spatial statistics for each channel. Then, it generates one-dimensional dynamic channel weights through Sigmoid normalization for feature enhancement. Simultaneously, the feature map is fed into a spatial compression stream, where a stride convolution with a stride of 2 downsamples the spatial height and width by half. The downsampled low-resolution feature stream is then fed into a cascaded Transformer block. This block replaces the traditional spatial token attention with a channel-dimensional cross-covariance multi-head self-attention mechanism. It calculates cross-head cross-covariance across channels to implicitly model global pixel associations, accurately identifying and capturing local high-frequency rain streaks with strong spatial directionality and linear distribution in complex nonlinear entanglement, and establishing long-distance spatial nonlocal geometric dependencies. The low-resolution feature stream, refined through Transformer deep modeling and global feature extraction, is spatially upsampled and restored using a transposed convolutional layer with a kernel size of 3×3 and a stride of 2. Finally, a long-range feature selection and fusion block performs element-wise pixel-level addition between the upsampled and restored feature stream and the previously enhanced original resolution residual features, resulting in a finely repaired preliminary rain-free background feature map in the spatial domain. .
[0011] S22: Global Frequency Domain Enhancement Background Recovery Branch, i.e., d2 branch: The input feature map stream is directly input into the global frequency domain enhancement module GFEM. The four-dimensional feature map tensor input in the spatial domain (where B is BatchSize, C is the number of feature channels, and H and W are the current spatial resolution height and width) is directly transformed to the complex frequency space through a two-dimensional fast real Fourier transform (2DRFFT) to obtain a complex frequency domain tensor with more global frequency representation. To enable efficient backpropagation with full differentiability on mainstream computing architectures, the module explicitly extracts the real and imaginary parts of the complex tensor. The extracted real and imaginary feature maps are then applied in parallel to each other using an adaptive pointwise convolutional layer with a kernel size of only 1×1 and a learnable non-linear activation function layer, PReLU. The real and imaginary outputs, after frequency-level convolution reconstruction and nonlinear gain, are recombined to generate a refined complex frequency domain tensor. The reconstructed complex frequency domain tensor is projected back into the spatial domain using a two-dimensional inverse fast real Fourier transform (2DiRFFT), revealing the target restored spatial dimensions of the specified image to ensure perfect alignment of height and width geometric resolutions. Finally, global residual jump connections are established between the features transformed back into the spatial domain by the inverse Fourier transform and the original spatial features input to the front end of the GFEM module. A preliminary rainless background feature map containing low-frequency macroscopic background information was reconstructed. .
[0012] S23: Spatial Cooperative Adaptive Seamless Fusion of Dual-Domain Recovery Features in SCIM Module: Integrating Preliminary Background Features Preliminary background features The input is shared into the Spatial Cross-Branch Interaction Module (SCIM). First, [the input is...] and Explicit cross-domain concatenation is performed along the feature channel dimension, resulting in a highly fused cross-domain joint feature flow with a shape of (2C, H, W). To capture the complex nonlinear entanglement of spatial and frequency domain representations, the joint feature flow is fed into a multi-domain feature spatial interaction flow network. This network consists of a first 3×3 spatial convolutional layer, an intermediate cascaded parameterized nonlinear activation layer (PReLU), and a second 3×3 spatial convolutional layer connected in series. A sigmoid-normalized activation layer is connected at the end of the two spatial convolutional layers to adaptively generate a dual-domain spatial joint decoupling weight mask within the [0, 1] interval for each spatial pixel and each feature channel mapping. The generated joint spatial attention mask is uniformly bisected along the channel dimension to precisely separate and peel off the first domain spatial adaptive gating mask specifically used for modulating the spatial background branch. And a second-domain spatially adaptive gated mask specifically for modulating the frequency domain background branch. Constructing a cross-domain spatial soft-gated dynamic cross-injection fusion flow: Utilizing two sets of generated dynamic masks, complementary adaptive multiply-accumulate logic is performed on the originally independent spatial and frequency domain features. The two sets of complementary interactive feature streams after fusion are fed into two completely independent 1×1 adaptive local projection layers for cross-channel feature alignment and linear dimension smoothing. Finally, the projected output features and the original input features fed into the SCIM module establish local residual shortcut connections, outputting highly decoupled and seamlessly fused refined spatial background output features and refined frequency background output features, completing the forward propagation of the entire single-layer FIRM module.
[0013] S3: Multi-scale recovery feature fusion and deep concatenation iterative refinement: After the FIRM module outputs of the 1 / 2 scale recovery branch and the 1 / 4 scale recovery branch have completed feature reconstruction, the system calls the multi-scale fusion hub.
[0014] Using an upsampling module, the dual-channel feature outputs of the 1 / 2-scale and 1 / 4-scale recovery branches are magnified by 2x and 4x in spatial height and width dimensions, respectively, to completely restore their spatial geometric resolution to the top-level master-scale resolution. The background stream output from the master-scale branch, the background stream restored from the 1 / 2-scale branch, and the background stream restored from the 1 / 4-scale branch are concatenated in a three-scale, three-channel cascade along the feature channel dimension to form a high-density scale-converged feature map with 7C channels. Two independent 1×1 adjusted point-to-point convolutional layers are then used to compress the number of channels back to the standard base dimension, completing the adaptive deep convergence of multi-scale rain removal information.
[0015] In order to perform final decoupling and refinement on the fused multi-scale features, the converged feature stream is fed into multiple cascaded FIRM structures at the tail for further spatial-frequency decoupling and SCIM adaptive collaboration, outputting a high-fidelity background feature stream that is highly pure and completely stripped of rain streaks.
[0016] S4: Multi-target image reconstruction and closed-loop collaborative loss optimization supervision: The background stream output by the cascaded FIRM is fed into the reconstruction: The refined feature map of the first channel stream is input into a set of 3×3 standard reconstruction convolutional layers that maintain the same spatial size, directly mapping the high-dimensional feature map back to the 3-channel standard color image space, and outputting the first predicted clear rainless background image. The refined feature map of the second channel stream is simultaneously input into another set of completely parallel 3×3 standard reconstruction convolutional layers. Without changing the spatial geometry, it is directly mapped back to the standard 3-channel color image space, outputting a second predicted clear, rain-free background image. .
[0017] To combine the detailed information captured in the spatial domain with the global contour captured in the frequency domain, the system uses the two preliminary rain-free background images recovered above. and The images are stitched together along the channel dimension to form a multi-channel joint image space. This joint image is then passed through a 1×1 reconstruction convolutional layer to adaptively fuse the pixel advantages of the two images, eliminating boundary artifacts and outputting the final, clear, rain-free background image predicted by the system. .
[0018] The formula for the total training loss function of the system is: This invention uniformly adopts a composite loss function composed of three-dimensional joint metrics for the above three output branches: It includes a Charbonnier pixel fitting term, which robustly fits the absolute pixel difference between the predicted background image and the real, clear, rain-free background to prevent over-smoothing and blurring; a structural similarity constraint term, which uses a negative structural similarity loss function to explicitly apply to the generated image and the target image to ensure that the restored background meets human visual perception expectations; and a Laplacian high-frequency edge term, which extracts the global high-frequency edge contour components of the predicted image and the real image respectively through a specific Laplacian operator kernel and applies L1 norm constraints to accurately eliminate periodic oscillation artifacts, ringing effects, or edge ghosting that are easily caused by inverse Fourier transform or transposed convolution during scale restoration.
[0019] The proposed single-image deraining method based on spatial-frequency co-modeling has the following advantages: This invention utilizes the two-dimensional fast real Fourier transform to open a frequency domain path GFEM, directly intercepting and capturing the macroscopic global low-frequency rainless background contour in the frequency domain. From a physical mechanism perspective, it avoids the interference of high-frequency components of rain marks in the spatial domain on background restoration, perfectly protecting the boundaries and high-frequency textures of objects. This invention obtains an infinite global receptive field covering the entire image with extremely low computational complexity through 1×1 pointwise convolution in the frequency domain, which significantly reduces the number of network parameters, computational overhead and inference latency. This invention abandons the rigid, pixel-level matrix addition and subtraction decoupling method of existing technologies, which lacks dynamic perception capabilities, and innovatively designs a spatial cross-branch interaction module. Through a spatial soft gating mechanism, a dual-domain joint weight mask is dynamically and adaptively generated, realizing a strong combination of fine features in the spatial domain and macroscopic structures in the frequency domain, as well as bidirectional repair of blind spots. This completely eliminates texture loss, tone distortion, and edge blur artifacts commonly found in heavy rain, highlights, or edge areas. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 This is a flowchart of a single-image rain removal method based on spatial-frequency co-modeling. Figure 2 A technical framework diagram of a single-image rain removal method based on spatial-frequency co-modeling; Figure 3 The flowchart for the global frequency domain enhancement module; Figure 4 The flowchart shows the method for cross-branch interaction modules in space. Detailed Implementation
[0022] To make the objectives, technical solutions, and beneficial effects of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described in this specification are merely for explaining the invention and are not intended to limit the invention.
[0023] Example: Please see Figures 1-4This invention provides a single-image rain removal method based on spatial-frequency co-modeling. In step S1, the input is a 256×256×3 rain-affected RGB image. First, a standard convolutional layer with a kernel size of 3×3, a stride of 1, and padding of 1 projects the 3-channel RGB image into a 32-channel high-dimensional feature space, outputting a feature map of size 256×256×32. This feature map is then fed into a built-in parallel dual-branch depthwise separable convolutional layer. The outputs of the two branches are concatenated along the channel dimension and compressed back to 32 channels through a 1×1 convolutional layer, completing shallow spatial detail encoding.
[0024] To adaptively address the degradation of macroscopic rain streaks and microscopic raindrops with varying densities and receptive field scales, the system subsequently constructed a three-scale parallel feature processing network topology: the main scale recovery branch directly preserves the original feature map resolution dimensions (H, W), with the channel capacity set to a fixed base value, corresponding to the core high-resolution recovery backbone of the network; the 1 / 2 scale recovery branch synchronously inputs the output feature stream into the first-level downsampling module, using bilinear interpolation to reduce the spatial height and width geometric dimensions of the feature map to 1 / 2 of the original image, followed by a 1×1 channel-level convolutional layer to expand the feature channel dimension to 2C; the 1 / 4 scale recovery branch synchronously inputs the feature stream into the second-level downsampling module, using bilinear interpolation to further downsample and shrink the spatial height and width geometric dimensions to 1 / 4 of the original image, and using a 1×1 convolutional layer to expand the feature channel dimension to 4C.
[0025] Three cascaded frequency domain heuristic interaction and representation modules (FIRMs) are deployed on three scale branches. Each FIRM module consists of three parts: a spatial domain background restoration branch (d1), a global frequency domain enhanced background restoration branch (d2), and a spatial cross-branch interaction module (SCIM). The spatial domain background restoration branch (d1) implements channel attention enhancement. The input feature map first enters the enhanced depthwise separable convolutional channel attention block (CAB), and the feature map is simultaneously fed into the spatial compression stream. It is downsampled to 128×128×32 through stride convolution with a stride of 2 and a kernel size of 3×3, and then fed into two cascaded Transformer blocks. Each Transformer block adopts an 8-head cross-covariance multi-head self-attention mechanism based on the channel dimension. It calculates the cross-head cross covariance across channels, implicitly models global pixel association, and accurately captures rain streaks with strong directionality. The Transformer output is upsampled back to 256×256×32 through a transposed convolution with a stride of 2 and a kernel size of 3×3. Then, through long-range feature selection and fusion blocks, the upsampled features are element-wise added to the original resolution residual features from the CAB output, producing a preliminary rain-free background feature map. (256×256×32).
[0026] The global frequency domain enhancement background recovery branch d2 implements frequency domain transformation, frequency domain feature reconstruction, inverse transformation, and residual connection. In the global frequency domain enhancement background recovery branch (d2 branch), the input feature map stream is directly input into the global frequency domain enhancement module (GFEM). The specific algorithm evolution flow is as follows: Figure 3 As shown, the four-dimensional feature map tensor input in the spatial domain is directly transformed to the complex frequency space through a two-dimensional fast real Fourier transform (2DRFFT), resulting in a complex frequency domain tensor with greater global frequency representation. To enable efficient backpropagation with full differentiability on mainstream computing architectures, the module explicitly extracts the real and imaginary parts of this complex tensor. Since each frequency response point in the complex frequency domain naturally aggregates the energy of all pixels in the original spatial domain image, the frequency domain branch naturally obtains an infinite global receptive field covering the entire image. Adaptive pointwise convolutional layers with a kernel size of only 1×1 are applied in parallel to the extracted real and imaginary feature maps. Utilizing the property that the multiplication in the frequency space is equivalent to the circular convolution in the spatial domain, linear adaptive adjustment and reconstruction of cross-channel frequency components are achieved. This is followed by a learnable nonlinear activation function layer, PReLU, which merges the real and imaginary outputs after frequency-level convolution reconstruction and nonlinear gain, and reconstructs a refined complex frequency domain tensor. The reconstructed complex frequency domain tensor is projected back into the spatial domain using a two-dimensional inverse fast real Fourier transform (2DiRFFT). During the inverse transform, the target spatial dimensions of the image are explicitly specified to ensure perfect alignment of the height and width geometric resolutions. Finally, global residual jump connections are established between the features transformed back into the spatial domain by the inverse Fourier transform and the original spatial features input to the front end of the GFEM module. Through this design, GFEM can adaptively filter high-frequency rain streak oscillations and reconstruct a preliminary rainless background feature map rich in low-frequency macroscopic background information.
[0027] See Figure 4 Next, the preliminary background features output from the high-path spatial domain are... Preliminary background characteristics of low-path frequency domain output The inputs are combined and integrated into the Spatial Cross-Branch Interaction Module (SCIM).
[0028] First, and Explicit cross-domain cascading is performed along the feature channel dimension to output a highly fused cross-domain joint feature flow with a shape of (B, 2C, H, W). To capture the complex nonlinear entanglement of spatial and frequency domain representations, the joint feature flow is fed into a multi-cross-domain feature spatial interaction flow network consisting of a first 3×3 spatial convolutional layer, an intermediate cascaded parameterized nonlinear activation layer PReLU, and a second 3×3 spatial convolutional layer. The two spatial convolutional layers fully exploit the correlation between surrounding local pixels and cross-domain channels. A Sigmoid normalized activation layer is connected at the end of the feature interaction flow, thereby adaptively generating a dual-domain spatial joint decoupling weight mask in the interval [0, 1] for each spatial pixel and each feature channel mapping.
[0029] The generated joint spatial attention mask is uniformly bisected along the channel dimension, thereby accurately separating and peeling off a first-domain spatial adaptive gating mask specifically for modulating the spatial background branch, and a second-domain spatial adaptive gating mask specifically for modulating the frequency background branch. Using these two sets of dynamic masks, complementary adaptive multiply-accumulate logic is performed on the originally independent spatial and frequency features. Through a spatial soft-gating mechanism, the spatial and frequency background features are strongly combined and mutually injected with complementary flows, achieving collaborative bidirectional repair of spatial and frequency blind spots. The two sets of complementary interactive feature flows are then fed into two completely independent 1×1 adaptive local projection layers for cross-channel feature alignment and linear dimension smoothing.
[0030] Finally, the projected output features and the original input features fed into the SCIM module establish local residual shortcut connections respectively, output highly decoupled and seamlessly integrated refined spatial background output features and refined frequency background output features, completing the forward propagation of the entire single-layer FIRM module.
[0031] Next, step S3 is executed: multi-scale feature fusion and deep concatenation iterative refinement. After the FIRM modules within the 1 / 2-scale and 1 / 4-scale recovery branches have completed feature reconstruction, the system calls the multi-scale fusion hub. Using the upsampling module, the dual-channel feature outputs of the 1 / 2-scale and 1 / 4-scale recovery branches are magnified by 2x and 4x respectively in the spatial height and width dimensions, so that their spatial geometric resolution is completely restored to the top-level main scale resolution. The background stream and rain streak stream output from the main scale branch, the background stream and rain streak stream restored from the 1 / 2-scale branch, and the background stream and rain streak stream restored from the 1 / 4-scale branch are concatenated and stitched together in three-scale, three-channel dimensions to form a high-density scale converged feature map with a channel number expanded to 7C. Two independent 1×1 adjusted point-to-point convolutional layers are used to compress the number of channels back to the standard base dimension, explicitly completing the adaptive deep convergence of multi-scale rain removal information while ensuring ultra-low parameter count. To perform final decoupling and refinement of the fused multi-scale features, the converged feature stream is fed into multiple cascaded FIRM structures at the tail. Within the cascaded FIRM structures, spatial-frequency domain decoupling and SCIM adaptive collaboration are used again to output a high-fidelity background feature stream that is highly pure and completely free of rain streaks.
[0032] Finally, step S4 is executed, involving supervised optimization of multi-target image reconstruction and loop closure collaborative loss. The background stream output from the cascaded FIRM volume is then fed into the reconstruction process: the refined feature map of the first channel stream is input into a set of standard 3×3 reconstruction convolutional layers that maintain spatial dimensions, directly mapping the high-dimensional feature map back to the standard 3-channel color image space, outputting the first predicted clear, rain-free background image, denoted as... The refined feature map of the second channel stream is simultaneously input into another set of completely parallel 3×3 standard reconstruction convolutional layers. Again, without changing the spatial geometry, it is directly mapped back to the standard 3-channel color image space, outputting a second predicted clear, rain-free background image, denoted as... The two preliminary background images generated and The composite background loss is calculated directly with a real, clear, rainless background. To allow them to play an auxiliary guiding role in training, each loss term is assigned a balancing weight of 0.2.
[0033] To combine the detailed information captured in the spatial domain with the global contour captured in the frequency domain, the system uses the two preliminary rain-free background images recovered above. and The images are stitched together along the channel dimension to form a multi-channel joint image space. This joint image is then passed through a 1×1 reconstruction convolutional layer to adaptively fuse the pixel advantages of the two images, eliminating boundary artifacts and outputting the final, clear, rain-free background image predicted by the system. .
[0034] As alternatives to this technical solution: In the GFEM module, the two-dimensional fast real Fourier transform can be replaced by the two-dimensional discrete cosine transform; in the spatial compression flow of the d1 branch, the cascaded cross-covariance multi-head self-attention mechanism based on the channel dimension can be replaced by a multi-scale large kernel attention network, a deformable convolutional network, or a selective state space model; in the SCIM module, the fusion of dual-domain features can be replaced by a dual-stream interaction body based on the cross-attention mechanism, or by an adaptive modulation structure based on the matrix Hadamard product and the frequency domain dynamic filter.
[0035] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A single-image deraining method based on spatial-frequency co-modeling, characterized in that: Includes the following steps: S1: Feature initialization and multi-scale parallel feature flow construction: The input two-dimensional rainy image is projected into a high-dimensional feature channel space through the image patch embedding module, and a three-scale parallel feature processing network topology containing a main scale recovery branch, a 1 / 2 scale recovery branch and a 1 / 4 scale recovery branch is constructed. S2: Parallel frequency-space rain removal recovery and dual-domain background feature reconstruction based on FIRM module: Frequency domain heuristic interaction and representation module FIRM is deployed in parallel on three-scale recovery branches for distributed iterative recovery. In each scale FIRM module, the feature map is simultaneously fed into the spatial domain background recovery branch d1 and the global frequency domain enhanced background recovery branch d2 for deep reconstruction. Then, the initial rain-free background feature map output by the two branches is spatially collaboratively and adaptively fused through the spatial cross-branch interaction module SCIM. S3: Multi-scale recovery feature fusion and deep concatenated iterative refinement: The output features of the 1 / 2 scale recovery branch and the 1 / 4 scale recovery branch are upsampled and restored to the main scale resolution. They are then concatenated with the output features of the main scale recovery branch in the channel dimension in three scales. After compression by the convolutional layer, they are fed into multiple concatenated FIRM structures at the end for final decoupling and refinement. S4: Multi-target image reconstruction and closed-loop collaborative loss optimization supervision: The feature stream output by the cascaded FIRM structure is mapped and restored through two sets of parallel reconstruction convolutional layers to output two clear and rain-free background images; the two prediction images are stitched together in the channel dimension and adaptively fused through the reconstruction convolutional layer to output the final clear and rain-free background image, and the network is trained and optimized using a composite loss function.
2. The single-image deraining method based on spatial-frequency co-modeling according to claim 1, characterized in that: In step S1, the image patch embedding module includes a standard convolutional layer with a kernel size of 3×3 and a built-in parallel dual-branch depth-separable convolutional layer. The 3×3 standard convolutional layer is used to project the image from the 3-channel RGB space to the high-dimensional feature channel space, and the parallel dual-branch depth-separable convolutional layer is used for shallow spatial detail encoding.
3. The single-image deraining method based on spatial-frequency co-modeling according to claim 1, characterized in that: In step S1, the specific construction method of the three-scale parallel feature processing network topology is as follows: The main scale recovery branch directly preserves the original feature map resolution size. The channel capacity is set to the base value C; The 1 / 2 scale recovery branch uses bilinear interpolation to reduce the spatial height and width of the feature map to half of the original image, and expands the feature channel dimension to 2C through a 1×1 convolutional layer; The 1 / 4 scale recovery branch uses bilinear interpolation to reduce the spatial height and width of the feature map to 1 / 4 of the original image, and expands the feature channel dimension to 4C through a 1×1 convolutional layer.
4. The single-image deraining method based on spatial-frequency co-modeling according to claim 1, characterized in that: In step S2, the output of each FIRM module includes refined spatial background output features and refined frequency background output features. The spatial domain background restoration branch d1 is used to capture local high-frequency rain streaks with spatial directionality and repair spatial details, while the global frequency domain enhanced background restoration branch d2 is used to reconstruct low-frequency macroscopic background information.
5. The single-image deraining method based on spatial-frequency co-modeling according to claim 4, characterized in that: The specific reconstruction process of the spatial domain background restoration branch d1 in step S2 is as follows: The input feature map stream first enters the Enhanced Depthically Separable Convolutional Channel Attention Block (CAB) for channel-dimensional weight recalibration and feature enhancement; The feature maps are simultaneously fed into the spatial compression stream. After the spatial height and width dimensions are downsampled by half through strided convolution with a stride of 2, they are fed into cascaded Transformer blocks to establish long-distance spatial nonlocal geometric dependencies. The low-resolution feature stream refined by the Transformer block is upsampled and restored in terms of spatial size through a transposed convolutional layer with a stride of 2; Finally, through long-range feature selection and fusion blocks, the upsampled and recovered feature stream is added to the original resolution residual features enhanced by CAB at the pixel level to output a preliminary rainless background feature map.
6. The single-image deraining method based on spatial-frequency co-modeling according to claim 5, characterized in that: The Transformer block employs a channel-dimensional cross-covariance multi-head self-attention mechanism, which implicitly models global pixel association by calculating cross-head cross-covariance across channels, and identifies and captures local high-frequency rain streaks with strong spatial directionality and linear distribution.
7. The single-image deraining method based on spatial-frequency co-modeling according to claim 1, characterized in that: In step S2, the specific reconstruction process of the global frequency domain enhanced background recovery branch d2 is as follows: The input four-dimensional feature map tensor is mapped to the complex frequency space through two-dimensional fast real Fourier transform (2DRFFT) to obtain a complex frequency domain tensor. Extract the real and imaginary feature maps of the complex tensor, apply them in parallel with an adaptive pointwise convolutional layer with a kernel size of 1×1 to adjust the frequency components across channels, and follow up with a parameterized nonlinear activation function layer PReLU. The processed real and imaginary outputs are merged again to reconstruct the refined complex frequency domain tensor. The reconstructed complex frequency domain tensor is projected back into the spatial domain by applying a two-dimensional inverse fast real Fourier transform (2DiRFFT), and the target restored spatial dimensions of the specified image are displayed to ensure geometric resolution alignment. Finally, the features transformed back to the spatial domain by the inverse Fourier transform are used to establish a global residual jump connection with the original spatial features input from the front end, and a preliminary rainless background feature map is output.
8. The single-image deraining method based on spatial-frequency co-modeling according to claim 1, characterized in that: In step S2, the specific fusion process of the Spatial Cross-Branch Interaction Module (SCIM) is as follows: The preliminary rainless background feature map output from the spatial domain branch and the preliminary rainless background feature map output from the frequency domain branch are concatenated and stitched across the feature channel dimension to generate a cross-domain joint feature flow. The joint feature flow is fed into a network consisting of a first 3×3 spatial convolutional layer, an intermediate cascaded PReLU activation layer, a second 3×3 spatial convolutional layer, and a tail Sigmoid normalized activation layer, and adaptively generates a dual-domain spatial joint decoupling weight mask in the interval [0, 1]. The dual-domain joint decoupling weight mask is uniformly bisected along the channel dimension to obtain a first-domain adaptive gated mask and a second-domain adaptive gated mask. By using two sets of dynamic masks to perform complementary adaptive multiply-accumulate logic on independent spatial and frequency domain features, a cross-domain spatial soft-gated dynamic mutual injection fusion flow is constructed. The two fused feature streams are fed into two independent 1×1 adaptive local projection layers for cross-channel feature alignment. The projected features and the original input features fed into the SCIM module are connected to establish local residual shortcut connections, and refined spatial background output features and refined frequency background output features are output.
9. A single-image deraining method based on spatial-frequency co-modeling according to claim 1, characterized in that: The specific process of step S3 is as follows: The upsampling module is used to amplify the output features of the 1 / 2 scale recovery branch by a factor of 2 and the output features of the 1 / 4 scale recovery branch by a factor of 4, so that their spatial resolution is consistent with that of the main scale recovery branch. The output features of the three scale recovery branches are concatenated and stitched together in the channel dimension to form a high-density scale converged feature map. By using 1×1 convolutional layers to compress the number of channels in the converged feature map back to the standard base dimension, adaptive deep convergence of multi-scale rain removal information is achieved. The compressed feature stream is fed into multiple cascaded FIRM structures at the tail, and spatial-frequency domain decoupling and SCIM adaptive collaboration are performed again to output the final high-fidelity background feature stream.
10. A single-image deraining method based on spatial-frequency co-modeling according to claim 1, characterized in that: In step S4, the composite loss function includes a Charbonnier pixel fitting term, a structural similarity constraint term, and a Laplacian high-frequency edge term. The Charbonnier pixel fitting term is used to robustly fit the absolute pixel difference between the generated predicted background image and the real, clear, rain-free background. The structural similarity constraint term uses a negative structural similarity loss function to force the network to highly align local contrast, brightness, and macroscopic topology. The Laplacian high-frequency edge term extracts the global high-frequency edge contour components of the predicted image and the real image respectively through the Laplacian operator kernel and applies L1 norm constraints.