Image processing method fusing frequency perception and MAMBA network

CN120387955APending Publication Date: 2025-07-29SHANGHAI UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510459674.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The existing physical model-based image rain removal method is poor in the face of different rainfall scenes, and it is difficult to adapt to complex and changeable real environments, resulting in poor quality of raindrop image recovery.

Method used

The image processing method of fusion frequency perception and Mamba network is adopted, and through wavelet transformation, convolution operation and high-frequency enhancement technology, combined with dual-branch feature fusion and prior guidance units, refined texture reconstruction is carried out to improve rain removal effect.

Benefits of technology

While keeping the controllable model parameters, the image recovery quality is significantly improved, high-quality rainwater removal effect is achieved, and the visual effect is better than existing methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120387955A_ABST
    Figure CN120387955A_ABST
Patent Text Reader

Abstract

The invention discloses an image processing method fusing frequency perception and a Mama network, and the method comprises the steps: collecting original data, carrying out the preprocessing, obtaining a sample set, training a constructed deep learning network, processing an input image in real time through the trained deep learning network at an online stage, and achieving a rainwater removal effect. On the basis of a fusion frequency perception and Mama network architecture, on one hand, discrete wavelet transform down-sampling operation, convolution operation, two wavelet domain recovery module feature extraction operation, convolution operation, two wavelet domain recovery module feature extraction operation, convolution operation and discrete wavelet inverse transform operation are sequentially carried out on a raindrop picture; and on the other hand, a high-frequency component is extracted from frequency information after wavelet transformation, and after frequency enhancement, a wavelet domain recovery module is sequentially guided to perform refined texture reconstruction, so that the rain removal effect is improved, and an image restoration result with higher quality is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a technology in the field of image processing, and specifically to a method for realizing the effect of removing rain in an image by integrating frequency perception and the MAMBA network. Background Art

[0002] Traditional image de-raining methods are mainly based on physical models. These methods usually utilize the physical characteristics in the imaging process to model the formation of raindrops or rain streaks, so as to separate and remove the raindrop components and restore a clear background image. Compared with data-driven methods, the image de-raining methods based on physical models are more theoretically interpretable, but their effects often depend on the accuracy of the raindrop model and the accuracy of the input parameters. In the face of different rainfall scenarios (such as raindrop morphology, density, lighting conditions, etc.), the robustness of such methods is poor and it is difficult to adapt to the complex and changeable real environment. Summary of the Invention

[0003] Aiming at the defects of poor restoration quality and weak robustness of raindrop images in the prior art, the present invention proposes an image processing method that integrates frequency perception and the Mamba network. In the architecture based on the integration of frequency perception and the Mamba network, on the one hand, the raindrop image is successively subjected to discrete wavelet transform downsampling operation, convolution operation, feature extraction operations of two wavelet domain restoration modules, convolution operation, feature extraction operations of two wavelet domain restoration modules, convolution operation, and inverse discrete wavelet transform operation; on the other hand, high-frequency components are extracted from the frequency information after wavelet transform, and after frequency enhancement, the wavelet domain restoration modules are successively guided to perform refined texture reconstruction so as to improve the de-raining effect and obtain a higher-quality image restoration result.

[0004] The present invention is realized through the following technical solutions:

[0005] The present invention relates to an image processing method that integrates frequency perception and the MAMBA network. By collecting and preprocessing the original data, a sample set is obtained for training the constructed deep learning network. In the online stage, the input image is processed in real time by the trained deep learning network to achieve the effect of rain removal.

[0006] The described deep learning network includes: a wavelet transform module, four wavelet domain restoration modules with high-frequency enhancement units, and a wavelet inverse transform module, which are arranged in sequence. Among them: the wavelet transform module performs wavelet transform downsampling processing on the read raindrop image information to obtain a low-frequency sub-image LL containing overall contour information and high-frequency sub-images LH, HL, and HH containing edge and texture information respectively. The wavelet domain restoration module performs global-local feature extraction processing on the low- and high-frequency sub-images after wavelet transform to obtain a global-local information fusion result. The high-frequency enhancement unit performs high-frequency enhancement processing on the high-frequency sub-image after wavelet transform to obtain an enhanced high-frequency reconstruction result. The wavelet inverse transform module performs wavelet inverse transform upsampling processing based on the feature extraction information of the wavelet domain restoration module guided by the enhanced high-frequency sub-image to obtain a clean image reconstruction result.

[0007] The described wavelet domain restoration module includes: a plurality of double-branch feature fusion units and a plurality of prior guidance units. Among them: the double-branch feature fusion unit processes global feature information and local feature information respectively based on the shallow feature information after wavelet transform; the prior guidance unit separates the high-frequency sub-image from the global-local information obtained by the double-branch feature fusion unit and the high-frequency sub-image obtained by the high-frequency enhancement module. Technical effects

[0008] Through double-branch feature fusion, 2D characteristic frequency scanning, and prior guidance processing, on the one hand, the model has a powerful frequency perception ability; on the other hand, comprehensive experiments on the raindrop dataset show that our method has achieved good performance compared with the state-of-the-art methods while keeping the number of model parameters efficiently controllable. And from the visual effect, it can be seen that our method can achieve high-quality degraded image reconstruction. Brief description of the drawings

[0009] Figure 1 It is the flowchart of the present invention;

[0010] Figure 2 It is the schematic diagram of the deep learning network structure of the present invention;

[0011] Figure 3 It is the schematic diagram of the double-branch feature fusion unit;

[0012] Figure 4 It is the schematic diagram of the prior guidance unit;

[0013] Figure 5 It is the 2D characteristic frequency scanning mechanism

[0014] Figure 6 It is the processing flowchart of the deep learning network;

[0015] Figure 7 It is the schematic diagram of the effect of the embodiment;

[0016] In the figure: (a) is the input raindrop picture, (b) is the clean result output by the TransWeather model, (c) is the clean result output by the WeatherDiff model, (d) is the clean result output by the present invention, and (e) is the ground truth picture. Specific implementation mode

[0017] As Figure 1 shown, this embodiment relates to an image processing method that combines frequency perception and the MAMBA network. After collecting and preprocessing the original data, a sample set is obtained for training the constructed deep learning network (FA-MAMBA). In the online stage, the trained deep learning network is used to process the input image in real time and achieve the effect of rain removal.

[0018] As Figure 2 shown, the deep learning network includes: a wavelet transform module (DWT), four wavelet domain restoration modules (WDRM) with high-frequency enhancement units (HFEM), and an inverse wavelet transform module (IWT) arranged in sequence. Among them: a first convolutional layer is provided between the wavelet transform module and the first wavelet domain restoration module, second to fourth convolutional layers are provided between the second wavelet domain restoration module and the third wavelet domain restoration module, and fifth and sixth convolutional layers are provided between the fourth wavelet domain restoration module and the inverse wavelet transform module. The output of the wavelet transform module is accumulated with the output of the third convolutional layer and then input to the fourth convolutional layer. The output of the first convolutional layer is accumulated with the output of the second convolutional layer and then input to the third convolutional layer. The output of the first wavelet domain restoration module is accumulated with the output of the second wavelet domain restoration module and then input to the second convolutional layer. The output of the fourth convolutional layer is accumulated with the output of the fifth convolutional layer and then input to the sixth convolutional layer. The output of the third wavelet domain restoration module is accumulated with the output of the fourth wavelet domain restoration module and then input to the fifth convolutional layer. The output of the wavelet transform module, the output of the third convolutional layer, and the output of the sixth convolutional layer are accumulated and then input to the inverse wavelet transform module.

[0019] The high-frequency enhancement units of the four wavelet domain restoration modules are connected in series in sequence and are connected to their respective wavelet domain restoration modules.

[0020] The described wavelet domain restoration module includes: a plurality of dual-branch feature fusion units (DFEB) and a plurality of prior guidance units (PGB). Among them: The DFEB processes the global feature information and local feature information respectively according to the shallow feature information obtained by 3×3 convolution after wavelet transform. In the extraction of local feature information, through convolution, pooling, and convolution operations in sequence, the result of local feature information is obtained; in the extraction of global feature information, the result of global feature information is obtained through a normalization layer and 2D characteristic frequency scanning; finally, the extracted global feature information and local feature information are fused by addition; The PGB separates the high-frequency subgraph from the global-local information processed by the dual-branch feature fusion unit (DFEB) and the high-frequency subgraph obtained by the high-frequency enhancement module, performs convolution and dot product interaction processing to obtain a high-frequency result, and performs residual connection output with the originally separated high-frequency subgraph, and finally connects and outputs with the low-frequency subgraph separated from the global-local information processed by the dual-branch feature fusion unit (DFEB).

[0021] As Figure 3 shown, the described DFEB includes: four convolutional layers, two pooling layers, a normalization layer, and a 2D characteristic frequency scanning module (AFSM). Among them: Four convolutional layers and two pooling layers form a branch to perform local feature extraction operations; a normalization layer and a 2D characteristic frequency scanning module (AFSM) form a branch to perform global feature extraction operations, enhancing the ability to recover the information lost due to raindrop occlusion.

[0022] As Figure 4 shown, the described PGB includes: a 3×3 convolutional layer and four 1×1 convolutional layers. Among them: The unit input 1 generates V I and Q I values through 1×1 convolution, and the unit input 2 generates K E values through 3×3 convolution. First, perform dot product interaction operation on Q I and K E and then perform softmax operation to obtain result 1, and then perform dot product interaction operation on result 1 and V I values to obtain the attention weight. Through the dot product interaction operation, the unit input 1 obtains a refined reconstruction result of texture details under the guidance of the unit input 2;

[0023] As Figure 5 shown, the described 2D characteristic frequency scanning module (AFSM) includes: two linear layers, a depthwise separable convolutional layer, two activation layers, and a 2D characteristic frequency scanning unit. Among them: Under the wavelet domain, enhance the detailed part of the image by performing 2D linear characteristic scanning processing on the unit input, while retaining important structures and features.

[0024] As Figure 6As shown, the real-time processing specifically includes:

[0025] Step 1: Wavelet transform processing: The wavelet transform module performs a first-level wavelet transform on the input using the Haar transform, converting the input single image into four sub-images, that is, decomposing the original image signal into different low- and high-frequency sub-images LL, HL, LH, and HH, so that the low-frequency sub-image LL retains the overall contour information, while the high-frequency sub-images HL, LH, and HH contain details such as edges and textures respectively. The transformation from the spatial domain to the wavelet domain also performs downsampling on the original image while preserving the original information, and the size of the sub-images obtained is 1 / 2H and 1 / 2W of the original input. Specifically: I LL ,I LH ,I HL ,I HH =f DWT (I in ), and then splice the information of the four sub-images, that is Among them:

[0026] Step 2: Shallow feature extraction, that is: F in =f conv (F e ), where: f conv (·) is a 3×3 convolutional layer for extracting shallow features, and then the rain droplet removal task is achieved through four wavelet domain restoration modules in sequence, that is: F out =f WDRM (F in )=f PGB (f DFEB (F in ),f HFEM (F HF )) specifically includes:

[0027] 2.1. The first wavelet domain restoration module performs 3*3 convolutional processing on the sub-images obtained in Step 1 to extract the shallow feature F in ;

[0028] 2.2. The high-frequency enhancement unit extracts the high-frequency sub-image F HF from the sub-images and performs high-frequency enhancement to obtain the enhanced high-frequency sub-image;

[0029] 2.3. The DFEB processes the global feature information and local feature information respectively according to the result F in in Step 2.1. On the one hand, in the extraction of local feature information, the local feature information result is obtained through convolutional, pooling, and convolutional operations in sequence; on the other hand, in the extraction of global feature information, the global feature information result is obtained through a normalization layer and 2D characteristic frequency scanning; finally, the extracted global feature information and local feature information are fused by addition;

[0030] 2.4. The PGB will extract high-frequency subgraphs as unit input 1 according to the result of step 2.3, and generate V I and Q I values through 1×1 convolution. Take the result of step 2.2 as unit input 2 and generate K E values through 3×3 convolution; first perform dot product interaction operation on Q I and K E , then perform softmax operation to get the result, and then perform dot product interaction operation on the result and V I to obtain attention weights. Add the attention weights to the high-frequency subgraphs of the result of step 2.3, and finally splice them with the low-frequency subgraphs of the result of step 2.3 to obtain the result F out .

[0031] Step 3. Inverse wavelet transform processing: Perform inverse wavelet transform on the output F out to obtain the final prediction graph F pred = f DWT (F out ), where: f DWT is the inverse wavelet transform.

[0032] Through specific actual experiments, it is implemented by constructing a neural network in the PyTorch framework. The training process iterates 350 epochs with the support of an NVIDIA RTX 3090 graphics processing unit. For the optimizer, we use the Adam optimizer and train the model with a learning rate of 0.0003. Read the picture through wavelet transform and convert it into four subgraphs, splice them to obtain a tensor with 12 channels, and further obtain high-dimensional features with 180 channels through convolution, and input them into four wavelet domain restoration modules in turn.

[0033] In the first and second wavelet domain restoration modules mentioned above, 6 double-branch feature fusion units and 6 PGBs are adopted; in the last two wavelet domain restoration modules, 4 double-branch feature fusion units and 4 PGBs are adopted.

[0034] The input channel number of the high-frequency enhancement unit is 9, and the output channel number is 9. In the PGM module, extract high-frequency components from the output of the double-branch feature fusion unit, and then perform convolution on it and the output of the high-frequency enhancement unit to obtain high-dimensional features with the same channel number of 270 for high-frequency guidance. Finally, obtain the predicted picture through inverse wavelet transform.

[0035] To comprehensively evaluate the effectiveness of existing methods, in this embodiment, the peak signal-to-noise ratio (PSNR) and structural similarity (SSIM) are used as evaluation metrics. PSNR measures the recovery quality of an image, and a higher value indicates a better recovery effect; SSIM evaluates the similarity of images by comparing luminance, contrast, and structural information, and can better reflect the visual perception quality. The closer the SSIM value is to 1, the more similar the two images are.

[0036] The peak signal-to-noise ratio mentioned above Where: MAX is the maximum pixel value of the image (for an 8-bit image, it is usually 255).

[0037] The mean squared error mentioned above Where: I orig (i,j) represents the pixel value of the original image, I rec (i,j) represents the pixel value of the original image, and H and W are the height and width of the image respectively.

[0038] The structural similarity mentioned above Where: μ x , μ y are the means of images x and y respectively, representing luminance information, are the variances of images x and y respectively, representing contrast information, σ xy is the covariance between images x and y, representing the similarity of structural information, and C1 and C2 are stable constants to prevent the denominator from being zero.

[0039] The present invention is compared with existing image restoration methods CCN, IDT, TransWeather, and the two latest image restoration methods PatchDM and WaveDM. The results are shown in Table 1 and Table 2 and Figure 7 as shown, from the enlarged view at the lower right corner of the visualization result Figure 7 it can be seen that at the texture details of the part occluded by raindrops, the restoration of this method is closest to the ground truth value. It can be seen that the result of this method is better than the above comparison methods.

[0040] Table 1: Quantitative comparison of different methods on the Raindrop dataset

[0041] Table 2: Quantitative analysis of ablation experiments on different modules

[0042] Compared with the prior art, through the dual-branch feature fusion unit, the PSNR of the present invention in the rain removal task is 32.41 dB, achieving a preliminary rain removal effect; through the prior guidance unit, the PSNR is 33.04 dB. Finally, relying on the texture distribution characteristics of the high-frequency subgraphs, the high-frequency subgraphs HL and LH are scanned horizontally and vertically, and the high-frequency subgraph HH is scanned in a zigzag pattern. Through the 2D characteristic frequency scanning module, the PSNR of the entire model in the rain removal task is as high as 33.18 dB. The evaluation results on the Raindrop dataset show that the image rain removal method based on the fusion of frequency perception and the Mamba network of the present invention is superior to other state-of-the-art methods.

[0043] The above specific embodiments can be locally adjusted in different ways by those skilled in the art without departing from the principles and purposes of the present invention. The protection scope of the present invention is subject to the claims and is not limited by the above specific embodiments, and all implementation solutions within its scope are subject to the constraints of the present invention.

Claims

1. An image processing method that integrates frequency perception and the MAMBA network, characterized in that, A sample set is obtained by collecting and preprocessing the original data and is used to train and construct the deep learning network. In the online stage, the trained deep learning network is used to process the input images in real time and achieve the effect of rain removal. The deep learning network includes: a wavelet transform module, four wavelet domain restoration modules with high-frequency enhancement units, and a wavelet inverse transform module, which are arranged in sequence. Among them: the wavelet transform module performs wavelet transform downsampling processing on the read raindrop image information to obtain a low-frequency sub-image LL containing overall contour information and high-frequency sub-images LH, HL, and HH containing edge and texture information respectively. The wavelet domain restoration module performs global-local feature extraction processing on the low-frequency and high-frequency sub-images after wavelet transform to obtain a global-local information fusion result. The high-frequency enhancement unit performs high-frequency enhancement processing on the high-frequency sub-image after wavelet transform to obtain an enhanced high-frequency reconstruction result. The wavelet inverse transform module performs wavelet inverse transform upsampling processing based on the feature extraction information of the wavelet domain restoration module guided by the enhanced high-frequency sub-image to obtain a clean image reconstruction result.

2. The image processing method integrating frequency perception and MAMBA network according to claim 1, characterized in that, The wavelet domain restoration module includes: a plurality of double-branch feature fusion units and a plurality of prior guidance units. Among them: the double-branch feature fusion unit processes global feature information and local feature information respectively according to the shallow feature information after wavelet transform; the prior guidance unit separates the high-frequency sub-image from the global-local information processed by the double-branch feature fusion unit and the high-frequency sub-image obtained by the high-frequency enhancement module. The high-frequency enhancement units of the four wavelet domain restoration modules are connected in series in sequence and are connected to their respective wavelet domain restoration modules.

3. The image processing method integrating frequency perception and MAMBA network according to claim 2, characterized in that, The double-branch feature fusion unit includes: four convolutional layers, two pooling layers, a normalization layer, and a 2D characteristic frequency scanning module. Among them: the four convolutional layers and the two pooling layers form one branch and perform local feature extraction operations; a normalization layer and a 2D characteristic frequency scanning module form one branch and perform global feature extraction operations to enhance the ability to restore the information lost due to raindrop occlusion.

4. The image processing method integrating frequency perception and MAMBA network according to claim 2, characterized in that, The prior guidance unit described above includes: a 3*3 convolutional layer and four 1*1 convolutional layers, where: the unit input 1 generates V through a 1*1 convolution I and Q I values, and the unit input 2 generates K through a 3*3 convolution E values. First, Q I and K E are subjected to a dot product interaction operation and then a softmax operation to obtain result 1. Then, result 1 and V I values are subjected to a dot product interaction operation to obtain attention weights, and through the dot product interaction operation, the unit input 1 obtains a refined reconstruction result of texture details under the guidance of the unit input 2.

5. The image processing method integrating frequency perception and MAMBA network according to claim 3, characterized in that, The 2D characteristic frequency scanning module includes: two linear layers, a depthwise separable convolutional layer, two activation layers, and a 2D characteristic frequency scanning unit. Among them: in the wavelet domain, the details of the image are enhanced by performing 2D linear characteristic scanning processing on the unit input, while important structures and features are retained.

6. The image processing method integrating frequency perception and MAMBA network according to claim 1, characterized in that, The real-time processing specifically includes: Step 1, Wavelet Transform Processing: The wavelet transform module performs a one-level wavelet transform on the input using the Haar transform, converting the input single image into four sub-images, that is, decomposing the original image signal into different low- and high-frequency sub-images LL, HL, LH, and HH, so that the low-frequency sub-image LL retains the overall contour information, while the high-frequency sub-images HL, LH, and HH contain details such as edges and textures respectively. The transformation from the spatial domain to the wavelet domain also performs a downsampling operation on the original image while preserving the original information, and the size of the sub-images obtained is 1 / 2H and 1 / 2W of the original input. Specifically: I LL ,I LH ,I HL ,I HH = f DWT (I in ), and then splice the information of the four sub-images, that is Among them: F LF = I LL ; Step 2, shallow feature extraction, i.e.: F in = f conv (F e ), where: f conv (·) is a 3×3 convolutional layer for extracting shallow features, and then the raindrop removal task is implemented through four wavelet domain restoration modules in sequence, i.e.: F out = f WDRM (F[[ID=...]] in ) = f PGB (f DFEB (F in ), f HFEM (F HF )); Step 3, inverse wavelet transform processing: Output F out Perform inverse wavelet transform to obtain the final prediction map F pred = f DWT (F out ), where: f DWT is the inverse wavelet transform.

7. The image processing method integrating frequency perception and the MAMBA network according to claim 6, characterized in that, Step 2 specifically includes: 2.

1. The first wavelet domain recovery module performs 3×3 convolution processing on the subgraph obtained in step 1 to extract the shallow feature F in ; 2.2 The high-frequency enhancement unit extracts the high-frequency subgraph F from the subgraph HF and performs high-frequency enhancement to obtain the enhanced high-frequency subgraph; 2.

3. DFEB acts according to the result F in step 2.1 in Perform global feature information and local feature information processing respectively. On the one hand, in the extraction of local feature information, successively through convolution, pooling and convolution operations to obtain the result of local feature information; on the other hand, in the extraction of global feature information, obtain the result of global feature information through a normalization layer and 2D characteristic frequency scanning; finally, fuse the extracted global feature information and local feature information by addition; 2.

4. PGB will extract high-frequency subgraphs as unit input 1 according to the result of step 2.3, and generate V I and Q I values. Take the result of step 2.2 as unit input 2 and generate K E values through 3×3 convolution. First, perform dot product interaction on Q I and K E , then perform softmax operation to get the result. Then perform dot product interaction on the result and V I to obtain attention weights. Add the attention weights to the high-frequency subgraph of the result of step 2.3, and finally concatenate with the low-frequency subgraph of the result of step 2.3 to get the result F out .

Citation Information

Cited By

  • Single-exposure HDR image reconstruction method based on high and low frequency double branches

    CN121190324A