Multi-feature enhancement fusion Mama image denoising method

By building a Mamba denoising model and combining it with the Mamba long-distance dependency capture unit and local residual module, we solve the problem in existing technologies that it is difficult to balance the recovery of global structure and local details in complex noisy environments, achieving efficient image denoising effects and suitable for resource-constrained devices.

CN120707868APending Publication Date: 2025-09-26NANJING UNIV OF INFORMATION SCI & TECH
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510820370.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing deep learning methods find it difficult to simultaneously take into account global structure and local detail recovery in image denoising tasks, especially in complex noisy environments. In addition, the Transformer model has high computational complexity when processing large-scale images, which limits its practical application.

Method used

The Mamba image denoising method with multi-feature enhancement fusion is adopted. By constructing a Mamba denoising model, which includes a shallow feature extraction module, a deep feature extraction module and an image reconstruction module, the Mamba long-distance dependency capture unit, the local residual module and the channel attention mechanism are used to achieve the coordinated expression of global and local features, and the robustness and computational efficiency of the model are improved through small-batch training.

Benefits of technology

It effectively restores the global structure and local details of the image, improves image quality, reduces computational complexity, and improves the adaptability and computational efficiency of the model in complex noisy environments. It is suitable for resource-constrained devices such as remote sensing platforms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120707868A_ABST
    Figure CN120707868A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of image processing, and discloses a multi-feature enhancement fusion Mama image denoising method, which comprises the following steps: constructing a Mama denoising model for image denoising processing, the Mama denoising model comprising a shallow feature extraction module, a depth feature extraction module and an image reconstruction module; the shallow layer feature extraction module is used for extracting shallow layer features of the input image and performing normalization processing on the shallow layer features; the depth feature extraction module comprises a plurality of DFEGs and a sixth convolutional layer, and feature extraction is performed on the normalized shallow layer features step by step through the plurality of DFEGs; according to the invention, through a plurality of DFEGs in the depth feature extraction module, gradual feature extraction from a shallow layer to a deep layer is realized, and through cooperation of the Mama long-distance dependence capture unit, the local residual error module LRB and the channel attention mechanism CAB, the problem that global structure and local detail recovery are difficult to consider at the same time in a complex noise environment in the prior art is effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and in particular relates to a Mamba image denoising method based on multi-feature enhancement and fusion. Background Art

[0002] Images are often affected by noise during acquisition, transmission, and storage, resulting in degraded image quality. This noise can originate from imaging sensor instability, electromagnetic interference during signal transmission, or information loss during compressed storage. Common noise types include Gaussian noise, salt and pepper noise, and speckle noise. Noise can degrade image visual quality, leading to loss of detail, reduced contrast, and blurred structures. This can severely impact the human eye's perception of image content and the performance of subsequent image processing tasks (such as object detection, recognition, and segmentation).

[0003] With the rise of deep learning technology, models such as convolutional neural networks (CNNs) and Transformers have been widely used in image denoising tasks. Deep learning methods can automatically learn multi-level features of images and implement complex nonlinear mappings, thereby significantly improving the performance of image denoising. However, existing deep learning denoising methods still have some shortcomings. On the one hand, convolutional neural networks have difficulty in simultaneously considering the global structure and local details of the image during feature extraction, resulting in a lack of detail recovery in the denoised image. On the other hand, the computational complexity of the Transformer model increases exponentially when processing large-scale images, which limits its promotion in practical applications. In addition, when faced with high noise levels, complex backgrounds, or a mixture of multiple noise types, the robustness and generalization capabilities of existing models are also challenged. Therefore, existing technologies have the problem of difficulty in simultaneously considering the recovery of global structure and local details in complex noise environments. Summary of the Invention

[0004] In view of the shortcomings of the existing technology, the purpose of the present invention is to provide a Mamba image denoising method with multi-feature enhancement fusion, which solves the problem in the existing technology that it is difficult to simultaneously restore global structure and local details in a complex noisy environment.

[0005] The purpose of the present invention can be achieved through the following technical solutions:

[0006] A multi-feature enhancement fusion Mamba image denoising method includes the following steps:

[0007] Construct a Mamba denoising model for image denoising, which includes a shallow feature extraction module, a deep feature extraction module, and an image reconstruction module;

[0008] The shallow feature extraction module is used to extract the shallow features of the input image and normalize the shallow features;

[0009] The deep feature extraction module includes multiple DFEGs and a sixth convolutional layer. The multiple DFEGs gradually extract the shallow features after normalization. The sixth convolutional layer integrates the features extracted by multiple DFEGs to generate deep features, where DFEG is the depth extraction component.

[0010] The image reconstruction module is used to fuse shallow features and deep features, and combines convolution mapping and residual enhancement to restore image details and textures, and finally outputs a denoised image through denormalization.

[0011] Input the image dataset as the training set to train the constructed Mamba denoising model;

[0012] The image to be denoised is input into the trained Mamba denoising model for denoising.

[0013] The shallow feature extraction module includes the first convolutional layer, which contains a 3×3 convolution kernel;

[0014] Extracting shallow features of the input image includes the following steps:

[0015] Let the input image be represented as I LQ , I LQ ∈R H×W×3 , where H and W represent the vertical height and horizontal width of the image respectively, and 3 represents the original number of channels of the input image;

[0016] The image is processed by 3×3 convolution through the shallow feature extraction module to generate shallow features Fs, which is expressed as F S ∈R H ×W×C , where C represents the number of channels output by the first convolutional layer;

[0017] The shallow features Fs generated are normalized by the first normalization layer, and the normalized shallow features Fs are input into the deep feature extraction module.

[0018] In the deep feature extraction module, multiple DFEGs are stacked sequentially;

[0019] Any DFEG includes multiple layers of DFEM and a fifth convolutional layer. The fifth convolutional layer is used to further refine the features extracted by the multiple layers of DFEM in the corresponding DFEG;

[0020] Any DFEM includes a first normalization, a Mamba long-distance dependency capture unit, a second normalization layer, a local residual module LRB and a channel attention mechanism unit CAB connected in sequence, where the input end of the first normalization layer is connected to the output end of the first convolutional layer.

[0021] The Mamba long-distance dependency capture unit includes the second convolutional layer, the third convolutional layer, the depth convolutional layer, the OSSMamba module and the fourth convolutional layer;

[0022] The input of the second convolutional layer is connected to the output of the first normalization layer. The output of the second convolutional layer is independently connected to the input of the depthwise convolutional layer and the input of the third convolutional layer. The output of the third convolutional layer is connected to the middle layer of the OSSMamba module. The output of the depthwise convolutional layer is connected to the input of the OSSMamba module. The output of the OSSMamba module is connected to the input of the fourth convolutional layer. The output of the fourth convolutional layer is connected to the input of the second normalization layer. The second convolutional layer, the depthwise convolutional layer, and the fourth convolutional layer all use 1×1 convolution kernels.

[0023] The OSSMamba module includes a SiLU activation function and two Mamba blocks, which are the first Mamba block and the second Mamba block respectively;

[0024] The local residual module LRB includes a seventh convolutional layer, a GELU activation function and an eighth convolutional layer connected in sequence. The input of the seventh convolutional layer is connected to the output of the second normalization layer, the output of the seventh convolutional layer is connected to the input of the GELU activation function, and the output of the GELU activation function is connected to the input of the eighth convolutional layer. Both the seventh and eighth convolutional layers have 3×3 convolution kernels.

[0025] The Mamba long-distance dependency capture unit performs the following processing steps on the normalized shallow features:

[0026] The shallow features after normalization are processed by the second convolutional layer with 1×1 convolution to generate two information streams F O1 and F O2 ;

[0027] Information Flow F O1 After 1×1 convolution processing through the depth convolution layer, the information flow F is obtained O1D , information flow F O1D Enter the OSSMamba module;

[0028] Information Flow F O2 After the third convolutional layer performs 1×1 convolution processing, the information flow F with the shape of B×C×H×W is obtained. O2D , and input it into the OSSMamba module;

[0029] OSSMamba modules respectively process information flow F O1DAfter bidirectional scanning in the longitudinal and transverse directions, the two-dimensional planar information of the features is captured, and the two-dimensional planar information of the features is stacked and reshaped to obtain four features with a shape of B×C×H×W. The feature is refined based on the SiLU activation function to obtain an information flow F with a shape of B×4C×H×W. O1S ;

[0030] Information Flow F O1S After the first Mamba block processing, the information flow F is generated O4D , where the principle formula of the Mamba block is in discrete form, as follows:

[0031] Discrete state update equation:

[0032] Discrete output equation: y t =Ch t-1 +Dx t

[0033] Discrete parameter calculation:

[0034]

[0035] where h t ∈R N Indicates status; h t-1 ∈R N Indicates the state at the previous moment; x t ∈R represents input; y t ∈R represents the output; Δt is the time scale parameter; A, B, and C are all continuous parameters; All are discrete parameters;

[0036] Information flow F O4D Perform a merge operation to merge the information flow F O4D Reshape back to the dimension of B×C×H×W to obtain the information flow F with complementary information OP ;

[0037] The information flow F is converted into OP With F O2D Fusion is performed to obtain the feature map F OM ;

[0038] For the feature map F OM Perform pooling operation to reduce the amount of calculation while retaining the main feature information, and obtain the first feature map with a shape of B×1×C;

[0039] The first feature map is scanned bidirectionally, i.e., the channel is scanned forward and the channel is scanned backward. The results of the bidirectional scanning are activated, and the information of different channels is integrated to obtain the second feature map F.OC1 ;

[0040] For the second feature map F OC1 The stacking and reshaping are performed, and the results of the stacking and reshaping are input into the second Mamba block for processing. The processed results are merged to obtain the third feature map F OC ;

[0041] Based on element-by-element multiplication, the third feature map F OC With the feature map F OM The fusion is performed and the related features are merged based on element-by-element addition. The merged result is input into the fourth convolutional layer for 1×1 convolution processing and the final output is F OSS ;

[0042] The second normalization layer is used to normalize the output F OSS Perform normalization processing;

[0043] Normalized F OSS Enter the local residual block LRB and perform the following steps:

[0044] The seventh convolutional layer receives the normalized F OSS , and perform 3×3 convolution processing to output the fourth feature map;

[0045] The GELU activation function receives the fourth feature map, applies the GELU activation function to the fourth feature map, increases the nonlinear expression capability of the fourth feature map, and inputs it into the eighth convolutional layer;

[0046] The eighth convolutional layer is used to restore local pixel similarity, receives the fourth feature map after activation and performs 3×3 convolution processing, and outputs the fifth feature map F LRB .

[0047] The channel attention mechanism unit CAB includes global average pooling, two fully connected layers and Sigmoid function. The two fully connected layers are the first fully connected layer and the second fully connected layer respectively.

[0048] The input of the global average pooling is connected to the output of the eighth convolutional layer, and the global average pooling is used to receive the fifth feature map F LRB , and the fifth feature map F LRB The spatial dimension information is compressed to the channel dimension to obtain the channel descriptor;

[0049] The input of the first fully connected layer is connected to the output of the global average pooling to reduce the channel dimension of the channel descriptor;

[0050] The input end of the second fully connected layer is connected to the output end of the first fully connected layer to restore the channel dimension of the channel descriptor;

[0051] The input end of the Sigmoid function is connected to the output end of the second fully connected layer to map the output of the second fully connected layer to the interval [0,1] to obtain the attention weight of each channel;

[0052] The attention weight is combined with the fifth feature map F LRB Multiply and perform weighted operations on different channels to obtain a weighted feature map.

[0053] The sixth and fifth convolutional layers both have 3×3 convolution kernels;

[0054] The output ends of the multi-layer DFEM in the first DFEG are all connected to the input ends of the corresponding fifth convolutional layer. The weighted feature maps output from the multi-layer DFEM in the DFEG are processed with 3×3 convolution to further refine the weighted feature maps.

[0055] After the weighted feature map refined by the first DFEG passes through each DFEG in turn, the weighted feature map output by the last DFEG is input into the sixth convolutional layer;

[0056] The sixth convolutional layer performs a 3×3 convolution on the input weighted feature map to generate deep features.

[0057] The image reconstruction module includes a feature fusion layer, a mapping convolution layer, a residual enhancement unit, and an anti-normalization unit;

[0058] The image reconstruction module fuses shallow features and deep features, and combines convolution mapping and residual enhancement to restore image details and textures. Finally, it outputs a denoised image through denormalization. The specific steps include:

[0059] The deep features Input the feature fusion layer with the shallow feature Fs and add the elements to get F R ,

[0060] F R Input to the mapping convolution layer for 3×3 convolution processing, F R The mapping is restored to an RGB image with a spatial resolution of H×W×3 of the original input image;

[0061] The RGB image obtained by the mapping convolution layer is combined with I LQ Input the residual enhancement unit and perform residual connection to obtain I′ HQ ;

[0062] Will I′ HQ Input the denormalization unit, perform denormalization to restore the image value range, and finally output the denoised image I HQ .

[0063] The image dataset is used as the training set to train the constructed Mamba denoising model. The specific steps include:

[0064] Input image dataset and divide it into training set, validation set and test set according to the proportion;

[0065] Perform noise synthesis processing on the high-definition images in the training set to generate Gaussian white noise images with noise levels σ = 25 and σ = 50, forming high-definition-noise image pairs;

[0066] Perform various transformations on the training set data to improve the model's generalization ability, including horizontal flipping, random rotation, and image cropping;

[0067] Set training parameters including batch size BatchSize and learning rate;

[0068] Set the training strategy including the loss function;

[0069] Set the number of iterations and iteratively train the Mamba denoising model until the set number of iterations is reached;

[0070] The Mamba denoising model is trained on the validation set and the optimal Mamba denoising model parameters are saved.

[0071] The performance of the model is evaluated on the test set, and the evaluation indicators include peak signal-to-noise ratio and structural similarity.

[0072] Beneficial effects of the present invention:

[0073] 1. The present invention provides a multi-feature enhanced fusion Mamba image denoising method, which realizes gradual feature extraction from shallow to deep layers through multiple DFEGs in the deep feature extraction module. The Mamba long-distance dependency capture unit can effectively capture the global features of the image, while the local residual module LRB focuses on restoring local pixel similarities and enhancing the expression ability of local features, realizing the coordinated expression of global and local features, so that the denoised image can not only retain the overall structural information but also finely restore the detailed texture of the image; the introduction of the channel attention mechanism CAB, by weighting the channel dimension of the feature map, can highlight key channel features and reduce channel redundancy, thereby improving the expression ability and computational efficiency of the model; in addition, the multi-directional expansion and modeling of features by the OSSMamba module further enhances the structure and detail expression ability of the image, enabling the model to better capture complex patterns in the image, effectively solving the problem in the prior art that it is difficult to simultaneously take into account the restoration of global structure and local details in a complex noisy environment;

[0074] 2. During the training process, multiple transformation operations are performed on the training data to increase data diversity, thereby improving the model's adaptability to different image changes and reducing the risk of overfitting. Even when small-batch training introduces high gradient noise, the model of the present invention can still maintain or exceed the performance of large-batch training, has excellent robustness to gradient fluctuations, and is not prone to falling into local poor solutions. This enables the model to converge more stably to a flat solution region when faced with high noise levels, complex backgrounds, or a mixture of multiple noise types, achieving better generalization performance and providing a reliable denoising solution for complex scenarios in practical applications. At the same time, the small-batch training method significantly reduces video memory overhead, enabling the model to be more efficiently trained and deployed on resource-constrained devices, such as in application scenarios with high computing resource requirements such as remote sensing platforms, which has obvious advantages. Compared with the traditional Transformer model, the present invention has lower computational complexity when processing large-scale images, can more efficiently complete image denoising tasks, and improves feasibility and practicality in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0075] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0076] Figure 1 It is a schematic diagram of the overall process of the present invention;

[0077] Figure 2 It is a schematic diagram of the structure of the OSSMamba module of the present invention;

[0078] Figure 3 Schematic diagram of the local residual module LRB structure of the present invention;

[0079] Figure 4 3 is a schematic diagram comparing verification results of the present invention when the noise level is 50;

[0080] Figure 5 3 is a schematic diagram comparing the verification results of the present invention when the noise level is 25. DETAILED DESCRIPTION

[0081] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.

[0082] like Figures 1 to 5 As shown, a multi-feature enhanced fusion Mamba image denoising method includes the following steps:

[0083] Construct a Mamba denoising model for image denoising, which includes a shallow feature extraction module, a deep feature extraction module, and an image reconstruction module;

[0084] The shallow feature extraction module is used to extract the shallow features of the input image and normalize the shallow features;

[0085] The deep feature extraction module includes multiple DFEGs and a sixth convolutional layer. The multiple DFEGs gradually extract the shallow features after normalization. The sixth convolutional layer integrates the features extracted by multiple DFEGs to generate deep features, where DFEG is the depth extraction component.

[0086] The image reconstruction module is used to fuse shallow features and deep features, and combines convolution mapping and residual enhancement to restore image details and textures, and finally outputs a denoised image through denormalization.

[0087] Input the image dataset as the training set to train the constructed Mamba denoising model;

[0088] The image to be denoised is input into the trained Mamba denoising model for denoising.

[0089] The shallow feature extraction module includes the first convolutional layer, which contains a 3×3 convolution kernel;

[0090] Extracting shallow features of the input image (low-quality image to be denoised) specifically includes the following steps:

[0091] Let the input image be represented as I LQ , I LQ ∈R H×W×3 , where H and W represent the vertical height and horizontal width of the image respectively, and 3 represents the original number of channels of the input image;

[0092] The image is processed by 3×3 convolution through the shallow feature extraction module to generate shallow features Fs, which is expressed as F S ∈R H ×W×C , where C represents the number of channels output by the first convolutional layer;

[0093] The shallow features Fs generated are normalized by the first normalization layer, and the normalized shallow features Fs are input into the deep feature extraction module.

[0094] In the deep feature extraction module, multiple DFEGs are stacked sequentially;

[0095] Any DFEG includes multiple layers of DFEM and a fifth convolutional layer. The fifth convolutional layer is used to further refine the features extracted by the multiple layers of DFEM in the corresponding DFEG;

[0096] Any DFEM includes a first normalization, a Mamba long-distance dependency capture unit, a second normalization layer, a local residual module LRB and a channel attention mechanism unit CAB connected in sequence, where the input end of the first normalization layer is connected to the output end of the first convolutional layer.

[0097] The Mamba long-distance dependency capture unit includes the second convolutional layer, the third convolutional layer, the depth convolutional layer, the OSSMamba module and the fourth convolutional layer;

[0098] The input of the second convolutional layer is connected to the output of the first normalization layer. The output of the second convolutional layer is independently connected to the input of the depthwise convolutional layer and the input of the third convolutional layer. The output of the third convolutional layer is connected to the middle layer of the OSSMamba module. The output of the depthwise convolutional layer is connected to the input of the OSSMamba module. The output of the OSSMamba module is connected to the input of the fourth convolutional layer. The output of the fourth convolutional layer is connected to the input of the second normalization layer. The second convolutional layer, the depthwise convolutional layer, and the fourth convolutional layer all use 1×1 convolution kernels.

[0099] The OSSMamba module includes a SiLU activation function and two Mamba blocks, which are the first Mamba block and the second Mamba block respectively;

[0100] The local residual module LRB includes a seventh convolutional layer, a GELU activation function and an eighth convolutional layer connected in sequence. The input of the seventh convolutional layer is connected to the output of the second normalization layer, the output of the seventh convolutional layer is connected to the input of the GELU activation function, and the output of the GELU activation function is connected to the input of the eighth convolutional layer. Both the seventh and eighth convolutional layers have 3×3 convolution kernels.

[0101] The Mamba long-distance dependency capture unit performs the following processing steps on the normalized shallow features:

[0102] The shallow features after normalization are processed by the second convolutional layer with 1×1 convolution to generate two information streams F O1 and F O2 ;

[0103] Information Flow F O1 After 1×1 convolution processing through the depth convolution layer, the information flow F is obtained O1D , information flow F O1DEnter the OSSMamba module;

[0104] Information Flow F O2 After the third convolutional layer performs 1×1 convolution processing, the information flow F with the shape of B×C×H×W is obtained. O2D , and input it into the OSSMamba module;

[0105] OSSMamba modules respectively process information flow F O1D After performing bidirectional scanning in the longitudinal and transverse directions (the bidirectional scanning in the longitudinal and transverse directions includes H Forward Scan, H Backward Scan and W forward scan, W Backward Scan), the two-dimensional planar information of the features is captured, and the two-dimensional planar information of the features is stacked and reshaped to obtain four features with a shape of B×C×H×W. Feature refinement is performed based on the SiLU activation function to obtain an information flow F with a shape of B×4C×H×W. O1S ;

[0106] Information Flow F O1S After the first Mamba block processing, the information flow F is generated O4D , where the principle formula of the Mamba block is in discrete form, as follows:

[0107] Discrete state update equation:

[0108] Discrete output equation: y t =Ch t-1 +Dx t

[0109] Discrete parameter calculation:

[0110]

[0111] where h t ∈R N Indicates status; h t-1 ∈R N Indicates the state at the previous moment; x t ∈R represents input; y t ∈R represents the output; Δt is the time scale parameter; A, B, and C are all continuous parameters; All are discrete parameters;

[0112] Information flow F O4D Perform a merge operation to merge the information flow F O4D Reshape back to the dimension of B×C×H×W to obtain the information flow F with complementary information OP ;

[0113] The information flow F is converted into OP With F O2D Fusion is performed to obtain the feature map F OM ;

[0114] For the feature map F OM Perform pooling operation to reduce the amount of calculation while retaining the main feature information, and obtain the first feature map with a shape of B×1×C;

[0115] The first feature map is scanned bidirectionally by channel forward scanning (C Forward Scan) and channel reverse scanning (C Backward Scan), and the results of the bidirectional scanning are activated to integrate the information of different channels to obtain the second feature map F OC1 ;

[0116] For the second feature map F OC1 The stacking and reshaping are performed, and the results of the stacking and reshaping are input into the second Mamba block for processing. The processed results are merged to obtain the third feature map F OC ;

[0117] Based on element-by-element multiplication, the third feature map F OC With the feature map F OM The fusion is performed and the related features are merged based on element-by-element addition. The merged result is input into the fourth convolutional layer for 1×1 convolution processing and the final output is F OSS ;

[0118] The second normalization layer is used to normalize the output F OSS Perform normalization, which can improve training stability, accelerate convergence, and enhance generalization ability, ensuring greater stability and training efficiency;

[0119] Normalized F OSS Enter the local residual block LRB and perform the following steps:

[0120] The seventh convolutional layer receives the normalized F OSS , and perform 3×3 convolution processing to output the fourth feature map;

[0121] The GELU activation function receives the fourth feature map, applies the GELU activation function to the fourth feature map, increases the nonlinear expression capability of the fourth feature map, and inputs it into the eighth convolutional layer;

[0122] The eighth convolutional layer is used to restore local pixel similarity, receives the fourth feature map after activation and performs 3×3 convolution processing, and outputs the fifth feature map F LRB .

[0123] The channel attention mechanism unit CAB includes global average pooling, two fully connected layers and Sigmoid function. The two fully connected layers are the first fully connected layer and the second fully connected layer respectively.

[0124] The input of the global average pooling is connected to the output of the eighth convolutional layer, and the global average pooling is used to receive the fifth feature map F LRB , and the fifth feature map F LRB The spatial dimension information is compressed to the channel dimension to obtain the channel descriptor;

[0125] The input of the first fully connected layer is connected to the output of the global average pooling to reduce the channel dimension of the channel descriptor;

[0126] The input end of the second fully connected layer is connected to the output end of the first fully connected layer to restore the channel dimension of the channel descriptor;

[0127] The Sigmoid function input is connected to the output of the second fully connected layer to map the output of the second fully connected layer to the interval [0, 1], obtain the attention weight of each channel, implement weighted operations on different channels, highlight key channel features, reduce channel redundancy, and improve the expressiveness of the model;

[0128] The attention weight is combined with the fifth feature map F LRB Multiply and perform weighted operations on different channels to obtain a weighted feature map.

[0129] The sixth and fifth convolutional layers both have 3×3 convolution kernels;

[0130] The outputs of the multi-layer DFEM in the first DFEG are all connected to the inputs of the corresponding fifth convolutional layer. The weighted feature maps output by the multi-layer DFEM in the DFEG are subjected to 3×3 convolution processing to further refine the weighted feature maps, compensate for the loss of local features, further solve the problem of local pixel forgetting, and restore the similarity between spatially adjacent pixels. This allows the Mamba denoising model to effectively capture local information, which is conducive to the model's accurate restoration of image details.

[0131] After the weighted feature map refined by the first DFEG passes through each DFEG in turn, the weighted feature map output by the last DFEG is input into the sixth convolutional layer;

[0132] The sixth convolutional layer performs a 3×3 convolution on the input weighted feature map to generate deep features.

[0133] The image reconstruction module includes a feature fusion layer, a mapping convolution layer, a residual enhancement unit, and an anti-normalization unit;

[0134] The image reconstruction module fuses shallow features and deep features, and combines convolution mapping and residual enhancement to restore image details and textures. Finally, it outputs a denoised image through denormalization. The specific steps include:

[0135] The deep features Input the feature fusion layer with the shallow feature Fs and add the elements to get F R ,

[0136] F R Input to the mapping convolution layer for 3×3 convolution processing, F R The mapping is restored to an RGB image with a spatial resolution of H×W×3 of the original input image;

[0137] The RGB image obtained by the mapping convolution layer is combined with I LQ Input the residual enhancement unit and perform residual connection to obtain I′ HQ ;

[0138] Will I′ HQ Input the denormalization unit, perform denormalization to restore the image value range, and finally output the denoised image I HQ ; Fuse features at different levels and make full use of the information extracted from the image at different stages to restore the details and texture of the image and improve the quality of the reconstructed image.

[0139] The image dataset is used as the training set to train the constructed Mamba denoising model. The specific steps include:

[0140] Input image dataset and divide it into training set, validation set and test set according to the proportion;

[0141] Perform noise synthesis processing on the high-definition images in the training set to generate Gaussian white noise images with noise levels σ = 25 and σ = 50, forming high-definition-noise image pairs;

[0142] Perform various transformations on the training set data to improve the model's generalization ability, including horizontal flipping, random rotation, and image cropping;

[0143] Horizontal flipping involves mirroring the image in the horizontal direction;

[0144] Random rotation includes rotating the image by 90°, 180°, and 270°;

[0145] Image cropping involves cropping the image into 128×128 image blocks;

[0146] Set training parameters including batch size BatchSize and learning rate;

[0147] Set the training strategy including the loss function;

[0148] Set the number of iterations and iteratively train the Mamba denoising model until the set number of iterations is reached;

[0149] The Mamba denoising model is trained on the validation set and the optimal Mamba denoising model parameters are saved.

[0150] The performance of the model is evaluated on the test set, and the evaluation indicators include peak signal-to-noise ratio (PSNR) and structural similarity (SSIM).

[0151] The present invention provides a multi-feature enhanced fusion Mamba image denoising method, which realizes gradual feature extraction from shallow to deep layers through multiple DFEGs in the deep feature extraction module. The Mamba long-distance dependency capture unit can effectively capture the global features of the image, while the local residual module (LRB) focuses on restoring local pixel similarities and enhancing the expression ability of local features, thereby realizing the coordinated expression of global and local features. The denoised image can not only retain the overall structural information but also finely restore the detailed texture of the image. The introduction of the channel attention mechanism (CAB) can highlight key channel features and reduce channel redundancy by weighting the channel dimensions of the feature map, thereby improving the expression ability and computational efficiency of the model. In addition, the multi-directional expansion and modeling of features by the OSSMamba module further enhances the image structure and detail expression ability, enabling the model to better capture complex patterns in the image, effectively solving the problem in the prior art that it is difficult to simultaneously take into account the restoration of global structure and local details in complex noisy environments. During the training process, multiple transformation operations are performed on the training data to increase the diversity of the data, thereby improving the model's adaptability to different image changes and reducing the risk of overfitting.

[0152] The experimental design uses public datasets (such as CBSD68 and Urban100) for testing. Gaussian noise of different intensities (σ = 25 and 50) is added. The number of iterations is set to 50,000 (lightweight), the learning rate is MultiStepLR, the initial learning rate is 1e-4, and the optimizer is Adam. The comparison methods include DnCNN, DRUNet, SCUNet, Restormer, and MambaIR. The evaluation metrics are Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity (SSIM).

[0153] like Figure 4 、 Figure 5 As shown, the present invention is compared with the existing method in an experiment: the denoising effect is compared under the same noise level (σ=25, σ=50); wherein, the image denoising method of the present invention is referred to as OSSMIR;

[0154] Results Analysis: The test results are shown in Table 1 below. Our architecture is compared with the following five existing advanced methods. Regardless of whether the noise level is 25 or 50, our proposed method performs best. In addition, our architecture uses a small batch size (i.e., 1), while other methods use batch sizes of 4 or 32 during training. Even in a higher gradient noise environment (small batch), our model can still maintain or exceed the performance of large batches, which shows that it has excellent robustness to gradient fluctuations and is not prone to falling into local suboptimal solutions. Despite the same training rounds and learning rate settings, our method still achieves higher PSNR and SSIM on the BSD68 and Urban100 datasets. This phenomenon shows that under the gradient noise introduced by small batches and the effects of random regularization, our network structure and optimization strategy can more stably converge to the flat solution region, thereby achieving better generalization performance.

[0155] At the same time, mini-batch training significantly reduces graphics memory overhead, enabling the model to be trained and deployed more efficiently on resource-constrained devices. This offers significant advantages in computationally demanding applications, such as remote sensing platforms. Compared to traditional Transformer models, this proposed method reduces computational complexity when processing large-scale images, enabling more efficient image denoising, and improving feasibility and practicality in practical applications.

[0156] In comparative experiments with various existing advanced denoising methods (such as DnCNN, DRUNet, SCUNet, Restormer and MambaIR), the method of the present invention achieved the best denoising effect under different noise levels (σ=25 and σ=50). Evaluation indicators such as peak signal-to-noise ratio (PSNR) and structural similarity (SSIM) were better than those of other methods, indicating that the present invention can more effectively remove noise from images while better preserving image details and texture information, providing higher quality image input for subsequent image processing tasks, and has important application value.

[0157] Table 1

[0158]

[0159] Throughout this specification, references to terms such as "one embodiment," "example," or "specific example" indicate that the specific features, structures, materials, or characteristics described in conjunction with that embodiment or example are included in at least one embodiment or example of the present invention. In this specification, schematic representations of these terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0160] The basic principles, main features, and advantages of the present invention are shown and described above. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions are merely illustrative of the principles of the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention, and such changes and modifications are intended to fall within the scope of the present invention.

Claims

1. A Mamba image denoising method based on multi-feature enhancement fusion, characterized in that: The following steps are involved: Construct a Mamba denoising model for image denoising, which includes a shallow feature extraction module, a deep feature extraction module, and an image reconstruction module; The shallow feature extraction module is used to extract the shallow features of the input image and normalize the shallow features; The deep feature extraction module includes multiple DFEGs and a sixth convolutional layer. The multiple DFEGs gradually extract the shallow features after normalization. The sixth convolutional layer integrates the features extracted by multiple DFEGs to generate deep features, where DFEG is the depth extraction component. The image reconstruction module is used to fuse shallow features and deep features, and combines convolution mapping and residual enhancement to restore image details and textures, and finally outputs a denoised image through denormalization. Input the image dataset as the training set to train the constructed Mamba denoising model; The image to be denoised is input into the trained Mamba denoising model for denoising.

2. The Mamba image denoising method based on multi-feature enhancement fusion according to claim 1 is characterized in that: The shallow feature extraction module includes the first convolutional layer, which contains a 3×3 convolution kernel; Extract shallow features of the input image, specifically The following steps are involved: Let the input image be represented as I LQ , I LQ ∈R H×W×3 , where H and W represent the vertical height and horizontal width of the image respectively, and 3 represents the original number of channels of the input image; The image is processed by 3×3 convolution through the shallow feature extraction module to generate shallow features Fs, which is expressed as F S ∈R H×W×C , where C represents the number of channels of the output of the first convolutional layer.

3. The Mamba image denoising method based on multi-feature enhancement fusion according to claim 2 is characterized in that: In the deep feature extraction module, multiple DFEGs are stacked sequentially; Any DFEG includes multiple layers of DFEM and a fifth convolutional layer. The fifth convolutional layer is used to further refine the features extracted by the multiple layers of DFEM in the corresponding DFEG; Any DFEM includes a first normalization, a Mamba long-distance dependency capture unit, a second normalization layer, a local residual module LRB and a channel attention mechanism unit CAB connected in sequence, where the input end of the first normalization layer is connected to the output end of the first convolutional layer.

4. The Mamba image denoising method based on multi-feature enhancement fusion according to claim 3 is characterized in that: The Mamba long-distance dependency capture unit includes the second convolutional layer, the third convolutional layer, the depth convolutional layer, the OSSMamba module and the fourth convolutional layer; The input of the second convolutional layer is connected to the output of the first normalization layer. The output of the second convolutional layer is independently connected to the input of the depthwise convolutional layer and the input of the third convolutional layer. The output of the third convolutional layer is connected to the middle layer of the OSSMamba module. The output of the depthwise convolutional layer is connected to the input of the OSSMamba module. The output of the OSSMamba module is connected to the input of the fourth convolutional layer. The output of the fourth convolutional layer is connected to the input of the second normalization layer. The second, third, depthwise convolutional layers, and fourth convolutional layers all use 1×1 convolution kernels. The OSSMamba module includes a SiLU activation function and two Mamba blocks, which are the first Mamba block and the second Mamba block respectively; The local residual module LRB includes a seventh convolutional layer, a GELU activation function and an eighth convolutional layer connected in sequence. The input of the seventh convolutional layer is connected to the output of the second normalization layer, the output of the seventh convolutional layer is connected to the input of the GELU activation function, and the output of the GELU activation function is connected to the input of the eighth convolutional layer. Both the seventh and eighth convolutional layers have 3×3 convolution kernels.

5. The Mamba image denoising method based on multi-feature enhancement fusion according to claim 4 is characterized in that: The Mamba long-distance dependency capture unit performs the following processing steps on the normalized shallow features: The shallow features after normalization are processed by the second convolutional layer with 1×1 convolution to generate two information streams F O1 and F O2 ; Information Flow F O1 After 1×1 convolution processing through the depth convolution layer, the information flow F is obtained O1D , information flow F O1D Enter the OSSMamba module; Information Flow F O2 After the third convolutional layer performs 1×1 convolution processing, the information flow F with the shape of B×C×H×W is obtained. O2D , and input it into the OSSMamba module; OSSMamba modules respectively process information flow F O1D After bidirectional scanning in the longitudinal and transverse directions, the two-dimensional planar information of the features is captured, and the two-dimensional planar information of the features is stacked and reshaped to obtain four features with a shape of B×C×H×W. The feature is refined based on the SiLU activation function to obtain an information flow F with a shape of B×4C×H×W. O1S ; Information Flow F O1S After the first Mamba block processing, the information flow F is generated O4D , where the principle formula of the Mamba block is in discrete form, as follows: Discrete state update equation: Discrete output equation: y t =Ch t-1 +Dx t Discrete parameter calculation: where h t ∈R N Indicates status; h t-1 ∈R N Indicates the state at the previous moment; x t ∈R represents input; y t ∈R represents the output; Δt is the time scale parameter; A, B, and C are all continuous parameters; All are discrete parameters; Information flow F O4D Perform a merge operation to merge the information flow F O4D Reshape back to the dimension of B×C×H×W to obtain the information flow F with complementary information OP ; The information flow F is converted into OP With F O2D Fusion is performed to obtain the feature map F OM ; For the feature map F OM Perform pooling operation to reduce the amount of calculation while retaining the main feature information, and obtain the first feature map with a shape of B×1×C; The first feature map is scanned bidirectionally, i.e., the channel is scanned forward and the channel is scanned backward. The results of the bidirectional scanning are activated, and the information of different channels is integrated to obtain the second feature map F. OC1 ; For the second feature map F OC1 The stacking and reshaping are performed, and the results of the stacking and reshaping are input into the second Mamba block for processing. The processed results are merged to obtain the third feature map F OC ; Based on element-by-element multiplication, the third feature map F OC With the feature map F OM The fusion is performed and the related features are merged based on element-by-element addition. The merged result is input into the fourth convolutional layer for 1×1 convolution processing and the final output is F OSS .

6. The Mamba image denoising method based on multi-feature enhancement fusion according to claim 5 is characterized in that: The second normalization layer is used to normalize the output F OSS Perform normalization processing; Normalized F OSS Enter the local residual block LRB and perform the following steps: The seventh convolutional layer receives the normalized F OSS , and perform 3×3 convolution processing to output the fourth feature map; The GELU activation function receives the fourth feature map, applies the GELU activation function to the fourth feature map, increases the nonlinear expression capability of the fourth feature map, and inputs it into the eighth convolutional layer; The eighth convolutional layer is used to restore local pixel similarity, receives the fourth feature map after activation and performs 3×3 convolution processing, and outputs the fifth feature map F LRB .

7. The Mamba image denoising method based on multi-feature enhancement fusion according to claim 6, characterized in that: The channel attention mechanism unit CAB includes global average pooling, two fully connected layers and Sigmoid function. The two fully connected layers are the first fully connected layer and the second fully connected layer respectively. The input of the global average pooling is connected to the output of the eighth convolutional layer, and the global average pooling is used to receive the fifth feature map F LRB , and the fifth feature map F LRB The spatial dimension information is compressed to the channel dimension to obtain the channel descriptor; The input of the first fully connected layer is connected to the output of the global average pooling to reduce the channel dimension of the channel descriptor; The input end of the second fully connected layer is connected to the output end of the first fully connected layer to restore the channel dimension of the channel descriptor; The input end of the Sigmoid function is connected to the output end of the second fully connected layer to map the output of the second fully connected layer to the interval [0,1] to obtain the attention weight of each channel; The attention weight is combined with the fifth feature map F LRB Multiply and perform weighted operations on different channels to obtain a weighted feature map.

8. The Mamba image denoising method based on multi-feature enhancement fusion according to claim 7 is characterized in that: The sixth and fifth convolutional layers both have 3×3 convolution kernels; The output ends of the multi-layer DFEM in the first DFEG are all connected to the input ends of the corresponding fifth convolutional layer. The weighted feature maps output from the multi-layer DFEM in the DFEG are processed with 3×3 convolution to further refine the weighted feature maps. After the weighted feature map refined by the first DFEG passes through each DFEG in turn, the weighted feature map output by the last DFEG is input into the sixth convolutional layer; The sixth convolutional layer performs a 3×3 convolution on the input weighted feature map to generate deep features.

9. The Mamba image denoising method based on multi-feature enhancement fusion according to claim 8, characterized in that: The image reconstruction module includes a feature fusion layer, a mapping convolution layer, a residual enhancement unit, and an anti-normalization unit; The image reconstruction module fuses shallow features and deep features, and combines convolution mapping and residual enhancement to restore image details and textures. Finally, it outputs a denoised image through denormalization. The specific steps include: The deep features Input the feature fusion layer with the shallow feature Fs and add the elements to get F R , F R Input to the mapping convolution layer for 3×3 convolution processing, F R The mapping is restored to an RGB image with a spatial resolution of H×W×3 of the original input image; The RGB image obtained by the mapping convolution layer is combined with I LQ Input the residual enhancement unit and perform residual connection to obtain I′ HQ ; Will I′ HQ Input the denormalization unit, perform denormalization to restore the image value range, and finally output the denoised image I HQ .

10. The Mamba image denoising method based on multi-feature enhancement fusion according to claim 9, characterized in that: The image dataset is used as the training set to train the constructed Mamba denoising model. The specific steps include: Input image dataset and divide it into training set, validation set and test set according to the proportion; Perform noise synthesis processing on the high-definition images in the training set to generate Gaussian white noise images with noise levels σ = 25 and σ = 50, forming high-definition-noise image pairs; Perform various transformations on the training set data to improve the model's generalization ability, including horizontal flipping, random rotation, and image cropping; Set training parameters including batch size BatchSize and learning rate; Set the training strategy including the loss function; Set the number of iterations and iteratively train the Mamba denoising model until the set number of iterations is reached; The Mamba denoising model is trained on the validation set and the optimal Mamba denoising model parameters are saved. The performance of the model is evaluated on the test set, and the evaluation indicators include peak signal-to-noise ratio and structural similarity.

Citation Information

Cited By

  • Image processing method, processing device and computer product

    CN121353749A

  • Generalized image restoration model based on contour prior guided mamba diffusion

    CN122391018A

  • Generalized image restoration apparatus based on contour prior guided mamba diffusion

    CN122391018B