Double-domain heterogeneous image denoising method
Through the dual-domain heterogeneous image denoising method, combined with frequency domain and spatial domain processing, and using Transformer and Mamba models, the shortcomings in the image denoising method in the prior art in terms of global structure and local details are solved, and an efficient and clear image denoising effect is achieved.
Patent Information
- Application Number
- CN202510444153.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-04-10
AI Technical Summary
When processing noise, existing image denoising methods are difficult to maintain the global structure and local details of the image at the same time, and the calculation complexity is high, resulting in poor denoising effect, especially in application scenarios that require high-quality image processing.
The dual-domain heterogeneous image denoising method is adopted, and the dual-domain parallel M-T Block processing is used in the shallow layer through a hierarchical dual-drive encoding and decoding architecture, and the deep switches to the lightweight T Block. Combining the frequency domain and spatial domain branches, the global frequency domain features and long-domain spatial dependence are captured respectively by the Transformer and Mamba models, and the feature fusion is optimized through the jump connection and the vertical stripe perception fusion attention mechanism module.
It significantly reduces computational redundancy, improves the efficiency and quality of image denoising, makes the denoised image clearer in global and local details, retains key information, and improves the overall quality of the image.
Smart Images

Figure CN120374438A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image denoising, and in particular to a dual-domain heterogeneous image denoising method. Background Art
[0002] At present, in the field of image processing, image denoising has always been a core and key research direction. In today's digital age, images are widely used in many fields, from daily photography to professional medical images, satellite remote sensing images, and industrial inspection images. However, in the image acquisition process, due to the physical characteristics of the sensor, environmental interference and other factors, noise is inevitably mixed into the image data; in the transmission process, network instability, signal interference, etc. may also cause noise in the image; during storage, the characteristics and potential damage of the storage medium may also introduce noise. Like the common Gaussian noise, it is usually caused by electronic circuit noise and sensor noise, showing a grayscale change that obeys the Gaussian distribution, which will make the image as a whole blurred; salt and pepper noise is manifested as black and white pixels that appear randomly in the image, like salt and pepper sprinkled on the image, which seriously damages the visual effect of the image. These noises greatly reduce the quality of the image, make the image details blurred, and greatly reduce the recognizability of the information. For subsequent image analysis tasks, such as feature extraction and classification of objects in the image, if there is noise interference in the image, the extracted features may be inaccurate, resulting in classification errors. In the field of target recognition, whether it is face recognition in security monitoring or road target recognition in autonomous driving, noise may cause the recognition algorithm to misjudge or fail to recognize the target, resulting in serious consequences. In medical image diagnosis, noise interference may mask the characteristics of lesions, leading to misdiagnosis or missed diagnosis by doctors, endangering the life and health of patients. The birth of the Transformer architecture has brought new opportunities to solve the problem of long-range dependencies. Based on the self-attention mechanism, it can effectively capture long-range dependencies when processing sequence data. It has achieved great success in the field of natural language processing and has gradually been introduced into image processing. However, Transformer is relatively weak in local feature extraction, has high computational complexity, and low processing efficiency. As an emerging model architecture, Mamba has a unique structure and efficient computing performance, providing a new way for image denoising. The image denoising method based on the Mamba+Transformer dual branch is expected to combine the advantages of both, give full play to the efficiency of Mamba and the ability of Transformer to process long-range dependencies, perform comprehensive feature extraction and denoising on images from different scales and domains, and open up a new way to improve the image denoising effect. It has broad application prospects and important research value in the field of image processing.
[0003] Traditional image denoising methods mainly rely on filtering techniques. For example, mean filtering replaces the central pixel value by calculating the average of neighboring pixels to smooth the noise; median filtering replaces the central pixel value with the median of pixel values in the neighborhood, which has a good effect on suppressing salt-and-pepper noise. However, these methods have obvious limitations. While smoothing the noise, they over-blur the image edges and details. For instance, when processing images containing fine textures, the texture information disappears after noise removal, resulting in image distortion and unable to meet application scenarios with high requirements for image details. With the booming development of deep learning technology, image denoising methods based on neural networks have shown great advantages. Convolutional neural networks (CNNs), with structures such as convolutional layers and pooling layers, can automatically learn the feature representations of images and have achieved remarkable results in image denoising tasks. However, the convolutional operation of CNNs is essentially local, with a limited receptive field size, and it is difficult to handle the long-range dependence relationships between pixels at long distances in the image, making it difficult to capture the global information of the image. As a result, the denoised image may have unreasonable global structures, affecting the integrity of the denoising effect.
[0004] Therefore, the present invention proposes a dual-domain heterogeneous image denoising method. Summary of the Invention
[0005] The present invention provides a dual-domain heterogeneous image denoising method. In the shallow layers (the first and second layers), a dual-domain parallel M-TBlock is adopted, and in the deep layers (the third and fourth layers), it switches to a lightweight T Block. Through hierarchical feature adaptation (the shallow layers take both details and local structures into account, while the deep layers focus on global semantics) and dynamic allocation of computing resources (high-resolution features in the shallow layers are processed by the M-T Block, and low-resolution features in the deep layers are handled by the lightweight T Block), the redundant computing burden is significantly reduced. The frequency domain (FreqFormer) and spatial domain (SpMamba) branches are combined in parallel, and global frequency domain features and long-range spatial dependencies are captured through Transformer and Mamba respectively, solving the limitations of single-domain processing in traditional methods. V-CBAM is embedded in the encoder-decoder skip connections to reconstruct the cross-level feature fusion process. Channel attention filters important feature channels, and vertical spatial attention adds a vertical Sobel convolution kernel to enhance the sensitivity to vertical stripe noise and specifically suppress the noise area.
[0006] The present invention provides a dual-domain heterogeneous image denoising method, including:
[0007] S1: Extract the preliminary feature map of the input image based on depthwise separable convolution;
[0008] S2: Use a hierarchical dual-driven encoder-decoder architecture to hierarchically process the preliminary features to obtain multiple layers of features, including:
[0009] In the shallow layer of the hierarchical dual-driven encoding and decoding architecture, an encoder of the dual-domain heterogeneous collaborative architecture is adopted for dual-domain heterogeneous collaborative processing, including:
[0010] Frequency domain processing branch: Reorganize the spectral bands and assign attention weights to the image frequency domain information to obtain the processed image frequency domain information;
[0011] Spatial domain processing branch: Capture the long-range spatial dependencies of the image, and combine the convolutional enhanced feed-forward network to suppress local pixel forgetting to obtain the processed image spatial domain information;
[0012] In the deep layer of the hierarchical dual-driven encoding and decoding architecture, switch to the T Block, and adopt the multi-head transposed attention mechanism and the deep spatial enhanced feed-forward network for global context modeling and local feature enhancement;
[0013] S3: Concatenate the processed image frequency domain information and the image spatial domain information to obtain the features of each layer after hierarchical processing;
[0014] S4: Based on the vertical stripe perception fusion attention mechanism module between the encoder and the decoder with skip connections, fuse the features of each layer after hierarchical processing and output them to the decoder to obtain the denoised image.
[0015] Preferably, for the dual-domain heterogeneous image denoising method, reorganize the spectral bands and assign attention weights to the image frequency domain information to obtain the processed image frequency domain information, including:
[0016] Use the frequency domain self-attention module to extract Q, K, and V through depthwise separable convolution, and use the fast Fourier transform to transform the spatial domain features to the spectral domain;
[0017] Classify the frequency domain elements into different frequency bands through spectral band reorganization, and assign different attention weights between different frequency bands. The attention in the low-frequency band is reduced, while the attention in the high-frequency band is enhanced to obtain the frequency domain attention features;
[0018] Convert the frequency domain attention features back to the spatial domain through the inverse fast Fourier transform, and output the spatial features through the convolutional layer;
[0019] Use the frequency domain enhanced feed-forward module to increase the channel dimension through the point convolutional layer, perform processing using parallel dilated convolutions, then convert the features of the two branches to the frequency domain, enhance the spectral domain information through learnable weights and bias variables respectively, and through the gating mechanism, use the activation output of one branch with a longer receptive field as the gating unit of the other branch to dynamically adjust feature enhancement and perform feature fusion, and obtain the processed image frequency domain information based on the output of the 1×1 convolution.
[0020] Preferably, the dual-domain heterogeneous image denoising method captures the long-range spatial dependence of the image, combines a convolutional enhanced feedforward network to suppress local pixel forgetting, and obtains the spatial domain information of the processed image, including:
[0021] The visual state space module scans the two-dimensional features in four diagonal directions to generate four 1D sequences in four directions, and then applies the discrete state space equation to each 1D sequence in each direction to obtain the results in four directions;
[0022] The results in four directions are combined through a summation operation, then the two-dimensional structure is restored through a reshape operation, and then a convolutional enhanced feedforward network is used to enhance the interaction ability of local features and suppress channel redundant information to obtain the spatial domain information of the processed image.
[0023] Preferably, the dual-domain heterogeneous image denoising method uses the encoder of the dual-domain heterogeneous collaborative architecture for dual-domain heterogeneous collaborative processing in the shallow layer of the hierarchical dual-driven encoding and decoding architecture, and further includes:
[0024] Calculate the adjusted number of channels for each layer based on the size of the feature map processed by each layer;
[0025] Calculate the computational balance error of the hierarchical dual-driven encoding and decoding architecture based on the adjusted number of channels for all layers and the size of the feature maps processed by all layers;
[0026] When the computational balance error of the hierarchical dual-driven encoding and decoding architecture is greater than the preset error threshold, the hierarchical channel dynamic allocation strategy is triggered.
[0027] Preferably, the dual-domain heterogeneous image denoising method calculates the adjusted number of channels for each layer based on the size of the feature map processed by each layer, including:
[0028]
[0029] where C is the adjusted number of channels for a single layer, C base is the basic number of channels, α is the expansion intensity coefficient, tanh() is the hyperbolic tangent function, log2() is the logarithm function with base 2, H is the height of the feature map processed by the corresponding layer, W is the width of the feature map processed by the corresponding layer, H0 is the height of the input image, and W0 is the width of the input image.
[0030] Preferably, the dual-domain heterogeneous image denoising method calculates the computational balance error of the hierarchical dual-driven encoding and decoding architecture based on the adjusted number of channels for all layers and the size of the feature maps processed by all layers, including:
[0031]
[0032] where δ is the computational balance error of the hierarchical dual-driven encoding and decoding architecture, n is the total number of layers of the hierarchical dual-driven encoding and decoding architecture, γi is the weight of the i-th layer, C i is the adjusted number of channels of the i-th layer, FLOPS0 is the computational amount of a single channel, H0 is the height of the input image, W0 is the width of the input image, H i is the height of the feature map processed by the i-th layer, W i is the width of the feature map processed by the i-th layer.
[0033] Preferably, in the dual-domain heterogeneous image denoising method, when the computational amount balance error of the hierarchical dual-driven encoding and decoding architecture is greater than a preset error threshold, a hierarchical channel dynamic allocation strategy is triggered, including:
[0034] When the computational amount balance error of the hierarchical dual-driven encoding and decoding architecture is greater than a preset error threshold, the size of the feature map processed by the corresponding layer, the size of the input image, and the computational amount balance error of the hierarchical dual-driven encoding and decoding architecture are input into the hierarchical channel dynamic allocation model to obtain the finally adjusted number of channels for each layer.
[0035] Preferably, in the dual-domain heterogeneous image denoising method, a multi-head transposed attention mechanism and a depth spatial enhancement feed-forward network are used for global context modeling and local feature enhancement, including:
[0036] Use multi-head transposed attention to extract Q, K, and V matrices through depthwise separable convolution and perform global context modeling;
[0037] Use a depth spatial enhancement feed-forward network and use depthwise separable convolution to extract local features, and screen key regions through spatial attention gating.
[0038] Preferably, in the dual-domain heterogeneous image denoising method, the vertical stripe perception fusion attention mechanism module optimizes cross-layer feature fusion through the synergistic effect of channel attention and vertically enhanced spatial attention.
[0039] Preferably, in the dual-domain heterogeneous image denoising method, the vertical stripe perception fusion attention mechanism module optimizes cross-layer feature fusion through the synergistic effect of channel attention and vertically enhanced spatial attention, including:
[0040] Perform global average pooling and max pooling on the input features, and generate a channel weight vector through two fully connected layers;
[0041] Use a vertical Sobel operator to extract vertical edge responses, and splice them with the horizontal convolution results to generate a spatial weight matrix;
[0042] At the same time, based on the enhancement mechanism, dynamically allocate the weights of the channel attention branch and the spatial attention branch until the output is the element-wise product of the channel weight vector and the spatial weight matrix.
[0043] The beneficial effects of the present invention compared with the prior art are as follows:
[0044] The initial feature map of the input image is extracted by using depthwise separable convolution. This convolution method can effectively capture the basic features of the image while reducing the computational complexity, laying a foundation for subsequent denoising processing. It can quickly perform preliminary processing on the image without losing too much feature information, improving the efficiency of the entire denoising process, and enabling the method to maintain good performance when processing large-scale image data.
[0045] The hierarchical dual-driven encoder-decoder architecture is adopted to hierarchically process the preliminary features, giving full play to the advantages of different processing stages and different branches.
[0046] In the shallow layer, an encoder with a dual-domain heterogeneous collaborative architecture is used, which processes in parallel through two branches in the frequency domain and the spatial domain. The frequency domain processing branch performs spectral band recombination and attention weight assignment on the frequency domain information of the image, which can specifically adjust the energy distribution of the image in the frequency domain, highlight key information, and remove frequency domain noise; the spatial domain processing branch captures the long-range spatial dependence of the image, combines convolution with an enhanced feedforward network to suppress the forgetting of local pixels, and effectively maintains the structure and detail information of the image in the spatial dimension, avoiding the loss of local information during the denoising process. The two work together to optimize the image from different dimensions.
[0047] In the deep layer, it switches to the T Block, and the multi-head transposed attention mechanism and the deep spatial enhanced feedforward network are used for global context modeling and local feature enhancement. The multi-head transposed attention mechanism can capture the dependence relationships between different regions of the image globally, and the deep spatial enhanced feedforward network further enhances the local features, making the denoised image not only consistent as a whole but also more clear in local details, effectively improving the image quality.
[0048] The processed image frequency domain information and image spatial domain information are concatenated, which can fully integrate the advantageous features of the two domains, making the features of each layer after hierarchical processing richer and more comprehensive. This fusion method retains the useful information extracted during the frequency domain and spatial domain processing respectively, providing a more powerful feature representation for subsequent generation of high-quality denoised images.
[0049] Based on the skip connection, a vertical stripe perception fusion attention mechanism module is set between the encoder and the decoder. This module can effectively fuse the features of each layer after hierarchical processing and output the fused features to the decoder. This fusion method can not only utilize the feature information of different layers of the encoder but also, through the vertical stripe perception attention mechanism, pay more attention to the image features in a specific direction, further improving the quality of the denoised image, making the denoised image perform better in terms of details and textures and being closer to the original clear image.
[0050] Other features and advantages of the present invention will be described in the following specification, and in part will become apparent from the specification, or will be understood by implementing the present invention. The objectives and other advantages of the present invention can be achieved and obtained by the structure specifically pointed out in this application document.
[0051] The technical solution of the present invention will be further described in detail below with reference to the drawings and embodiments. Description of the Drawings
[0052] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention. In the drawings:
[0053] Figure 1 It is a schematic flow chart of a dual-domain heterogeneous image denoising method in an embodiment of the present invention;
[0054] Figure 2 It is an execution logic flow chart of steps S2 and S3 in an embodiment of the present invention;
[0055] Figure 3 It is an execution logic flow chart of step S4 in an embodiment of the present invention. Detailed Embodiments
[0056] The following describes the preferred embodiments of the present invention with reference to the drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention and are not used to limit the present invention.
[0057] Embodiment 1:
[0058] The present invention provides a dual-domain heterogeneous image denoising method, referring to Figures 1 to 3 , including:
[0059] S1: Extract a preliminary feature map of the input image based on depthwise separable convolution;
[0060] S2: Hierarchically process the preliminary features using a hierarchical dual-driven encoder-decoder architecture to obtain multiple layers of features, including:
[0061] In the shallow layer of the hierarchical dual-driven encoder-decoder architecture, an encoder of the dual-domain heterogeneous collaboration architecture is used for dual-domain heterogeneous collaboration processing, including:
[0062] Frequency domain processing branch: Reorganize the spectral bands and assign attention weights to the frequency domain information of the image to obtain the processed frequency domain information of the image; Reorganize the spectral bands and assign attention weights to the frequency domain information of the image. First, the image is transformed from the spatial domain to the frequency domain (e.g., through Fourier transform) to obtain the frequency domain representation of the image. Then, different spectral bands in the frequency domain information are recombined to highlight or adjust the information within a specific frequency range. At the same time, attention weights are assigned to different spectral bands, and according to the characteristics of the image, the model is made to pay more attention to the spectral bands that are important for denoising. For example, if the noise in the image is mainly concentrated in the high-frequency part, then a higher attention weight is assigned to the high-frequency spectral band to better process the noise. Through spectral band reorganization and attention weight assignment, the frequency domain information of the image can be optimized in a targeted manner, removing noise while retaining the important features of the image, making the processed frequency domain information of the image more conducive to subsequent processing and denoising.
[0063] Spatial domain processing branch: Capture the long-range spatial dependencies of the image, and combine a convolutional enhanced feed-forward network to suppress the forgetting of local pixels to obtain the processed spatial domain information of the image; Capture the long-range spatial dependencies of the image, and combine a convolutional enhanced feed-forward network to suppress the forgetting of local pixels. Long-range spatial dependencies refer to the relationships between pixels that are far apart in the image. Some techniques (such as variants of the self-attention mechanism) are used to capture this long-range dependency and understand the interconnections between different regions in the image. At the same time, a convolutional enhanced feed-forward network is used to process local pixels to prevent the information of local pixels from being forgotten during the processing. For example, in a large image area, by capturing the long-range spatial dependencies, it can be known how a certain local area is related to other far-away areas, and the convolutional enhanced feed-forward network further enhances the feature representation of local pixels to avoid information loss. The image is comprehensively processed in the spatial domain, taking into account both the overall structural relationship of the image (long-range spatial dependencies) and ensuring the integrity of local pixel information, so as to obtain the processed high-quality spatial domain information of the image, providing richer spatial information support for denoising.
[0064] In the deep layer of the hierarchical dual-drive encoding and decoding architecture, it switches to the T Block, which adopts the multi-head transposed attention mechanism and the depth spatial enhancement feed-forward network for global context modeling and local feature enhancement. In the deep layer of the hierarchical dual-drive encoding and decoding architecture, the T Block is adopted, which includes the multi-head transposed attention mechanism and the depth spatial enhancement feed-forward network. The multi-head transposed attention mechanism is used for global context modeling. It can calculate the attention distributions in multiple different representation subspaces in parallel, so as to capture the global information of the image more comprehensively. The depth spatial enhancement feed-forward network focuses on local feature enhancement. Through a series of convolutions and non-linear transformations, it strengthens the features of local regions. For example, the multi-head transposed attention mechanism can simultaneously focus on different regions of the image, comprehensively consider the relationships between them, while the depth spatial enhancement feed-forward network deeply processes a certain local region to enhance the feature expressiveness of that region. In the deep layer of the architecture, through global context modeling and local feature enhancement, the feature representation of the image is further optimized. Global context modeling helps the model understand the image as a whole, while local feature enhancement enables the model to process local details more meticulously. The combination of the two helps to more accurately remove noise and retain the detail information of the image.
[0065] S3: Concatenate the processed image frequency domain information and the image spatial domain information to obtain the features of each layer after hierarchical processing. Concatenate the processed image frequency domain information obtained from the frequency domain processing branch and the processed image spatial domain information obtained from the spatial domain processing branch along the channel dimension. For example, if the processed image frequency domain information has C1 channels and the processed image spatial domain information has C2 channels, the resulting feature map after concatenation will have C1 + C2 channels. Through concatenation, the processing results in the frequency domain and the spatial domain are fused, making full use of the information advantages of the two domains, so that the features of each layer after hierarchical processing contain richer and more comprehensive image information, providing a more powerful feature representation for subsequent fusion and denoising.
[0066] S4: Based on the vertical stripe perception fusion attention mechanism module between the encoder and the decoder with skip connections, fuse the features of each layer after hierarchical processing and output them to the decoder to obtain the denoised image.
[0067] Construct a vertical stripe-aware fusion attention mechanism module between the encoder and the decoder using skip connections. Skip connections allow the shallow features of the encoder to be directly passed to the decoder and fused with the deep features, avoiding information loss during the encoding-decoding process. The vertical stripe-aware fusion attention mechanism module focuses on fusing the hierarchically processed features at each layer. It assigns attention weights to different features according to their different characteristics, and focuses on the feature information related to denoising. For example, a lower weight is assigned to the feature regions containing noise information, and a higher weight is assigned to the key structure and texture feature regions of the image. Then the fused features are input into the decoder, and the decoder restores the low-resolution feature map to a denoised image of the same size as the original image through a series of transposed convolution or upsampling operations. In this way, features at different levels are effectively fused, important features are highlighted and noise-related features are suppressed, and finally a high-quality denoised image is output, improving the effect and quality of image denoising.
[0068] In this embodiment, the model adopts a 4-layer encoder-decoder structure. First, initial features are extracted through depthwise separable convolution (DS-Conv), and then enter the encoder (M-T Block), which consists of two branches: one is the frequency domain processing branch (FreqFormer) based on Transformer, and the other is the spatial domain processing branch (SpMamba) based on Mamba. These two branches process the frequency domain information and spatial information of the image respectively. Since the image size decreases after downsampling, the T Block (based on Transformer) is selected as the spatial domain processing module to extract deeper features while reducing the computational complexity. Skip connections are embedded with a vertical stripe-aware fusion attention mechanism (V-CBAM), which highlights the key regions and suppresses the irrelevant parts through channel attention and vertical enhanced spatial attention, further improving the image processing ability.
[0069] Depthwise separable convolution is an efficient convolution method that decomposes the standard convolution into depthwise convolution and pointwise convolution. In this step, depthwise separable convolution is applied to the input image to extract the initial feature map of the image. Depthwise convolution is responsible for performing convolution operations independently on each channel to capture local spatial features, while pointwise convolution fuses and integrates the features between channels through 1×1 convolution. For example, for an RGB three-channel image, depthwise convolution will perform convolution on the red, green, and blue channels respectively to extract the local features of each channel, and then pointwise convolution fuses these channel features to obtain the initial feature map. Depthwise separable convolution greatly reduces the computational complexity compared with traditional convolution, and at the same time can effectively extract the underlying features of the image, providing a basis for subsequent processing. These initial feature maps contain some basic information of the image, such as edges, textures, etc., which are important bases for further analyzing and processing the image.
[0070] The M-T Block consists of parallel FreqFormer frequency-domain processing branches and SpMamba spatial-domain processing branches. The frequency-domain information and spatial-domain information of the images processed by the two branches are then concatenated.
[0071] SBSA first extracts Q, K, and V through depthwise separable convolutions. Then, it uses the FFT to transform the spatial-domain features into the spectral domain. Next, through spectral band recombination (SBR), the frequency-domain elements are classified into different frequency bands, and different attention weights are assigned between these bands. The attention in the low-frequency bands is reduced, while the attention in the high-frequency bands is enhanced. Finally, the frequency-domain attention features are transformed back to the spatial domain through the IFFT, and the spatial features are output through a convolutional layer.
[0072] SEFF increases the channel dimension through a point convolutional layer and then processes it through two parallel branches, each using 3×3 and 3×3 dilated depth convolutions. Then, the features of the two branches are transformed into the frequency domain, and the spectral-domain information is enhanced through learnable weight and bias variables respectively. Through a gating mechanism, the activation output of one branch with a longer-range receptive field acts as the gating unit of the other branch, dynamically adjusting feature enhancement.
[0073] SpMamba consists of a Visual State Space Module (VSSM) and a Convolutional Augmented Feed-Forward Network (CAE).
[0074] The VSSM is a Mamba-based state space model used to capture long-range dependencies in images. Among them, the 2D-SSM is used to flatten the two-dimensional image features into a 1D sequence so that Mamba can process them. The image features are scanned in four different directions (from top left to bottom right, from bottom right to top left, from bottom left to top right, from top right to bottom left), generating four 1D sequences. And the discrete state space equations are applied to each sequence respectively. It is best to merge the results of all sequences through a summation operation and then restore the two-dimensional structure through a reshape operation.
[0075] The CAE is a convolutional augmented feed-forward network. Since Mamba processes 1D sequences, local pixels in a two-dimensional image may be far apart after flattening, resulting in the problem of local pixel forgetting. The CAE not only enhances the interaction ability of local features but also suppresses channel redundant information.
[0076] The T Block selects the multi-head transposed attention mechanism (MDTA) in Restormer. MDTA uses depthwise separable convolutions to extract Q, K, and V, which can capture local features while maintaining computational efficiency and model global context in attention calculations. The depth spatial enhancement feed-forward network (DSFN) uses depth convolutions to reduce computational complexity while maintaining the effectiveness of feature extraction. Through the spatial attention mechanism, key regions are highlighted, features with less information are suppressed, and only useful information is allowed to propagate further.
[0077] The skip connection embeds the vertical stripe-aware fusion attention mechanism (V-CBAM), which optimizes cross-level feature fusion through the synergistic effect of channel attention and vertically enhanced spatial attention. Channel attention filters important feature channels, and the spatial attention branch adds a vertical Sobel convolution kernel to enhance the response to vertical stripe noise and suppress redundant information. This design ensures that features in the noise region are accurately suppressed during the decoding process while preserving image details.
[0078] The frequency domain (FreqFormer) and spatial domain (SpMamba) branches are combined in parallel. They capture global frequency domain features and long-range spatial dependencies through Transformer and Mamba respectively, solving the limitation of single-domain processing in traditional methods.
[0079] The dual-domain parallel M-T Block is used in the shallow layers (the first and second layers), and switches to the lightweight T Block in the deep layers (the third and fourth layers). Through hierarchical feature adaptation (paying attention to details and local structures in the shallow layers and focusing on global semantics in the deep layers) and dynamic allocation of computing resources (high-resolution features in the shallow layers are processed by the M-T Block, while low-resolution features in the deep layers are handled by the lightweight T Block), the redundant computational burden is significantly reduced.
[0080] Targeted feature enhancement mechanisms are introduced in different modules. The convolutional enhancement feed-forward network (CAE) in SpMamba solves the problem of local pixel forgetting in two-dimensional image processing by Mamba, enhances the local feature interaction ability, and suppresses channel redundant information. The depth spatial enhancement feed-forward network (DSFN) in the T Block uses depthwise separable convolutions and the spatial attention mechanism to solve the problem that the local feature extraction ability of the transformer is inferior to that of the CNN and there is overprocessing of irrelevant regions. While maintaining computational efficiency, it enhances the model's local feature extraction ability, focuses on key regions, and suppresses features in irrelevant regions.
[0081] V-CBAM is embedded in the encoder-decoder skip connection to reconstruct the cross-level feature fusion process. Channel attention filters important feature channels, and the vertically enhanced spatial attention adds a vertical Sobel convolution kernel to enhance the sensitivity to vertical stripe noise and specifically suppress the noise region.
[0082] Example 2:
[0083] Based on Example 1, the spectral band recombination and attention weight allocation are performed on the image frequency domain information to obtain the processed image frequency domain information, including:
[0084] Using the frequency domain self-attention module, through depthwise separable convolution, Q, K, and V are extracted, and the spatial domain features are transformed into the spectral domain using the fast Fourier transform; with the help of the depthwise separable convolution in the frequency domain self-attention module. Depthwise separable convolution splits the conventional convolution into depthwise convolution and pointwise convolution, which can effectively reduce the computational amount. In this step, the query (Q), key (K), and value (V) are extracted from the image frequency domain information through this convolution method. These three features are extremely crucial in the self-attention mechanism. Q is used to determine the focus of attention, K is used to match relevant information, and V provides the actual feature representation. For example, when processing an image containing multiple textures, Q may focus on a certain type of texture, K searches for the relevant parts in the image, and V carries the specific feature information of these textures. The fast Fourier transform (FFT) is used to transform the spatial domain features into the spectral domain. FFT can efficiently convert the spatial domain signal into a frequency domain representation, revealing the components of the signal at different frequencies. The features that originally described the pixel distribution of the image in the spatial domain are transformed into the representation of the intensities of different frequency components in the frequency domain after FFT. For example, when a natural scenery image is transformed from the spatial domain to the frequency domain, it can be seen that the low-frequency part reflects the general outline of the image, and the high-frequency part reflects the details and textures. The Q, K, and V extracted by the depthwise separable convolution provide the basic data for subsequent attention calculation, and the transformation to the spectral domain enables subsequent processing to operate on the image information in the frequency dimension, which helps to analyze and process the image frequency domain information more effectively and prepares for tasks such as denoising.
[0085] Through spectral band recombination, the frequency domain elements are classified into different frequency bands, and different attention weights are allocated between different frequency bands. The attention in the low-frequency band is reduced, while the attention in the high-frequency band is enhanced to obtain the frequency domain attention features; the frequency domain elements are spectrally band-recombined according to frequency and classified into different frequency bands. Given that noise is mostly concentrated in the high-frequency part, when allocating attention weights to different frequency bands, the attention in the low-frequency band is deliberately reduced and the attention in the high-frequency band is enhanced. In this way, on the basis of retaining the low-frequency information of the main structure of the image, the high-frequency information related to noise and details can be captured, making the processed frequency domain features more conducive to denoising.
[0086] Convert the frequency-domain attention features back to the spatial domain through the inverse fast Fourier transform, and output the spatial features through the convolutional layer; convert the frequency-domain attention features after spectral band recombination and weight assignment back to the spatial domain through the inverse fast Fourier transform to match the spatial dimension of the original image. Then, use the convolutional layer to process and output these spatial-domain features to further optimize the feature representation, preparing for subsequent fusion with other spatial-domain features or more in-depth processing.
[0087] Use the frequency-domain enhancement feed-forward module to increase the channel dimension through the point convolutional layer, process it using parallel dilated convolutions, then convert the features of the two branches to the frequency domain, enhance the spectral domain information through learnable weights and bias variables respectively, and through the gating mechanism, use the activation output of one branch with a longer receptive field as the gating unit for the other branch to dynamically adjust the feature enhancement and perform feature fusion, and obtain the processed image frequency-domain information based on the 1×1 convolution output. The frequency-domain enhancement feed-forward module first increases the channel dimension through the point convolutional layer, enabling the model to learn richer features. Then it uses parallel dilated convolutions to extract features from different scales and expand the receptive field. After that, it converts the features of the two branches to the frequency domain and enhances the spectral domain information using learnable weights and bias variables. Through the gating mechanism, according to the perception of long-range information by one branch, it dynamically adjusts the degree of feature enhancement of the other branch to achieve feature fusion. Finally, it outputs the processed image frequency-domain information through 1×1 convolution, comprehensively improving the quality of the frequency-domain information and providing strong support for image denoising.
[0088] Embodiment 3:
[0089] Based on Embodiment 1, capture the long-range spatial dependence of the image, and combine the convolutional enhanced feed-forward network to suppress the forgetting of local pixels to obtain the processed image spatial-domain information, including:
[0090] The visual state space module scans the two-dimensional features along four diagonal directions to generate four 1D sequences in four directions, and then applies the discrete state space equation to each 1D sequence in each direction to obtain the results in four directions;
[0091] The visual state space module processes the two-dimensional image features, scans along four diagonal directions (from top left to bottom right, from top right to bottom left, from bottom left to top right, from bottom right to top left), and converts the two-dimensional features into one-dimensional (1D) sequences in four directions. This is because the discrete state space equation to be applied later is more suitable for processing 1D sequences. After that, the discrete state space equation is applied to each 1D sequence separately. The discrete state space equation can capture the long-range dependencies in the sequence. For an image, in this way, the correlation information of pixels at different positions in the image over a long distance can be captured, and the processing results in four directions can be obtained. For example, in an image containing a complex scene, through this operation, the potential connections between distant regions can be discovered, such as a certain visual long-range dependence that may exist between a distant building and a nearby road.
[0092] The results in four directions are combined through a summation operation, and then the two-dimensional structure is restored through a reshape operation. That is, the results obtained after applying the discrete state space equation in the above four directions are subjected to a summation operation. The summation operation can synthesize the long-range dependence information captured in four directions, avoiding ignoring the important connections in other directions by only focusing on a single direction. Then, through the reshape operation, the combined one-dimensional data is restored to a two-dimensional structure to match the dimension of the original image. This step prepares for further processing at the two-dimensional image level, integrating the information processed by capturing long-range dependencies back into the two-dimensional space of the image.
[0093] Then, a convolutional enhanced feed-forward network is used to enhance the interaction ability of local features and suppress channel redundant information, obtaining the processed image spatial domain information. After restoring the two-dimensional structure of the feature information in the previous steps, there will be problems such as local pixel information being forgotten during the processing and channel information redundancy. Therefore, a convolutional enhanced feed-forward network is used to process it. Through a series of convolutional operations, the convolutional enhanced feed-forward network can enhance the interaction ability between local features, making the model pay more attention to the connections between local pixels, thereby suppressing the phenomenon of local pixel forgetting. At the same time, it can also screen and integrate channel information, suppressing channel redundant information, making the processed image spatial domain information more refined and effective, providing high-quality spatial domain feature support for image denoising. For example, in a local area of an image, there may be rich detail information, and the convolutional enhanced feed-forward network can better retain and highlight these details while removing some unnecessary channel information, improving the accuracy of image denoising.
[0094] Example 4:
[0095] Based on Example 1, for the dual-domain heterogeneous image denoising method, an encoder of the dual-domain heterogeneous collaborative architecture is used in the shallow layer of the hierarchical dual-driven encoding and decoding architecture for dual-domain heterogeneous collaborative processing, and it further includes:
[0096] Calculate the adjusted number of channels for each layer based on the size of the feature map processed by each layer; in the shallow layer of the hierarchical dual-driven encoding and decoding architecture, due to the differences in the sizes of the feature maps processed by different layers, in order to make more reasonable use of computing resources, it is necessary to determine the adjusted number of channels for each layer according to the size of the feature map processed by each layer. The feature map size includes height and width information. Through a specific calculation method, the relationship between the input image size and the feature map sizes of each layer is comprehensively considered to adjust the number of channels. This adjustment helps the model to extract and represent features more effectively at different layers, avoiding waste of computing resources or insufficient feature extraction caused by unreasonable number of channels. For example, if the size of the feature map of a certain layer is small, the number of channels may be appropriately reduced to reduce the computational amount; if the feature map size is large, the number of channels can be increased to capture more features.
[0097] Calculate the computational balance error of the hierarchical dual-driven encoding and decoding architecture based on the adjusted number of channels of all layers and the sizes of the feature maps processed by all layers; based on the adjusted number of channels of all layers obtained from the previous calculation, and the sizes of the feature maps processed by all layers, further calculate the computational balance error of the hierarchical dual-driven encoding and decoding architecture. This error reflects the degree of balance in the distribution of computing resources across all layers of the entire architecture. Factors such as the weights of each layer and the computational amount per single channel are involved in the calculation. By combining these factors with the number of channels and the feature map sizes, it is evaluated whether the distribution of the computational amount among the layers is reasonable. If the computational balance error is large, it indicates that the distribution of computing resources among the layers is uneven, and there may be situations where the computational amount of some layers is too large or too small.
[0098] When the computational balance error of the hierarchical dual-driven encoding and decoding architecture is greater than the preset error threshold, the hierarchical channel dynamic allocation strategy is triggered. When the calculated computational balance error of the hierarchical dual-driven encoding and decoding architecture is greater than the pre-set error threshold, it indicates that the current channel number allocation cannot achieve an ideal balance state of computing resources across all layers. At this time, it is necessary to trigger the hierarchical channel dynamic allocation strategy. This strategy will re-adjust the number of channels for each layer according to information such as the size of the feature map processed by the corresponding layer, the size of the input image, and the computational balance error, so as to optimize the allocation of computing resources and improve the overall efficiency and performance of the model. For example, through the dynamic allocation strategy, the computing resources are adjusted from the layer with too large computational amount to the layer with too small computational amount, making the computing of the entire architecture more balanced and efficient.
[0099] Example 5:
[0100] Based on Example 4, for the dual-domain heterogeneous image denoising method, calculating the adjusted number of channels for each layer based on the size of the feature map processed by each layer includes:
[0101]
[0102] Where C is the number of channels after single-layer adjustment;
[0103] C base is the basic number of channels, which is the starting benchmark value for channel number adjustment, the number of channels initially set in the model, and provides a basic framework for feature extraction;
[0104] α is the expansion intensity coefficient, which determines the intensity of channel number adjustment according to the change in the size of the feature map. The larger the value, the greater the amplitude of the channel number change with the change in the feature map size; conversely, the smaller it is;
[0105] tanh() is the hyperbolic tangent function that maps the output value of the logarithmic function to between -1 and 1. Its role is to perform a non-linear transformation on the result of the logarithmic function, making the adjustment of the channel number more flexible and adaptable, and avoiding the limitations brought by simple linear adjustment. Because the complexity of image features is not a linear relationship, the hyperbolic tangent function can better fit this complex change;
[0106] log2() is the logarithmic function with base 2;
[0107] H is the height of the feature map processed by the corresponding layer, and W is the width of the feature map processed by the corresponding layer;
[0108] H0 is the height of the input image, and W0 is the width of the input image.
[0109] This formula is used in the Mamba+Transformer dual-branch image denoising method to accurately calculate the number of channels after adjustment for each layer according to the size of the feature map processed by each layer. It enables the model to dynamically adjust the number of channels according to the size of the feature map and reasonably allocate computing resources.
[0110] Example 6:
[0111] Based on Example 4, for the dual-domain heterogeneous image denoising method, the computational load balance error of the hierarchical dual-drive encoding and decoding architecture is calculated based on the number of channels after adjustment for all layers and the size of the feature maps processed by all layers, including:
[0112]
[0113] Where δ is the computational load balance error of the hierarchical dual-drive encoding and decoding architecture, which is a key indicator to measure whether the computing resource allocation of the entire architecture is balanced;
[0114] n is the total number of layers of the hierarchical dual-drive encoding and decoding architecture, covering all layers participating in the calculation in the architecture;
[0115] γ i is the weight of the i-th layer, which reflects the relative importance of this layer in the entire architecture. The weights of different layers can be set according to their roles in feature extraction, processing, etc.;
[0116] C i is the adjusted number of channels of the i-th layer, which is calculated based on the feature map size in Example 5. The number of channels affects the ability and amount of calculation of the layer to process features;
[0117] FLOPS0 is the computational capacity of a single channel, which is the basic unit for measuring the computational resources required for each channel to process features;
[0118] H0 is the height of the input image, W0 is the width of the input image, representing the size information of the input image;
[0119] H i is the height of the feature map processed by the i-th layer, W i The width of the feature map processed by the i-th layer reflects the size of the feature map of this layer.
[0120] The formula calculates the balance error by comprehensively considering the factors related to the amount of computation of each layer. For each layer, the contribution value of the layer to the balance error of computation is obtained by using its weight, the adjusted number of channels, the amount of computation per channel, and the size ratio of the input image to the feature map of the layer. Then the contribution values of all layers are accumulated and summed to finally obtain the balance error of computation of the entire hierarchical dual-drive codec architecture. For example, if a layer has a large weight, a large number of channels, and a large difference between the feature map size and the input image size, then the contribution of the layer to the balance error of computation is large, which may mean that the allocation of computational resources of the layer is unbalanced with other layers, thereby increasing the overall balance error of computation.
[0121] Embodiment 7:
[0122] On the basis of Example 4, the dual-domain heterogeneous image denoising method, when the computational balance error of the hierarchical dual-drive encoding and decoding architecture is greater than a preset error threshold, triggers a hierarchical channel dynamic allocation strategy, including:
[0123] When the computational balance error of the hierarchical dual-drive codec architecture is greater than the preset error threshold, the feature map size processed by the corresponding layer, the size of the input image, and the computational balance error of the hierarchical dual-drive codec architecture are input into the hierarchical channel dynamic allocation model to obtain the final adjusted number of channels for each layer.
[0124] When the computational balance error of the hierarchical dual-drive encoding and decoding architecture calculated through computation is greater than the preset error threshold, it means that the computational resource allocation of each layer of the architecture is unbalanced, affecting the model efficiency. At this time, the hierarchical channel dynamic allocation strategy will be triggered. The size of the feature map (including height and width) processed by the corresponding layer, the size of the input image (height and width), and the calculated computational balance error are input into the pre-trained hierarchical channel dynamic allocation model together. This model will comprehensively analyze these data, considering the relationship between the feature maps of each layer and the size of the input image, as well as the degree of current computational imbalance. The model outputs the finally adjusted number of channels for each layer. In this way, the number of channels for each layer is reasonably adjusted, the allocation of computational resources among layers is optimized, the model runs more efficiently, and the overall performance of image denoising is improved.
[0125] Example 8:
[0126] Based on Example 1, the dual-domain heterogeneous image denoising method uses a multi-head transposed attention mechanism and a deep spatial enhancement feed-forward network for global context modeling and local feature enhancement, including:
[0127] Use multi-head transposed attention to extract Q, K, and V matrices through depthwise separable convolution and perform global context modeling; in the dual-domain heterogeneous image denoising method, use the multi-head transposed attention mechanism to extract Q (query), K (key), and V (value) matrices from image features through depthwise separable convolution. This convolution method can effectively reduce the computational amount. Then use the multi-head transposed attention mechanism to calculate the attention distribution in multiple subspaces in parallel, enabling the model to comprehensively capture the long-distance connections between different regions of the image, and then constructing a global context model that reflects the overall layout and relationships of the image. For example, when processing an urban landscape image, it can be used to establish an association between the distant high-rise buildings and the nearby streets at the global level, enabling the model to understand the entire scene.
[0128] Use a deep spatial enhancement feed-forward network and use depthwise separable convolution to extract local features, and screen key regions through spatial attention gating. It uses depthwise separable convolution to capture local details of the image, such as features like textures and small patterns. Then, through the spatial attention gating mechanism, according to the importance of the features for image denoising and feature expression, key regions are screened out, and unimportant parts are suppressed. For example, in an image, it can highlight the regions where the key contours and detailed textures of the object are located, weaken the irrelevant background, make the model more targeted in processing local features, improve the quality of local features, and assist in image denoising.
[0129] Example 9:
[0130] Based on Embodiment 1, for the dual-domain heterogeneous image denoising method, the vertical stripe-aware fusion attention mechanism module optimizes cross-level feature fusion through the collaborative action of channel attention and vertically enhanced spatial attention.
[0131] Channel attention focuses on different channels of features. Each channel of an image contains different types of information, such as color, texture, etc. By analyzing the importance of each channel, higher weights are assigned to key channels to highlight important information, suppress redundant channels, and improve the denoising and feature expression effects.
[0132] Vertically enhanced spatial attention focuses on the spatial dimension, especially the vertical direction. By means of a special method, the sensitivity to vertical stripe features is enhanced, and both vertical noise stripes and important vertical structures can be accurately captured and processed, enhancing the role of vertical direction features in fusion.
[0133] The synergistic effect of these two attentions dynamically adjusts the fusion weights for features with different scales and abstraction levels at different levels in the encoder-decoder structure, effectively combines shallow details with deep semantics, optimizes cross-level feature fusion, improves the image denoising quality, and makes the processed image better in terms of details, structure, and overall quality.
[0134] Embodiment 10:
[0135] Based on Embodiment 1, for the dual-domain heterogeneous image denoising method, the vertical stripe-aware fusion attention mechanism module optimizes cross-level feature fusion through the collaborative action of channel attention and vertically enhanced spatial attention, including:
[0136] Perform global average pooling and max pooling on the input features, and generate a channel weight vector through two fully connected layers; perform global average pooling and max pooling on the input features to obtain channel global information. Then, through the processing of two fully connected layers, a channel weight vector is generated to determine the importance of each channel in feature representation, preparing for channel weighting during cross-level feature fusion.
[0137] Use a vertical Sobel operator to extract the vertical edge response, and splice it with the horizontal convolution result to generate a spatial weight matrix; use a vertical Sobel operator to extract the vertical edge response, combine it with the horizontal convolution result for splicing, and generate a spatial weight matrix. This matrix emphasizes vertical direction features, reflects the spatial position weight distribution of the image, and helps to focus on the importance of specific spatial positions for denoising and feature fusion.
[0138] Meanwhile, based on the enhancement mechanism, the weights of the channel attention branch and the spatial attention branch are dynamically allocated until the output is the element-wise product of the channel weight vector and the spatial weight matrix. Based on the enhancement mechanism, the weights of the channel and spatial attention branches are dynamically adjusted, and finally the element-wise product of the two is output. In this way, by integrating the channel and spatial attention information, the weights during cross-level feature fusion are accurately regulated, and the fusion effect is optimized to improve the denoising quality.
[0139] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the present invention and its equivalent technologies, the present invention is also intended to include these changes and modifications.
Claims
1. A dual-domain heterogeneous image denoising method, characterized in that Including: S1: Extract a preliminary feature map of the input image based on depthwise separable convolution; S2: Use a hierarchical dual-driven encoder-decoder architecture to hierarchically process the preliminary features to obtain multiple layers of features, including: In the shallow layer of the hierarchical dual-driven encoder-decoder architecture, use the encoder of the dual-domain heterogeneous collaborative architecture to perform dual-domain heterogeneous collaborative processing, including: Frequency domain processing branch: Recombine spectral bands and assign attention weights to the image frequency domain information to obtain the processed image frequency domain information; Spatial domain processing branch: Capture the long-range spatial dependence of the image, and combine a convolutional enhanced feedforward network to suppress local pixel forgetting to obtain the processed image spatial domain information; In the deep layer of the hierarchical dual-driven encoder-decoder architecture, switch to the T Block, and use the multi-head transposed attention mechanism and the depth spatial enhanced feedforward network for global context modeling and local feature enhancement; S3: Concatenate the processed image frequency domain information and the image spatial domain information to obtain the features of each layer after hierarchical processing; S4: Based on the vertical stripe perception fusion attention mechanism module between the encoder and the decoder with skip connections, fuse the features of each layer after hierarchical processing and output them to the decoder to obtain the denoised image.
2. The dual-domain heterogeneous image denoising method according to claim 1, wherein Recombine spectral bands and assign attention weights to the image frequency domain information to obtain the processed image frequency domain information, including: Use the frequency domain self-attention module to extract Q, K, and V through depthwise separable convolution, and use the fast Fourier transform to transform the spatial domain features to the spectral domain; Classify the frequency domain elements into different frequency bands through spectral band recombination, and assign different attention weights between different frequency bands. The attention in the low-frequency band is reduced, while the attention in the high-frequency band is enhanced to obtain the frequency domain attention features; Convert the frequency domain attention features back to the spatial domain through the inverse fast Fourier transform, and output spatial features through the convolutional layer; Use the frequency domain enhanced feedforward module to increase the channel dimension through a point convolutional layer, perform processing using parallel dilated convolutions, then transform the features of the two branches to the frequency domain, enhance the spectral domain information through learnable weights and bias variables respectively, and through the gating mechanism, use the activation output of one branch with a longer receptive field as the gating unit of the other branch to dynamically adjust feature enhancement and perform feature fusion, and output based on 1×1 convolution to obtain the processed image frequency domain information.
3. The dual-domain heterogeneous image denoising method according to claim 1, wherein Capture the long-range spatial dependence of the image, and combine a convolutional enhanced feedforward network to suppress local pixel forgetting to obtain the processed image spatial domain information, including: The visual state space module scans the two-dimensional features along four diagonal directions to generate four 1D sequences in four directions, and then applies the discrete state space equation to each 1D sequence in each direction to obtain the results in four directions; Merge the results in four directions through a summation operation, then restore the two-dimensional structure through a reshape operation, and then use a convolutional enhanced feedforward network to enhance the interaction ability of local features and suppress channel redundant information to obtain the processed image spatial domain information.
4. The dual-domain heterogeneous image denoising method according to claim 1, wherein In the shallow layer of the hierarchical dual-driven encoder-decoder architecture, use the encoder of the dual-domain heterogeneous collaborative architecture to perform dual-domain heterogeneous collaborative processing, and also include: Calculate the adjusted number of channels for each layer based on the size of the feature map processed by each layer; Calculate the computational balance error of the hierarchical dual-drive encoding and decoding architecture based on the adjusted number of channels of all layers and the size of the feature maps processed by all layers; When the computational balance error of the hierarchical dual-drive encoding and decoding architecture is greater than the preset error threshold, trigger the hierarchical channel dynamic allocation strategy.
5. The dual-domain heterogeneous image denoising method according to claim 4, wherein, Calculate the adjusted number of channels for each layer based on the size of the feature maps processed by each layer, including: Where C is the number of channels after single-layer adjustment, C base is the basic number of channels, α is the expansion intensity coefficient, tanh() is the hyperbolic tangent function, log2() is the logarithm function with base 2, H is the height of the feature map processed by the corresponding layer, W is the width of the feature map processed by the corresponding layer, H0 is the height of the input image, and W0 is the width of the input image.
6. The dual-domain heterogeneous image denoising method according to claim 4, wherein Calculate the computational balance error of the hierarchical dual-drive encoding and decoding architecture based on the adjusted number of channels of all layers and the size of the feature maps processed by all layers, including: In the formula, δ is the computational load balance error of the hierarchical dual-drive encoding and decoding architecture, n is the total number of layers of the hierarchical dual-drive encoding and decoding architecture, γ i is the weight of the i-th layer, C i is the adjusted number of channels of the i-th layer, FLOPS0 is the computational load of a single channel, H0 is the height of the input image, W0 is the width of the input image, H i is the height of the feature map processed by the i-th layer, W i is the width of the feature map processed by the i-th layer.
7. The dual-domain heterogeneous image denoising method according to claim 4, wherein When the computational balance error of the hierarchical dual-drive encoding and decoding architecture is greater than the preset error threshold, trigger the hierarchical channel dynamic allocation strategy, including: When the computational balance error of the hierarchical dual-drive encoding and decoding architecture is greater than the preset error threshold, input the size of the feature maps processed by the corresponding layer, the size of the input image, and the computational balance error of the hierarchical dual-drive encoding and decoding architecture into the hierarchical channel dynamic allocation model to obtain the finally adjusted number of channels for each layer.
8. The dual-domain heterogeneous image denoising method according to claim 1, characterized in that Adopt the multi-head transposed attention mechanism and the depth spatial enhancement feed-forward network for global context modeling and local feature enhancement, including: Adopt the multi-head transposed attention to extract the Q, K, and V matrices through depthwise separable convolution and perform global context modeling; Adopt the depth spatial enhancement feed-forward network and use depthwise separable convolution to extract local features, and screen key regions through spatial attention gating.
9. The dual-domain heterogeneous image denoising method according to claim 1, characterized in that The vertical stripe perception fusion attention mechanism module optimizes cross-layer feature fusion through the synergistic effect of channel attention and vertically enhanced spatial attention.
10. The dual-domain heterogeneous image denoising method according to claim 9, characterized in that, The vertical stripe perception fusion attention mechanism module optimizes cross-layer feature fusion through the synergistic effect of channel attention and vertically enhanced spatial attention, including: Perform global average pooling and max pooling on the input features, and generate the channel weight vector through two fully connected layers; Adopt the vertical Sobel operator to extract the vertical edge response, and splice it with the horizontal convolution result to generate the spatial weight matrix; At the same time, dynamically allocate the weights of the channel attention branch and the spatial attention branch based on the enhancement mechanism until the output is the element-wise product of the channel weight vector and the spatial weight matrix.
Citation Information
Patent Citations
Moving image deblurring model based on Fourier transform
CN117115040A
Underwater image enhancement method of Mama hybrid architecture based on space-frequency fusion
CN118710507A
Deep hash image retrieval method based on frequency domain decoupling and visual Mamba
CN118820508A
Underwater image enhancement method based on frequency domain analysis and visual Mama
CN119784598A
Cited By
Underwater image enhancement method and system based on double-domain collaboration
CN121053048A
Financial bill information automatic input system based on artificial intelligence
CN121074935A
Artificial intelligence-based automatic financial document information input system
CN121074935B
FFA image generation system and method based on double-domain constraint Mama diffusion model, medium and device
CN121095098A
Transform-based GPR image clutter removal method
CN121120425A