Image Denoising Method Based on Multi-Order Local Attention and Hybrid Attention of Deep Learning
By constructing multi-order local attention and hybrid attention modules, the problem of high computational complexity of Transformer model and inability to focus on spatial information is solved, and efficient image denoising and detail recovery are achieved.
Patent Information
- Application Number
- CN202310723004.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-19
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2043-06-19
AI Technical Summary
The existing Transformer model has high computational complexity in image denoising tasks and high hardware requirements. The channel attention mechanism cannot focus on spatial information in the image spatial dimension, resulting in excessive smoothing of the image edge contour information after denoising and loss of local details.
A multi-order local attention module and a hybrid attention module are adopted, combined with channel index grouping strategy and gating mechanism, a hybrid attention module is designed to perform multi-order interactions through multi-neighborhood information, and the model's local modeling ability and feature extraction ability are improved.
It significantly improves the local details and edge texture information recovery ability after image denoising, reduces the computational complexity and memory requirements, and achieves efficient image denoising effect.
Smart Images

Figure CN116757953B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image processing technology, and further relates to an image denoising method based on deep learning multi-order local attention and mixed attention in the field of image restoration technology. The present invention can be used for denoising noisy images, and while removing the noise information of the image, the original information of the image can be better restored. Background Art
[0002] Due to the limitations of the equipment, noise will inevitably be generated during the image acquisition and transmission process, especially in the case of insufficient light at night, insufficient exposure time will cause a lot of noise in the image. Therefore, the image denoising task has become one of the most basic tasks in the field of image processing. The quality of the image algorithm is directly related to the effect of subsequent image processing, such as image segmentation, target recognition, edge extraction, etc. In order to obtain high-quality digital images, it is often necessary to denoise the image to keep the integrity of the original information (i.e., the main features) as much as possible while removing useless information from the signal. With the continuous reform of parallel computing of hardware devices such as GPU, the field of artificial intelligence has achieved new development, and image processing technology based on deep learning has also become an important method for image denoising. Convolutional neural networks (CNNs) perform well in learning generalizable image priors from large-scale data, so these models have been widely used in image denoising and related image restoration tasks. Recently, another type of neural network architecture, Transformer, has shown significant performance improvements in the field of image restoration technology. Although the Transformer model alleviates the shortcomings of CNN (i.e., limited receptive field and inadaptability to input content), its computational complexity grows quadratically with the spatial resolution, making it unsuitable for most image denoising tasks involving high-resolution images.
[0003] Syed Waqas Zamir et al. proposed a high-resolution image denoising method for efficient Transformers in their published paper "Restormer: Efficient Transformer for High Resolution Image Restoration" (Published as a conference paper at CVPR 2022). This method constructs an efficient Transformer module, specifically including: depthwise separable multi-head shifted attention MDTA (Multi-Dconv Head Transposed Attention) and depthwise separable gated feed-forward network GDFN (Gated-Dconv Feed-Forward Network). An important feature of the proposed MDTA is the local context mixing before feature covariance calculation, which is achieved through pixel-level aggregation of cross-channel context using 1*1 convolution and channel-level aggregation of local context using efficient depthwise separable convolution. The proposed GDFN uses a gating mechanism to reformulate the conventional feed-forward network to improve the information flow through the network. The gating layer is designed as the element-wise product of two linear projection layers, one of which is non-linearly activated by an activation function. The gating mechanism in GDFN controls which specific complementary features should flow forward and allows subsequent layers in the network hierarchy to focus specifically on finer image attributes, thus producing high-quality outputs. However, the drawback of this method is that the self-attention mechanism SA (self-attention) of Transformer is adopted in the network, and its computational complexity grows quadratically with the image spatial resolution. Therefore, the computational complexity of this method is still high, and it has high requirements for the computational power and memory of the hardware.
[0004] Soochow University disclosed an image denoising method based on a Transformer with channel attention in its patent document "A Swin-Transformer Image Denoising Method and System Based on Channel Attention" (Application No.: 2021114146259; Application Date: November 25, 2021; Publication No.: CN 114140353A). The implementation steps of this method are as follows: Step 1: Input a noisy image into the shallow feature extraction network in the denoising network; Step 2: Use the shallow feature extraction network to extract the noise and channel shallow feature information of the noisy image; Step 3: Input the channel shallow feature information into the deep feature extraction network in the denoising network model for deep feature extraction; Step 4: Input the shallow feature information and deep feature information into the reconstruction network of the denoising network model for feature fusion; Step 5: Output the denoised clean image. The deficiency of this method is that the proposed channel attention mechanism only obtains the importance of each channel feature at the channel dimension level and cannot focus on the spatial information in the image spatial dimension, resulting in the over-smoothing of the edge contour information of the denoised image and the loss of local detail information. Summary of the Invention
[0005] The object of the present invention is to provide an image denoising method based on deep learning multi-order local attention and hybrid attention for the deficiencies of the above-mentioned existing technologies, to solve the problems of high computational complexity of the self-attention mechanism using Transformer, high requirements for the computing power and memory of hardware, and the fact that the channel attention mechanism cannot focus on the spatial information in the image spatial dimension, resulting in the over-smoothing of the edge contour information of the denoised image and the loss of local detail information.
[0006] The idea of realizing the object of the present invention is that since spatial neighborhood information can better restore local details, the present invention adopts a channel exponential grouping strategy and a gating mechanism to design a multi-order local attention module, and uses multi-neighborhood information for multi-order interaction to improve the local modeling ability of the model; the present invention combines the advantages of multiple attentions, adopts a channel average grouping strategy and a mechanism of parallel multiple attention mechanisms to design a hybrid attention module, which can improve the feature extraction ability of the model. Combining the above two points, a denoising network is constructed to solve the problems of insufficient denoising ability and over-smoothing in the prior art and the problems of large computational amount and high memory occupancy in the prior art.
[0007] To achieve the above object, the specific implementation steps of the present invention are as follows:
[0008] Step 1, construct a multi-order local attention branch:
[0009] Build and set up a multi-stage local attention branch connected in series, with the following structure: a first convolutional layer, a channel exponential grouping layer, a depthwise separable convolution module group, and a second convolutional layer; among them, the depthwise separable convolution module group is composed of four depthwise separable convolution modules with the same structure in parallel, namely the first, second, third, and fourth; each depthwise separable convolution module is composed of a depthwise separable convolution layer, a matrix multiplier, and a convolutional layer connected in series; set the parameters of the multi-stage local attention branch;
[0010] Step 2: Build a hybrid attention module connected in series, with the following structure: a channel average grouping layer, an attention parallel group, a channel splicing layer, and a 1*1 convolutional layer; among them, the attention parallel group is composed of a self-attention branch, a multi-stage local attention branch, and a channel attention branch in parallel;
[0011] Step 3: Construct a hybrid attention Transformer module:
[0012] Build a hybrid attention Transformer module composed of a first normalization layer, a hybrid attention module, a second normalization layer, and a feedforward neural network module connected in series;
[0013] Step 4: Construct a deep feature extraction sub-network:
[0014] Build a deep feature extraction sub-network composed of a downsampling module group, a bottom layer feature extraction module, and an upsampling module group connected in series;
[0015] Step 5: Connect a first 3*3 convolutional layer, a deep feature extraction sub-network, and a second 3*3 convolutional layer in series to form a Unet denoising network;
[0016] Step 6: Generate a training set:
[0017] Select at least 300 pairs of natural images, each pair consisting of a real noisy image and a labeled clean image; crop each image into image patches of size 256*256, and perform data augmentation on each image patch; form a training set from all the augmented image patches;
[0018] Step 7: Train the Unet denoising network:
[0019] Shuffle the image patches of the training set, randomly extract noisy image patches and the corresponding manually annotated clean image patches, input the extracted noisy image patches into the Unet denoising network, use the Charbonnier_L1loss function of the mean absolute error loss with penalty to calculate the loss value between the output image patches and the clean image patches, adopt the gradient descent method, and iteratively update the parameters of each layer of the Unet denoising network until the Charbonnier_L1loss function of the mean absolute error loss with penalty converges, so as to obtain the trained Unet denoising network;
[0020] Step 8, process the natural image to be denoised:
[0021] Input the natural image to be denoised into the trained Unet denoising network, and output the denoised image.
[0022] The present invention has the following advantages compared with the existing technologies:
[0023] First, the present invention constructs a multi-order local attention module, constructs a new type of dynamic attention by adopting a channel exponential grouping strategy and a gating mechanism, uses convolutional layers with multiple receptive fields to capture feature information in different neighborhood ranges, and then performs multi-order interaction of local information on the captured multi-neighborhood features through the gating mechanism. The multi-order local attention module constructed by the present invention significantly improves the local modeling ability of the network, and can better recover the local detail information and edge texture information of the image from the noisy image.
[0024] Second, the present invention designs a hybrid attention module, combines three types of attention, including self-attention, multi-order local attention, and channel attention, by adopting a channel average grouping strategy and a mechanism of parallel multiple attention mechanisms. The three types of attention control the forward flow of specific feature information in parallel. This module not only retains the ability of the Transformer self-attention to capture global features, but also retains the strong local modeling ability of the CNN, and also retains the ability of the channel attention to capture the importance of each channel. Finally, the feature information captured by the three types of attention is fused. The hybrid attention module constructed by the present invention significantly improves the feature extraction ability of the network, can better denoise, and can also better recover the texture information and local detail information of the image. Brief Description of the Drawings
[0025] Figure 1 is the flow chart of the present invention;
[0026] Figure 2 is the structural schematic diagram of the multi-order local attention branch of the present invention;
[0027] Figure 3 is the structural schematic diagram of the hybrid attention module of the present invention;
[0028] Figure 4 is a schematic structural diagram of the hybrid attention Transformer module of the present invention;
[0029] Figure 5 is a schematic structural diagram of the deep feature extraction sub-network of the present invention;
[0030] Figure 6 is a schematic structural diagram of the upsampling module of the present invention;
[0031] Figure 7 is a schematic structural diagram of the downsampling module of the present invention;
[0032] Figure 8 is a schematic structural diagram of the Unet denoising network of the present invention. Detailed implementation manners
[0033] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments.
[0034] Refer to Figure 1 to further describe in detail the implementation steps of the embodiments of the present invention.
[0035] Step 1: Construct a multi-stage local attention branch.
[0036] Refer to Figure 2 to further describe the constructed multi-stage local attention branch structure.
[0037] Build and set a sequentially connected multi-stage local attention branch, the structure of which is: a channel index grouping layer, a first convolutional layer, a depthwise separable convolution module group, and a second convolutional layer; wherein, the depthwise separable convolution module group is composed of four depthwise separable convolution modules with the same structure in parallel, namely the first, second, third, and fourth; each depthwise separable convolution module is composed of a depthwise separable convolution layer, a matrix multiplier, and a convolutional layer in series; set the parameters of the multi-stage local attention branch.
[0038] The setting of the parameters of the multi-stage local attention branch means that the kernel sizes of the first and second convolutional layers in the multi-stage local attention branch are both set to 1*1; the kernel sizes of the depthwise separable convolution layers in the first to fourth depthwise separable convolution modules are sequentially set to 9*9, 7*7, 5*5, 3*3; the kernel sizes of the convolutional layers are both set to 1*1.
[0039] The said channel index grouping layer is implemented by the following formula:
[0040] Value, Querys = split(input)
[0041] Q1, Q2, Q3, Q4 = expsplit(Querys)
[0042] Among them, Value and Querys respectively represent the two outputs of the channel grouping function split(), and the number of channels is C represents the number of channels of the input feature vector input; Q1, Q2, Q3, and Q4 respectively represent the four outputs of the channel exponential grouping function expsplit(), and the corresponding numbers of channels are
[0043] The implementation process of the multi-order local attention branch is as follows: Pass the input feature vector through the first convolutional layer and the channel exponential grouping layer to obtain Value, Q1, Q2, Q3, and Q4; Pass Q1, Q2, Q3, and Q4 through the depthwise separable convolutional layers of four parallel depthwise separable convolutional modules respectively; Use Value as the input of the matrix multiplier in the first depthwise separable convolutional module to perform matrix multiplication with the output of the 9*9 depthwise separable convolutional layer, and then pass through a 1*1 convolutional layer; Sequentially use the output of the previous depthwise separable convolutional module as the input of the matrix multiplier in the current depthwise separable convolutional module to perform matrix multiplication with the depthwise separable convolutional layer in the current depthwise separable convolutional module, and then pass through a 1*1 convolutional layer; Finally, pass the output of the fourth depthwise separable convolutional module through the second convolutional layer.
[0044] Refer to Figure 3 , and make a further description of the constructed hybrid attention module structure.
[0045] Step 2, construct a sequentially connected hybrid attention module, whose structure consists of a channel average grouping layer, an attention parallel group, a channel concatenation layer, and a 1*1 convolutional layer; Among them, the attention parallel group is composed of a self-attention branch, a multi-order local attention branch, and a channel attention branch in parallel.
[0046] The described channel average grouping layer is implemented by the following formula:
[0047] F1, F2, F3 = avgsplit(input)
[0048] Among them, F1, F2, and F3 respectively represent the three outputs of the channel average grouping avgsplit() function, and the number of channels is C represents the number of channels of the input feature vector input.
[0049] The implementation process of the hybrid attention module is as follows: After performing channel average grouping on the input feature vector, pass through three parallel attention branches respectively, then concatenate the outputs of the three attention branches in the channel dimension, and finally pass through a 1*1 convolutional layer.
[0050] Reference Figure 4 , a further description of the constructed hybrid attention Transformer module structure is given.
[0051] Step 3, construct a hybrid attention Transformer module.
[0052] Build a hybrid attention Transformer module by sequentially connecting a first layer normalization layer, a hybrid attention module, a second layer normalization layer, and a feed-forward neural network module in series.
[0053] The feed-forward neural network module is composed of a first convolutional layer, a Relu bias activation layer, and a second convolutional layer connected in series. Set the convolutional kernel sizes of the first and second convolutional layers in the feed-forward neural network module to 1*1.
[0054] Reference Figure 5 , a further description of the constructed deep feature extraction sub-network structure is given.
[0055] Step 4, construct a deep feature extraction sub-network.
[0056] Build a deep feature extraction sub-network by sequentially connecting a down-sampling module group, a bottom layer feature extraction module, and an up-sampling module group in series.
[0057] The up-sampling module group is composed of four up-sampling modules with the same structure at the first, second, third, and fourth levels connected in series.
[0058] Reference Figure 6 , a further description of the constructed up-sampling module structure is given.
[0059] Each up-sampling module is composed of 2 hybrid attention Transformer modules, a 1*1 convolutional layer, and a pixel recombination up-sampling layer connected in series.
[0060] The down-sampling module group is composed of four down-sampling modules with the same structure at the first, second, third, and fourth levels connected in series.
[0061] Reference Figure 7 , a further description of the constructed down-sampling module structure is given.
[0062] Each down-sampling module is composed of 2 hybrid attention Transformer modules, a 1*1 convolutional layer, and a pixel recombination down-sampling layer connected in series.
[0063] The bottom layer feature extraction module is composed of 8 hybrid attention Transformer modules with the same structure connected in series.
[0064] Reference Figure 8, a further description of the constructed Unet denoising network structure is given.
[0065] Step 5, connect the first 3*3 convolutional layer, the deep feature extraction sub-network, and the second 3*3 convolutional layer in series to form the Unet denoising network.
[0066] The implementation process of the Unet denoising network is as follows: the input image passes through the first 3*3 convolutional layer, the deep feature extraction sub-network, and the second 3*3 convolutional layer in sequence; the input of the first 3*3 convolutional layer and the output of the second 3*3 convolutional layer are connected using a skip connection; the output of each downsampling module and the input of the corresponding upsampling module are also connected using a skip connection.
[0067] Step 6, generate a training set.
[0068] Select at least 300 pairs of natural images, each pair consisting of a real noisy image and an annotated clean image; crop each image into image patches of size 256*256, and perform data augmentation on each image patch; form the training set from all the augmented image patches.
[0069] Step 7, train the Unet denoising network.
[0070] Shuffle the order of the image patches in the training set, randomly select noisy image patches and the corresponding manually annotated clean image patches, input the selected noisy image patches into the Unet denoising network, use the Charbonnier_L1loss function of the mean absolute error loss function with penalty to calculate the loss value between the output image patch and the clean image patch, adopt the gradient descent method, and iteratively update the parameters of each layer of the Unet denoising network until the Charbonnier_L1loss function of the mean absolute error loss function with penalty converges, obtaining the trained Unet denoising network.
[0071] The Charbonnier_L1loss function of the mean absolute error loss function with penalty is as follows:
[0072]
[0073] where H and W represent the width and height of the denoised image and the clean image respectively, represents the pixel value at the i-th row and j-th column in the denoised image, I(i,j) represents the pixel value at the i-th row and j-th column in the clean image, and ε is the penalty coefficient.
[0074] Step 8, process the natural image to be denoised:
[0075] Input the natural image to be denoised into the trained Unet denoising network, and output the denoised image.
Claims
1. An image denoising method based on multi - order local attention and hybrid attention of deep learning, characterized in that, Construct multi - order local attention branches and a hybrid attention module respectively; the specific steps of this image denoising method are as follows: Step 1, construct multi - order local attention branches: Build and set up a sequentially cascaded multi - order local attention branch, whose structure is: the first convolutional layer, the channel exponent grouping layer, the depth - separable convolution module group, the second convolutional layer; among them, the depth - separable convolution module group is composed of four depth - separable convolution modules with the same structure in parallel, namely the first, the second, the third, and the fourth; each depth - separable convolution module is composed of a depth - separable convolutional layer, a matrix multiplier, and a convolutional layer in series; set the parameters of the multi - order local attention branch; Step 2, build a sequentially cascaded hybrid attention module, whose structure is composed of a channel average grouping layer, an attention parallel group, a channel splicing layer, and a 1*1 convolutional layer; among them, the attention parallel group is composed of a self - attention branch, a multi - order local attention branch, and a channel attention branch in parallel; Step 3, construct a hybrid attention Transformer module: Build a hybrid attention Transformer module composed of a first normalization layer, a hybrid attention module, a second normalization layer, and a feed - forward neural network module in series; Step 4, construct a deep feature extraction sub - network: Build a deep feature extraction sub - network composed of a down - sampling module group, a bottom - layer feature extraction module, and an up - sampling module group in series; Step 5, compose a Unet denoising network by cascading a first 3*3 convolutional layer, a deep feature extraction sub - network, and a second 3*3 convolutional layer in series; Step 6, generate a training set: Select at least 300 pairs of natural images, each pair consisting of a real noisy image and an annotated clean image; crop each image into image patches of size 256*256, and perform data augmentation on each image patch; form a training set with all the augmented image patches; Step 7, train the Unet denoising network: Shuffle the image patches in the training set, randomly extract noisy image patches and the corresponding artificially annotated clean image patches, input the extracted noisy image patches into the Unet denoising network, use the Charbonnier_L1loss function of the mean absolute error loss with penalty to calculate the loss value between the output image patch and the clean image patch, adopt the gradient descent method, and iteratively update the parameters of each layer of the Unet denoising network until the Charbonnier_L1loss function of the mean absolute error loss with penalty converges, obtaining a trained Unet denoising network; Step 8, process the natural image to be denoised: Input the natural image to be denoised into the trained Unet denoising network, and output the denoised image.
2. The image denoising method based on deep learning multi-order local attention and hybrid attention according to claim 1, wherein, The parameter setting of the multi - order local attention branch in Step 1 refers to setting the convolution kernel sizes of the first and second convolutional layers in the multi - order local attention branch to 1*1; setting the convolution kernel sizes of the depth - separable convolutional layers in the first to fourth depth - separable convolution modules to 9*9, 7*7, 5*5, and 3*3 in sequence; and setting the convolution kernel sizes of the convolutional layers to 1*1.
3. The image denoising method based on deep learning multi-order local attention and hybrid attention according to claim 1, wherein The channel index grouping layer described in step 1 is implemented by the following formula: Value, Querys = split(input) Q1, Q2, Q3, Q4 = expsplit(Querys) Among them, Value and Querys respectively represent the two outputs of the channel grouping function split(), and the number of their channels is C represents the number of channels of the input feature vector input; Q1, Q2, Q3, and Q4 respectively represent the four outputs of the channel index grouping function expsplit(), and the corresponding numbers of channels are 4. The image denoising method based on deep learning multi-order local attention and hybrid attention according to claim 1, characterized in that, The channel average grouping layer described in step 2 is implemented by the following formula: F1, F2, F3 = avgsplit(input) Among them, F1, F2, and F3 respectively represent the three outputs of the channel average grouping avgsplit() function, and the number of channels is C represents the number of channels of the input feature vector input.
5. The image denoising method based on deep learning multi - order local attention and hybrid attention according to claim 1, characterized in that, The feed-forward neural network module described in step 3 is composed of a first convolutional layer, a Relu bias activation layer, and a second convolutional layer in series; the convolutional kernel sizes of the first and second convolutional layers in the feed-forward neural network module are both set to 1*1.
6. The image denoising method based on deep learning multi-order local attention and hybrid attention according to claim 1, characterized in that, The downsampling module group described in step 4 is composed of four downsampling modules with the same structure at the first, second, third, and fourth levels in series; each downsampling module is composed of 2 hybrid attention Transformer modules, a 1*1 convolutional layer, and a pixel reorganization downsampling layer in series.
7. The image denoising method based on deep learning multi-order local attention and hybrid attention according to claim 1, wherein The bottom layer feature extraction module described in step 4 is composed of 8 hybrid attention Transformer modules with the same structure in series.
8. The image denoising method based on deep learning multi-order local attention and hybrid attention according to claim 1, wherein, The upsampling module group described in step 4 is composed of four upsampling modules with the same structure at the first, second, third, and fourth levels in series; each upsampling module is composed of 2 hybrid attention Transformer modules, a 1*1 convolutional layer, and a pixel reorganization upsampling layer in series.
9. The image denoising method based on deep learning multi-stage local attention and hybrid attention according to claim 1, wherein, The penalized mean absolute error loss function Charbonnier_L1loss described in step 7 is as follows: where H and W represent the width and height of the denoised image and the clean image, respectively, represents the pixel value at the i-th row and j-th column in the denoised image, I(i, j) represents the pixel value at the i-th row and j-th column in the clean image, and ε is the penalty coefficient.
Citation Information
Patent Citations
Ultrasonic image denoising model building method and ultrasonic image denoising method
CN112200750A
Swin-Transform image denoising method and system based on channel attention
CN114140353A