Real image denoising method based on noise correction and guided residual estimation
By constructing a noise predictor and corrector network and combining it with a feature domain residual estimation network, the problem of complex noise distribution in real images is solved and high-quality image denoising effect is achieved.
Patent Information
- Application Number
- CN202210385679.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-13
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2042-04-13
AI Technical Summary
Existing technologies have difficulty in effectively processing the complex noise distribution in real images, resulting in poor image denoising effects, especially when the noise levels in different parts are inconsistent.
By constructing a noise predictor network to predict the initial noise level map, designing a non-blind feature domain residual estimation denoising network module, and continuously correcting the noise level map through the noise corrector network, combined with deep network iterative training, finally restoring high-quality images.
The subjective and objective effects of image denoising are significantly improved, especially in real noisy environments. The restored image has clear edges and textures, good details, and reduced artificial traces.
Smart Images

Figure BDA0003594903340000032 
Figure BDA0003594903340000033 
Figure BDA0003594903340000044
Abstract
Description
Technical Field
[0001] The present invention relates to image denoising technology, and in particular to real image denoising based on noise correction and guided residual estimation, and belongs to the image restoration direction in the field of digital image processing. Background Art
[0002] Due to imperfections in imaging systems, transmission media, and recording devices, noise is ubiquitous in the images we use daily. It can affect visual quality and even hinder accurate image recognition. Image denoising is an effective solution to this problem. Its goal is to restore high-quality images from noisy images, laying the foundation for subsequent operations such as image recognition and segmentation. This technology is a hot topic of extensive research in digital image processing and has significant practical applications.
[0003] In recent years, with the development of deep learning, many deep neural network-based methods have achieved remarkable success in removing specific levels of additive white Gaussian noise (AWGN). However, in real life, the noise level of an image is usually not a specific value, and the noise distribution in each part of the image may be different. Moreover, due to the influence of the image signal processing (ISP) module in the camera imaging system, the noise distribution of real images is often difficult to predict. Image denoising is a typical ill-posed inverse problem. Therefore, it is particularly critical to obtain accurate prior knowledge of image noise and fully utilize it in the denoising process. Summary of the Invention
[0004] The purpose of the present invention is to introduce noise level prediction and correction, and integrate them into a deep network to guide the estimation of denoising residuals, ultimately decomposing the originally complex real blind denoising problem into two easier-to-solve parts: noise level prediction / correction and non-blind denoising, thereby achieving the suppression of real noise. First, a preliminary noise level map of the noisy image is predicted and used to guide the initial non-blind denoising of the network. Secondly, the noise level map is corrected using the denoising features of the intermediate stage of the network, making the noise prior information obtained by the network more accurate. Finally, under the guidance of more accurate prior information, the non-blind network is able to restore a higher quality image.
[0005] The real image denoising method based on noise correction and guided residual estimation proposed in the present invention mainly includes the following steps:
[0006] (1) Construct a noise predictor network to predict the initial noise level map of the noisy image;
[0007] (2) Based on the guidance of the noise level map in step (1), a non-blind feature domain residual estimation denoising network module is designed to estimate the denoising residual of the noisy image to obtain a preliminary denoising result;
[0008] (3) Construct a noise corrector network to use the denoising results of the previous stage to correct the predicted noise level map to make it more accurate;
[0009] (4) Based on the feature domain residual estimation network constructed in step (2), the corrected noise level map and the denoising result of the previous stage are simultaneously input into the new residual estimation network by sharing the weights of its network modules to update the denoising residual and obtain a better denoising result;
[0010] (5) Repeat steps (3) to (4) until the specified number of iterations k is reached, and finally a complete real image denoising network based on noise correction and guided residual estimation is constructed to output the final denoising result;
[0011] (6) Using a publicly available training image dataset, the deep network constructed in step (5) is trained by minimizing the loss function;
[0012] (7) Finally, the noisy image is input into the deep network trained in step (6) to obtain the restored clean image. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 It is a block diagram of the real image denoising method based on noise correction and guided residual estimation of the present invention.
[0014] Figure 2 It is a network structure diagram of the noise predictor and noise corrector of the present invention.
[0015] Figure 3 It is a structural diagram of the feature domain residual estimation network of the present invention.
[0016] Figure 4 This is a comparison chart of the recovery results of an image block in the SIDD validation set of the test library using the present invention and different methods.
[0017] Figure 5 This is a comparison chart of the restoration results of an image block in the test image library DnD using the present invention and different methods. DETAILED DESCRIPTION
[0018] The present invention will be further described below in conjunction with the accompanying drawings:
[0019] Figure 1 In [1], a real image denoising method based on noise correction and guided residual estimation includes the following steps:
[0020] (1) Construct a noise predictor network to predict the initial noise level map of the noisy image;
[0021] (2) Based on the guidance of the noise level map in step (1), a non-blind feature domain residual estimation denoising network module is designed to estimate the denoising residual of the noisy image to obtain a preliminary denoising result;
[0022] (3) Construct a noise corrector network to use the denoising results of the previous stage to correct the predicted noise level map to make it more accurate;
[0023] (4) Based on the feature domain residual estimation network constructed in step (2), the corrected noise level map and the denoising result of the previous stage are simultaneously input into the new residual estimation network by sharing the weights of its network modules to update the denoising residual and obtain a better denoising result;
[0024] (5) Repeat steps (3) to (4) until the specified number of iterations k is reached, and finally a complete real image denoising network based on noise correction and guided residual estimation is constructed to output the final denoising result;
[0025] (6) Using a publicly available training image dataset, the deep network constructed in step (5) is trained by minimizing the loss function;
[0026] (7) Finally, the noisy image is input into the deep network trained in step (6) to obtain the restored clean image.
[0027] Specifically, in the step (1), the raw data (RawData) output by the sensor in the imaging system inside the camera is processed by the ISP and converted into the final color image. Due to the complex noise distribution of the real noise image and the ISP processing flow inside the camera, it is not easy to extract features from the noise image and then estimate the noise level, so the present invention models the noise in the Raw space rather than the color space. Specifically, the method proposed by Guo et al. is used, reference "Shi Guo, Zifei Yan, Kai Zhang, Wangmeng Zuo, and Lei Zhang. Toward convolutional blind denoising of real photographs. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 1712–1722, Jun. 2019." Let y represent the input noise image, σ represent the noise level map of the noise image in the raw space, and a noise predictor P is used to preliminarily predict the noise level map of the noise image. The present invention optimizes the predictor by minimizing the L1 distance:
[0028]
[0029] Among them, θ P are the parameters of the predictor P. The predictor is implemented using a convolutional neural network, using a module consisting of four convolutional layers interspersed with activation layers to estimate the noise level map. To further improve the network's expressiveness and more effectively estimate the noise level map, we introduce parameter-free attention before the last convolutional layer, specifically using the method proposed by Yang et al. (referenced in "Lingxiao Yang, Ru-Yuan Zhang, Lida Li, and Xiaohua Xie. Simam: A simple, parameter-free attention module for convolutional neural networks. In International Conference on Machine Learning (ICLR), pages 11863–11874. PMLR, 2021.").
[0030] In step (2), the feature mapping function M is used to transform the noise image from the pixel space to the feature space to extract the higher-level feature information I0 of the image, I0 = M(y). Correspondingly, the image reconstruction function Q is used to transform the restored clean image features from the feature space to the pixel space. We use the "convolution layer + activation layer + convolution layer" structure to achieve feature extraction and image reconstruction. The initial feature I0 is passed through the noise level map Guide, estimate the feature domain residual R1. The preliminary feature domain denoising result is I1=I0+R1.
[0031] In step (3), a noise corrector C is constructed to use the denoising features of the previous stage to correct the noise level map predicted in the previous stage to make it closer to the ground truth. The present invention optimizes the corrector by minimizing the L1 distance:
[0032]
[0033] Among them, θ C is the parameter of the corrector C. The corrector is similar to the predictor, but the input is replaced by the output feature I of the previous stage i and noise level after channel expansion The correction amount of the noise level map obtained by the corrector is Corrected noise level map
[0034] In the step (4), the denoising residual R of the feature domain i+1 It is realized through the feature domain residual estimation network F.
[0035] The feature-domain residual estimation network F employs a common residual architecture. However, conventional residual learning-based denoising networks process images in the pixel space, while our network performs restoration in the feature space. Network F consists of two parts: a fused denoising network and a noise feature-driven reconstruction controller. The fused denoising network implements image denoising and is primarily composed of multiple dynamic joint attention modules D. The noise feature-driven reconstruction controller encodes the estimated noise level map and extracts key features to guide the reconstruction of the denoised residual.
[0036] Three inputs of the feature domain residual estimation network F: I0, I i , each undergoes feature adjustment through a convolutional layer, and the three adjusted features are cascaded and fused:
[0037]
[0038] Among them, f0 represents the fused feature, C1, C2, and C3 all represent a convolution layer with a convolution kernel size of 3*3. Indicates cascade, C fus Represents the convolutional layer for feature fusion. The fused features are fed into n series-connected dynamic joint attention modules D, which guide the fused features through residual learning.
[0039] The dynamic joint attention module D consists of "multiple convolutional layers + channel-space joint attention", but in the channel-space joint attention, both channel features and spatial features are modulated by dynamic convolution. Finally, the output f of the nth dynamic joint attention module is n for:
[0040] f n =D n (D n-1 (…D2(D1(f0))…))
[0041] After fine-tuning, the output features are multiplied with the noise features output by the controller, using this prior information to guide feature learning. The controller uses a combination of three "1*1 convolution + activation layers" to perform spatial and channel transform encoding on the noise information. Then, a "convolution layer + activation layer + convolution layer" is used to convert the encoded information into the feature space and enhance it. Finally, a sigmoid function is used to normalize the enhanced features and generate feature weights.
[0042] In order to better preserve the output feature information of the previous stage, the convolutional I iPerform residual learning. Finally, a convolutional layer is used to output the feature residual R i+1 :
[0043]
[0044] Among them C R Represents a "convolution layer + activation layer" combination for fine-tuning, F en represents the reconstruction controller driven by noise characteristics, C tail Represents the convolutional layer at the end of the feature domain residual estimation network. The updated feature space denoising result is I i+1 =I0+R i+1 .
[0045] In step (5), after the iteration reaches the specified number k, the output feature space denoising result is I k+1 The denoising result I of the feature space is converted into k+1 Restore to pixel space, This is the final restored high-quality image.
[0046] In step (6), the parameters θ of the feature domain residual estimation network F, the feature mapping function M and the image reconstruction function Q are F ,θ M ,θ Q Both are optimized by minimizing L1 distance and structural similarity (SSIM):
[0047]
[0048] Where SSIM is the structural similarity calculation formula, and λ is the trade-off hyperparameter.
[0049] To better illustrate the effectiveness of this invention, we will use a comparative experiment to demonstrate the denoising effect. The experiment uses two commonly used real-world noise datasets: SIDD and DnD. Six representative image denoising methods are selected for comparison with the experimental results of this invention. These six representative image denoising methods are:
[0050] Method 1: The method proposed by Dabov et al., reference “Kostadin Dabov, Alessandro Foi, Vladimir Katkovnik, and Karen Egiazarian. Image denoising by sparse 3-d transform-domain collaborative filtering. IEEE Transactions on Image Processing, 16(8): 2080-2095, Aug. 2007.”
[0051] Method 2: The method proposed by Zhang et al., reference “Kai Zhang, Wangmeng Zuo, Yunjin Chen, Deyu Meng, and Lei Zhang. Beyond a gaussian denoiser: residual learning of deep cnn for image denoising. IEEE Transactions on Image Processing, 26(7): 3142-3155, Jul. 2016.”
[0052] Method 3: The method proposed by Guo et al., reference "Shi Guo, Zifei Yan, Kai Zhang, Wangmeng Zuo, and Lei Zhang. Toward convolutional blind denoising of real photographs. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 1712–1722, Jun. 2019."
[0053] Method 4: The method proposed by Anwar et al., reference "Saeed Anwar and Nick Barnes. Realimage denoising with feature attention. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 3155–3164, Oct. 2019."
[0054] Method 5: The method proposed by Kim et al., reference "Yoonsik Kim, Jae Woong Soh, Gu YongPark, and Nam Ik Cho. Transfer learning from synthetic to real-noise denoising with adaptive instance normalization. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 3482-3492, Jun. 2020."
[0055] Method 6: The method proposed by Ren et al., reference "Chao Ren, Xiaohai He, Chuncheng Wang, and Zhibo Zhao. Adaptive consistency prior based deep network for image denoising. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 8596–8606, Jun. 2021."
[0056] For the proposed method of correcting noise prior information, we used a mixture of synthetic and real noise data for training and tested it in a real-noise experimental setting. We used Peak Signal to Noise Ratio (PSNR, dB) and Structure Similarity Index (SSIM) as objective evaluation metrics. The higher the PSNR value and the closer the SSIM value is to 1, the better the denoising effect.
[0057] The contents of the comparative experiment are as follows:
[0058] Experiment 1 uses the SIDD dataset for testing. It includes two test sets: the SIDD validation set and the SIDD benchmark, both containing 1280 image blocks of size 256*256. Methods 1, 2, 3, 4, 5, 6, and the proposed method are used to denoise the noisy images. Table 1 shows the average results of each method on these test images. In addition, for visual comparison, Figure 4 The restoration results of an image in the SIDD validation set of the test gallery are given when the noise standard deviation is 50.
[0059] Table 1
[0060]
[0061] Experiment 2 corresponds to the DnD dataset, which contains 50 high-resolution test image pairs. However, due to the large size of the images, DnD cropped all images into 512*512 image blocks, resulting in a total of 1000 image blocks for testing. The noisy images were denoised using methods 1, 2, 3, 4, 5, 6, and the method of the present invention. Table 2 shows the average results of each method on the test image library. In addition, for visual comparison, Figure 5 The denoising results of each method on an image in the test gallery DnD are given.
[0062] Table 2
[0063]
[0064] Comparative experiments show that the proposed method achieves significantly higher PSNR and SIDD values across different datasets. Furthermore, the proposed method achieves superior denoising results on the DnD test set, which lacks a matching training set. This further demonstrates the method's significant advantages in removing real noise and restoring detail.
[0065] from Figure 4 and Figure 5 The experimental results show that the edges and textures of the images denoised by methods 1, 2, 3, 4, 5, and 6 are too smooth or blurred, losing many details. Some methods even produce some false textures, which are misleading to visual observation. However, the denoised images of the present invention have clearer edges and textures, produce fewer artifacts, are closer to the original image, and have better visual effects.
[0066] In summary, compared with the comparative methods, the denoising effect of the present invention has obvious advantages in both subjective and objective evaluation. Therefore, the present invention is an effective image denoising method.
Claims
1. A real image denoising method based on noise correction and guided residual estimation, characterized by The following steps are involved: Step 1: Build a noise predictor network to predict the noise level map of the noisy image; Step 2: Using the guidance of the noise level map, a feature domain residual estimation denoising network module is designed to estimate the denoising residual of the noisy image to obtain a preliminary denoising result; Step 3: Construct a noise corrector network to correct the noise level map using the denoising results from the previous stage; Step 4: Using the weight sharing method, the corrected noise level map and the denoising result of the previous stage are simultaneously input into the new residual estimation network to update the denoising residual and obtain a better denoising result; Step 5: Repeat steps 3 and 4 until the specified number of iterations k is reached, and finally a complete real image denoising network based on noise correction and guided residual estimation is constructed to output the final denoising result; Step 6: Use the public training image dataset to train the constructed deep network by minimizing the loss function; Step 7: Input the noisy image into the deep network trained in step 6 to obtain the restored clean image.
2. The method according to claim 1, characterized in that In step 1, a noise predictor P is constructed to preliminarily predict the noise level map of the noise image y. The predictor is optimized by minimizing the L1 distance. The specific formula is as follows: Among them, σ represents the noise level map of the noise image in the raw space, θ P is the parameter of the predictor P; the predictor is implemented through a convolutional neural network, which adopts a structure of four convolutional layers interspersed with activation layers, and introduces parameter-free attention before the last convolutional layer.
3. The method according to claim 1, characterized in that The feature domain residual proposed in step 2 uses the feature mapping function M to transform the noise image from the pixel space to the feature space to extract the higher-level feature information I0 of the image, I0 = M(y); Correspondingly, the image reconstruction function Q is used to transform the restored clean image features from the feature space to the pixel space; both feature extraction and image reconstruction are implemented using the "convolution layer + activation layer + convolution layer" structure; the initial feature I0 is passed through the noise level map Guide, estimate the feature domain residual R1, and finally obtain the preliminary feature domain denoising: I1=I0+R1.
4. The method according to claim 1, characterized in that In step 3, a noise corrector C is constructed to correct the noise level map predicted in the previous step to make it closer to the ground truth. The present invention optimizes the corrector by minimizing the L1 distance: Among them, θ C is the parameter of the corrector C; the corrector has a similar structure to the predictor, but the input is replaced by the output feature I of the previous stage i and noise level after channel expansion The correction amount of the noise level map obtained by the corrector is Corrected noise level map 5. The method according to claim 1, characterized in that The feature domain residual estimation network constructed in step 4; the denoising residual R of the feature domain i+1 It is realized through the feature domain residual estimation network F. The network F adopts a residual structure to perform restoration in the feature space; the three inputs of the network F: I0, I i and Each of them is subjected to feature adjustment through a convolutional layer, and the three adjusted features are cascaded and fused: Among them, f0 represents the fused feature, C1, C2, and C3 all represent a convolution layer with a convolution kernel size of 3*3. Indicates cascade, C fus Represents the convolution layer for feature fusion; the fused features are sent to n series-connected dynamic joint attention modules D, and residual learning is performed by guiding the fusion features through attention; module D is composed of "multiple convolution layers + channel-space joint attention", but in the channel-space joint attention, both channel features and spatial features are modulated by dynamic convolution; finally, the output f of the nth dynamic joint attention module is n for: f n =D n (D n-1 (...D2(D1(f0))...)) After fine-tuning the output features, they are multiplied with the noise features output by the controller, and the noise features are used as prior information to guide feature learning. The controller uses a combination of three "1*1 convolution + activation layers" to perform spatial and channel transformation encoding on the noise information, and then uses "convolution layer + activation layer + convolution layer" to convert the encoded information into feature space and enhance it. Finally, the Sigmoid function is used to normalize the enhanced features and generate feature weights. In order to better retain the output feature information of the previous stage, the I after convolution is added at the end of the network. i Perform residual learning; finally, output the feature residual R through a convolutional layer i+1 : Among them C R Represents a "convolution layer + activation layer" combination for fine-tuning, F en represents the reconstruction controller driven by noise characteristics, C tail Represents the convolutional layer at the end of the feature domain residual estimation network; the updated feature space denoising result is I i+1 =I0+R i+1 .
6. The method according to claim 1, characterized in that The deep network model constructed in step 5; after the specified number of iterations k, the output feature space denoising result is I k+1 ; The denoising result I of the feature space is reconstructed by the image reconstruction function Q k+1 Restore to pixel space, This is the final restored high-quality image.
Citation Information
Patent Citations
CNN medical CT image denoising method based on noise prior
CN112419169A
Convolutional blind denoising method containing noise estimation
CN112837231A