Image denoising method and device
By training the self-supervised denoising framework on noise data, combining random blind spots and incremental blind spot strategies, using cascade networks and frequency domain enhancement expansion convolution modules, the problem of poor performance in real noise processing is solved, and an efficient and highly adaptable image denoising effect is achieved.
Patent Information
- Application Number
- CN202411995492.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-12-31
AI Technical Summary
Existing deep learning-based image denoising methods have poor performance when processing noise in real scenes, especially in fields such as medical and military. Due to the complex noise distribution and spatial correlation, existing methods are difficult to effectively denoising.
A self-supervised denoising framework based on noise data training is proposed, combining random blind spot strategy and incremental blind spot training strategy, and image denoising is realized through cascading network module and frequency domain enhancement expansion convolution module.
Maintaining high-quality image reconstruction in complex real noise environments significantly improves the denoising performance, is suitable for image processing tasks in various complex noise scenarios, without paired data support.
Smart Images

Figure CN120047341A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and in particular, relates to a method and device for image denoising. Background Art
[0002] Noise is an important factor affecting image quality. With the development of cameras, the resolution has been greatly improved. However, due to factors such as low-light conditions, sensor heating, and channel interference, noise is still inevitably introduced. Noise will destroy the signal distribution and texture features of image pixels, thus affecting the performance of tasks such as object detection, semantic segmentation, and object tracking. Images are widely used in modern society. In addition to aesthetics and clarity, they are also widely used in industrial, military, medical and other fields. The quality of pictures directly affects the success or failure of tasks such as defect detection, medical diagnosis, deployment arrangement, and information collection. Therefore, it is very important to recover clear images from damaged images.
[0003] Image denoising technology is a basic technology in low-level vision tasks, aiming to recover clean images from noisy observations, and is widely used in camera signal processing pipelines, industrial, medical, military and other fields. With the development of deep neural networks, learning-based image denoising methods have made remarkable progress. Compared with traditional denoising methods, they have obvious advantages in terms of denoising speed, denoising performance, and large-scale denoising.
[0004] Currently, most deep learning-based denoising algorithms are supervised learning methods, and such methods need to be trained on a large number of noise-clean image pairs. The most common method for constructing a training dataset is to add additive white Gaussian noise (AWGN) to clean images to artificially synthesize noisy images, so as to obtain a large number of clean-noise image pairs. However, compared with synthetic noise, the noise distribution in real scenes is often more complex, with signal dependence and spatial correlation. This results in poor performance or even failure of the trained denoiser when applied to real noise images. To solve this problem, some researchers have tried to capture clean-noise image pairs in real scenes to form a training dataset, such as the SIDD dataset. However, collecting real datasets requires long exposure or multiple shootings, which cannot be achieved in some special fields such as medical and military.
[0005] To overcome the dependence on large-scale paired datasets, self-supervised learning denoising methods that do not require clean images have received increasing attention. In 2019, Alexander Krull et al. proposed the Noise2Void method. This method is based on the following assumptions: pixel signals are spatially correlated, while noise signals are spatially independent and have a mean of zero. On this basis, they proposed the Blind Spot Network (BSN) denoising method. Since then, many researchers have improved the blind spot network proposed in Noise2Void and achieved remarkable results when dealing with synthetic noise (such as AWGN). The blind spot network has become one of the most representative methods in the field of self-supervised image denoising.
[0006] Each pixel generated by the blind spot network is estimated from the noise pixels within its receptive field and does not include the pixel itself. This design enables the network to be trained through a self-supervised loss function, making it possible to train the network based only on noisy images, that is, both the input and target of the network are the same noisy image. In addition, the design of the blind spot can effectively prevent the network from converging to the identity mapping. Under the assumption that the noise has zero mean and is independent and identically distributed (IID), BSN can theoretically converge to a noise-free clean image. To achieve the blindness of the network, researchers usually adopt two strategies: one is to mask a large number of pixels in the image, and the other is to mask the central pixel through the design of the convolution kernel. However, the idealized noise assumption of this method limits its practical application performance. Since the noise in the real world usually does not satisfy the assumptions of independent and identically distributed and zero mean, such methods perform poorly when dealing with real image noise.
[0007] To address the above problems, researchers have proposed a variety of self-supervised real image denoising methods. Among them, Lee et al. proposed the APBSN method, which introduced the Pixel Shuffling Down-sampling (PD) operation into the blind spot network. Specifically, this method first downsamples the input noisy image to break the correlation between noise pixels, thus meeting the requirements of the blind spot network for the independence of noise signals. Subsequently, the downsampled image after recombination is input into the blind spot network for denoising processing. This method has successfully achieved self-supervised real image denoising and obtained good results.
[0008] On this basis, many subsequent studies have further improved such methods. However, the masking mechanism of most methods only masks the central pixel within the receptive field and has limited effectiveness when dealing with large-area spatially correlated noise. Some methods, such as LG-BPN, MM-BSN, and AT-BSN, increase the number of blind spots to address the problem of high noise correlation. However, the blind spots in these methods are usually set at fixed positions, which easily leads to overfitting of the network to specific types of noise, resulting in poor model generalization.
[0009] In addition, self-supervised real-image denoising methods mainly break the correlation of noise through downsampling or neighborhood masking. However, according to the Nyquist-Shannon sampling theory, the downsampling process will destroy the spatial structure of the image, reduce the sampling density, and cause the loss of high-frequency details. At the same time, the neighborhood masking method may lead to the loss of key information in the network input, further affecting the denoising effect. Therefore, it is often difficult for these methods to effectively recover high-frequency details in the denoised image, limiting their performance in high-precision applications. Summary of the Invention
[0010] The present invention provides a method and device for image denoising to solve the above problems or at least partially solve the above problems.
[0011] In a first aspect, a method for image denoising is disclosed. The method includes:
[0012] Step S1: Obtain an image to be denoised and input the image into a trained cascaded network module;
[0013] Step S2: The cascaded network module performs denoising processing on the image and outputs the denoised image;
[0014] Wherein, the cascaded network module includes a pixel reorganization downsampling module, a shallow feature extraction module, a first branch and a second branch connected in sequence. The output of the first branch and the output of the second branch are fused to obtain a fused feature. The fused feature is input into a reconstruction module. The reconstruction module maps the fused feature to a three-dimensional RGB space to obtain a reconstructed image. The reconstructed image is input into a pixel reorganization upsampling module, and the pixel reorganization upsampling module outputs the denoised image;
[0015] The first branch includes a first random blind spot module and a first deep feature extraction module connected in sequence. The second branch includes a second random blind spot module and a second deep feature extraction module connected in sequence. The first random blind spot module generates a first blinded feature and uses the first blinded feature as the input of the first deep feature extraction module, and outputs a first feature. The second random blind spot module generates a second blinded feature and uses the second blinded feature as the input of the second deep feature extraction module, and outputs a second feature.
[0016] Preferably, the first random blind spot module receives the shallow features of the image, performs a convolution operation on the shallow features of the image using a first random blind spot convolution kernel, and generates first blinded features; wherein, the first random blind spot convolution kernel is obtained by element-wise multiplication of a convolution kernel and a first noise mask matrix, the first noise mask matrix is randomly generated, has the same dimension as the convolution kernel size, and is a matrix composed of 0s and 1s, and the elements with 0 in the noise mask matrix represent blind spots; the size of the convolution kernel corresponding to the first random blind spot convolution kernel is 3×3; that is:
[0017] k1 binld = k1 conv ⊙ k1 mask
[0018] k1 binld is the first random blind spot convolution kernel, k1 conv is the convolution kernel corresponding to the first random blind spot convolution kernel, k1 mask is the first noise mask matrix;
[0019] The second random blind spot module receives the shallow features of the image, performs a convolution operation on the shallow features of the image using a second random blind spot convolution kernel, and generates second blinded features; wherein, the second random blind spot convolution kernel is obtained by element-wise multiplication of a convolution kernel and a second noise mask matrix, the second noise mask matrix is randomly generated, has the same dimension as the convolution kernel size, and is a matrix composed of 0s and 1s, and the elements with 1 in the noise mask matrix represent blind spots; the size of the convolution kernel corresponding to the second random blind spot convolution kernel is 5×5; that is:
[0020] k2 binld = k2 conv ⊙ k2 mask
[0021] k2 binld is the second random blind spot convolution kernel, k2 conv is the convolution kernel corresponding to the second random blind spot convolution kernel, k2 mask is the second noise mask matrix.
[0022] Preferably, during the training process of the first random blind spot module and the second random blind spot module, the number of elements with 0 in their respective corresponding first noise mask matrix and second noise mask matrix gradually increases.
[0023] Preferably, the first deep feature extraction module has the same structure as the second deep feature extraction module, and both include 8 frequency domain enhanced dilated convolution sub-modules connected in sequence. Each frequency domain enhanced dilated convolution sub-module includes a first sub-branch and a second sub-branch connected in parallel, and a multi-scale feature fusion sub-module for fusing the output of the first sub-branch and the output of the second sub-branch;
[0024] The first sub-branch performs discrete wavelet transform on the input to obtain a frequency feature map, then performs 3×3 dilated convolution on the frequency feature map, and processes the high-frequency components in the first feature map obtained by the dilated convolution by the first ReLU layer to obtain a frequency domain feature map, and performs inverse discrete wavelet transform on the frequency domain feature map to obtain a frequency enhanced feature; the high-frequency components refer to the features in the frequency domain whose frequencies exceed a preset threshold after the input is transformed from the spatial domain to the frequency domain by discrete wavelet transform;
[0025] The second sub-branch performs 3×3 dilated convolution on the input row by row, processes the second feature map obtained by the dilated convolution by the second ReLU layer, then performs 1×1 convolution on the second feature map to obtain a pointwise convolution feature map, and processes the pointwise convolution feature map by the third ReLU layer to obtain a second feature;
[0026] The frequency enhanced feature and the second feature are respectively input into the multi-scale feature fusion sub-module;
[0027] The multi-scale feature fusion sub-module adds the frequency enhanced feature and the second feature pointwise to obtain a first fusion feature, and inputs the first fusion feature into a third sub-branch and a fourth sub-branch connected in parallel; the third sub-branch compresses the first fusion feature to one dimension through 1×1 convolution to obtain a global feature, inputs the global feature into the fourth ReLU layer, and performs 1×1 convolution processing on the processed global feature to obtain a first global feature; the fourth sub-branch extracts the local feature of the first fusion feature through 1×1 convolution, inputs the local feature into the fifth ReLU layer, and performs 1×1 convolution processing on the processed local feature to obtain a first local feature; adds the first global feature and the first local feature, and then activates the added result by the Sigmoid activation function to obtain an attention feature map; then fuses the attention feature map, the frequency enhanced feature and the second feature to obtain a second fusion feature.
[0028] Preferably, the reconstruction module maps the fusion feature to a three-dimensional RGB space to obtain a reconstructed image, and inputs the reconstructed image into a pixel reorganization upsampling module, and the pixel reorganization upsampling module outputs a denoised image, where:
[0029] The reconstruction module is five consecutive convolutional layers with a convolutional kernel size of 1×1, which maps the fusion feature to a three-dimensional RGB space to obtain a reconstructed image.
[0030] In a second aspect, an image denoising device is disclosed. The device includes:
[0031] A feature acquisition module: configured to acquire an image to be denoised and input the image into a trained cascaded network module;
[0032] A denoising module: configured to perform denoising processing on the image by the cascaded network module and output a denoised image;
[0033] Wherein, the cascaded network module includes a pixel reorganization downsampling module, a shallow feature extraction module, a first branch and a second branch connected in sequence. The output of the first branch and the output of the second branch are fused to obtain a fusion feature. The fusion feature is input into a reconstruction module. The reconstruction module maps the fusion feature to a three-dimensional RGB space to obtain a reconstructed image. The reconstructed image is input into a pixel reorganization upsampling module, and the pixel reorganization upsampling module outputs a denoised image;
[0034] The first branch includes a first random blind spot module and a first deep feature extraction module connected in sequence. The second branch includes a second random blind spot module and a second deep feature extraction module connected in sequence. The first random blind spot module generates a first blind feature, uses the first blind feature as the input of the first deep feature extraction module, and outputs a first feature. The second random blind spot module generates a second blind feature, uses the second blind feature as the input of the second deep feature extraction module, and outputs a second feature.
[0035] In a third aspect, an electronic device is disclosed. The electronic device includes:
[0036] At least one processor; and
[0037] A memory communicatively connected to the at least one processor; wherein,
[0038] The memory stores instructions executable by the at least one processor. The instructions are executed by the at least one processor so that the at least one processor can execute the method as described above.
[0039] In a fourth aspect, a non-transitory computer-readable storage medium storing computer instructions is disclosed. The computer instructions are used to cause the computer to execute the method as described above.
[0040] The present invention has the following technical effects:
[0041] The present invention aims to address the problem that the real noise distribution is more complex and spatially correlated than synthetic noise, and to break through the limitation of the existing blind spot network on the assumption of independent and identically distributed (IID) noise. In view of the difficulty in obtaining paired data of noise and clean images in fields such as medical and military, the present invention proposes a self-supervised denoising framework based on noise data training. At the same time, a random blind spot strategy and an incremental blind spot training strategy are designed to solve the overfitting problem caused by fixed blind spot positions and to improve the generalization ability of the model in multi-noise environments. In addition, the present invention introduces a frequency domain enhanced dilated convolution module, which reduces the information loss caused by downsampling and the blind spot mechanism while enhancing the detail recovery effect, achieving the goal of significantly improving the denoising performance without increasing the number of parameters.
[0042] Through the collaborative work of multiple modules such as a self-supervised learning framework, a random blind spot strategy, and multi-scale feature fusion, the present invention realizes the improvement of the denoising performance without increasing the number of parameters. It can maintain high-quality image reconstruction in complex real noise environments and achieve efficient denoising without the support of paired data, with strong adaptability and excellent generalization ability.
[0043] The technical solution of the present invention breaks through the bottleneck of existing self-supervised denoising methods in real noise processing, and through a series of innovative designs, improves the robustness, generalization ability, and detail recovery effect of the model, and is applicable to image processing tasks under various complex noise scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 It is a schematic flow diagram of the method for image denoising according to the present invention;
[0045] Figure 2 It is a schematic architecture diagram of the cascade network module according to the present invention;
[0046] Figure 3 It is a schematic diagram of the downsampling operation with a stride of 2 according to the present invention;
[0047] Figure 4 It is a schematic diagram of random blind spot processing using a convolution kernel of size 3×3 according to the present invention;
[0048] Figure 5 It is a schematic diagram of random blind spot processing using a convolution kernel of size 5×5 according to the present invention;
[0049] Figure 6 It is a schematic architecture diagram of the frequency domain enhanced dilated convolution sub-module according to the present invention;
[0050] Figure 7 It is a schematic diagram of the principle of the multi-scale feature fusion sub-module according to the present invention;
[0051] Figure 8 It is a schematic structure diagram of the device for image denoising according to the present invention. Detailed implementation manners
[0052] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0053] As Figure 1 shown, the present invention provides a method for image denoising, and the method includes:
[0054] Step S1: Obtain an image to be denoised, and input the image into a trained cascaded network module;
[0055] Step S2: The cascaded network module performs denoising processing on the image and outputs a denoised image;
[0056] Wherein, the cascaded network module includes a pixel reorganization downsampling module, a shallow feature extraction module, a first branch and a second branch connected in sequence. The output of the first branch and the output of the second branch are fused to obtain a fused feature. The fused feature is input into a reconstruction module. The reconstruction module maps the fused feature to a three-dimensional RGB space to obtain a reorganized image. The reorganized image is input into a pixel reorganization upsampling module, and the pixel reorganization upsampling module outputs a denoised image;
[0057] The first branch includes a first random blind spot module and a first deep feature extraction module connected in sequence. The second branch includes a second random blind spot module and a second deep feature extraction module connected in sequence. The first random blind spot module generates a first blinded feature, takes the first blinded feature as the input of the first deep feature extraction module, and outputs a first feature. The second random blind spot module generates a second blinded feature, takes the second blinded feature as the input of the second deep feature extraction module, and outputs a second feature.
[0058] In the present invention, first, the pixel reorganization downsampling module breaks the spatial correlation of the noise to make it conform to the independence assumption of the blind spot network. Then, the shallow feature extraction module maps the downsampled image to a high-dimensional feature map and extracts shallow features. Next, the random blind spot module realizes blindization by dynamically shielding pixels, enabling the network to learn in a self-supervised learning framework and further breaking the correlation of the noise. The deep feature extraction module further extracts global and complex feature information. Finally, the reconstruction module restores the high-dimensional feature map to an RGB image, and restores the original resolution through the upsampling module to output the final denoising result.
[0059] As Figure 3As shown, the pixel recombination downsampling module adopts pixel recombination downsampling operation, i.e., PD operation (this operation has been disclosed in APBSN). This operation reorganizes the pixel arrangement to break the spatial correlation of real noise, making it more in line with the independence assumption of the blind spot network for noisy images, and laying a foundation for subsequent blind spot network denoising.
[0060] The shallow feature extraction module is a convolutional neural network with a convolutional kernel size of 1×1, which is used to extract the shallow features of the image, that is, map each pixel of the image to be denoised into a high-dimensional feature space. For example, given a mixed degraded image X, the feature size of X is 3×H×W, where H is the height of the image and W is the width of the image. After inputting into the 1×1 convolutional layer, a high-dimensional shallow feature map of 128×H×W is obtained.
[0061] Furthermore, the first random blind spot module receives the shallow features of the image and performs a convolutional operation on the shallow features of the image using the first random blind spot convolutional kernel to generate the first blinded feature; wherein, the first random blind spot convolutional kernel is obtained by element-wise multiplication of the convolutional kernel and the first noise mask matrix, and the first noise mask matrix is randomly generated, has the same dimension as the convolutional kernel size, and is a matrix composed of 0s and 1s. The elements of the noise mask matrix that are 0 represent blind spots; the size of the convolutional kernel corresponding to the first random blind spot convolutional kernel is 3×3; that is:
[0062] k1 binld =k1 conv ⊙k1 mask
[0063] k1 binld is the first random blind spot convolutional kernel, k1 conv is the convolutional kernel corresponding to the first random blind spot convolutional kernel, k1 mask is the first noise mask matrix.
[0064] The second random blind spot module receives the shallow features of the image and performs a convolutional operation on the shallow features of the image using the second random blind spot convolutional kernel to generate the second blinded feature; wherein, the second random blind spot convolutional kernel is obtained by element-wise multiplication of the convolutional kernel and the second noise mask matrix, and the second noise mask matrix is randomly generated, has the same dimension as the convolutional kernel size, and is a matrix composed of 0s and 1s. The elements of the noise mask matrix that are 1 represent blind spots; the size of the convolutional kernel corresponding to the second random blind spot convolutional kernel is 5×5; that is:
[0065] k2 binld =k2 conv ⊙k2 mask
[0066] k2 binld is the second random blind spot convolutional kernel, k2conv is the convolution kernel corresponding to the second random blind spot convolution kernel, k2 mask is the second noise mask matrix.
[0067] Furthermore, during the training process of the first random blind spot module and the second random blind spot module, the number of zero elements in their respective corresponding first noise mask matrix and second noise mask matrix gradually increases.
[0068] In the present invention, as Figures 4 - 5 shown, the key of the random blind spot module lies in dynamically masking some pixels during the convolution process, enabling the network to robustly process the noise in the input data under the self-supervised learning framework. The masked pixel positions are regarded as "blind spots", that is, the values at these positions are ignored or randomly replaced during training. This can prevent the model from relying on noise information and enable it to learn to infer features based on the context.
[0069] The mathematical expression and operation formula of the random blind spot module are as follows:
[0070] k binld = k conv ⊙ k mask
[0071] where k binld represents the convolution kernel of the random blind spot module, k conv represents an ordinary 3×3 or 5×5 convolution kernel, k mask represents the noise mask matrix and represents the blind spot position. The symbol "⊙" represents element-wise multiplication. The product of the convolution kernel k conv and the mask k mask forms the random blind spot convolution kernel k binld . During the training process, the mask matrix k mask is randomly generated on the premise of ensuring that the center point is a blind spot and the number of blind spots is determined, so as to realize the random blindening operation of pixel positions.
[0072] The function of the random blind spot module is:
[0073] Break the correlation of noise: Each time during training, the mask matrix will change randomly, ensuring that the model does not memorize specific spatial patterns. This random blind spot mechanism effectively breaks the spatial dependence of noise, enabling the network to better process complex noise.
[0074] Prevent overfitting: Through the random blindening operation, the network cannot rely on the input noise features, thereby improving the generalization ability of the model and avoiding overfitting on the training data.
[0075] Self-supervised optimization: The masked blind spots are completed by the network through context information, and the true values of these pixels are restored in the reconstruction module. Finally, the reconstruction loss is calculated with the original noisy image to optimize the network performance.
[0076] Furthermore, the first deep feature extraction module and the second deep feature extraction module have the same structure, and both include 8 frequency domain enhanced dilated convolution sub-modules connected in sequence. Each frequency domain enhanced dilated convolution sub-module includes a first sub-branch and a second sub-branch connected in parallel, and a multi-scale feature fusion sub-module that fuses the output of the first sub-branch and the output of the second sub-branch;
[0077] The first sub-branch performs discrete wavelet transform on the input to obtain a frequency feature map, then performs 3×3 dilated convolution on the frequency feature map, and processes the high-frequency components in the first feature map obtained by the dilated convolution by the first ReLU layer to obtain a frequency domain feature map, and performs inverse discrete wavelet transform on the frequency domain feature map to obtain a frequency enhanced feature; the high-frequency components refer to the features whose frequencies exceed a preset threshold in the frequency domain after the input is transformed from the spatial domain to the frequency domain by the discrete wavelet transform.
[0078] The second sub-branch performs 3×3 dilated convolution on the input row by row, processes the second feature map obtained by the dilated convolution by the second ReLU layer, then performs 1×1 convolution on the second feature map to obtain a pointwise convolution feature map, and processes the pointwise convolution feature map by the third ReLU layer to obtain a second feature;
[0079] The frequency enhanced feature and the second feature are respectively input into the multi-scale feature fusion sub-module;
[0080] The multi-scale feature fusion sub-module adds the frequency enhanced feature and the second feature pointwise to obtain a first fusion feature, and inputs the first fusion feature into a third sub-branch and a fourth sub-branch connected in parallel; the third sub-branch compresses the first fusion feature to one dimension through 1×1 convolution to obtain a global feature, inputs the global feature into the fourth ReLU layer, and performs 1×1 convolution processing on the processed global feature to obtain a first global feature; the fourth sub-branch extracts the local feature of the first fusion feature through 1×1 convolution, inputs the local feature into the fifth ReLU layer, and performs 1×1 convolution processing on the processed local feature to obtain a first local feature; adds the first global feature and the first local feature, and then activates the added result by the Sigmoid activation function to obtain an attention feature map; then fuses the attention feature map, the frequency enhanced feature and the second feature to obtain a second fusion feature.
[0081] The second feature map is a large receptive field feature map.
[0082] Furthermore, the reconstruction module maps the fused features to a three-dimensional RGB space to obtain a reconstructed image, and inputs the reconstructed image into a pixel shuffle upsampling module, which outputs a denoised image, where:
[0083] The reconstruction module is composed of five consecutive convolutional layers with a convolutional kernel size of 1×1, which maps the fused features to a three-dimensional RGB space to obtain a reconstructed image.
[0084] The pixel shuffle upsampling module restores the reconstructed image to the same resolution as the image to be denoised to obtain a denoised image.
[0085] In the present invention, as Figures 6 - 7 shown, a deep feature extraction module is designed to solve the problem of loss of detailed information caused by PD operations and blind spot settings. The core of the design of this module lies in combining frequency enhancement and dilated convolution techniques to highlight the high-frequency details in the image. This module is composed of 8 frequency domain enhanced dilated convolution modules, each module includes a frequency branch and a dilated convolution branch, and they are fused through a multi-scale feature fusion module.
[0086] Specifically, the frequency branch first uses the discrete wavelet transform (DWT) to transform the image from the spatial domain to the frequency domain, then processes the high-frequency components (Rule layer) through a 3×3 dilated convolution, and finally uses the inverse discrete wavelet transform (IDWT) to transform the processed image back to the spatial domain. The dilated convolution branch uses operations such as 3×3 dilated convolution, ReLU activation function, and pointwise convolution.
[0087] (1) Frequency branch design: As Figure 6 shown in the upper branch in, the input shallow feature F is transformed to the frequency domain using the discrete wavelet transform (DWT) to obtain a frequency feature map F freq . Then, the high-frequency components are processed sequentially through a 3×3 dilated convolution and a ReLU layer to highlight and enhance the high-frequency details in the image. The frequency domain feature map F h after high-frequency processing is inversely transformed back to the spatial domain to obtain the final frequency enhanced feature F 1 . These operations are expressed as:
[0088] F 1 = IDWT(ReLU(Conv(DWT(F))))
[0089] (2) Dilated convolution branch design: As Figure 6 shown in the lower branch in, the input shallow feature F is processed through a 3×3 dilated convolution, ReLU, 1×1 convolution, and ReLU layers respectively to obtain the extracted feature F 2 . These operations are expressed as:
[0090] F 2 = ReLU(Conv 3 (ReLU(Conv 1 (F))))
[0091] (3) Multi-scale feature fusion module: As Figure 7 shown, first, the frequency-enhanced feature F 1 and the feature F 2 are added pointwise to obtain the fused feature. The fused feature is then input into two branches respectively. In the first branch, the fused feature is first compressed from high-dimensional to one-dimensional through a 1×1 convolution to obtain a feature map of 1×H×W to acquire global features. Then it is input into ReLU and a 1×1 convolution for further processing, and finally a global feature of 1×H×W is obtained; in the second branch, the dimension of the fused feature remains unchanged, and local features are mainly extracted through a 1×1 convolution, ReLU, and a 1×1 convolution, and a local feature of C×H×W is output. After adding the feature 1×H×W and C×H×W, the Sigmoid activation function is used to generate the attention feature map F att (C×H×W). Finally, the attention feature map F att is combined with the frequency-enhanced feature F 1 and the feature F 2 to obtain the fused feature F fuse . These operations are expressed as:
[0092]
[0093] In the present invention, the high-dimensional feature map after deep feature extraction is transmitted to the reconstruction module. This module maps the deep features back to the three-dimensional RGB space to generate a preliminarily denoised image. The reconstruction module consists of 5 1×1 convolutions, gradually mapping the high-dimensional feature map to a three-dimensional RGB image.
[0094] Finally, the reconstructed image is restored to the original resolution through the pixel recombination upsampling module, and the denoising result is output. This module ensures the restoration of image details while retaining the denoising effect of the network during the multi-scale processing.
[0095] To address the challenges brought by the random blind spot module during the training process, the present invention uses an incremental blind spot training strategy. As the training progresses, the number of blind spots gradually increases, and the amount of effective information that the network can directly utilize in each convolution operation decreases accordingly, which makes the network face greater difficulties in processing complex tasks. To avoid the network falling into too high a learning difficulty in the initial stage, this strategy adopts a step-by-step progressive method to ensure that the network can fully learn the basic features and global information of the image in the early stage.
[0096] In the initial stage of training, the number of blind spots is small, and the network can access more pixel information, enabling it to quickly learn the basic patterns and local features in the data. The training in this stage is relatively simple, which helps the network establish a good foundation and ensures a relatively stable convergence effect in the initial stage. As the training progresses, the number of blind spots gradually increases, and the local information directly available to the network becomes sparser. This forces the model to rely more on global context information for prediction, thereby improving its adaptability to complex scenarios. This progressive training process effectively increases the difficulty of the task, enabling the network to gradually learn to handle more challenging scenarios and enhancing the generalization ability of the model on unknown data.
[0097] In addition, the strategy of incrementally increasing blind spots helps prevent the network from overfitting simple data in the early stage, avoiding the model relying too much on local features while ignoring global structural information. By gradually increasing the number of blind spots, the network benefits from a more hierarchical feature learning approach, enabling it to extract features at different scales. This multi-scale feature extraction ability makes the model more robust in image denoising tasks under various different noise patterns and complex environments. The specific training arrangement is shown in Table 1.
[0098] Table 1: Incremental Blind Spot Training Plan during Training
[0099]
[0100] The present invention has the following characteristics:
[0101] 1. Self-supervised learning denoising method
[0102] The present invention designs a self-supervised learning denoising framework that only needs to train the model on a noisy dataset without the need for paired data of clean images and noisy images, solving the problem of difficult acquisition of large-scale paired data in special fields such as medical and military. This method reduces the data collection cost and improves the applicability of the model in practical applications.
[0103] 2. Random blind spot module
[0104] Aiming at the problem that the real noise distribution in real-world scenarios is more complex and has spatial correlation, which does not meet the independent and identically distributed (IID) noise assumption of the blind spot network, the present invention proposes a random blind spot module. By dynamically and randomly masking pixel positions, this strategy further reduces noise correlation and improves the robustness and denoising performance of the model in complex real noise environments.
[0105] 3. Incremental blind spot training strategy
[0106] Aiming at the problem that the blind spot positions in the existing blind spot network are fixed, which easily leads to the model overfitting to specific noise types, the present invention designs an incremental blind spot training strategy. By gradually increasing the number of blind spots during the training process, the model can gradually adapt to more complex noise patterns, improving the generalization ability of the network. In addition, this strategy also optimizes the training process, making the model easier to converge.
[0107] 4. Frequency Domain Enhanced Dilated Convolution Module
[0108] The present invention designs a frequency domain enhanced dilated convolution module to solve the problem of information loss caused by downsampling and the blind spot mechanism during the denoising process. This module enhances the feature extraction ability in the frequency domain by expanding the receptive field of the convolution kernel, thereby improving the detail recovery effect of the model without increasing the number of parameters, especially performing well in the reconstruction of high-frequency information.
[0109] 5. Multi-scale Feature Fusion Module
[0110] To further improve the denoising ability of the model, the present invention introduces a multi-scale feature fusion module. This module integrates image features of different scales to achieve efficient fusion of global information and local details, thereby improving the processing ability of complex image structures and multi-level noises during the denoising process. The fusion of multi-scale features ensures that the model has higher robustness and adaptability when dealing with noises of different sizes and shapes.
[0111] 6. Achievement of High-efficiency Denoising Performance
[0112] The present invention realizes the improvement of denoising performance without increasing the number of parameters through the collaborative work of multiple modules such as a self-supervised learning framework, a random blind spot strategy, and multi-scale feature fusion. The system can maintain high-quality image reconstruction in a complex real noise environment and achieve efficient denoising without the support of paired data, with strong adaptability and excellent generalization ability.
[0113] The technical solution of the present invention breaks through the bottleneck of existing self-supervised denoising methods in real noise processing, and through a series of innovative designs, improves the robustness, generalization ability, and detail recovery effect of the model, and is applicable to image processing tasks under various complex noise scenarios.
[0114] As Figure 8 shown, the present invention provides an image denoising device, and the device includes:
[0115] Feature acquisition module: configured to acquire the image to be denoised and input the image into the trained cascaded network module;
[0116] Denoising module: configured to perform denoising processing on the image by the cascaded network module and output the denoised image;
[0117] Among them, the cascaded network module includes a pixel reorganization downsampling module, a shallow feature extraction module, a first branch and a second branch connected in sequence. The output of the first branch and the output of the second branch are fused to obtain a fused feature. The fused feature is input into a reconstruction module. The reconstruction module maps the fused feature to a three-dimensional RGB space to obtain a reconstructed image. The reconstructed image is input into a pixel reorganization upsampling module, and the pixel reorganization upsampling module outputs a denoised image;
[0118] The first branch includes a first random blind spot module and a first deep feature extraction module connected in sequence. The second branch includes a second random blind spot module and a second deep feature extraction module connected in sequence. The first random blind spot module generates a first blinded feature, uses the first blinded feature as the input of the first deep feature extraction module, and outputs a first feature. The second random blind spot module generates a second blinded feature, uses the second blinded feature as the input of the second deep feature extraction module, and outputs a second feature.
[0119] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that it is still possible to modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features, and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for denoising an image, characterized in that: The method comprises the following steps: Step S1: obtaining an image to be denoised, and inputting the image into a trained cascade network module; Step S2: the cascade network module performs denoising on the image and outputs the denoised image; The cascade network module includes a pixel reorganization downsampling module, a shallow feature extraction module, a first branch and a second branch connected to the shallow feature extraction module, the output of the first branch is fused with the output of the second branch to obtain a fused feature, the fused feature is input into a reconstruction module, the reconstruction module maps the fused feature to a three-dimensional RGB space, and the obtained reorganized image is input into a pixel reorganization upsampling module, and the pixel reorganization upsampling module outputs a denoised image; The first branch includes a first random blind spot module and a first deep feature extraction module connected in sequence, and the second branch includes a second random blind spot module and a second deep feature extraction module connected in sequence; the first random blind spot module generates a first blinding feature, uses the first blinding feature as the input of the first deep feature extraction module, and outputs a first feature; the second random blind spot module generates a second blinding feature, uses the second blinding feature as the input of the second deep feature extraction module, and outputs a second feature.
2. The method according to claim 1, characterized in that The first random blind spot module receives the shallow features of the image, and uses the first random blind spot convolution kernel to perform a convolution operation on the shallow features of the image to generate a first blind feature; wherein the first random blind spot convolution kernel is obtained by multiplying the convolution kernel by the first noise mask matrix element by element, the first noise mask matrix is a randomly generated matrix composed of 0 and 1 with the same dimension as the convolution kernel size, and the elements of the noise mask matrix with the element of 0 represent blind spots; the size of the convolution kernel corresponding to the first random blind spot convolution kernel is 3×3; that is: k1 binld =k1 conv ⊙k1 mask k1 binld is the first random blind spot convolution kernel, k1 conv is the convolution kernel corresponding to the first random blind spot convolution kernel, k1 mask is the first noise mask matrix; The second random blind spot module receives the shallow features of the image, and uses the second random blind spot convolution kernel to perform a convolution operation on the shallow features of the image to generate a second blind feature; wherein the second random blind spot convolution kernel is obtained by multiplying the convolution kernel by the second noise mask matrix element by element, the second noise mask matrix is a randomly generated matrix composed of 0 and 1 with the same dimension as the convolution kernel size, and the elements of the noise mask matrix with an element of 1 represent blind spots; the size of the convolution kernel corresponding to the second random blind spot convolution kernel is 5×5; that is: k2 binld =k2 conv ⊙k2 mask k2 binld is the second random blind spot convolution kernel, k2 conv is the convolution kernel corresponding to the second random blind spot convolution kernel, k2 mask is the second noise mask matrix.
3. The method according to claim 2, characterized in that During the training process of the first random blind spot module and the second random blind spot module, the number of 0 elements in the first noise mask matrix and the second noise mask matrix corresponding to each of them gradually increases.
4. The method according to claim 2, characterized in that The first deep feature extraction module and the second deep feature extraction module have the same structure, and both include 8 frequency domain enhancement dilated convolution submodules connected in sequence, each frequency domain enhancement dilated convolution submodule includes a first sub-branch and a second sub-branch connected in parallel, and a multi-scale feature fusion submodule that fuses the output of the first sub-branch with the output of the second sub-branch; The first sub-branch performs a discrete wavelet transform on the input to obtain a frequency feature map, then performs a 3×3 dilated convolution on the frequency feature map, processes the high-frequency components in the first feature map obtained by the dilated convolution by the first ReLU layer to obtain a frequency domain feature map, and performs an inverse discrete wavelet transform on the frequency domain feature map to obtain a frequency enhancement feature; the high-frequency component refers to a feature in the frequency domain where the frequency exceeds a preset threshold value when the input is converted from the spatial domain to the frequency domain by discrete wavelet transform; The second sub-branch performs a 3×3 dilated convolution on the input, processes the dilated convolution by the second ReLU layer to obtain a second feature map, performs a 1×1 convolution on the second feature map to obtain a point-by-point convolution feature map, and processes the point-by-point convolution feature map by the third ReLU layer to obtain a second feature; Inputting the frequency enhancement feature and the second feature into the multi-scale feature fusion submodule respectively; The multi-scale feature fusion submodule adds the frequency enhancement feature and the second feature point by point to obtain a first fused feature, and inputs the first fused feature into the third sub-branch and the fourth sub-branch connected in parallel respectively; the third sub-branch compresses the first fused feature to one dimension through 1×1 convolution to obtain a global feature, inputs the global feature into the fourth ReLU layer, and performs 1×1 convolution on the processed global feature to obtain a first global feature; the fourth sub-branch extracts the local feature of the first fused feature through 1×1 convolution, inputs the local feature into the fifth ReLU layer, and performs 1×1 convolution on the processed local feature to obtain a first local feature; the first global feature and the first local feature are added, and the result of the addition is activated by a Sigmoid activation function to obtain an attention feature map; the attention feature map, the frequency enhancement feature and the second feature are fused to obtain a second fused feature.
5. The method according to claim 2, characterized in that The reconstruction module maps the fusion features to a three-dimensional RGB space to obtain a reconstructed image, and the reconstructed image is input into a pixel reconstructing up-sampling module, and the pixel reconstructing up-sampling module outputs a denoised image, wherein: The reconstruction module is five sequentially connected convolution layers with a convolution kernel size of 1×1, which maps the fusion features to a three-dimensional RGB space to obtain a reconstructed image.
6. A device for denoising an image, characterized in that: The device comprises: Feature acquisition module: configured to acquire the image to be denoised and input the image into the trained cascade network module; Denoising module: configured as the cascade network module to perform denoising on the image and output a denoised image; The cascade network module includes a pixel reorganization downsampling module, a shallow feature extraction module, a first branch and a second branch connected to the shallow feature extraction module, the output of the first branch is fused with the output of the second branch to obtain a fused feature, the fused feature is input into a reconstruction module, the reconstruction module maps the fused feature to a three-dimensional RGB space, and the obtained reorganized image is input into a pixel reorganization upsampling module, and the pixel reorganization upsampling module outputs a denoised image; The first branch includes a first random blind spot module and a first deep feature extraction module connected in sequence, and the second branch includes a second random blind spot module and a second deep feature extraction module connected in sequence; the first random blind spot module generates a first blinding feature, uses the first blinding feature as the input of the first deep feature extraction module, and outputs a first feature; the second random blind spot module generates a second blinding feature, uses the second blinding feature as the input of the second deep feature extraction module, and outputs a second feature.
7. A computer-readable storage medium, wherein a plurality of instructions are stored in the storage medium; the plurality of instructions are used for a processor to load and execute the method as claimed in any one of claims 1 to 5.
8. An electronic device, characterized in that: The electronic device comprises: A processor, which is used to execute multiple instructions; A memory for storing a plurality of instructions; The plurality of instructions are used to be stored in the memory and loaded and executed by the processor according to any one of claims 1 to 5.
Citation Information
Patent Citations
Real image denoising method based on multi-scale fusion and edge enhancement
CN112233038A
Ocean image denoising method based on SwinIR
CN117670714A
Self-supervised image denoising method, system and device and readable storage medium
CN117710240A
Self-supervised image denoising method based on three-stage feature extraction
CN118097159A