Method and device for denoising real image
By designing a denoising network framework combining blind-spot branches and non-blind-spot branches, the noise correlation and information loss problems of traditional blind-spot networks when processing real noisy images are solved, and a more efficient image denoising effect is achieved.
Patent Information
- Application Number
- CN202411978130.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-12-31
AI Technical Summary
When traditional blind spot networks process real noisy images, there is pixel-level correlation between noise, resulting in poor denoising effect, and masking operations may lead to the loss of image key information, especially when restoring high-frequency information.
A denoising network framework combining blind-spot branches and non-blind-spot branches is designed. Blind-spot branches break the spatial correlation of noise through dense sampling block blind-spot convolution module and U-Net network. Non-blind-spot branches avoid information loss through gradient-free training and revisible loss functions, and weighted the denoising results of both.
It effectively solves the noise correlation problem when denoising real noise images, avoids information loss, improves image denoising effect, and achieves higher denoising accuracy and image quality.
Smart Images

Figure CN120047340A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and in particular, relates to a method and device for denoising real images. Background Art
[0002] In the field of image processing, image denoising is an important link to improve image quality and enhance image details. Traditional denoising methods rely on the statistical characteristics or prior knowledge of images. In recent years, deep learning-based denoising methods have been widely used. In particular, the Blind-Spot Network can perform self-supervised learning without clean labels and shows certain denoising capabilities. However, the Blind-Spot Network has two main defects: First, for real noise images, the noise often has pixel-level correlations. Based on the assumption of independent noise, the denoising effect of the Blind-Spot Network is poor. Second, during the denoising process, the Blind-Spot Network may cause the loss of key image information through masking operations, especially when restoring high-frequency information such as edges and textures, the effect is not good. Therefore, there is an urgent need for a new image denoising method that can solve the problem of real correlation noise and avoid information loss.
[0003] The Blind-Spot Network is a variant of the convolutional neural network and is an image denoising method based on self-supervised learning. When the Blind-Spot Network is applied to the image denoising task of real noise, there are the following two problems:
[0004] (1) During the image denoising process of the Blind-Spot Network, the central pixel x RF(i) of each neighborhood region x in the input image will be masked, and the pixel x i contains the true image information of the corresponding pixel in the output image i . The loss of valuable information will reduce the upper limit of denoising and thus affect the performance of the denoising algorithm.
[0005] (2) The Blind-Spot Network predicts the information of the central pixel point through a decentralized neighborhood region. When the noise signal in the image shows spatial correlation, that is, the noise signals of neighboring pixel points are correlated with each other, the noise signal will inevitably be introduced into the information of the central pixel point predicted by the Blind-Spot Network, resulting in the inability to effectively remove the noise. In the real world, noise images often show spatial correlation, which does not conform to the assumption of independent and zero-mean distribution of noise in the Blind-Spot Network. Therefore, the Blind-Spot Network often performs poorly in the real noise removal task. Summary of the Invention
[0006] The present invention provides a method and device for denoising real images to solve the above problems or at least partially solve the above problems.
[0007] First aspect, a method for denoising a real image is disclosed. The method includes:
[0008] Step S1: Obtain training samples, use the training samples as input to train a denoising module, and obtain a trained denoising module.
[0009] Step S2: Obtain a real image to be denoised, input the real image into the trained denoising module, and the blind spot branch of the trained denoising module performs denoising processing on the real image and outputs a denoised image.
[0010] Wherein, the denoising module includes a non-blind spot branch and a blind spot branch connected in parallel.
[0011] The non-blind spot branch is a non-blind spot branch without gradients, including a downsampling module, a convolutional denoising network module, and an upsampling module connected in sequence. The non-blind spot branch receives an input, and the downsampling module performs downsampling on the input in the spatial direction; the convolutional denoising network module is a U-Net network for denoising the downsampled result to obtain denoised features, and the upsampling module restores the denoised features to the original resolution through transposed convolution operations and outputs the first denoising result of the non-blind spot branch.
[0012] Wherein, the downsampling module includes a convolutional layer with a kernel size of 1×1 and a pooling layer connected in sequence. The convolutional layer with a kernel size of 1×1 performs dimensionality increase on the training samples in the channel direction, and then the pooling layer performs max-pooling operations to achieve downsampling in the spatial direction.
[0013] The blind spot branch includes a dense sampling block blind spot convolutional module and a convolutional denoising network module connected in sequence. The dense sampling block blind spot convolutional module receives an input, uses a blind spot convolutional kernel to perform a masking operation on the input. When performing the masking operation, the masked pixel points are restored based on the pixel information in the neighborhood interval of the masked pixel points, and a masked feature is generated based on the restored pixel points and the unmasked pixel points; wherein, the blind spot convolutional kernel is obtained by element-wise multiplication of a convolutional kernel and a noise masking matrix, and the noise masking matrix is randomly generated, has the same dimension as the convolutional kernel size, and is a matrix composed of 0s and 1s. The elements of the noise masking matrix that are 0 represent blind spots; the convolutional denoising network module denoises the masked feature through convolutional operations to obtain denoised features, which are used as the second denoising result of the blind spot branch.
[0014] Fuse the first denoising result and the second denoising result as the output of the denoising module.
[0015] Preferably, the formula for fusing the first denoising result and the second denoising result is:
[0016]
[0017] Among them, a is a hyperparameter, and x b is the first denoising result, and x unb is the second denoising result. When the number of training epochs epoch ≤ 8, a = 2.0; when the number of training epochs epoch > 8,
[0018] Preferably, the loss function of the non-blind spot branch during training is
[0019] L reg = ‖x b - y‖ 1
[0020] Among them, L reg is the loss function value of the non-blind spot branch, y is the input noisy image; ||·|| 1 is the L1 loss function.
[0021] Preferably, the loss function of the blind spot branch during training is
[0022]
[0023] Among them, L rev is the loss function value of the blind spot branch, ||·|| 2 is the mean square error loss function.
[0024] Preferably, the total loss function L is:
[0025]
[0026] Preferably, in step S2, obtain the real image to be denoised, input the features of the real image into the trained denoising module, and the blind spot branch of the trained denoising module performs denoising processing on the real image, where:
[0027] Input the features of the real image into the blind spot convolution module of the dense sampling block, and use the blind spot convolution kernel to perform a masking operation on the input. When performing the masking operation, restore the masked pixel points based on the pixel information in the neighborhood interval of the masked pixel points, and generate a masked feature based on the restored pixel points and the unmasked pixel points; among them, the blind spot convolution kernel is obtained by multiplying the convolution kernel element by element with the noise mask matrix, and the noise mask matrix is randomly generated, has the same dimension as the convolution kernel size, and is a matrix composed of 0s and 1s. The elements of the noise mask matrix that are 0 represent blind spots; the convolutional denoising network module denoises the masked feature through convolutional operations to obtain the denoised feature, which is used as the second denoising result of the blind spot branch;
[0028] Use the second denoising result as the denoised image.
[0029] In a second aspect, a device for denoising a real image is disclosed. The device includes:
[0030] A training module: configured to obtain training samples, use the training samples as input to train a denoising module, and obtain a trained denoising module;
[0031] A real image denoising module: configured to obtain a real image to be denoised, input the real image into the trained denoising module, and perform denoising processing on the real image by a blind spot branch of the trained denoising module, and output the denoised image;
[0032] Wherein, the denoising module includes a non-blind spot branch and a blind spot branch connected in parallel;
[0033] The non-blind spot branch is a non-blind spot branch without gradients, including a downsampling module, a convolutional denoising network module, and an upsampling module connected in sequence. The non-blind spot branch receives an input, and the downsampling module performs downsampling on the input in the spatial direction; the convolutional denoising network module is a U-Net network, which is used to denoise the downsampled result to obtain denoised features, and the upsampling module restores the denoised features to the original resolution through transposed convolution operations and outputs the first denoising result of the non-blind spot branch;
[0034] Wherein, the downsampling module includes a convolutional layer with a convolution kernel size of 1×1 and a pooling layer connected in sequence. The convolutional layer with a convolution kernel size of 1×1 performs dimensionality increase on the training samples in the channel direction, and then the pooling layer performs a max pooling operation to achieve downsampling in the spatial direction;
[0035] The blind spot branch includes a dense sampling block blind spot convolution module and a convolutional denoising network module connected in sequence. The dense sampling block blind spot convolution module receives an input, uses a blind spot convolution kernel to perform a masking operation on the input. When performing the masking operation, the masked pixel points are restored based on the pixel information in the neighborhood interval of the masked pixel points, and masked features are generated based on the restored pixel points and the unmasked pixel points; wherein, the blind spot convolution kernel is obtained by element-wise multiplication of a convolution kernel and a noise masking matrix, and the noise masking matrix is randomly generated, has the same dimension as the convolution kernel size, and is a matrix composed of 0s and 1s. The elements of the noise masking matrix that are 0 represent blind spots; the convolutional denoising network module denoises the masked features through convolution operations to obtain denoised features, which are used as the second denoising result of the blind spot branch;
[0036] Fuse the first denoising result and the second denoising result as the output of the denoising module.
[0037] In a third aspect, an electronic device is disclosed, the electronic device comprising:
[0038] at least one processor; and
[0039] a memory communicatively connected to the at least one processor; wherein,
[0040] the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method as described above.
[0041] In a fourth aspect, a non-transitory computer-readable storage medium storing computer instructions is disclosed, the computer instructions being used to cause the computer to execute the method as described above.
[0042] The present invention has the following technical effects:
[0043] By designing a blind spot network framework for removing real correlation noise, the framework adopts a combined structure of a blind spot branch and a non-blind spot branch. The blind spot branch uses a dense sampling block blind point convolution module (DSPMC, which has been disclosed in the LG-BPN method) and a U-Net network, and breaks the spatial correlation of noise through local decentralized convolution operations, thus solving the problem of poor performance of traditional blind spot networks in processing real noise images.
[0044] By introducing a gradient-free non-blind spot training branch and a re-visibility loss function, the present invention solves the problem of key information loss caused by masking operations in blind spot networks. The gradient-free non-blind spot branch reintroduces the pixel or feature information masked by the blind spot network, and performs weighted fusion of the denoising results of the blind spot branch and the results of the non-blind spot branch, ensuring that high-frequency information, especially textures and edges, is not lost during the denoising process, thereby improving the denoising effect of the image.
[0045] The present invention also maintains the original loss function of the blind spot network as a regularization term, while introducing a new loss function to maintain the stability of the training process. Through this technical means, the denoising network further improves the image detail restoration ability without increasing excessive computational complexity, achieving higher denoising accuracy and image quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 is a schematic flowchart of the method for denoising real images according to the present invention;
[0047] Figure 2 is a schematic diagram of the training process architecture of the present invention;
[0048] Figure 3 is a schematic diagram of the inference process architecture of the present invention;
[0049] Figure 4 Schematic diagram of the blind spot convolution module of the present invention;
[0050] Figure 5 Schematic diagram of the downsampling module of the present invention;
[0051] Figure 6 Schematic structural diagram of the device for denoising real images according to the present invention. Detailed implementation manners
[0052] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0053] As Figures 1 - 3 shown, the present invention provides a method for denoising real images, and the method includes:
[0054] Step S1: Obtain training samples, use the training samples as input to train the denoising module, and obtain a trained denoising module;
[0055] Step S2: Obtain a real image to be denoised, input the real image into the trained denoising module, and the blind spot branch of the trained denoising module performs denoising processing on the real image, and outputs a denoised image;
[0056] Wherein, the denoising module includes a non-blind spot branch and a blind spot branch connected in parallel;
[0057] The non-blind spot branch is a non-blind spot branch without gradient, and includes a downsampling module, a convolutional denoising network module, and an upsampling module connected in sequence. The non-blind spot branch receives an input, and the downsampling module performs downsampling on the input in the spatial direction; the convolutional denoising network module is a U-Net network, which is used to denoise the downsampled result to obtain denoised features, and the upsampling module restores the denoised features to the original resolution through transposed convolution operations, and outputs the first denoising result of the non-blind spot branch;
[0058] Wherein, the downsampling module includes a convolutional layer with a convolution kernel size of 1×1 and a pooling layer connected in sequence. The convolutional layer with a convolution kernel size of 1×1 performs dimension elevation on the training samples in the channel direction, and then the pooling layer performs a maximum pooling operation to achieve downsampling in the spatial direction;
[0059] The blind spot branch includes a dense sampling block blind spot convolution module and a convolution denoising network module connected in sequence. The dense sampling block blind spot convolution module receives an input, uses a blind spot convolution kernel to perform a masking operation on the input. When performing the masking operation, the masked pixel points are restored based on the pixel information in the neighborhood interval of the masked pixel points, and masked features are generated based on the restored pixel points and the unmasked pixel points. Among them, the blind spot convolution kernel is obtained by element-wise multiplication of a convolution kernel and a noise masking matrix. The noise masking matrix is randomly generated, has the same dimension as the convolution kernel size, and is a matrix composed of 0s and 1s. The elements of the noise masking matrix that are 0 represent blind spots. The convolution denoising network module denoises the masked features through convolution operations to obtain denoised features, which are used as the second denoising result of the blind spot branch.
[0060] Fuse the first denoising result and the second denoising result as the output of the denoising module.
[0061] In the present invention, generally, the calculations of a convolutional neural network have gradients and will automatically update network parameters. In the code implementation of the present invention, the values of the parameters involving gradients are changed from the default True to False, so that it becomes gradient-free.
[0062] Further, the formula for fusing the first denoising result and the second denoising result is:
[0063]
[0064] where a is a hyperparameter, x b is the first denoising result, x unb is the second denoising result. When the number of training epochs epoch ≤ 8, a = 2.0; when the number of training epochs epoch > 8, The number of training epochs is the number of times the entire training data set is traversed during the training process.
[0065] In the present invention, a is a hyperparameter that controls the fusion ratio of two images. During the training process, the value of a will gradually increase as epoch increases. When the denoising network starts training, the convolution denoising network module does not yet have the ability to denoise. Therefore, it is more dependent on the blind spot network branch to train the network to have a preliminary denoising ability. As the training epoch increases, the convolution denoising network module gradually has the ability to denoise. Since the information input by the non-blind spot branch does not lack key information and the output denoising effect is better, the network becomes more and more dependent on the training process of the non-blind spot branch.
[0066] In the present invention, the convolutional denoising network module adopts a U-Net network. As the main framework for image denoising, the U-Net network aims to utilize its symmetric encoder-decoder structure to effectively remove noise through multi-scale feature extraction and restoration capabilities, while preserving the details and global information in the image. The unique skip connections in the U-Net enable the network to retain the high-resolution details of the image during the denoising process, making the denoising effect more natural and suitable for various complex noise environments.
[0067] The U-Net network consists of two parts: the encoder gradually downsamples the image through multiple convolutional and max-pooling operations to extract multi-scale features; the decoder then gradually upsamples the features through transposed convolutions to restore the original resolution of the image. At the same time, skip connections fuse the features in the encoder with the decoder to ensure that sufficient detail information is retained during the restoration process, and finally, a denoised image is output through 1×1 convolution.
[0068] The advantage of the U-Net framework lies in its powerful multi-scale feature extraction ability and skip connection mechanism, which enables it not only to effectively remove noise but also to maintain the details and global structure of the image. Especially when dealing with images with complex noise distributions, the U-Net can accurately capture noise features and effectively remove them, while retaining the texture and edge information of the image, making the denoised image clearer and more natural.
[0069] The upsampling module mainly restores the denoised image to its original resolution through transposed convolution operations. Transposed convolution (also known as deconvolution) expands the image in the spatial dimension to reversely restore the resolution reduction caused by the convolution operation. The role of this module is to gradually restore the spatial resolution of the image while maintaining or enhancing the feature information, so that important details are not lost during the restoration process, providing a high-quality denoising result for the final output.
[0070] In the present invention, the noisy image or feature map undergoes a masking operation through the dense sampling block blind spot convolution module and then is input into the convolutional denoising network module for denoising. The role of the dense sampling block blind spot convolution module is to mask part of the pixel information of the input image or feature map and restore the pixel points masked by the mask with the pixel information in the decentralized neighborhood interval. This process is mainly to avoid converging to the identity mapping during the training process. At the same time, the dense sampling block blind spot convolution module can also play a role in breaking the spatial correlation of the noise, and the functions of different blind spot convolution modules are not the same. In the present invention, the DSPMC module (which has been disclosed in the LG-BPN method) is adopted, that is, a convolutional kernel with a rhombus area centered on the center point is used to perform convolution operations on the input feature map, thereby realizing the masking of the pixel point information of the entire image. The dense sampling block blind spot convolution module is as Figure 4 shown.
[0071] The blind-spot convolution module of the dense sampling block in the present invention enables the network to ignore local noise points during processing through a specific convolution structure, thereby reducing the impact of noise on the model. Its main function is to further break the correlation of noise, enabling the network to better extract clean image features without relying on noise points. This blind-spot convolution strategy effectively enhances the denoising effect and is particularly suitable for scenarios with high noise correlation.
[0072] The present invention adds a non-blind-spot branch without gradients, which together with the blind-spot branch constitutes a non-blind-spot training strategy. In the non-blind-spot branch without gradients, the image does not pass through the blind-spot convolution module of the dense sampling block, but is directly input into the convolutional denoising network module after downsampling to obtain the denoised image. Since in this branch, the image or features do not need to pass through the blind-spot convolution module of the dense sampling block, the image information of this branch is not masked and lost. And the denoised image x ub retains relatively complete information. At the same time, this branch is set as a process without gradients, that is, when the image passes through the convolutional denoising network module, it will not affect the parameters of the network, thus avoiding the possibility of the network converging to the identity mapping.
[0073] As Figure 5 shown, the convolutional layer of the downsampling module increases the number of channels without changing the spatial dimension of the image, enhancing the feature expression ability. The max pooling reduces the spatial resolution of the image by extracting local maxima and retains the key features.
[0074] Furthermore, the loss function of the non-blind-spot branch during the training process is
[0075] L reg =‖x b -y‖ 1
[0076] where L reg is the loss function value of the non-blind-spot branch, y is the input noisy image; ||·|| 1 is the L1 loss function (Absolute Loss).
[0077] The loss function of the blind-spot branch during the training process is
[0078]
[0079] where L rev is the loss function value of the blind-spot branch, ||·|| 2 is the mean squared error loss function (Mean Squared Error, MSE).
[0080] The total loss function L is:
[0081]
[0082] Further, in step S2, a real image to be denoised is obtained, and the features of the real image are input into the trained denoising module. The blind spot branch of the trained denoising module performs denoising processing on the real image, where:
[0083] The features of the real image are input into the blind spot convolution module of the dense sampling block. A blind spot convolution kernel is used to perform a masking operation on the input. When performing the masking operation, the masked pixel points are restored based on the pixel information in the neighborhood interval of the masked pixel points. Mask features are generated based on the restored pixel points and the unmasked pixel points. Among them, the blind spot convolution kernel is obtained by element-wise multiplication of the convolution kernel and a noise mask matrix. The noise mask matrix is randomly generated, has the same dimension as the convolution kernel size, and is a matrix composed of 0s and 1s. The elements of the noise mask matrix that are 0 represent blind spots. The convolutional denoising network module denoises the mask features through convolutional operations to obtain denoised features, which are used as the second denoising result of the blind spot branch.
[0084] The second denoising result is used as the denoised image.
[0085] In the present invention, during the training process, the noisy images pass through the blind spot network branch network and the non-blind spot branch network respectively, and the denoised images x b , and image x unb . are obtained respectively. Finally, the two images are weighted and averaged to obtain the denoised image x. During the inference process, the noisy image directly passes through the trained blind spot branch network for denoising, and finally the denoised image x is obtained.
[0086] Among them, the non-blind spot branch part during the training process proposes a training strategy for the problem of key information loss in the blind spot network during the denoising process. The blind spot mechanism is a method proposed for the deficiencies of the blind spot network in the denoising task of real noisy images.
[0087] The present invention can solve the problem of key information loss in the prediction process of the blind spot network, and overcome the problem of the blind spot network failing in the denoising task of real noisy images.
[0088] The present invention designs a network framework for removing real correlation noise. The pixel information in an image generally has spatial correlation, that is, the information of the central pixel point can be predicted through the pixel information in the neighborhood. The blind spot network utilizes this characteristic and through the decentralized neighborhood interval x RF(i)To predict the information of the central pixel. When the noise signal in the image is spatially independent and has a zero mean, this prediction process can remove the noise information in the image and generate the true-value pixels in the image. However, the noise in real noise images often has pixel-level correlations and does not conform to the assumptions of the blind spot network. Therefore, the blind spot network performs poorly in the task of denoising real noise images.
[0089] The present invention designs a blind spot network framework for removing real correlated noise. This framework consists of a blind spot branch and a non-blind spot branch. The blind spot branch is mainly used to break the noise correlation and obtain a preliminary denoised image. The non-blind spot branch is mainly used to solve the problem of the decline in denoising performance caused by the loss of key information in the blind spot network.
[0090] The non-blind spot branch of the present invention uses the re-visibility loss and the gradient-free non-blind spot training branch to solve the problem of the decline in denoising performance caused by the loss of key information in the blind spot network.
[0091] The blind spot network belongs to the self-supervised learning method and can be trained in the case of only noise images without corresponding clean images. Since the input and the target of the blind spot network are the same noise image, in order to prevent the network from learning the identity mapping, the blind spot network introduces a mask blind spot mechanism in the traditional convolutional neural network, performs a masking operation on some pixel points of the input image or features, and then achieves the purpose that the predicted pixel points learn through their neighborhood intervals. However, the masking operation on the input image or features by the blind spot network results in the loss of key information, reduces the upper limit of denoising, and causes the loss of high-frequency information such as textures and edges during the denoising process.
[0092] Based on the blind spot denoising network, the present invention proposes a gradient-free non-blind spot training branch. This method reintroduces the pixel information or feature information masked by the blind spot network into the network while avoiding non-identity mapping, and performs weighted fusion on the denoised image obtained by this branch and the denoised image obtained by the blind spot network branch to obtain a new denoised image. At the same time, a re-visibility loss function is introduced, which calculates the loss between the newly obtained denoised image and the noise image, and then guides the update of the network parameters. At the same time, in order to ensure the stability of training, the original blind spot network loss is retained as a regularization term.
[0093] The blind spot branch of the present invention combines the blind spot mechanism and the U-Net network to solve the problem that the blind spot network cannot effectively denoise real noise.
[0094] The pixel information in the image generally has spatial correlation, that is, the information of the central pixel point can be predicted through the pixel information in the neighborhood. The blind spot network utilizes this characteristic and predicts the central pixel point through the decentralized neighborhood interval x RF(i)To predict the information of the central pixel. When the noise signal in the image is spatially independent and has a zero mean, this prediction process can remove the noise information in the image and generate the true-value pixels in the image. However, the noise in real noise images often has pixel-level correlations and does not conform to the assumptions of the blind spot network. Therefore, the blind spot network performs poorly in the task of denoising real noise images.
[0095] The present invention designs a dense sampling block blind spot convolution module (DSPMC), and uses this module as a preprocessing module, and then combines it with the U-Net network to form a blind spot network branch that can remove correlated noise.
[0096] As Figure 6 shown, the present invention provides a device for denoising real images, and the device includes:
[0097] A training module: configured to obtain training samples, use the training samples as inputs to train the denoising module, and obtain a trained denoising module;
[0098] A real image denoising module: configured to obtain a real image to be denoised, input the real image into the trained denoising module, and the blind spot branch of the trained denoising module performs denoising processing on the real image and outputs a denoised image;
[0099] Wherein, the denoising module includes a non-blind spot branch and a blind spot branch connected in parallel;
[0100] The non-blind spot branch is a non-blind spot branch without gradients, and includes a downsampling module, a convolutional denoising network module, and an upsampling module connected in sequence. The non-blind spot branch receives an input, and the downsampling module performs downsampling on the input in the spatial direction; the convolutional denoising network module is a U-Net network, which is used to denoise the downsampled result to obtain denoised features, and the upsampling module restores the denoised features to the original resolution through transposed convolution operations and outputs the first denoising result of the non-blind spot branch;
[0101] Wherein, the downsampling module includes a convolutional layer with a convolution kernel size of 1×1 and a pooling layer connected in sequence. The convolutional layer with a convolution kernel size of 1×1 performs dimensionality increase on the training samples in the channel direction, and then the pooling layer performs a max pooling operation to achieve downsampling in the spatial direction;
[0102] The blind spot branch includes a dense sampling block blind spot convolution module and a convolution denoising network module connected in sequence. The dense sampling block blind spot convolution module receives an input, uses a blind spot convolution kernel to perform a masking operation on the input. When performing the masking operation, the masked pixel points are restored based on the pixel information in the neighborhood interval of the masked pixel points, and masked features are generated based on the restored pixel points and the unmasked pixel points. Among them, the blind spot convolution kernel is obtained by element-wise multiplication of a convolution kernel and a noise masking matrix. The noise masking matrix is randomly generated, has the same dimension as the convolution kernel size, and is a matrix composed of 0s and 1s. The elements of the noise masking matrix that are 0 represent blind spots. The convolution denoising network module denoises the masked features through convolution operations to obtain denoised features, which are used as the second denoising result of the blind spot branch.
[0103] Fuse the first denoising result and the second denoising result as the output of the denoising module.
[0104] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that it is still possible to modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features, and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for denoising a real image, characterized in that: The method comprises the following steps: Step S1: Obtain training samples, and use the training samples as input to train a denoising module to obtain a trained denoising module; Step S2: obtaining a real image to be denoised, inputting the real image into a trained denoising module, performing denoising processing on the real image by a blind spot branch of the trained denoising module, and outputting a denoised image; Wherein, the denoising module comprises a non-blind spot branch and a blind spot branch connected in parallel; The non-blind spot branch is a non-blind spot branch without gradient, and includes a downsampling module, a convolutional denoising network module, and an upsampling module connected in sequence. The non-blind spot branch receives an input, and the downsampling module downsamples the input in the spatial direction; the convolutional denoising network module is a U-Net network, which is used to denoise the down-sampled result to obtain the denoised feature. The upsampling module restores the denoised feature to the original resolution through a transposed convolution operation, and outputs the first denoising result of the non-blind spot branch; The downsampling module includes a convolution layer and a pooling layer connected in sequence, wherein the convolution layer with a convolution kernel size of 1×1 performs dimension increase in the channel direction on the training sample, and then the pooling layer performs a maximum pooling operation to achieve downsampling in the spatial direction; The blind spot branch includes a dense sampling block blind spot convolution module and a convolution denoising network module connected in sequence. The dense sampling block blind spot convolution module receives input and uses a blind spot convolution kernel to perform a mask operation on the input. When performing the mask operation, the masked pixel is restored based on the pixel information of the masked pixel neighborhood interval, and a mask feature is generated based on the restored pixel and the unmasked pixel. The blind spot convolution kernel is obtained by element-by-element multiplication of the convolution kernel and the noise mask matrix. The noise mask matrix is a randomly generated matrix composed of 0 and 1 with the same dimension as the convolution kernel size. The elements of the noise mask matrix with an element of 0 represent blind spots. The convolution denoising network module denoises the mask feature through a convolution operation to obtain a denoised feature as a second denoising result of the blind spot branch. The first denoising result and the second denoising result are fused as the output of the denoising module.
2. The method according to claim 1, characterized in that The formula for fusing the first denoising result and the second denoising result is: Among them, a is a hyperparameter, x b is the first denoising result, x unb is the second denoising result. When the number of training rounds epoch ≤ 8, a = 2.0; when the number of training rounds epoch > 8, 3. The method according to claim 2, characterized in that The loss function of the non-blind spot branch during training is L reg =||xb-y||1 Among them, L reg is the loss function value of the non-blind spot branch, y is the input noise image; ||·||1 is the L1 loss function.
4. The method according to claim 3, characterized in that The loss function of the blind spot branch during training is Among them, L rev is the loss function value of the blind spot branch, and ||·||2 is the mean square error loss function.
5. The method according to claim 4, characterized in that The total loss function L is:
6. The method according to claim 1, characterized in that In step S2, a real image to be denoised is obtained, and features of the real image are input into a trained denoising module, and the blind spot branch of the trained denoising module performs denoising on the real image, wherein: Input the features of the real image into the blind spot convolution module of the dense sampling block, use the blind spot convolution kernel to perform a mask operation on the input, and when performing the mask operation, restore the masked pixel based on the pixel information of the neighborhood interval of the masked pixel, and generate mask features based on the restored pixel and the unmasked pixel; wherein the blind spot convolution kernel is obtained by element-by-element multiplication of the convolution kernel and the noise mask matrix, the noise mask matrix is a randomly generated matrix composed of 0 and 1 with the same dimension as the convolution kernel size, and the elements of the noise mask matrix with the element being 0 represent blind spots; the convolution denoising network module denoises the mask features through a convolution operation to obtain denoised features as the second denoising result of the blind spot branch; The second denoising result is used as the denoised image.
7. A device for denoising a real image, characterized in that: The device comprises: Training module: configured to obtain training samples, use the training samples as input to train the denoising module, and obtain a trained denoising module; A real image denoising module: configured to obtain a real image to be denoised, input the real image into a trained denoising module, perform denoising on the real image by a blind spot branch of the trained denoising module, and output a denoised image; Wherein, the denoising module comprises a non-blind spot branch and a blind spot branch connected in parallel; The non-blind spot branch is a non-blind spot branch without gradient, and includes a downsampling module, a convolutional denoising network module, and an upsampling module connected in sequence. The non-blind spot branch receives an input, and the downsampling module downsamples the input in the spatial direction; the convolutional denoising network module is a U-Net network, which is used to denoise the down-sampled result to obtain the denoised feature. The upsampling module restores the denoised feature to the original resolution through a transposed convolution operation, and outputs the first denoising result of the non-blind spot branch; The downsampling module includes a convolution layer and a pooling layer connected in sequence, wherein the convolution layer with a convolution kernel size of 1×1 performs dimension increase in the channel direction on the training sample, and then the pooling layer performs a maximum pooling operation to achieve downsampling in the spatial direction; The blind spot branch includes a dense sampling block blind spot convolution module and a convolution denoising network module connected in sequence. The dense sampling block blind spot convolution module receives input and uses a blind spot convolution kernel to perform a mask operation on the input. When performing the mask operation, the masked pixel is restored based on the pixel information of the masked pixel neighborhood interval, and a mask feature is generated based on the restored pixel and the unmasked pixel. The blind spot convolution kernel is obtained by element-by-element multiplication of the convolution kernel and the noise mask matrix. The noise mask matrix is a randomly generated matrix composed of 0 and 1 with the same dimension as the convolution kernel size. The elements of the noise mask matrix with an element of 0 represent blind spots. The convolution denoising network module denoises the mask feature through a convolution operation to obtain a denoised feature as a second denoising result of the blind spot branch. The first denoising result and the second denoising result are fused as the output of the denoising module.
8. A computer-readable storage medium, wherein a plurality of instructions are stored in the storage medium; the plurality of instructions are used for a processor to load and execute the method as claimed in any one of claims 1 to 6.
9. An electronic device, characterized in that: The electronic device comprises: A processor, which is used to execute multiple instructions; A memory for storing a plurality of instructions; The plurality of instructions are used to be stored in the memory and loaded and executed by the processor according to any one of claims 1 to 6.
Citation Information
Patent Citations
Dynamic scene blind deblurring method based on asymmetric U-Net network
CN116188313A
Image denoising method and device for multi-scale complementary learning
CN117474797A
Self-supervised image denoising method, system and device and readable storage medium
CN117710240A
Self-supervised image denoising method based on three-stage feature extraction
CN118097159A
Cited By
Magnetic resonance diffusion weighted image denoising method and system, computer equipment and medium
CN120471800A
Self-supervised image denoising method and system for multi-angle sequence
CN120543415A