Transform blind spot network and knowledge distillation-based image denoising method

By using Transformer blind spot network and knowledge distillation method in image denoising technology, the problem of noise correlation difference and excessive calculation amount is solved, and efficient image denoising and detail utilization is achieved.

CN120219216APending Publication Date: 2025-06-27NANJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510276674.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art ignores the difference in the influence of noise correlation in the flat area and texture area of ​​the image in image denoising, and the model calculation amount is too large, making it difficult to make full use of image details.

Method used

Using image denoising method based on Transformer blind spot network and knowledge distillation, by constructing multiple blind spot network models and lightweight denoising network models, global and local features are extracted using networks of different blind spot sizes, and the model is optimized through knowledge distillation to reduce the calculation amount.

Benefits of technology

Effectively denoising images improves the visual quality and information fidelity of the image, while reducing the calculation amount of the model and improving the efficiency and effect of image denoising.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219216A_ABST
    Figure CN120219216A_ABST
Patent Text Reader

Abstract

The invention discloses an image denoising method based on a transform blind spot network and knowledge distillation. The method comprises the following steps: obtaining a training set; constructing a blind spot denoising network model based on a transform, and carrying out the training of the blind spot denoising network model; changing the size of a blind convolution kernel of the blind spot denoising network model to obtain a group of pre-trained blind spot denoising network models; respectively inputting the noise images into the group of pre-training blind spot denoising network models to obtain pseudo clean images with different weights; constructing a lightweight denoising network model based on the improved U-net, and training the lightweight denoising network model by using the noise image and the pseudo clean images with different weights; and processing a noise image by using the trained lightweight denoising network model to obtain a final denoised image. According to the method, image details are reserved, meanwhile, the complexity of the network is reduced, and the visual performance, the peak signal-to-noise ratio (PSNR) and the structural similarity (SSIM) of the image are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image information processing, and particularly relates to an image denoising method based on a transformer blind spot network and knowledge distillation. Background Art

[0002] In the process of generating, transmitting, and storing image information, due to the influence of various factors such as the environment and equipment, images are often interfered by various noises. These noises may come from channel noise during transmission, thermal noise in storage media, inherent noise of shooting equipment, etc., resulting in a decline in image quality and affecting the visual effect and information fidelity of the image. Image denoising technology has broad application prospects in the fields of medical treatment, aerospace, security monitoring, etc. In the medical field, denoised medical images can help doctors diagnose diseases more accurately; in the aerospace field, denoised remote sensing images can improve the accuracy of ground object recognition; in the security monitoring field, denoised video images can improve the reliability and stability of the monitoring system. With the rapid development and innovation of emerging information technologies such as 5G, the Internet of Things, and artificial intelligence, social informatization is rapidly developing globally. These technologies provide more powerful computing capabilities and richer data resources for image denoising, promoting the continuous progress of image denoising technology.

[0003] The core goal of image denoising technology is to recover clear and real image content from images contaminated by noise. By removing the noise in the image, the visual quality and information fidelity of the image can be significantly improved, bringing a better visual experience to people. Traditional methods usually use additive white Gaussian noise (AWGN) and perform supervised learning on a large-scale training data set by synthesizing noise-free image pairs. Since the characteristics of actual noise are very different from those of AWGN, the models learned on synthetic noise cannot generalize well in practice. To overcome this limitation, attempts have been made to construct real-world data set pairs, such as SIDD and NIND.

[0004] Using real-world training pairs, supervised denoising methods can be trained to recover clean images from noisy real-world inputs. For example, the Noise2Noise algorithm introduces a self-supervised method that does not rely on paired training data. Under the assumption that the noise signal is pixel-wise independent and has a zero mean, the blind spot network reconstructs a clean pixel from adjacent noisy pixels without referring to the corresponding input pixels, but ignores the impact of noise correlation on denoising. The LG-BPN algorithm proposes to suppress the spatial correlation of noise by masking the central region of a large convolutional kernel and proposes an extended Transformer block to extract global information. However, the introduction of large kernels will bring greater computational overhead. It can be seen that although existing methods consider eliminating the impact of noise correlation on image denoising, they ignore the difference in the impact of noise correlation in the flat and texture regions of the image. Summary of the Invention

[0005] The present invention aims to solve at least one of the technical problems existing in the related art to a certain extent.

[0006] The purpose of the present invention is to provide an image denoising method based on a transformer blind spot network and knowledge distillation, which solves the problems of denoising of related noise, insufficient utilization of original image details, and excessive computational amount of the model in the prior art.

[0007] To achieve the above object, on the one hand, the present invention provides an image denoising method based on a transformer blind spot network and knowledge distillation, including:

[0008] Obtain a training set containing only noisy images;

[0009] Construct a blind spot denoising network model based on a transformer and train the blind spot denoising network model using the noisy images in the training set;

[0010] Based on the trained blind spot denoising network model, by changing the size of the blind convolution kernel therein, a group of pre-trained blind spot denoising network models are obtained;

[0011] Input the noisy images into the group of pre-trained blind spot denoising network models respectively to obtain pseudo-clean images with different weights;

[0012] Construct a lightweight denoising network model based on an improved U-net and train the lightweight denoising network model using the noisy images and the pseudo-clean images with different weights;

[0013] Process the noisy images with the trained lightweight denoising network model to obtain the final denoised images.

[0014] A further preferred technical solution of the present invention is that the constructed blind spot denoising network model based on Transformer includes a first blind spot network and a second blind spot network;

[0015] Both the first blind spot network and the second blind spot network are composed of a first-layer Conv layer, a DSPMC layer, several LGFEB layers, and three consecutive Conv layers; the noisy image is input into the first-layer Conv layer, the output of the first-layer Conv layer is input into the DSPMC layer, the output of the DSPMC layer is input into several LGFEB layers, and the output of the LGFEB layer passes through three Conv layers in sequence to obtain the network output;

[0016] Among them, the convolution kernel size of the DSPMC layer of the first blind spot network is 9×9, the stride is (1,1), the padding size is (4,4), the padding mode is reflect, and the total number of feature maps is 128;

[0017] The convolution kernel size of the DSPMC layer of the second blind spot network is 3×3, the stride is (1,1), the padding size is (1,1), the padding mode is reflect, and the total number of feature maps is 128.

[0018] Preferably, the first-layer Conv layer of the first blind spot network is used to increase the number of channels of the input image, the convolution kernel size is 1×1, the stride is (1,1), the activation function uses the ReLU activation function, and the total number of feature maps is 24;

[0019] The DSPMC layer of the first blind spot network is used to perform masked convolution operation on the output features of the first-layer Conv layer, the convolution kernel size is 9×9, the stride is (1,1), the padding size is (4,4), the padding mode is reflect, the total number of feature maps is 128, and the masked part, that is, the part with an Euclidean distance from the central pixel not much different from 5 pixel points, is 0, and the rest is 1;

[0020] Each LGFEB layer of the first blind spot network is composed of a DC layer and an SSTB module in parallel;

[0021] The DC layer is composed of two Conv layers in series. Among them, the convolution kernel size of the first Conv layer is 3×3, the stride is (1,1), the padding size is (2,2), the dilation rate is (2,2), the activation function uses the ReLU activation function, the total number of feature maps is 128, and the convolution kernel size of the second Conv layer is 1×1, the stride is (1,1), and the total number of feature maps is 128;

[0022] The SSTB module is used to normalize the input feature map, and then process it through three parallel DC layers to obtain Q, K, and V. The DC layer consists of two Conv layers. The first Conv layer has a kernel size of 1×1, a stride of (1,1), and a total of 384 feature maps. The second Conv layer has a kernel size of 3×3, a stride of (1,1), a padding size of (3,3), a dilation rate of (3,3), a group size of 384, and a total of 384 feature maps. Then, calculate the dot product of Q and K, normalize the dot product and then take the dot product with K to obtain the attention feature map. Input the attention feature map into a Conv layer with a kernel size of 1×1, a stride of (1,1), and a total of 128 feature maps. Then add and fuse it with the input feature map to obtain the output feature map of the attention. Normalize the output feature map of the attention mechanism and input it into two identical DC layers and a Conv layer in sequence. The DC layer has a kernel size of 3×3, a stride of (1,1), a padding size of (3,3), a dilation rate of (3,3), uses the ReLU activation function, and has a total of 680 feature maps. The Conv layer has a kernel size of 1×1, a stride of (1,1), and a total of 128 feature maps. The obtained output is added and fused with the input feature map to obtain the final output feature map.

[0023] The three consecutive Conv layers of the first blind spot network are used to reduce the number of channels of the feature extraction map and fuse the extracted features. The kernel sizes of the three Conv layers are all 1×1, the strides are all (1,1), the activation functions all use the ReLU activation function, and the total number of feature maps are 128, 64, and 3 respectively.

[0024] In the second blind spot network, the DSPMC layer has a kernel size of (1,1), a stride of (1,1), a padding size of (1,1), a padding mode of reflect, and a total of 128 feature maps. The rest is the same as the first blind spot network.

[0025] Preferably, the blind spot denoising network model is trained using the noisy images in the training set. The specific method is as follows:

[0026] Take the noisy image as the input image x and input it into the first blind spot network to obtain the output image Use the loss function L flat To constrain the difference between the input image x and the output image Between them;

[0027] Then take the noisy image as the input image x and input it into the second blind spot network to obtain the output image Use the loss function L texture To constrain the difference between the input image x and the denoised image Between them;

[0028] Add the output images of the first blind spot network and the second blind spot network and to obtain a pseudo-clean image Utilize the loss function L self to constrain the difference between the input image x and the pseudo-clean image and continuously adjust the parameters of the model until the model converges, completing the training of the blind spot denoising network model.

[0029] Preferably, the loss function L of the first blind spot network flat is expressed as:[[]]

[0030]

[0031] where P represents all pixel points of an image, p represents the intensity value of the p-th pixel of the image, x p represents the intensity value of the p-th pixel of the input image, represents the intensity value of the p-th pixel of the output image passing through the first blind spot network;

[0032] The loss function L of the second blind spot network texture is expressed as:[[]]

[0033]

[0034] where represents the intensity value of the p-th pixel of the output image passing through the second blind spot network;

[0035] The loss function L of the blind spot denoising network model self is expressed as:[[]]

[0036]

[0037] where θ p is the texture degree of the p-th pixel of the image.

[0038] Preferably, the calculation formula for the texture degree θ of the p-th pixel of the image p is:[[]]

[0039]

[0040] where p is the pixel at the i-th row and j-th column of the image; S(·) represents the Sigmoid function; σ(i,j) is the standard deviation of the input image x and the output image of the first blind spot network and its calculation formula is:[[]]

[0041]

[0042] Wherein, n represents the local window size, and (a, b) are the pixels at the a-th row and the b-th column of the image.

[0043] Preferably, based on the trained blind spot denoising network model, when obtaining a group of pre-trained blind spot denoising network models by changing the size of the blind convolution kernel therein, the specific steps are determined according to the Euclidean distance n of the pixel where the blind spot is located from the central pixel, and include:

[0044] (1) When n < 3, use the trained blind spot denoising network model as the meta-teacher network, set the convolution kernel size of the DSPMC layer of the second blind spot network to (2n + 1)×(2n + 1), the stride to (1, 1), the padding size to (n, n), the padding mode to reflect, the total number of feature maps to 128, the masked part, that is, the part where the Euclidean distance of the blind spot pixel from the central pixel is not much different from n pixels, is 0, and the rest is 1; input the noisy image as the input image into the modified meta-teacher network to obtain i groups of pseudo-clean images wherein, n = 0, 1, and i = 1, 2;

[0045] (2) When n ≥ 3, use the trained blind spot denoising network model as the meta-teacher network, set the convolution kernel size of the DSPMC layer of the second blind spot network to (2n - 1)×(2n - 1), the stride to (1, 1), the padding size to (n - 1, n - 1), the padding mode to reflect, the total number of feature maps to 128, the masked part, that is, the part where the Euclidean distance of the blind spot pixel from the central pixel is not much different from n pixels, is 0, and the rest is 1; input the noisy image as the input image into the modified meta-teacher network to obtain i groups of pseudo-clean images wherein, n = 3, 5, 7, 9, 11, …, and i = 3, 4, 5, ….

[0046] Preferably, the constructed lightweight denoising network model based on the improved U-net is composed of a first-layer Conv layer, i encoders, i - 1 decoders, and two consecutive Conv layers;

[0047] The noisy image is input into the first-layer Conv layer, the output of the first-layer Conv layer is sequentially input into i encoders to extract feature maps, the extracted feature maps are sequentially input into i - 1 decoders, and finally pass through two Conv layers in sequence to obtain the final output image.

[0048] Preferably, the first-layer Conv layer of the lightweight denoising network model is used to increase the number of channels of the input image, the convolution kernel size is 3×3, the stride is (1, 1), the padding size is (2, 2), the activation function uses the ReLU activation function, and the total number of feature maps is 16;

[0049] The encoder of the lightweight denoising network model consists of a DC layer and a pooling layer. The convolutional kernel of the DC layer has a size of 3×3, a stride of (1,1), a padding size of (2,2), a dilation rate of (2,2), uses the ReLU activation function, and the total number of feature maps is 96. The window size of the pooling layer is 2, the stride is 2, and the element spacing in the window is 1.

[0050] The decoder of the lightweight denoising network model consists of a DC layer and an upsampling layer. The DC layer consists of DilatedConv and Conv. The convolutional kernel of Dilated Conv has a size of 3×3, a stride of (1,1), a padding size of (2,2), a dilation rate of (2,2), uses the ReLU activation function, and the total number of feature maps is 96. The convolutional kernel of Conv has a size of 3×3, a stride of (1,1), a padding size of (2,2), uses the ReLU activation function, and the total number of feature maps is 64. The upsampling layer specifies the output to be twice the output.

[0051] The convolutional kernels of two consecutive Conv layers in the lightweight denoising network model both have a size of 1×1, a stride of (1,1), use the ReLU activation function, and the total number of feature maps is 16 and 3 respectively.

[0052] Preferably, the lightweight denoising network model is trained with the noise image and the pseudo-clean images with different weights. The specific method is as follows:

[0053] The noise image is used as the input image x and input into the lightweight denoising network model to obtain the denoised image. Using the loss function L distill constrains the difference between the pseudo-clean image and the denoised image, and continuously adjusts the parameters of the model until the model converges, completing the training of the blind spot network.

[0054] Preferably, the loss function L of the lightweight denoising network model distill , is expressed as:

[0055]

[0056] In the formula, P represents all the pixel points of an image, p represents the intensity value of the p-th pixel of the image, represents the intensity value of the p-th pixel of the denoised image, represents the intensity value of the p-th pixel of the i-th group of pseudo-clean images obtained through the blind spot denoising network model.

[0057] Preferably, when obtaining the training set containing only noise images, preprocess the noise images by data augmentation.

[0058] On the other hand, the present invention provides a non-transitory computer-readable storage medium storing computer instructions that cause a computer to execute the above-mentioned image denoising method based on the transformer blind spot network and knowledge distillation.

[0059] On yet another aspect, the present invention provides an electronic device, comprising: a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus, and the processor calls logic instructions in the memory to execute the above-mentioned image denoising method based on the transformer blind spot network and knowledge distillation.

[0060] On still another aspect, the present invention provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer executes the above-mentioned image denoising method based on the transformer blind spot network and knowledge distillation.

[0061] Beneficial effects: The image denoising method based on the transformer blind spot network and knowledge distillation of the present invention uses a feature extraction module that combines a dilated convolutional neural network and a Transformer, enabling more sufficient extraction of global information and local features and more comprehensive utilization of features. At the same time, the present invention extracts the features of the texture region as an additional branch, and utilizes the difference in the feature extraction capabilities of the blind spot network in the texture region and the texture region to make more full use of the correlation of pixel information. The present invention uses blind spot networks with different blind spot sizes as the teacher network, and utilizes the differences in the processing of smooth regions and texture regions of different teacher networks, enabling the lightweight network to learn the features of the original image more fully and further reducing the computational amount of the model. Description of the Drawings

[0062] Figure 1 It is a processing flow block diagram of the blind spot denoising network model constructed for the present invention;

[0063] Figure 2 It is a processing flow block diagram of the first blind spot network constructed for the present invention;

[0064] Figure 3 It is a processing flow block diagram of the second blind spot network constructed for the present invention;

[0065] Figure 4 It is a demonstration diagram of the operation process of the DSPMC layer in the blind spot denoising network model of the present invention;

[0066] Figure 5 It is a processing flow block diagram of the SSTB module in the blind spot denoising network model of the present invention;

[0067] Figure 6It is the flow chart of the pseudo-clean image generation process of the present invention;

[0068] Figure 7 It is the processing flow chart of the lightweight denoising network model constructed by the present invention;

[0069] Figure 8 It is the comparison chart of the denoising effects of the method of the present invention and other methods on the SIDD dataset in Example 1;

[0070] Figure 8 In (a) is the input noisy image, Figure 8 In (b) is the denoised image of the Noise2Void algorithm, Figure 8 In (c) is the denoised image of the AP-BSN algorithm, Figure 8 In (d) is the pseudo-clean image obtained after processing by the blind spot denoising network model in Example 1, Figure 8 In (e) is the denoised image obtained after processing by the lightweight denoising network model in Example 1, Figure 8 In (f) is the real image. Detailed implementation manners

[0071] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention, and they should not be construed as limiting the present invention. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the scope of protection of the present invention. In the description of the present invention, it should be understood that the terms used are only for the purpose of description and cannot be construed as indicating or implying relative importance.

[0072] Before elaborating on the embodiments of the present invention in detail, it is necessary to explain or define some technical terms or abbreviations involved in the model constructed by the present invention.

[0073] Transformer: It is a deep learning model architecture based on the self-attention mechanism. Through the encoder-decoder structure and the multi-head attention mechanism, it can efficiently capture the global dependencies in the input sequence.

[0074] Blind spot denoising network: It is an algorithm model based on a neural network, aiming to remove noise from a noise-contaminated image and restore the original clean image. It usually adopts an end-to-end architecture, which can directly take the noisy image as the input and output the denoised image after a series of processes. The "blind spot" of this network is an actively designed computational blind area, which realizes the ignoring of the central pixel by modifying the convolution kernel structure.

[0075] Blind convolution kernel: In convolution operations, it refers to a convolution kernel whose specific values or parameters are unknown and need to be estimated or learned through certain methods or algorithms.

[0076] Knowledge distillation: It is a technical method that transfers the knowledge learned by a complex and high-performance teacher model to a relatively simple student model. Its purpose is to enable the student model to mimic the behavior and performance of the teacher model as much as possible while maintaining a small model size and computational cost, thus achieving a good balance between efficiency and effectiveness.

[0077] Teacher Model: In deep learning, it is a relatively complex and high-performance model that usually has more parameters, a deeper network structure, or stronger feature extraction capabilities. It is used as a knowledge provider to transfer the knowledge it has learned to other models. For example, in knowledge distillation, it transfers knowledge to the student model to help the student model improve its performance and generalization ability.

[0078] Student Model: In deep learning methods such as knowledge distillation, it is a model used to receive and learn the knowledge transferred by the teacher model. It is usually a relatively simple model with a small number of parameters and low computational complexity. Its purpose is to approach the performance of the teacher model as much as possible while minimizing the consumption of model resources, thereby enabling efficient model deployment and application.

[0079] Conv layer (convolution layer) is the core component in a convolutional neural network (CNN). It is mainly used to extract features from input data. It extracts local features of the input data through a convolution kernel to generate a feature map, and significantly reduces the number of parameters through parameter sharing and local receptive fields, making it suitable for processing high-dimensional data.

[0080] DSPMC (Densely-Sampled Patch-Masked Convolution), densely sampled patch-masked convolution, is a convolution method that combines masking technology during convolution operations on densely sampled patch data. It obtains a large number of patches by densely sampling the input data, and then applies convolution operations with masking to each patch, enabling the model to capture local features of the data more finely. At the same time, it uses the masking mechanism to control the flow and utilization of information, thus achieving better results in some tasks (such as image generation, video processing, natural language processing, etc.). For example, in image inpainting tasks, it can better utilize surrounding known information to repair damaged areas.

[0081] LGFEB (Local and Global Feature Extractor Block), the local and global feature extraction block, is a core module in deep learning for fusing local details and global context information. Its design goal is to enhance the model's comprehensive understanding ability of data by integrating feature representations at different scales.

[0082] The DC layer (Dilated Conv layer), the dilated convolutional layer, is a core module in deep learning for increasing the receptive field of the convolutional layer. It efficiently captures multi-scale context information by introducing holes (dilation) in the traditional convolutional kernel.

[0083] SSTB (Self-Similarity Transformer Block), the self-similarity Transformer block, is a deep learning module that combines the self-similarity mechanism and the Transformer architecture. It is mainly used to capture long-range dependencies in data (such as images, sequences) and enhance feature representations through self-similarity.

[0084] The encoder is a core module in a deep learning model responsible for converting input data into a compact and abstract feature representation. Its core goal is to extract the key information of the data through hierarchical transformations and provide effective input for subsequent tasks (such as classification, generation, inference).

[0085] The decoder is a core module in a deep learning model responsible for converting the abstract feature representation into an understandable output (such as raw data, natural language, image). Its core goal is to restore or generate data through hierarchical transformations and is often used in pairs with the encoder.

[0086] The following combines Figures 1 - 8 Describe the image denoising method, non-transitory computer-readable storage medium, electronic device, and computer program product provided by the present invention based on the transformer blind spot network and knowledge distillation.

[0087] Embodiment 1: This embodiment provides an image denoising method based on a transformer blind spot network and knowledge distillation. The main steps of this method include: obtaining a training set containing only noisy images; constructing a blind spot denoising network model based on a transformer, and training the blind spot denoising network model with the noisy images in the training set; based on the trained blind spot denoising network model, by changing the size of the blind convolution kernel therein, obtaining a group of pre-trained blind spot denoising network models; inputting the noisy images into this group of pre-trained blind spot denoising network models respectively to obtain pseudo-clean images with different weights; constructing a lightweight denoising network model based on an improved U-net, and training the lightweight denoising network model with the noisy images and the pseudo-clean images with different weights; using the trained lightweight denoising network model to process the noisy images to obtain the final denoised images.

[0088] In this embodiment, the SIDD dataset is used. There are a total of 640 noisy images, among which 608 noisy images are used as the training set and 32 noisy images are used as the validation set. Since the SIDD dataset collects pictures in real scenes and is not synthetic, it can better verify the effectiveness of the present invention.

[0089] As Figure 1 shown, first, the noisy images are randomly rotated for data augmentation, and the processed noisy images are used as the input image x and input into the first blind spot network to obtain the output image Calculate the output image of the first blind spot network The standard deviation σ(i,j) and texture degree θ of the input image x p and the loss function L flat ; then, the noisy images are used as the input image x and input into the second blind spot network to obtain the output image Calculate the loss function L between the input image x and the denoised image ; add the output images texture of the first blind spot network and the second blind spot network and to obtain the pseudo-clean image Use the loss function L self to constrain the difference between the input image x and the pseudo-clean image , and continuously adjust the parameters of the model until the model converges, thus completing the training of the blind spot denoising network model.

[0090] As Figure 2 shown, the first blind spot network consists of a first-layer Conv layer, a DSPMC layer, 5 LGFEB layers, and three consecutive Conv layers.

[0091] Among them, the first-layer Conv layer is used to increase the number of channels of the input image. The convolution kernel size is 1×1, the stride is (1,1), the ReLU activation function is adopted, and the total number of feature maps is 24.

[0092] As Figure 4 shown, the DSPMC layer is used to perform masked convolution operations on the output features of the first-layer Conv layer. The convolution kernel size is 9×9, the stride is (1,1), the padding size is (4,4), the padding mode is reflect, the total number of feature maps is 128. For the masked part, that is, the part with an Euclidean distance from the central pixel not more than 5 pixel points is 0, and the rest is 1.

[0093] The LGFEB layer is composed of a DC layer and an SSTB module in parallel. The DC layer is composed of two Conv layers in series. Among them, the convolution kernel size of the first Conv layer is 3×3, the stride is (1,1), the padding size is (2,2), the dilation rate is (2,2), the ReLU activation function is adopted, and the total number of feature maps is 128. The convolution kernel size of the second Conv layer is 1×1, the stride is (1,1), and the total number of feature maps is 128.

[0094] As Figure 5 shown, the SSTB module is used to normalize the input feature map, and then process it through three parallel DC layers to obtain Q, K, and V. Among them, the DC layer is composed of two Conv layers. The convolution kernel size of the first Conv layer is 1×1, the stride is (1,1), and the total number of feature maps is 384. The convolution kernel size of the second Conv layer is 3×3, the stride is (1,1), the padding size is (3,3), the dilation rate is (3,3), the group size is 384, and the total number of feature maps is 384. Then calculate the dot product of Q and K, normalize the dot product and then dot it with K to obtain the attention feature map. Input the attention feature map into a Conv layer with a convolution kernel size of 1×1, a stride of (1,1), and a total number of feature maps of 128. Then add and fuse it with the input feature map to obtain the output feature map of the attention. Normalize the output feature map of the attention mechanism and input it into two identical DC layers and a Conv layer in sequence. Among them, the convolution kernel size of the DC layer is 3×3, the stride is (1,1), the padding size is (3,3), the dilation rate is (3,3), the ReLU activation function is adopted, and the total number of feature maps is 680. The convolution kernel size of the Conv layer is 1×1, the stride is (1,1), and the total number of feature maps is 128. The obtained output is added and fused with the input feature map to obtain the final output feature map.

[0095] Three consecutive Conv layers are used to reduce the number of channels of the feature extraction map and fuse the extracted features. The convolution kernel sizes of the three Conv layers are all 1×1, the strides are all (1,1), and the ReLU activation function is used for all activation functions. The total number of feature maps is 128, 64, and 3 respectively.

[0096] The calculation formula for the standard deviation σ(i,j) of the first blind spot network is:

[0097]

[0098] In the formula, n represents the local window size, (i,j) represents the pixel point at the i-th row and j-th column, and (a,b) is the pixel at the a-th row and b-th column of the image.

[0099] The calculation formula for the texture degree θ(i,j) of the image is:

[0100]

[0101] In the formula, S(·) represents the Sigmoid function.

[0102] The loss function L in the training process of the first blind spot network flat is:

[0103]

[0104] In the formula, P represents all pixel points of an image, p represents the intensity value of the pixel at point p of the image, x p represents the intensity value of the pixel at point p of the input image, represents the intensity value of the pixel at point p of the output image after passing through the first blind spot network.

[0105] As Figure 3 shown, in the second blind spot network, the convolution kernel size of the DSPMC layer is (1,1), the stride is (1,1), the padding size is (1,1), the padding mode is reflect, the total number of feature maps is 128, and the rest is the same as the first blind spot network.

[0106] The loss function L in the training process of the second blind spot network texture is:

[0107]

[0108] In the formula, represents the intensity value of the pixel at point p of the output image after passing through the second blind spot network.

[0109] The loss function L of the blind spot denoising network model self , is expressed as:

[0110]

[0111] where θ p is the texture degree of the pixel at point p in the image.

[0112] As Figure 6 shown, the generation process of the pseudo-clean image in this embodiment is determined according to the Euclidean distance n between the pixel where the blind spot is located and the central pixel, and is divided into two cases:

[0113] (1) When n < 3, the trained blind-spot denoising network model is used as the meta-teacher network. Set the convolution kernel size of the DSPMC layer of the second blind-spot network to (2n + 1)×(2n + 1), the stride to (1,1), the padding size to (n,n), the padding mode to reflect, the total number of feature maps to 128. The masked part, that is, the part where the Euclidean distance between the blind-spot pixel and the central pixel is approximately the same as n pixels, is 0, and the rest is 1. The noisy image is used as the input image and input into the modified meta-teacher network to obtain i groups of pseudo-clean images where n = 0, 1, and i = 1, 2;

[0114] (2) When n ≥ 3, the trained blind-spot denoising network model is used as the meta-teacher network. Set the convolution kernel size of the DSPMC layer of the second blind-spot network to (2n - 1)×(2n - 1), the stride to (1,1), the padding size to (n - 1,n - 1), the padding mode to reflect, the total number of feature maps to 128. The masked part, that is, the part where the Euclidean distance between the blind-spot pixel and the central pixel is approximately the same as n pixels, is 0, and the rest is 1. The noisy image is used as the input image and input into the modified meta-teacher network to obtain i groups of pseudo-clean images where n = 3, 5, 7, 9, 11, and i = 3, 4, 5, 6, 7.

[0115] As Figure 7 shown, the knowledge distillation model is constructed in this embodiment, that is, a lightweight denoising network model based on the improved U-net, which consists of the first-layer Conv layer, 6 encoders, 5 decoders, and two consecutive Conv layers.

[0116] The first-layer Conv layer is used to increase the number of channels of the input image. The convolution kernel size is 3×3, the stride is (1,1), the padding size is (2,2), the activation function uses the ReLU activation function, and the total number of feature maps is 16;

[0117] The encoder consists of a DC layer and a pooling layer. The convolution kernel size of the DC layer is 3×3, the stride is (1,1), the padding size is (2,2), the dilation rate is (2,2), the activation function uses the ReLU activation function, and the total number of feature maps is 96; the window size of the pooling layer is 2, the stride is 2, and the element spacing in the window is 1;

[0118] The decoder consists of a DC layer and an upsampling layer. The DC layer consists of Dilated Conv and Conv. The convolution kernel of Dilated Conv has a size of 3×3, a stride of (1,1), a padding size of (2,2), a dilation rate of (2,2), and the activation function is the ReLU activation function. The total number of feature maps is 96. The convolution kernel of Conv has a size of 3×3, a stride of (1,1), a padding size of (2,2), and the activation function is the ReLU activation function. The total number of feature maps is 64. The upsampling layer specifies the output to be twice the output.

[0119] For two consecutive Conv layers, the convolution kernel sizes are both 1×1, the strides are both (1,1), the activation functions are both the ReLU activation function, and the total numbers of feature maps are 16 and 3 respectively.

[0120] The training process of the lightweight denoising network model is as follows: The noisy image is used as the input image x and input into the lightweight denoising network model to obtain the denoised image Using the loss function L distill constrains the difference between the pseudo-clean image and the denoised image, and continuously adjusts the parameters of the model until the model converges, completing the training of the blind spot network.

[0121] The loss function L of the lightweight denoising network model distill , is expressed as:

[0122]

[0123] In the formula, P represents all the pixel points of an image, p represents the intensity value of the pixel at point p in the image, represents the intensity value of the pixel at point p in the denoised image, represents the intensity value of the pixel at point p in the i-th group of pseudo-clean images obtained through the blind spot denoising network model.

[0124] In this embodiment, PSNR and SSIM are used as evaluation metrics. The method of this embodiment is compared with other existing methods, where our is the method for obtaining the pseudo-clean image after being processed by the blind spot denoising network model in this embodiment; our(D) is the method for obtaining the denoised image after being processed by the lightweight denoising network model in this embodiment. The comparison results are shown in Table 1, Table 2, and Figure 8 as shown.

[0125] Table 1 Quantitative test comparison on the SIDD dataset

[0126] Method PSNR SSIM Noise2Void 27.68 0.668 AP - BSN 35.97 0.925 LG - BPN 37.28 0.936 Spatially - Adaptive (U - Net) 37.41 0.934 Our 36.54 0.927 Our (D) 37.52 0.940

[0127] As can be seen from Table 1, the results of the method of this embodiment are significantly better than other methods in terms of PSNR and SSIM.

[0128] Parameter Quantity Comparison of Table 2 on SIDD Dataset

[0129] Method Number of parameters (M) AP - BSN 3.66 SDAP (E) 3.66 LGBPN 4.56 Our 4.21 Our (D) 1.20

[0130] As can be seen from Table 2, the parameter quantity of the lightweight denoising network model of this embodiment is much lower than other methods.

[0131] As Figure 8 shown, it is a comparison of an image of the method of this embodiment and other methods on the SIDD dataset. Figure 8 Among them, (a) is the input noisy image. Figure 8 Among them, (b) is the denoised image of the Noise2Void algorithm. Figure 8 Among them, (c) is the denoised image of the AP-BSN algorithm. Figure 8 Among them, (d) is the pseudo-clean image obtained after processing by the blind spot denoising network model in this embodiment. Figure 8 Among them, (e) is the denoised image obtained after processing by the lightweight denoising network model in this embodiment. Figure 8 Among them, (f) is the real image. It can be seen that the blind spot denoising network model and the lightweight denoising network model of this embodiment are superior to other methods in terms of color and details, etc., and are closer to the real image.

[0132] The method of this embodiment retains the image details, reduces the complexity of the network, and improves the visual performance of the image as well as the peak signal-to-noise ratio (PSNR) and structural similarity (SSIM).

[0133] Embodiment 2: This embodiment provides a non-transitory computer-readable storage medium, on which computer instructions are stored, and the computer instructions cause the computer to execute an image denoising method based on a transformer blind spot network and knowledge distillation. The method includes the following steps:

[0134] Obtain a training set containing only noisy images;

[0135] Construct a blind spot denoising network model based on a transformer, and use the noisy images in the training set to train the blind spot denoising network model;

[0136] Based on the trained blind spot denoising network model, by changing the size of the blind convolution kernel therein, a group of pre-trained blind spot denoising network models are obtained;

[0137] Input the noisy images into the group of pre-trained blind spot denoising network models respectively to obtain pseudo-clean images with different weights;

[0138] Construct a lightweight denoising network model based on the improved U-net, and train the lightweight denoising network model with the noisy image and the pseudo-clean images with different weights;

[0139] Process the noisy image with the trained lightweight denoising network model to obtain the final denoised image.

[0140] Embodiment 3: This embodiment provides an electronic device, which may include: a processor, a communications interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus. The processor can call the logical instructions in the memory to execute an image denoising method based on the transformer blind spot network and knowledge distillation. The method includes the following steps:

[0141] Obtain a training set containing only noisy images;

[0142] Construct a blind spot denoising network model based on the transformer, and train the blind spot denoising network model with the noisy images in the training set;

[0143] Based on the trained blind spot denoising network model, obtain a group of pre-trained blind spot denoising network models by changing the size of the blind convolution kernel therein;

[0144] Input the noisy image into the group of pre-trained blind spot denoising network models respectively to obtain pseudo-clean images with different weights;

[0145] Construct a lightweight denoising network model based on the improved U-net, and train the lightweight denoising network model with the noisy image and the pseudo-clean images with different weights;

[0146] Process the noisy image with the trained lightweight denoising network model to obtain the final denoised image.

[0147] In addition, when the logical instructions in the above-mentioned memory are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0148] Embodiment 4: This embodiment provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute an image denoising method based on a transformer blind spot network and knowledge distillation. The method includes the following steps:

[0149] Obtain a training set that only contains noisy images;

[0150] Construct a blind spot denoising network model based on a transformer, and use the noisy images in the training set to train the blind spot denoising network model;

[0151] Based on the trained blind spot denoising network model, by changing the size of the blind convolution kernel therein, obtain a set of pre-trained blind spot denoising network models;

[0152] Input the noisy images into this set of pre-trained blind spot denoising network models respectively to obtain pseudo-clean images with different weights;

[0153] Construct a lightweight denoising network model based on an improved U-net, and use the noisy images and pseudo-clean images with different weights to train the lightweight denoising network model;

[0154] Use the trained lightweight denoising network model to process the noisy images to obtain the final denoised images.

[0155] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative work.

[0156] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0157] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An image denoising method based on transformer blind spot network and knowledge distillation, characterized in that: include: Get a training set containing only noise images; Construct a transformer-based blind spot denoising network model and train it using noisy images in the training set; Based on the trained blind spot denoising network model, a set of pre-trained blind spot denoising network models are obtained by changing the size of the blind convolution kernel; The noisy images are respectively input into the group of pre-trained blind spot denoising network models to obtain pseudo clean images with different weights; Construct a lightweight denoising network model based on the improved U-net, and train the lightweight denoising network model with noisy images and pseudo-clean images with different weights; The noisy image is processed with the trained lightweight denoising network model to obtain the final denoised image.

2. The image denoising method based on transformer blind spot network and knowledge distillation according to claim 1, characterized in that: The transformer-based blind spot denoising network model constructed includes the first blind spot network and the second blind spot network; The first blind spot network and the second blind spot network are both composed of a first Conv layer, a DSPMC layer, several LGFEB layers, and three consecutive Conv layers; the noise image is input to the first Conv layer, the output of the first Conv layer is input to the DSPMC layer, the output of the DSPMC layer is input to several LGFEB layers, and the output of the LGFEB layer passes through three Conv layers in sequence to obtain the network output; Among them, the convolution kernel size of the DSPMC layer of the first blind spot network is 9×9, the stride is (1,1), the padding size is (4,4), the padding mode is reflect, and the total number of feature maps is 128; The DSPMC layer of the second blind spot network has a convolution kernel size of 3×3, a stride of (1,1), a padding size of (1,1), a padding mode of reflect, and a total number of 128 feature maps.

3. The image denoising method based on transformer blind spot network and knowledge distillation according to claim 2, characterized in that: The blind spot denoising network model is trained using the noisy images in the training set, and the specific method is as follows: The noise image is input as the input image x to the first blind spot network, and the output image is obtained. Using the loss function L f l at Constrain the input image x and the output image The difference between Then the noise image is input as the input image x to the second blind spot network to obtain the output image Using the loss function L te xt ure Constrain the input image x and the denoised image The difference between The output images of the first blind spot network and the second blind spot network are and Add together to get a pseudo-clean image Using the loss function L self Constrain the input image x and the pseudo-clean image The difference between the two is calculated, and the parameters of the model are continuously adjusted until the model converges, completing the training of the blind spot denoising network model.

4. The image denoising method based on transformer blind spot network and knowledge distillation according to claim 3, characterized in that: The loss function L of the first blind spot network f l at , expressed as: In the formula, P represents all pixels of an image, p represents the intensity value of the pixel at point p in the image, and x p represents the intensity value of the pixel at point p of the input image, represents the intensity value of the pixel at point p of the output image after the first blind spot network; The loss function L of the second blind spot network te xt ure , expressed as: In the formula, represents the intensity value of the pixel at point p of the output image after the second blind spot network; Loss function L of the blind spot denoising network model self , expressed as: In the formula, θ p is the texture degree of pixel p in the image.

5. The image denoising method based on transformer blind spot network and knowledge distillation according to claim 4, characterized in that: The texture degree θ of the pixel at point p in the image p The calculation formula is: Where p is the pixel in the i-th row and j-th column of the image; S(·) represents the Sigmoid function; σ(i,j) is the input image x and the output image of the first blind spot network The standard deviation of is calculated as: Where n represents the local window size, and (a, b) is the pixel in the ath row and bth column of the image.

6. The image denoising method based on transformer blind spot network and knowledge distillation according to claim 3, characterized in that: Based on the trained blind spot denoising network model, by changing the size of the blind convolution kernel, a set of pre-trained blind spot denoising network models is obtained. The specific steps are determined according to the Euclidean distance n between the pixel where the blind spot is located and the center pixel, including: (1) When n<3, the trained blind spot denoising network model is used as the meta-teacher network, and the convolution kernel size of the DSPMC layer of the second blind spot network is set to (2n+1)×(2n+1), the step size is (1,1), the padding size is (n,n), the padding mode is reflect, the total number of feature maps is 128, the mask part, that is, the part where the Euclidean distance between the blind spot pixel and the center pixel is not more than n pixels, is 0, and the rest is 1; the noisy image is input as the input image to the modified meta-teacher network to obtain i groups of pseudo-clean images Where n = 0, 1, i = 1, 2; (2) When n ≥ 3, the trained blind spot denoising network model is used as the meta-teacher network, and the convolution kernel size of the DSPMC layer of the second blind spot network is set to (2n-1) × (2n-1), the step size is (1, 1), the padding size is (n-1, n-1), the padding mode is reflect, the total number of feature maps is 128, the mask part, that is, the part where the Euclidean distance between the blind spot pixel and the center pixel is not more than n pixels, is 0, and the rest is 1; the noisy image is input as the input image to the modified meta-teacher network to obtain i groups of pseudo-clean images. Among them, ==3,5,7,9,11,…, i=3,4,5,….

7. The image denoising method based on transformer blind spot network and knowledge distillation according to claim 1, characterized in that: The constructed lightweight denoising network model based on the improved U-net consists of the first Conv layer, i encoders, i-1 decoders and two consecutive Conv layers; The noise image is input to the first Conv layer, and the output of the first Conv layer is input to i encoders in turn to extract feature maps. The extracted feature maps are input to i-1 decoders in turn, and finally pass through two Conv layers in turn to obtain the final output image.

8. The image denoising method based on transformer blind spot network and knowledge distillation according to claim 7, characterized in that: The lightweight denoising network model is trained using noisy images and pseudo-clean images with different weights. The specific method is as follows: The noisy image is input as the input image x to the lightweight denoising network model to obtain the denoised image Using the loss function L dis t i l Constrain the difference between the pseudo-clean image and the denoised image, and continuously adjust the model parameters until the model converges to complete the training of the blind spot network.

9. The image denoising method based on transformer blind spot network and knowledge distillation according to claim 8, characterized in that: The loss function L of the lightweight denoising network model dis t i l, expressed as: In the formula, P represents all pixels of an image, p represents the intensity value of the pixel at point p of the image, represents the intensity value of the pixel at point p of the denoised image, Represents the intensity value of the pixel at point p of the i-th group of pseudo-clean images obtained through the blind spot denoising network model.

10. The image denoising method based on transformer blind spot network and knowledge distillation according to claim 8, characterized in that: When obtaining a training set containing only noise images, the noise images are preprocessed by data augmentation.

Citation Information

Cited By

  • Third harmonic microscopic imaging device with three-dimensional rapid intelligent noise reduction function

    CN121559729A