An edge extraction image denoising method based on convolutional neural network
By combining the Canny operator and convolutional neural network, image edge details are extracted and enhanced, and a fusion network of channels and spatial attention mechanisms is used to solve the problem of loss of edge details after image denoising in the prior art, and a higher quality image denoising effect is achieved.
Patent Information
- Application Number
- CN202211175527.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-26
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-09-26
AI Technical Summary
Existing image denoising methods ignore edge details of the image, resulting in a degradation in image visual quality after denoising.
The edge extraction image denoising method based on convolutional neural network is adopted, combined with the edge extraction module of the Canny operator and the convolutional neural network, through technologies such as dual threshold detection and residual learning, the edge details of the image are extracted and enhanced, and through the fusion network of the channel and the spatial attention mechanism, the greater weight is adaptively allocated to important image edge details.
While improving the denoising efficiency, it realizes that the edge details of the image are retained and enhanced, and the image visual quality after denoising is significantly improved. The PSNR is 0.13dB to 0.82dB higher than the traditional method, and the denoising speed is faster.
Smart Images

Figure CN115496687B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image processing and relates to an edge extraction image denoising method based on a convolutional neural network. Background Art
[0002] With the wide application of digital devices in the fields of military reconnaissance, security detection, medical imaging, satellite remote sensing, etc., the underlying computer vision task of image denoising has become more important. Current image denoising methods can be divided into four types: filtered image denoising, sparse representation image denoising, low-rank image denoising, and deep learning image denoising. The most classic filtered image denoising method is the block matching and 3D filtering (BM3D) method. This method combines non-local self-similarity with sparse representation in the transform domain, matches and groups images according to block similarity, and then performs frequency domain filtering denoising within each similar group. The denoising method based on sparse representation utilizes the characteristic that noise cannot be sparsely represented in the model to achieve image denoising by constraining the sparsity of natural images. Since noise will destroy the low-rank property of natural images, the low-rank based image denoising method often reduces noise by restricting the low-rank matrix.
[0003] The denoising method based on deep learning utilizes the powerful representation ability of deep networks and external information from large datasets to reconstruct feature images, greatly improving the denoising performance and efficiency. The prior art proposes a feedforward denoising convolutional neural network (DnCNN), which combines residual learning and batch normalization, and can significantly accelerate the training of the model while improving the denoising performance. Another person proposes a convolutional neural network image denoising method guided by an attention mechanism, which uses the attention mechanism to fully extract the noise information hidden in the complex background and effectively improves the processing of images with complex noise information. Although the above methods have achieved remarkable results in the denoising task, they all ignore the edge details of the image, which, as an important factor for evaluating the performance of denoising algorithms, is directly related to the visual quality of the image. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide an edge extraction image denoising method based on a convolutional neural network, which integrates an edge extraction module based on the Canny operator and a convolutional neural network in one framework to obtain a clear image with more edge details.
[0005] To achieve the above purpose, the present invention provides the following technical solutions:
[0006] An edge extraction image denoising method based on a convolutional neural network, comprising the following steps:
[0007] S1: Input the original image into an edge extraction network based on the Canny operator to extract the edge information of the image;
[0008] S2: Input the original image into a preprocessing network based on residual learning for preliminary denoising;
[0009] S3: Input the preliminarily denoised image and the edge information image into a fusion network based on channel and spatial attention mechanisms to obtain a denoised image with clear edges.
[0010] Furthermore, the edge extraction network based on the Canny operator locates the image edges through a double-threshold detection method, specifically including the following steps:
[0011] S11: Smooth the image using a Gaussian filter to remove noise; the size of the Gaussian filter kernel is (2k + 1)×(2k + 1), and the expression is as follows:
[0012]
[0013] where 1 ≤ i, j ≤ (2k + 1), (i, j) is the pixel point coordinate, and σ is the standard deviation;
[0014] S12: Calculate the gradient and gradient direction of each pixel;
[0015] S13: Use the non-maximum suppression algorithm to suppress pseudo-boundary points;
[0016] S14: Use the double-threshold detection method to determine real edges and potential edges;
[0017] S15: Complete edge extraction by suppressing isolated weak edges.
[0018] Furthermore, the preprocessing network based on residual learning contains a Conv+Relu layer, four residual units, and a convolutional layer. The outputs of three residual skip connection residual units are concatenated together through the Cat function and input into the convolutional layer. Finally, the learned noise map is subtracted through the input-end to output-end residual skip connection to obtain a preliminarily denoised clean image.
[0019] Furthermore, the fusion network based on channel and spatial attention mechanisms contains a Conv+Relu+BN layer, a Conv layer, and channel and spatial attention mechanisms. The preliminarily denoised image obtained by preprocessing and the edge information image obtained by edge extraction are aggregated through the Cat function, the number of channels is compressed through the Conv+Relu+BN convolutional layer, and the channel and spatial attention network CSAN adaptively assigns greater weights to important image edge details, thereby obtaining a denoised image with clear edges.
[0020] Furthermore, the working process of the fusion network based on channel and spatial attention mechanisms is as follows:
[0021] The upper branch is the channel attention network. Through average pooling and max pooling operations, the channels of the input feature map are compressed into two points, and these two points are aggregated with the noise level as the input;
[0022] Then, two convolutional layers are used to learn the relationships between channels. Assuming C is the number of channels, H and W are the width and height of the image respectively, then the size of the input feature map is C×H×W;
[0023] The first convolutional layer is a 3×1 convolution with an output channel of C / 16, and the second convolutional layer is a 1×1 convolution with an output channel number of C;
[0024] The lower branch is a spatial attention network. This network concatenates the maximum value, average value, and noise level map at each spatial position as the input, and uses two convolutional layers with a convolutional kernel size of 3×3 to learn the relationships between spatial positions;
[0025] Finally, each element of the input is scaled by multiplying it with the corresponding amplification factor; the process of the scaling operation is defined as the following formula:
[0026] SA k,i,j = input k,i,j ×B k ×S i,j
[0027] where k is the channel, (i, j) is the spatial position, SA is the output of the fusion network, B is the output of the channel attention network, and S is the output of the spatial attention network.
[0028] Furthermore, the number of channels of all convolutional layers in the preprocessing network and the fusion network is set to 64, the convolutional kernel size is 3×3, the stride is 1, and zero padding is performed before convolution for all network layers.
[0029] Furthermore, the L2 norm is used as the loss function to adjust the parameters of the overall network model. Let the output image estimation be f(y), the input noisy image be y, the original clean noise-free image be x, and the extracted edge information image be e. Then the loss function is as follows:
[0030]
[0031] where θ1 are the trainable parameters of the proposed denoising model, and N represents the total number of training samples.
[0032] The beneficial effects of the present invention are as follows: The present invention only needs to train a preprocessing network and a fusion network based on a convolutional neural network, and the edge extraction network does not need to be trained. This not only reduces the complexity of the model, but also greatly improves the denoising efficiency. Through experimental evaluation, the average PSNR of this method is 0.13 dB and 0.29 dB higher than that of DnCNN respectively, and 0.76 dB and 0.82 dB higher than that of BM3D respectively. At the same time, the denoising speed is also faster than that of DnCNN and BM3D.
[0033] Other advantages, objectives and features of the present invention will be described to some extent in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following specification. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be described in detail preferably with reference to the accompanying drawings, where:
[0035] Figure 1 is the flowchart of model training;
[0036] Figure 2 is the visualization result of different Gaussian filter kernels;
[0037] Figure 3 in (a) is the basic residual unit, and (b) is the residual unit of the present invention;
[0038] Figure 4 is the channel and spatial attention network;
[0039] Figure 5 is the overall network structure diagram of the present invention;
[0040] Figure 6 is the denoising result of different methods on the test045 image in the Set68 dataset when σ = 15;
[0041] Figure 7 is the denoising result of different methods on the Peppers image in the Set12 dataset when σ = 25. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0042] The following specific examples are used to illustrate the embodiments of the present invention. Those skilled in the art can easily understand the other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the drawings provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0043] Among them, the drawings are only for illustrative purposes, showing only schematic diagrams, not physical diagrams, and should not be construed as a limitation to the present invention; in order to better illustrate the embodiments of the present invention, some components in the drawings will be omitted, enlarged or reduced, which do not represent the dimensions of the actual product; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.
[0044] In the drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "rear", etc. indicating the orientation or positional relationship, they are based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the terms describing the positional relationship in the drawings are only for illustrative purposes and should not be construed as a limitation to the present invention. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.
[0045] An edge extraction image denoising method based on a convolutional neural network provided by the present invention is based on a dual-branch denoising network. The upper branch is an edge extraction network based on the Canny operator, which is used to make up for the defect that the lower branch preprocessing denoising network cannot fully extract the high-frequency details of the image and is used to enhance the image details after the initial denoising is completed. Finally, through a fusion network based on the channel and spatial attention mechanism, larger weights are adaptively assigned to the relatively more important features captured by the channel to obtain a better denoising effect.
[0046] To effectively solve the problem of edge detail loss in denoised images, the proposed edge extraction image denoising model based on a convolutional neural network is as Figure 5 shown. This method consists of three parts: a preprocessing network based on residual learning, an edge extraction network based on the Canny operator, and a fusion network.
[0047] The preprocessing network contains a Conv+Relu layer, four residual units (ResUnit), and a Conv layer. To avoid the loss of hierarchical feature information in the preprocessing module, the outputs of the three residual skip connection residual units are concatenated through the Cat function and input into the last convolutional layer. Finally, the learned noise map is subtracted through the input-to-output residual skip connection to obtain a preliminarily denoised clean image. The Gaussian filter kernel is 3×3, which can extract more image edge information. Therefore, the edge extraction network uses the Canny operator with a Gaussian filter kernel size of 3×3 to extract the edge information of the image to enhance the edge details of the image. The fusion network contains a Conv+Relu+BN layer, a Conv layer, and channel and spatial attention mechanisms. The preliminarily denoised image obtained by simply aggregating the preprocessing module and the edge information image obtained by the edge extraction module through the Cat function are compressed in the number of channels through the Conv+Relu+BN convolutional layer, and then the CSAN adaptively assigns larger weights to the important image edge details, thereby obtaining a denoised image with clear edges.
[0048] In the present invention, the Canny operator adopted by the edge extraction network locates the image edges through the double-threshold detection method, is not easily "filled" by noise, can detect real weak edges, and realizes the accurate positioning of the image edges. This method mainly includes the following five steps:
[0049] (1) Smooth the image using a Gaussian filter to filter out noise.
[0050] (2) Calculate the gradient and gradient direction of each pixel.
[0051] (3) Adopt the non-maximum suppression algorithm to suppress pseudo-boundary points.
[0052] (4) Adopt the double-threshold detection method to determine real edges and potential edges.
[0053] (5) Complete the edge extraction by suppressing isolated weak edges.
[0054] In the edge extraction network based on the Canny operator, noise will affect the accuracy of edge extraction. Therefore, the Canny algorithm uses a Gaussian filter to filter out noise, reduce the influence of noise on the edge extraction result, and prevent error detection. The size of the Gaussian filter kernel is generally designed as (2k+1)×(2k+1), and the expression is as follows:
[0055]
[0056] In the formula, 1≤i,j≤(2k+1). (i, j) is the pixel point coordinate, and σ is the standard deviation. Generally speaking, the larger the size of the Gaussian filter kernel, the lower the noise sensitivity of the Canny operator and the greater the error of edge detection.Figure 2 The size of the Gaussian filter kernel is 3×3 and 5×5, and the edge information images are extracted. It can be clearly seen that when the size of the Gaussian filter kernel is 3×3, the edge detection network retains more edge details.
[0057] Theoretically, the deeper the convolutional neural network, the better the denoising effect. However, due to the vanishing gradient caused by the increase in network depth, deepening the convolutional layer will not improve the accuracy of the image processing task, but instead lead to performance degradation. To solve this problem, through the idea of residual learning, information is directly passed from the input end to the output end. In this way, even if the network depth increases, the output end will still retain the information at the input end, thus avoiding the vanishing gradient. As Figure 3 (a) shows, when the input is x i , assuming that y is the output after passing through two layers of neural networks, this process can be described by the following formula:
[0058] x i+2 = y - x i (2)
[0059] If the x in formula (2) i is fitted to the noisy image, y is fitted to the noise mapping learned by the model, and x i+2 is the restored clear image. By applying the residual skip connection, the preprocessing network can learn the relatively less informative noise information, thus reducing the learning difficulty of the model. At the same time, based on the idea of residual learning, the present invention designs Figure 3 (b) shows a residual unit containing three convolutional layers and a residual connection to improve the training stability.
[0060] In most denoising methods based on convolutional neural networks, all channel features are processed equally without being adjusted according to their importance. However, different feature channels capture different types of information in all regions of a single noisy image, and some information is more significant than others and should be given more weights.
[0061] The present invention designs a Figure 4The shown channel and space attention network (CSAN) assigns greater weights to important feature information in the image to highlight significant useful features and suppress and ignore irrelevant features. The upper branch is the channel attention network, which compresses the channels of the input feature map into two points through average pooling and max pooling operations, and aggregates these two points and the noise level as the input. Then, two convolutional layers are used to learn the relationships between channels. Assuming C is the number of channels, and H and W are the width and height of the image respectively, the size of the input feature map is C×H×W. The first convolutional layer is a 3×1 convolution with an output channel of C / 16, and the second convolutional layer is a 1×1 convolution with an output channel number of C. The lower branch is a spatial attention network, which concatenates the maximum value, average value, and noise level map at each spatial position as the input, and uses two convolutional layers with a convolutional kernel size of 3×3 to learn the relationships between spatial positions. Finally, each element of the input is scaled by multiplying it by the corresponding scaling factor. If the outputs of the channel and spatial attention networks are processed as B and S respectively, the scaling operation process can be defined as the following formula:
[0062] SA k,i,j = input k,i,j ×B k ×S i,j
[0063] where k is the channel, (i, j) is the spatial position, and SA is the output of this attention network.
[0064] The number of channels of all convolutional layers in the preprocessing network and the fusion network is set to 64, the convolutional kernel size is 3×3, the stride is 1, and all network layers perform simple zero-padding before convolution.
[0065] The denoising method proposed by the present invention only needs to train the preprocessing network and the fusion network based on the convolutional neural network, and the edge extraction network does not need to be trained. The training process is as Figure 1 shown. To optimize the proposed denoising method and make the model reach the best convergence state, the L2 norm, which is more sensitive to outliers, is used as the loss function to adjust the model parameters. Assuming the output image estimate of the denoising method is f(y), the input noisy image is y, the original clean noise-free image is x, and the edge information image extracted is e, the loss function of the denoising module is as follows:
[0066]
[0067] where θ1 are the trainable parameters of the proposed denoising model, and N represents the total number of training samples.
[0068] The effectiveness of denoising methods is usually measured from two dimensions: objective evaluation metrics and subjective vision. Subjective vision involves directly observing the image with the human eye. In this invention, the peak signal-to-noise ratio (PSNR) and running time are used to evaluate the denoising effect of the image. The running time assesses the efficiency of the denoising method. A good denoising method needs to ensure denoising performance while minimizing the running time as much as possible. PSNR evaluates the performance of the denoising method and is defined by the mean squared error (MSE). The higher the PSNR, the better the denoising effect, indicating that the denoised image is closer to the original image. Assuming I is the input clean image and K is the predicted image of the denoising method, then PSNR can be defined by the following formula:
[0069]
[0070]
[0071] where M and N are the height and width of the image respectively, and the value of n is 8.
[0072] The proposed denoising method is compared with traditional denoising methods (BM3D, EPLL, WNNM) and the most commonly used methods based on deep learning in recent years (TNRD, DnCNN, IRCNN). Experiments are conducted on the test sets BSD68 and Set12 when the noise levels are 15, 25, and 50 respectively.
[0073] Table 1 shows the average PSNR of different methods on the BSD68 dataset. Compared with the traditional classic denoising method BM3D, the peak signal-to-noise ratio of the denoising method of this invention is about 0.8 dB higher than that of BM3D, and on average, it is about 0.3 dB higher than that of the methods based on deep learning. Tables 2 - 4 show the peak signal-to-noise ratios of different denoising methods on 12 images in the Set12 dataset. It can be seen that except for the images "Barbara" and "Man" which have many similar internal structures and are more advantageous for non-local self-similarity-based denoising methods, the denoising method of this article has the highest peak signal-to-noise ratio on the remaining images, and the average peak signal-to-noise ratio is 0.24 dB, 0.31 dB, and 0.33 dB higher than that of DnCNN respectively. In Tables 1 - 4, the highest peak signal-to-noise ratios at different noise levels are highlighted in bold.
[0074] Table 1
[0075]
[0076] Table 2
[0077]
[0078] Table 3
[0079]
[0080]
[0081] Table 4
[0082]
[0083] In terms of visual effects, to conveniently and intuitively observe the difference in the results between the denoising method of this article and other methods, Figures 6 - 7 the specific area in the image is magnified in the [document]. Figure 6 It is a comparison of the denoising results on the test045 image in the BSD68 dataset when σ = 15, Figure 7 and it is a comparison of the denoising results on the Peppers image in the Set12 dataset when σ = 25. It can be seen from the figure that the edge details of the images restored by the WNNM, BM3D, and DnCNN methods are relatively smooth. At the Figure 6 tip of the triangular part of the bottle body, the texture is relatively blurred, Figure 7 and there are a large number of artifacts in the edge of the pepper stalk, and the spots on the stalk almost disappear. In contrast, the denoising method of the present invention can restore more edge texture details of the bottle body part and the pepper stalk part without many image artifacts.
[0084] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the purpose and scope of the present technical solution, and they should all be covered within the scope of the claims of the present invention.
Claims
1. An edge extraction image denoising method based on a convolutional neural network, characterized in that: It includes the following steps: S1: Input the original image into an edge extraction network based on the Canny operator to extract the edge information of the image; S2: Input the original image into a preprocessing network based on residual learning for preliminary denoising; S3: Input the preliminarily denoised image and the edge information image into a fusion network based on channel and spatial attention mechanisms to obtain a denoised image with clear edges; The fusion network based on channel and spatial attention mechanisms includes a Conv+Relu+BN layer, a Conv layer, and channel and spatial attention mechanisms. The preliminarily denoised image obtained by preprocessing and the edge information image obtained by edge extraction are aggregated through the Cat function, the number of channels is compressed by the Conv+Relu+BN convolutional layer, and then the channel and spatial attention network CSAN adaptively assigns greater weights to important image edge details, thereby obtaining a denoised image with clear edges; The working process of the fusion network based on channel and spatial attention mechanisms is as follows: The upper branch is a channel attention network. The channels of the input feature map are compressed into two points through average pooling and max pooling operations, and these two points and the noise level are aggregated as the input; Then, two convolutional layers are used to learn the relationship between channels. Assuming C is the number of channels, H and W are the width and height of the image respectively, then the size of the input feature map is C×H×W; The first convolutional layer is a 3×1 convolution, the output channel is C / 16, and the second convolutional layer is a 1×1 convolution, and the output number of channels is C; The lower branch is a spatial attention network. The network concatenates the maximum value, average value, and noise level map at each spatial position as the input, and uses two convolutional layers with a convolutional kernel size of 3×3 to learn the relationship between spatial positions; Finally, each element of the input is scaled by multiplying by the corresponding scaling factor; the process of the scaling operation is defined as the following formula: SA k,i,j = input k,i,j × B k × S i,j where k is the channel, (i, j) is the spatial position, SA is the output of the fusion network, B is the output of the channel attention network, and S is the output of the spatial attention network.
2. The edge extraction image denoising method based on a convolutional neural network according to claim 1, characterized in that: The edge extraction network based on the Canny operator locates image edges through a double-threshold detection method, specifically including the following steps: S11: Smooth the image using a Gaussian filter to remove noise; the size of the Gaussian filter kernel is (2k+1)×(2k+1), and the expression is as follows: In the formula, 1≤i,j≤(2k+1), (i, j) is the pixel point coordinate, and σ is the standard deviation; S12: Calculate the gradient and gradient direction of each pixel; S13: Use the non-maximum suppression algorithm to suppress pseudo-edge points; S14: Use the double-threshold detection method to determine real edges and potential edges; S15: Complete edge extraction by suppressing isolated weak edges.
3. The edge extraction image denoising method based on a convolutional neural network according to claim 1, characterized in that: The preprocessing network based on residual learning includes a Conv+Relu layer, four residual units, and a convolutional layer. The outputs of three residual skip connection residual units are concatenated through the Cat function and input into the convolutional layer, and finally, the learned noise map is subtracted through the residual skip connection from the input end to the output end to obtain a preliminarily denoised clean image.
4. The edge extraction image denoising method based on a convolutional neural network according to claim 1, characterized in that: The number of channels in all convolutional layers of the preprocessing network and the fusion network is set to 64, the convolutional kernel size is 3×3, the stride is 1, and zero-padding is performed before convolution for all network layers.
5. The edge extraction image denoising method based on a convolutional neural network according to claim 1, characterized in that: The L2 norm is used as the loss function to adjust the parameters of the overall network model. Let the estimated output image be f(y), the input noisy image be y, the original clean noise-free image be x, and the extracted edge information image be e. Then the loss function is as follows: where θ1 are the trainable parameters of the proposed denoising model, and N represents the total number of training samples.