Apparatus and method for automatic convolutional coding for noise reduction

DE602022015949T2Active Publication Date: 2025-06-18ACER INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE602022015949
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-01-24
Publication Date
2025-06-18
Estimated Expiration
2042-01-24

AI Technical Summary

Technical Problem

Existing image denoising technologies require multiple training iterations for each type of image distortion, making them inefficient for handling diverse and random types of image noise and blur.

Method used

A noise reduction convolutional auto-encoding device and method that uses a single model architecture with multi-stride encoding and decoding convolutional layers, skip-connections, and a combined loss function to adapt to different types of image distortions, allowing for efficient denoising and deblurring of various distortion types with a single training iteration.

Benefits of technology

The proposed solution enables simultaneous denoising and deblurring of multiple types of distorted images, achieving performance comparable to specialized technologies for Gaussian noise, while also effectively reducing speckle noise, salt and pepper noise, Gaussian blur, and motion blur, with improved accuracy in image recognition tasks.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUND OF THE INVENTION Field of the Invention

[0001] The present invention relates to an image noise reduction device, and in particular to a noise reduction convolutional auto-encoding device and a noise reduction convolutional auto-encoding method.Description of the Related Art

[0002] Due to the vigorous development of deep learning, more and more documents use deep learning methods for image denoising. In the early days, Viren Jain and H. Sebastian Seung were pioneers in proposing the use of convolutional neural networks as image denoising architecture, and provided performance comparable to traditional image processing methods, but the structure was too simple. More recent technologies include Residual Encoder-Decoder Networks ("RED-Net"). However, RED-Net also has the disadvantage that it needs to use a very deep network to train to improve the denoising performance, resulting in the need for a lot of calculations.

[0003] The paper "A Perceptually Inspired New Blind Image Denoising Method Using L1 and Perceptual Loss" by A. F. M. Shahab Uddin et al., published in IEEE Access vol. 7, July 24, 2019, on pages 90538-90549, relates to discriminative learning-based denoising methods which have received much attention and have been studied to a large extent because of their high denoising performance with significantly shorter inference time compared to model based denoising methods. In this paper, the authors consider a perceptually motivated blind image denoising problem, which removes various levels of noise from an observed noisy image with a single model and produces visually pleasant images. It is well known that very few blind image denoising methods are available and they have shown the limited quality of restored images because they sacrifice the fine image details during the denoising process due to over-smoothing, resulting in visually unpleasant images. To overcome these problems, there is proposed a novel loss function that encourages the network to restore noise free images by focusing on the perceived visual quality. Then, the proposed loss function is adopted to an encoder-decoder network with skip connections that produces visually pleasant images with highly preserved fine image details. Further, this method is robust and constantly performs well for several unknown noise levels. The authors' extensive experimental results show the superiority of the proposed method both quantitatively and qualitatively.

[0004] WO 2020 / 162196 A1 discloses an imaging device with is provided with: a semiconductor substrate which has a first surface and a second surface opposite to each other and to which a plurality of pixels are provided; a wiring layer which is provided to the second surface side of the semiconductor substrate and to which a signal to each of the pixels is transmitted; a light shielding film which is disposed opposite to the wiring layer with the semiconductor substrate therebetween and which has openings satisfying expression B<A; and waveguides that are provided to the first surface side of the semiconductor substrate for the pixels and are directed to the respective openings of the light shielding film.

[0005] The paper "Connecting Image Denoising and High-Level Vision Tasks via Deep Learning" published on website in ARVIX.ORG, Cornell University Library, Ithaca, NY 14853, on Sep. 6, 2018, relates to image denoising and high-level vision tasks which are usually handled independently in the conventional practice of computer vision, and their connection is fragile. In this paper, the authors cope with the two jointly and explore the mutual influence between them with the focus on two questions, namely (1) how image denoising can help improving high-level vision tasks, and (2) how the semantic information from high-level vision tasks can be used to guide image denoising. First for image denoising there is proposed a convolutional neural network in which convolutions are conducted in various spatial resolutions via downsampling and upsampling operations in order to fuse and exploit contextual information on different scales. Second there is proposed a deep neural network solution that cascades two modules for image denoising and various high-level tasks, respectively, and use the joint loss for updating only the denoising network via back-propagation. The authors experimentally show that on one hand, the proposed denoiser has the generality to overcome the performance degradation of different high-level vision tasks. On the other hand, with the guidance of high-level vision information, the denoising network produces more visually appealing results. Extensive experiments demonstrate the benefit of exploiting image semantics simultaneously for image denoising and high-level vision tasks via deep learning. In addition, most methods that use deep learning methods to remove noise or deblurring need to modify or retrain the model architecture according to different types of distortion, and after the model is trained, only a single type of distortion can be processed. However, in the real world, it is impossible to determine what kind of distortion will appear in the input picture, and the types of distortion are random and diverse.

[0006] Therefore, how to construct and train a single denoising convolutional auto-encoder model architecture has become one of the problems to be solved in the field, wherein the model only needs a few training iterations before it can remove multiple types of image distortion and achieve the same level as the current technology-not just for a single type of noise or blurry distortion.BRIEF SUMMARY OF THE INVENTION

[0007] The present invention provides a noise reduction convolutional auto-encoding device. The noise reduction convolutional auto-encoding device includes a processor and a storage device. The processor is configured to access a noise reduction convolutional auto-encoding model stored in the storage device, in order to execute the noise reduction convolutional auto-encoding model. The processor executes the following operations: A distorted image is received and then input into the noise reduction convolutional auto-encoding model. In the noise reduction convolutional auto-encoding model, an image feature of the distorted image is transferred to a first deconvolution layer through skip-connection. The invention is characterized in that the processor further executes the following operations: A plurality of multi-stride encoding convolutional layers are performed for the distortion image to reduce a dimension, and then a same-dimensional encoding convolutional layer is performed. according to the corresponding multi-stride encoding convolutional layers and same-dimensional encoding convolutional layers, the corresponding a plurality of decoding multi-stride convolutional layers and a same-dimensional decoding convolutional layers are upgraded. The first deconvolution layer inputs the result of up-scaled dimension to the same-dimensional decoding convolutional layer of a balanced channel. A reconstructed image is output by the same-dimensional decoding convolutional layer of the balanced channel.

[0008] The present invention also provides a noise reduction convolutional auto-encoding method. The noise reduction convolutional auto-encoding method includes the following operations. A distorted image is received and input into a noise reduction convolutional auto-encoding model. In the noise reduction convolutional auto-encoding model, an image feature of the distorted image is transferred to a first deconvolution layer through skip-connection. The method is characterized by further operations as follows: A plurality of multi-stride encoding convolutional layers are performed for the distortion image to reduce a dimension. A same-dimensional encoding convolutional layer is performed. According to the corresponding multi-stride encoding convolutional layers and same-dimensional encoding convolutional layers, the corresponding a plurality of decoding multi-stride convolutional layers and a same-dimensional decoding convolutional layers are upgraded. The first deconvolution layer inputs the result of up-scaled dimension to the same-dimensional decoding convolutional layer of a balanced channel. A reconstructed image is output by the same-dimensional decoding convolutional layer of the balanced channel.

[0009] The noise reduction convolutional auto-encoding device and the noise reduction convolutional auto-encoding method described in the present invention can simultaneously denoise or eliminate blurring of different types of distorted images, and the noise reduction convolutional auto-encoding device has a simple structure, only needing to be trained once. By using skip-connection to transfer image features to the deconvolution layer, it aids in deconvolution to restore a better, clearer image. Using multi-stride encoding convolutional layers and multi-stride decoding convolutional layers to increase and decrease the dimensionality of the image can achieve better denoising performance. Moreover, adding the same-dimensional encoding convolutional layer and same-dimensional decoding convolutional layer after multi-stride encoding convolutional layers and multi-stride decoding convolutional layers can prevent the blocking effect. The blocking effect is caused by multi-stride encoding convolutional layers and multi-stride decoding convolutional layers. The blocking effect can affect the restored feature image. Then, performing the feature extraction again can reduce the noise transmitted by skip-connection.

[0010] In addition, the noise reduction convolutional auto-encoding device and the noise reduction convolutional auto-encoding method design a loss function by combining the mean square error and structural similarity index measurement to update the weights. When finally performing collaborative training on pictures with different distortion types, the weights can adapt to different distortions. The foregoing experimental results show that the noise reduction convolutional auto-encoding device used for noise reduction not only achieves the same results as the known technology in reducing Gaussian noise, but can also reduce speckle noise and salt and pepper noise at the same time. Moreover, the noise reduction convolutional auto-encoding device and noise reduction convolutional auto-encoding method have excellent performance. The noise reduction convolutional auto-encoding device is used to eliminate blur. It can eliminate Gaussian blur and motion blur at the same time with the same performance as other documents.

[0011] In addition, though the noise reduction convolutional auto-encoding device and the noise reduction convolutional auto-encoding method combine a noise reduction convolutional auto-encoding model with a deep convolutional neural network (VGG-16), the application of the noise reduction convolutional auto-encoding device can significantly improve the accuracy of the convolutional network's recognition of distorted images, and prove that a single noise reduction convolutional auto-encoding model only needs to be trained once to be able to reconstruct images from a variety of different types of distortion.BRIEF DESCRIPTION OF THE DRAWINGS

[0012] FIG. 1 is a block diagram of a noise reduction convolutional auto-encoding device in accordance with one embodiment of the present disclosure; FIG. 2 is a flowchart of a noise reduction convolutional auto-encoding method in accordance with one embodiment of the present disclosure; FIG. 3 is a schematic diagram illustrating a noise reduction convolutional auto-encoding method in accordance with one embodiment of the present disclosure; FIGS. 4A-4E are schematic diagrams illustrating different distortion types in accordance with one embodiment of the present disclosure; FIG. 5A is a schematic diagram of the noise reduction convolutional auto-encoding model and VGG-16 in the training stage according to an embodiment of the present invention; FIG. 5B is a schematic diagram of the noise reduction convolutional auto-encoding model and VGG-16 in the testing stage in accordance with one embodiment of the present disclosure; FIGS. 6A-6E are schematic diagrams illustrating the comparison of the results of the noise reduction convolutional auto-encoding model and VGG-16 combined with the classification accuracy of the original VGG-16 model in various types and levels of the distorted image according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0013] The following description is of the best-contemplated mode of carrying out the invention. This description is made for the purpose of illustrating the general principles of the invention and should not be taken in a limiting sense. The scope of the invention is best determined by reference to the appended claims.

[0014] The present invention will be described with respect to particular embodiments and with reference to certain drawings, but the invention is not limited thereto and is only limited by the claims. It will be further understood that the terms "comprises," "comprising," "comprises" and / or "including," when used herein, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0015] Use of ordinal terms such as "first", "second", "third", etc., in the claims to modify a claim element does not by itself connote any priority, precedence, or order of one claim element over another or the temporal order in which acts of a method are performed, but are used merely as labels to distinguish one claim element having a certain name from another element having the same name (but for use of the ordinal term) to distinguish the claim elements.

[0016] Please refer to FIGS. 1-2, FIG. 1 is a block diagram of a noise reduction convolutional auto-encoding device 100 in accordance with one embodiment of the present disclosure. FIG. 2 is a flowchart of a noise reduction convolutional auto-encoding method 200 in accordance with one embodiment of the present disclosure. In one embodiment, noise reduction convolutional auto-encoding method 200 can be implemented by the noise reduction convolutional auto-encoding device 100.

[0017] As shown in FIG. 1, the noise reduction convolutional auto-encoding device 100 can be a desktop computer, a mobile phone, or a virtual machine built on a host operation system.

[0018] In one embodiment, the function of the noise reduction convolutional auto-encoding device 100 can be implemented by a hardware circuit, chip, firmware, or software.

[0019] In one embodiment, the noise reduction convolutional auto-encoding device 100 includes a processor 10 and a storage device 20. In one embodiment, the noise reduction convolutional auto-encoding device 100 further includes a display.

[0020] In one embodiment, the processor 10 can be a microcontroller, a microprocessor, a digital signal processor, or an application specific integrated circuit (ASIC), or a logic circuit to implement it.

[0021] In one embodiment, the storage device 20 can be implemented by a read-only memory, a flash memory, a floppy disk, a hard disk, a compact disk, a flash drive, a magnetic tape, a network accessible database, or a storage medium having the same function by those skilled in the art.

[0022] In one embodiment, the processor 10 is configured to access programs stored in the storage device 20 to implement the noise reduction convolutional auto-encoding method 200.

[0023] In one embodiment, a noise reduction convolutional coding model 30 generated by the noise reduction convolutional auto-encoding device 100 can be implemented by hardware, software, or firmware.

[0024] In one embodiment, the noise reduction convolutional coding model 30 implements its functions by software or firmware, and is stored in the storage device 20. The noise reduction convolutional auto-encoding device 100 accesses the noise reduction convolutional coding model 30 stored in the storage device 20 through the processor 10 to realize the function of the noise reduction convolutional auto-encoding device 100.

[0025] Please refer to FIGS. 2 to 3 together. FIG. 3 is a schematic diagram illustrating a noise reduction convolutional auto-encoding method in accordance with one embodiment of the present disclosure.

[0026] In step 210, the processor 10 receives a distorted image IT, and inputs the distorted image IT into the noise reduction convolutional auto-encoding model 30.

[0027] In step 220, the processor 10 transfers an image feature of the distorted image to a deconvolution layer Lu7 through skip-connection in the noise reduction convolutional auto-encoding model 30.

[0028] As shown in FIG. 3, the architecture of the noise reduction convolutional auto-encoding model 30 can be divided into two parts: dimensionality reduction operation and dimensionality upscaling operation. The distorted image IT input into the noise reduction convolutional auto-encoding model 30, and the dimensionality reduction is firstly performed. The dimensionality reduction operation involves sequentially executing the convolutional layers Ll1-Ll7, capturing the image features of the distorted image IT. Then, perform the dimension upscaling operation. The dimension upscaling operation includes the sequential execution of the convolutional layers Lu1-Lu7. Finally, a convolution is performed after the convolution layer Lu7, and then the reconstructed image OT is outputted.

[0029] In one embodiment, skip-connection refers to transferring the result of the dimensionality reduction operation in the convolutional layers L11-L17 during the dimensionality reduction operation to the corresponding convolutional layers Lu1-Lu7 of the dimensional upscaling computing layer during the dimensionality upscaling operation. For example, as shown in FIG. 3, the convolutional layer Ll1 skip-connection to the convolutional layer Lu7, the convolutional layer Ll2 skip-connection to the convolutional layer Lu6, the convolutional layer Ll4 skip-connection to the convolutional layer Lu4, and the convolutional layer Ll6 skip-connection to the convolutional layer Lu2.

[0030] In one embodiment, after distorted image IT inputs into the noise reduction convolutional auto-encoding model 30, each multi-stride encoding convolutional layer and same-dimensional encoding convolutional layer in noise reduction convolutional auto-encoding model 30 outputs their respective images features, and the respective image features are connected to the corresponding multi-stride decoding convolutional layers and same-dimensional decoding convolutional layers through skip-connection.

[0031] In one embodiment, the image features transmitted to these multi-stride decoding convolutional layers or same-dimensional decoding convolutional layers are used as reference features during decoding.

[0032] Therefore, the application of skip-connection can transmit the image features calculated by the dimensionality reduction of each layer to the corresponding convolutional layer in the dimensionality upscaling process.

[0033] In step 230, the processor 10 executes a plurality of multi-stride encoding convolutional layers for the distorted image IT to reduce the dimensionality, and then executes a same-dimensional encoding convolutional layer.

[0034] In one embodiment, as shown in FIG. 3, the processor 10 executes a plurality of multi-stride encoding convolutional layers Ll1- Ll2 for the distorted image IT to reduce dimensionality (referring to the dimensionality reduction operation relationship between the data of the convolutional layer Ll1 and Ll2, the same hereinafter). Then the same-dimensional encoding convolutional layer Ll2-Ll3 is executed (referring to the same-dimensional operation relationship between the data of the convolutional layer Ll2 and Ll3, the same hereinafter). Then, a plurality of multi-stride encoding convolutional layers Ll3-Ll4 is executed to reduce the dimensionality, and then the same-dimensional encoding convolutional layer Ll4-Ll5 is executed. Then, a plurality of multi-stride encoding convolutional layers Ll5-Ll6, Ll6-Ll7 are performed to reduce the dimensionality.

[0035] In one embodiment, the convolutional layers L17 and Lu1 can be regarded as the turning layers of dimensionality upscaling and dimensionality reduction in this architecture.

[0036] In other words, in the example in FIG. 3, the multi-stride encoding convolutional layers are the convolutional layers L12, L14, L16, and L17. The same-dimensional encoding convolutional layers are the convolutional layer Ll3 and Ll5. The convolutional layer Ll1 is the initial single-stride convolutional layer with the same dimension. It can be seen that there will be one or more multi-stride encoding convolutional layers before one or more same-dimensional encoding convolutional layers will be performed.

[0037] In one embodiment, after these multi-stride encoding convolutional layers are continuously sorted, the processor 10 sequentially sorts the last multi-stride encoding convolutional layers of these multi-stride encoding convolutional layers, and the image features outputted by the last multi-stride encoding convolutional layers are input to the same-dimensional encoding convolutional layer to avoid blocking effect. The blocking effect is a visual defect. The main reason for this effect is that the blocking artifacts caused by the block-based codec make the coding of adjacent blocks in the image are similar, and the visual effect looks like a mosaic.

[0038] In one embodiment, multi-stride encoding convolutional layers are used to define the amount of kernel pixel shift when performing convolution operations.

[0039] In one embodiment, when the amount of kernel pixel shift defined by the multi-stride encoding convolutional layers for convolution operation is higher, the higher the dimensionality reduction effect, the fewer the captured image features, and the more obvious the captured image features.

[0040] In step 240, according to the corresponding multi-stride encoding convolutional layers, same-dimensional encoding convolutional layers, the processor 10 performs upscaling dimension on the corresponding plural multi-stride decoding convolutional layers and a same-dimensional decoding convolutional layer.

[0041] In one embodiment, when the multi-stride encoding convolutional layers are a plurality of 2 stride encoding convolutional layers, it means that the amount of kernel pixel shift of the convolution operation is 2 when reducing the dimension. Corresponding to these multi-stride encoding convolutional layers, when these multi-stride decoding convolutional layers are a plurality of 2 stride encoding convolutional layers, it means that the amount of kernel pixel shift of the convolution operation is 2 when upscaling the dimension.

[0042] In the example in FIG. 3, the multi-stride decoding convolutional layers are the convolutional layers Lu1, Lu3, Lu5, and Lu7. The same-dimensional decoding convolutional layers are the convolutional layer Lu2, Lu4, and Lu6. It can be seen that there will be one or more multi-stride decoding convolutional layers first, and then one or more same-dimensional decoding convolutional layers will be performed.

[0043] In one embodiment, the number of dimensionality reduction is the same as the number of dimensionality upscaling.

[0044] In one embodiment, the convolutional layer used for decoding can be referred to as a deconvolution layer.

[0045] In step 250, the deconvolution layer Lu7 inputs a result after the dimension upscale is completed to the same-dimensional decoding convolutional layer of a balanced channel, and the same-dimensional decoding convolutional layer of the balanced channel outputs a reconstructed image OT.

[0046] In one embodiment, the deconvolution layer Lu7 refers to the image feature from the convolutional layer Ll1 for the result after the dimension upscale is completed, and then enters the same-dimensional decoding convolutional layer (not shown) of the balanced channel. The same-dimensional decoding convolutional layer of the balanced channel is used to output the reconstructed image OT with the three primary colors of color light.

[0047] In one embodiment, the same-dimensional encoding convolutional layer is regarded as a 1 stride encoding convolutional layer. The 1-stride encoding convolutional layer is arranged after a plurality of dimensionality reduction operations. The same-dimensional decoding convolutional layer is regarded as a 1-stride decoding convolutional layer for feature extraction and removal of noise carried in the process of transmitting the distorted image through skip-connection. The 1-stride decoding convolutional layer is arranged after a plurality of dimensionality upscaling operations.

[0048] In an embodiment, the processor 10 regards the 1-stride encoding convolutional layer executed y times after executing a plurality of x-stride encoding convolutional layers as a dimensionality reduction operation unit. The dimensionality reduction operation is completed by repeatedly executing the dimensionality reduction operation unit N times. The processor 10 regards the 1-stride decoding convolutional layer y times after executing a plurality of x-stride decoding convolutional layers as an upscaling operation unit. The upscaling operation is completed by repeatedly executing the upscaling operation unit N times. N and y are positive integers, and x is a positive integer greater than 2.

[0049] In one embodiment, as in the example in FIG. 3, in the encoder used for dimensionality reduction operations, there are a total of 7 convolutional layers in the process of dimensionality reduction operations, 3 same-dimensional encoding convolutional layers (64, 128, 256) for performing feature extraction, and 4 multi-stride encoding convolutional layers for performing four downsampling and feature extraction. In the feature extraction of the same-dimensional encoding convolutional layer (64, 128, 256), take 64 as an example, 64 represents the number of convolution kernels used. The two-dimensional size of the convolution kernel is 3*3, and 64 represents the depth of the convolution layer (using several convolution kernels), so 3*3*64 is generally used to represent the convolution layer. The meaning of 128 and 256 means that 128 convolution kernels and 256 convolution kernels are used, respectively.

[0050] Each dimensionality reduction feature map will be passed to the corresponding transposed convolutional layer (i.e., multi-stride decoding convolutional layers) through skip connections. Therefore, in the decoder used for the upscaling operation, there are provided corresponding four transposed convolutional layers (i.e., multi-stride decoding convolutional layers).

[0051] Since the amount of kernel pixel shift (step) defined by these multi-stride encoding convolutional layers for convolution operations is higher, the dimensionality reduction effect is higher, and the captured image features are more obvious (for example, when the resolution of the distorted image is large enough, a step of 3 or more can be used). Therefore, through multi-stride decoding of the convolutional layer, obvious image features can be captured faster, and the number of layers for dimensionality reduction operations can be greatly reduced, which can reduce the calculation time required for the noise reduction convolutional auto-encoding model 30 in the training stage. Moreover it can at least achieve the accuracy equivalent to or better than the traditional one-stride encoding convolutional layer to train the model. In addition, the processor 10 executes a 1-stride encoding convolutional layer, which can be used for feature extraction and remove noise transmitted by skip-connection.

[0052] In one embodiment, the noise reduction convolutional auto-encoding model 30 is used to restore a variety of distorted images (such as distorted image IT) to generate a reconstructed image OT. In one embodiment, the distorted image IT is, for example, a Gaussian noise image, a Gaussian blur image, a motion blur image, a speckle noise image, and / or a salt and pepper noise image.

[0053] In one embodiment, noise reduction convolutional auto-encoding model 30 uses batch normalization after each convolutional layer to reduce overfitting and gradient vanishing. LeakyReLU is used as the activation function to reduce the appearance of dead neurons.

[0054] In one embodiment, in the architecture of the noise reduction convolutional auto-encoding model 30, convolutional layers and transposed convolutional layers controlled by strides (i.e., multi-stride decoding convolutional layers) are used to replace traditional max-pooling and upsampling. The up-and-down dimension of the image can be used to obtain better distortion removal performance. Adding a convolutional layer (i.e., same-dimensional decoding convolutional layer) after each transposed convolutional layer of the decoder (i.e., multi-stride decoding convolutional layers) is to prevent transposed convolutional layers (i.e., multi-stride decoding convolutional layers) from causing the blocking effect that affects the restored feature map. Moreover, feature extraction is performed to reduce the noise transmitted by skip-connections.

[0055] In one embodiment, the training samples used by the noise reduction convolutional auto-encoding model 30 in the training stage include: multiple Gaussian noise images, multiple Gaussian blur images, multiple motion blur images, multiple speckle noise images, and multiple salt and pepper noise images.

[0056] In one embodiment, the noise reduction convolutional auto-encoding model 30 is used in the application stage to receive and restore multiple Gaussian noise images, multiple Gaussian blur images, multiple motion blur images, multiple speckle noise images, and multiple salt and pepper noise images to generate a reconstructed image.

[0057] In one embodiment, noise reduction convolutional auto-encoding model 30 is used in the training stage to combine mean-square error (MSE) and structural similarity index measurement (SSIM) to design a loss function, to update the weights or converge, and input images with different distortion types into noise reduction convolutional auto-encoding model 30 for training.

[0058] In one embodiment, in order to adapt the noise reduction convolutional auto-encoding model 30 to different types of distortion patterns, combine mean-square error (MSE) and structural similarity index measurement (SSIM) are used as the loss function to train the model. The direction of the gradient update can be affected by structural similarity index measurement and mean-square error at the same time, so that the reconstructed picture is closer to human vision and maintains a smaller numerical error with the original picture, which can handle various distortions. Combined loss function formula: Loss x y = α ⋅ 1 − SSIM x y + β ⋅ MSE x y ; where x, y represent the input and output images, and α and β are hyper parameters. During the training process, before inputting the distorted image IT into the noise reduction convolutional auto-encoding model 30, the processor 10 normalizes the distorted image IT and finds that the loss of SSIM drops faster than the loss of MSE, which means that the contribution of SSIM is greater than MSE in each epoch, and the value ratio is about 2: 1. Therefore, the processor 10 sets α=0.33 and β=0.66 to balance the contributions of SSIM and MSE. The value here is only an example and is not limited thereto.

[0059] In one embodiment, the noise reduction convolutional auto-encoding model 30 can also use peak signal-to-noise ratio (PSNR), structural similarity index (SSIM index), objective average score of the human eye, or other algorithms to design the loss function in the training stage. The design of the loss function is not limited thereto.

[0060] Please refer to FIGS. 4A to 4E. FIGS. 4A to 4E are schematic diagrams illustrating different distortion types in accordance with one embodiment of the present disclosure. In FIGS 4A to 4E, the larger σ, the more noise, and the larger the SNR, the less distortion.

[0061] FIG. 4A is the Guassian noise image, the first horizontal row is the original image, the second row is the distorted image input to noise reduction convolutional auto-encoding model 30, and the third row is the noise image output by the noise reduction convolutional auto-encoding model 30.

[0062] FIG. 4B is the speckle noise image, and the first row is the original image, the second row is the distorted image input to noise reduction convolutional auto-encoding model 30, and the third row is the noise image output by the noise reduction convolutional auto-encoding model 30.

[0063] FIG. 4C is the salt and pepper noise image, the first horizontal row is the original image, the second row is the distorted image input to noise reduction convolutional auto-encoding model 30, and the third row is the noise image output by the noise reduction convolutional auto-encoding model 30.

[0064] FIG. 4D is motion blur image, the first horizontal row is the original image, the second row is the distorted image input to noise reduction convolutional auto-encoding model 30, the third row is the noise image output by the noise reduction convolutional auto-encoding model 30.

[0065] FIG. 4E is a Gaussian blur image he first horizontal row is the original image, the second row is the distorted image input to noise reduction convolutional auto-encoding model 30, the third row is the noise image output by the noise reduction convolutional auto-encoding model 30.

[0066] It can be seen from FIGS. 4A to 4E above that after inputting distorted images with different distortion types into noise reduction convolutional auto-encoding model 30, noise reduction convolutional auto-encoding model 30 can output noise reduction and restored reconstructed images.

[0067] In one embodiment, three data sets Set5, Set12, and BSD68 are used to objectively compare the performance of the noise reduction convolutional auto-encoding model 30 with other technologies that specifically remove a single noise (Gaussian noise).

[0068] For Gaussian noise, three different levels of tests are performed, σ = {15, 25, 30}, and the results of the denoising report of Gaussian noise under the same test conditions are cited in the document [3], and BM3D [1], DnCNN-B [2] and FC-AIDE [3] perform performance comparison, and the results are shown in Table 1. Table 1data setnoise levelBM3D [1](PSNR / SSIM)DnCNN-B [2](PSNR / S SIM)FC-AIDE [3](PSNR / SSIM)noise reduction convoluti onal auto-encoding model 30(PSNR / SSIM)Set 5σ = 1529.64 / 0.898328.76 / 0.936430.69 / 0.951431.14 / 0.8890σ = 2526.47 / 0.898326.14 / 0.894827.83 / 0.918229.84 / 0.8324σ = 3025.32 / 0.876425.15 / 0.873526.80 / 0.903029.45 / 0.8156Set 12σ = 1532.15 / 0.885632.50 / 0.889932.91 / 0.899532.48 / 0.9121σ = 2529.67 / 0.832730.15 / 0.843530.51 / 0.854530.74 / 0.8549σ = 3028.74 / 0.808529.30 / 0.823329.63 / 0.835329.94 / 0.8287BSD 68σ = 1531.07 / 0.871731.40 / 0.880431.71 / 0.889731.68 / 0.8633σ = 2528.56 / 0.801328.99 / 0.813229.26 / 0.826729.68 / 0.8286σ = 3027.74 / 0.772728.17 / 0.784728.44 / 0.799528.84 / 0.8063 BM3D [1], DnCNN-B [2] and FC-AIDE [3] only process the Gaussian noise image and output the PSNR / SSIM value. The noise reduction convolutional auto-encoding model 30 can handle Gaussian noise image, Gaussian blur image, motion blur image, speckle noise image, and salt and pepper noise image.

[0069] The document [1] is "K. Dabov, A. Foi, V. Katkovnik, and K. Egiazarian. Image denoising by sparse 3-d transform-domain collaborative filtering. IEEE Trans. Image Processing, 16(8)): 2080-2095, 2007".

[0070] The document [2] is "K. Zhang, W. Zuo, Y. Chen, D. Meng, and L. Zhang. Beyond a Gaussian denoiser: Residual learning of deep CNN for image denoising. IEEE Trans. Image Processing, 26(7):3142 - 3155, 2017".

[0071] The document [3] is "S. Cha and T. Moon, "Fully convolutional pixel adaptive image denoiser," in Proc. IEEE Int. Conf. Comput. Vis., Seoul, South Korea, Oct. 2019, pp. 4160-4169".

[0072] For salt and pepper noise and speckle noise, compare the output of the noise reduction convolutional auto-encoding model 30 with BM3D [1] and DnCNN [2], both are only designed to remove Gaussian noise, and their source code can be obtained online for simulation. Salt and pepper noise and speckle noise are also evaluated at three different levels. Salt and pepper noise: SNR = {0.95, 0.90, 0.80}, speckle noise: σ = {0.1, 0.2, 0.3}, the results are shown in Table 2 and Table 3, respectively. Table 2datasetnoise levelBM3D [1](PSNR / S SIM)DnCNN [2](PSNR / SS IM)noise reduction convolutional auto-encoding model 30(PSNR / SSI M)Set 5SNR = 0.9524.84 / 0.630519.68 / 0.558531.95 / 0.9097SNR = 0.9022.46 / 0.516416.49 / 0.366531.03 / 0.9036SNR = 0.8016.86 / 0.264113.30 / 0.212130.10 / 0.8567Set 12SNR = 0.9525.32 / 0.659120.86 / 0.611033.60 / 0.9507SNR = 0.9023.67 / 0.568917.42 / 0.404132.24 / 0.9419SNR = 0.8017.96 / 0.281614.03 / 0.234730.73 / 0.8864BSD 68SNR = 0.9523.80 / 0.572320.06 / 0.561632.72 / 0.9302SNR = 0.9022.10 / 0.485816.92 / 0.374331.07 / 0.9013SNR = 0.8017.19 / 0.272713.73 / 0.220130.93 / 0.8802 Table 3 datasetnoise levelBM3D [1](PSN R / SSIM)DnCNN [2](PSNR / SSIM)noise reduction convolutional auto-encoding model 30(PSNR / SSIM)Set 5σ = 0.129.15 / 0.829530.99 / 0.921132.47 / 0.8939σ = 0.228.28 / 0.797026.91 / 0.812431.37 / 0.8591σ = 0.325.29 / 0.664520.96 / 0.662430.76 / 0.8313Set 12σ = 0.128.38 / 0.822130.32 / 0.915233.39 / 0.9356σ = 0.227.85 / 0.786126.23 / 0.804132.57 / 0.9013σ = 0.324.77 / 0.634319.75 / 0.574331.53 / 0.8693BSD 68σ = 0.125.05 / 0.696628.65 / 0.871533.19 / 0.8997σ = 0.224.89 / 0.687526.48 / 0.824531.43 / 0.8631σ = 0.324.18 / 0.620721.22 / 0.660530.64 / 0.8354 BM3D [1] and DnCNN [2] in Table 2 and Table 3 can only remove Gaussian noise and output the PSNR / SSIM value. The noise reduction convolutional auto-encoding model 30 can handle Gaussian noise image, Gaussian blur image, motion blur image, speckle noise image, and salt and pepper noise image.

[0073] Experiments show that noise reduction convolutional auto-encoding model 30 can use a single model to remove three types of distortion at the same time. From Table 1 to Table 3, it can be seen that noise reduction convolutional auto-encoding model 30 achieves the same level of performance in removing Gaussian noise as the known technology. Moreover, it also significantly improves the ability to remove other types of distortion such as salt and pepper noise and speckle noise. It can simultaneously reduce speckle noise and salt and pepper noise, and has excellent performance, which is better than other technologies that specifically remove single noise (Gaussian noise).

[0074] In FIGS. 4A-4C, noise reduction convolutional auto-encoding model 30 also shows the visualization results in the CBSD68 data set. Through a single model, all types and levels of noise can be removed.

[0075] In one embodiment, the noise reduction convolutional auto-encoding model 30 can be used in combination with the deep convolutional neural network VGG-16. In order to verify the effectiveness of the auto-encoder in improving the accuracy of image recognition, the noise reduction convolutional auto-encoding model 30 is combined with the deep convolutional neural network VGG-16. It means that the distorted image will go through the noise reduction convolutional auto-encoding model before entering the deep convolutional neural network (DCNN) for classification. The calculation framework for training and inference is shown in FIGS. 5A-5B.

[0076] FIG. 5A is a schematic diagram of the noise reduction convolutional auto-encoding model 30 and VGG-16 in the training stage according to an embodiment of the present invention. In FIG. 5A, the noise reduction convolutional auto-encoding model 30 and the deep convolutional neural network VGG-16 are trained separately in the training stage. More specifically, in the training stage, the distorted image is input to the noise reduction convolutional auto-encoding model 30, and the original image (the undistorted image) is input to the deep convolutional neural network VGG-16.

[0077] FIG. 5B is a schematic diagram of the noise reduction convolutional auto-encoding model 30 and VGG-16 in the testing stage in accordance with one embodiment of the present disclosure. In FIG. 5B, the distorted image is input to noise reduction convolutional auto-encoding model 30. The noise reduction convolutional auto-encoding model 30 inputs the output reconstructed image to the deep convolutional neural network VGG-16 to output a classification result.

[0078] In one embodiment, the processor 10 compares the combined result of the noise reduction convolutional auto-encoding model 30 and the VGG-16 with the classification accuracy of the original VGG-16 model under various types and levels of the distorted image. Taking the data set CIFAR-10 as an example, the processor 10 tests 5 increasing levels of corresponding types of noise and blur distortion (Gaussian noise image, Gaussian blur image, motion blur image, speckle noise image, and salt and pepper noise image). As shown in FIGS. 6A to 6E, FIGS. 6A to 6E are schematic diagrams illustrating the comparison of the results of the noise reduction convolutional auto-encoding model 30 and VGG-16 combined with the classification accuracy of the original VGG-16 model in various types and levels of the distorted image according to an embodiment of the present invention. It can be seen from FIGS. 6A to 6E that the classification accuracy of the combination of the noise reduction convolutional auto-encoding model 30 and the deep convolutional neural network VGG-16 is much higher than that of only the deep convolutional neural network VGG-16.

[0079] Experiments show that the combined noise reduction convolutional auto-encoding model 30 and VGG-16 have significantly higher accuracy rates in five different types and levels of distortion than the original VGG-16 model, and it only needs to be trained once. The noise reduction convolutional auto-encoding model 30 can simultaneously remove noise and blurring with 5 different distortion types for small-size images (CIFAR-10 / 100).

[0080] The method of the present invention, or a specific type or part thereof, may exist in the form of code. The program code can be included in physical media, such as floppy disks, CDs, hard disks, or any other machine-readable (such as computer-readable) storage media, or computer program products that are not limited to external forms. When the program code is loaded and executed by a machine, such as a computer, the machine becomes a device for participating in the present invention. The code can also be transmitted through some transmission media, such as wire or cable, optical fiber, or any transmission type. When the code is received, loaded and executed by a machine, such as a computer, the machine becomes used to participate in this Invented device. When implemented in the same-dimensional purpose processing unit, the program code combined with the processing unit provides a unique device that operates similar to the application of a specific logic circuit.

[0081] The noise reduction convolutional auto-encoding device and the noise reduction convolutional auto-encoding method described in present invention can simultaneously denoise or eliminate blurring of different types of distorted images, and the noise reduction convolutional auto-encoding device has a simple structure, only need to train once. By using skip-connection to transfer image features to the deconvolution layer, it helps deconvolution to restore a better, clearer image. Using multi-stride encoding convolutional layers and multi-stride decoding convolutional layers to increase / decrease the dimensionality of the image can achieve better denoising performance. Moreover, adding the same-dimensional encoding convolutional layer / same-dimensional decoding convolutional layer after multi-stride encoding convolutional layers / multi-stride decoding convolutional layers can prevent blocking effect caused by multi-stride encoding convolutional layers / multi-stride decoding convolutional layers, due to the blocking effect will affect the restored feature image. Then, performing the feature extraction again can reduce the noise transmitted by skip-connection.

[0082] In addition, noise reduction convolutional auto-encoding device and noise reduction convolutional auto-encoding method design a loss function by combining the mean square error and structural similarity index measurement to update the weights. When finally performing collaborative training on pictures with different distortion types, the weights can adapt to different distortions. The foregoing experimental results show that the noise reduction convolutional auto-encoding device used for noise reduction not only achieves the same level as the known technology in reducing Gaussian noise, but can also reduce noise for speckle noise and salt and pepper noise at the same time,. Moreover, the noise reduction convolutional auto-encoding device and noise reduction convolutional auto-encoding method have excellent performance. The noise reduction convolutional auto-encoding device is also a noise reduction convolutional auto-encoding device used to eliminate blur. It can eliminate Gaussian blur and motion blur at the same time with the same performance as other documents.

[0083] In addition, through noise reduction convolutional auto-encoding device and noise reduction convolutional auto-encoding method combine noise reduction convolutional auto-encoding model and deep convolutional neural network (VGG-16), the application of the noise reduction convolutional auto-encoding device can significantly improve the accuracy of convolutional network recognition of distorted images, and proved that a single noise reduction convolutional auto-encoding model only needs to be trained once, so as to achieve the effect of reconstructing images from a variety of different types of distortion images.

Claims

1. A noise reduction convolutional auto-encoding device (100), comprising: a processor (10); and a storage device (20), wherein the processor is configured to access a noise reduction convolutional auto-encoding model (30) stored in the storage device to execute the noise reduction convolutional auto-encoding model, wherein the processor (10) executes: receiving a distorted image (IT), and inputting the distorted image into the noise reduction convolutional auto-encoding model (30); in the noise reduction convolutional auto-encoding model (30), an image feature of the distorted image (IT) is transferred to a first deconvolution layer (Ll1) through skip-connection; characterized in that the processor further executes: performing a plurality of multi-stride encoding convolutional layers (Ll2, Ll4, Ll6, Ll7) for the distortion image (IT) to reduce a dimension, and then performing a same-dimensional encoding convolutional layer (L13, Ll5); according to the corresponding multi-stride encoding convolutional layers and same-dimensional encoding convolutional layers, a plurality of decoding multi-stride convolutional layers (Lu1, Lu3, Lu5, Lu7) and a same-dimensional decoding convolutional layers (Lu2, Lu4, Lu6) are upgraded; and inputting a result of up-scaled dimension to the same-dimensional decoding convolutional layer of a balanced channel using the first deconvolution layer (Ll1), and outputting a reconstructed image (OT) via the same-dimensional decoding convolutional layer of the balanced channel.

2. The noise reduction convolutional auto-encoding device (100) as claimed in claim 1, wherein after the distorted image (IT) is input to the noise reduction convolutional auto-encoding model (30), each of the multi-stride encoding convolutional layers (Ll2, Ll4, Ll6, Ll7) and the same-dimensional encoding convolutional layer (L13, Ll5) in the noise reduction convolutional auto-encoding model (30) respectively outputs the image feature, and transmits each image feature to the corresponding multi-stride decoding convolutional layers (Lu1, Lu3, Lu5, Lu7) and the same-dimensional decoding convolutional layer (Lu2, Lu4, Lu6) through skip-connection.

3. The noise reduction convolutional auto-encoding device (100) as claimed in claim 1 or 2, wherein the image feature transmitted to the multi-stride decoding convolutional layers or the same-dimensional decoding convolutional layer is used as a reference feature during decoding.

4. The noise reduction convolutional auto-encoding device (100) as claimed in one of the preceding claims, wherein the multi-stride encoding convolutional layers are used to define an amount of kernel pixel shift when performing convolution operations.

5. The noise reduction convolutional auto-encoding device (100) as claimed in claim 4, wherein when the multi-stride encoding convolutional layers (Ll2, Ll4, Ll6, Ll7) are a plurality of 2 stride encoding convolutional layers, it means that the amount of kernel pixel shift of the convolution operation is 2 when reducing the dimension; corresponding to the multi-stride encoding convolutional layers (Ll2, Ll4, Ll6, Ll7), when the multi-stride decoding convolutional layers (Lu1, Lu3, Lu5, Lu7) are a plurality of 2 stride encoding convolutional layers, it means that the amount of kernel pixel shift of the convolution operation is 2 when upscaling the dimension.

6. The noise reduction convolutional auto-encoding device (100) as claimed in claim 4 or 5, wherein when the amount of kernel pixel shift defined by the multi-stride encoding convolutional layers (Ll2, Ll4, Ll6, Ll7) for convolution operation is higher, the higher the dimensionality reduction effect, the fewer the captured image features, and the more obvious the captured image features.

7. The noise reduction convolutional auto-encoding device (100) as claimed in one of the preceding claims, wherein after the multi-stride encoding convolutional layers (Ll2, Ll4, Ll6, Ll7) are continuously sorted, the processor sequentially sorts the last multi-stride encoding convolutional layers of the multi-stride encoding convolutional layers, and inputs the output image features to the same-dimensional encoding convolutional layer to avoid blocking effects.

8. A noise reduction convolutional auto-encoding method (200), comprising: (210): receiving a distorted image (IT), and inputting the distorted image (IT) into a noise reduction convolutional auto-encoding model (30); (220): in the noise reduction convolutional auto-encoding model (30), an image feature of the distorted image (IT) is transferred to a first deconvolution layer (Ll1) through skip-connection; characterized in that the method further comprises: (230): performing a plurality of multi-stride encoding convolutional layers (Ll2, Ll4, Ll6, Ll7) for the distortion image (IT) to reduce a dimension, and then performing a same-dimensional encoding convolutional layer (L13, Ll5); (240): according to the corresponding multi-stride encoding convolutional layers and same-dimensional encoding convolutional layers, a plurality of decoding multi-stride convolutional layers (Lu1, Lu3, Lu5, Lu7) and a same-dimensional decoding convolutional layers (Lu2, Lu4, Lu6) are upgraded; and (250): inputting the result of up-scaled dimension to the same-dimensional decoding convolutional layer of a balanced channel using the first deconvolution layer (Ll1), and outputting a reconstructed image (OT) using the same-dimensional decoding convolutional layer of the balanced channel.

9. The noise reduction convolutional auto-encoding method (200) as claimed in claim 8, wherein after the distorted image (IT) is input to the noise reduction convolutional auto-encoding model (30), each of the multi-stride encoding convolutional layers (Ll2, Ll4, Ll6, Ll7) and the same-dimensional encoding convolutional layer (L13, Ll5) in the noise reduction convolutional auto-encoding model (30) respectively outputs the image feature, and transmits each image feature to the corresponding multi-stride decoding convolutional layers (Lu1, Lu3, Lu5, Lu7) and the same-dimensional decoding convolutional layer (Lu2, Lu4, Lu6) through skip-connection.

10. The noise reduction convolutional auto-encoding method (200) as claimed in claim 8 or 9, wherein the image feature transmitted to the multi-stride decoding convolutional layers or the same-dimensional decoding convolutional layer is used as a reference feature during decoding.

11. The noise reduction convolutional auto-encoding method (200) as claimed in one of claims 8-10, wherein the multi-stride encoding convolutional layers are used to define the amount of kernel pixel shift when performing convolution operations.

12. The noise reduction convolutional auto-encoding method (200) as claimed in one of claims 8-11, wherein when the multi-stride encoding convolutional layers (Ll2, Ll4, Ll6, Ll7) are a plurality of 2 stride encoding convolutional layers, it means that the amount of kernel pixel shift of the convolution operation is 2 when reducing the dimension; corresponding to the multi-stride encoding convolutional layers (Ll2, Ll4, Ll6, Ll7), when the multi-stride decoding convolutional layers (Lu1, Lu3, Lu5, Lu7) are a plurality of 2 stride encoding convolutional layers, it means that the amount of kernel pixel shift of the convolution operation is 2 when upscaling the dimension.

13. The noise reduction convolutional auto-encoding method (200) as claimed in claim 11, wherein when the amount of kernel pixel shift defined by the multi-stride encoding convolutional layers (Ll2, Ll4, Ll6, Ll7) for convolution operation is higher, the higher the dimensionality reduction effect, the fewer the captured image features, and the more obvious the captured image features.

14. The noise reduction convolutional auto-encoding method (200) as claimed in one of claims 8-13 wherein after the multi-stride encoding convolutional layers (Ll2, Ll4, Ll6, Ll7) are continuously sorted, the noise reduction convolutional auto-encoding method further comprises: sequentially sorting the last multi-stride encoding convolutional layers of the multi-stride encoding convolutional layers, and inputting the output image features to the same-dimensional encoding convolutional layer to avoid blocking effects.