Image noise reduction method and device, electronic equipment and readable storage medium
By performing feature extraction, downsampling, upsampling and decoding on the image, the target noise reduction processing model is used to achieve effective noise reduction of the image, solving the problem of poor noise reduction effect in the prior art, and achieving good noise reduction effect.
Patent Information
- Application Number
- CN202311697556.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-11
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art has poor results in the image noise reduction process, making it difficult to effectively remove noise in the image, resulting in unsatisfactory noise reduction effect.
By extracting the features of the image to be processed for downsampling, a first target feature map is obtained, and then upsampling is performed to obtain the second target feature map. Finally, the decoder unit of the target noise reduction processing model is used to decode the second target feature map to obtain the target image.
Effective noise reduction processing on the image is realized, the characteristics of the image to be processed are clearly represented, and the good noise reduction effect is achieved, which solves the problem of poor noise reduction effect in the prior art.
Smart Images

Figure CN120147167A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of computers, and particularly relates to an image denoising method, device, electronic device, and readable storage medium. Background Art
[0002] With the development of technology, people's requirements for the photographed photo images are getting higher and higher. Image denoising has become a hot research field.
[0003] In order to improve the effect and speed of image denoising, related technologies often use ordinary machine learning models to process images for denoising. However, the methods of image denoising in related technologies often have the problem of poor denoising effect. Summary of the Invention
[0004] Embodiments of this application provide an image denoising method, device, electronic device, and readable storage medium, thereby solving the problem of poor denoising effect in related technologies.
[0005] In a first aspect, embodiments of this application provide an image denoising method, including: obtaining a to-be-processed image, where the to-be-processed image is a noisy image; extracting features of the to-be-processed image and performing downsampling processing to obtain a first target feature map; performing upsampling processing on the first target feature map to obtain a second target feature map; using a decoder unit of a target denoising processing model to perform a decoding operation on the second target feature map to obtain a target image; where the target image is a denoised image predicted by inputting the to-be-processed image into a target denoising processing model including an encoder unit and a decoder unit.
[0006] In a second aspect, embodiments of this application provide an image denoising device, including: an obtaining module, configured to obtain a to-be-processed image, where the to-be-processed image is a noisy image; a processing module, configured to extract features of the to-be-processed image and perform downsampling processing to obtain a first target feature map; perform upsampling processing on the first target feature map to obtain a second target feature map; use a decoder unit of a target denoising processing model to perform a decoding operation on the second target feature map to obtain a target image; where the target image is a denoised image predicted by inputting the to-be-processed image into a target denoising processing model including an encoder unit and a decoder unit.
[0007] In a third aspect, embodiments of this application provide an electronic device, including: a memory and a processor, where the memory stores a computer program, and when the computer program is executed, the steps of the method described in the first aspect are implemented.
[0008] Fourthly, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored, and when the program or instruction is executed, the steps of the method described in the first aspect are implemented.
[0009] In an embodiment of the present application, a to-be-processed image is obtained, and the to-be-processed image is a noisy image; features of the to-be-processed image are extracted and downsampled to obtain a first target feature map; the first target feature map is upsampled to obtain a second target feature map; a decoder unit of a target noise reduction processing model is used to perform a decoding operation on the second target feature map to obtain a target image; wherein the target image is a noise reduction image predicted by inputting the to-be-processed image into the target noise reduction processing model including an encoder unit and a decoder unit. In this way, the target image after being processed by the target noise reduction processing model can clearly represent the features of the to-be-processed image, achieving a good noise reduction effect and solving the problem of poor noise reduction effect in the related art. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required to be used in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0011] Figure 1 is a flowchart of an image noise reduction method provided by an embodiment of the present application;
[0012] Figure 2 is a conceptual diagram of the structure of a target noise reduction processing model provided by an embodiment of the present application;
[0013] Figure 3 is a flowchart of an image noise reduction method provided by an embodiment of the present application;
[0014] Figure 4 is a conceptual diagram of a target convolution kernel provided by an embodiment of the present application;
[0015] Figure 5 is a conceptual diagram of the structure of the M-th downsampling layer of the target noise reduction processing model provided by an embodiment of the present application;
[0016] Figure 6 A conceptual diagram of the structure of the (M - 1)-th downsampling layer of the target noise reduction processing model provided by an embodiment of the present application;
[0017] Figure 7 A conceptual diagram of the structure of the first upsampling layer of the target noise reduction processing model provided by an embodiment of the present application;
[0018] Figure 8 It is a flowchart of an image denoising method provided by an embodiment of the present application;
[0019] Figure 9 It is a conceptual diagram of an image denoising method provided by an embodiment of the present application;
[0020] Figure 10 It is a conceptual diagram of the M-th downsampling layer structure of the initial denoising processing model provided by an embodiment of the present application;
[0021] Figure 11 It is a flowchart of the training process of the initial denoising processing model provided by an embodiment of the present application;
[0022] Figure 12 It is a structural block diagram of an image denoising device provided by an embodiment of the present application;
[0023] Figure 13 It is a structural block diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0024] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. The components of the embodiments of the present invention described and illustrated here usually can be arranged and designed in various different configurations.
[0025] Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0026] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order different from those illustrated or described here, and the objects distinguished by "first", "second", etc. generally belong to the same category, and the number of objects is not limited. For example, the first object can be one or multiple. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally represents an "or" relationship between the associated objects before and after.
[0027] The embodiment of the present application provides an image denoising method, which obtains an image to be processed, wherein the image to be processed is a noisy image; extracts the features of the image to be processed and performs downsampling processing to obtain a first target feature map; performs upsampling processing on the first target feature map to obtain a second target feature map; uses a decoder unit of a target denoising processing model to decode the second target feature map to obtain a target image; wherein the target image is a denoised image predicted by inputting the image to be processed into a target denoising processing model including an encoder unit and a decoder unit. In this way, the target image after denoising by the target denoising processing model can clearly represent the features of the image to be processed, achieve a good denoising effect, and solve the problem of poor denoising effect in related technologies.
[0028] In an embodiment of the present application, the target noise reduction processing model includes an encoder unit and a decoder unit; the encoder unit includes a convolution layer and M downsampling layers; the decoder unit includes M upsampling layers and convolution layers, where M≥2. In the Mth downsampling layer, a feature enhancement convolution layer is also included. In the feature enhancement convolution layer, blind spot denoising enhancement processing and pixel independence feature enhancement processing are performed on the input feature map of the feature enhancement convolution layer to obtain a first processing feature map. Through the processing of the feature enhancement convolution layer, the first processing feature map obtained can better retain the image features of the image to be processed, so that the image noise reduction processing effect is better.
[0029] In an embodiment of the present application, the Mth downsampling layer also includes a first edge enhancement convolution layer, which performs image gradient feature enhancement processing on the input feature map of the edge enhancement convolution layer. After processing by the edge enhancement convolution layer, not only the central features of the image can be retained, but also the edge features will be retained, so that the target image output by the target denoising processing model can better restore the image to be processed and achieve a good denoising effect.
[0030] In an embodiment of the present application, before obtaining the image to be processed, the image denoising method further includes training an initial denoising processing model. During the training process of the initial denoising processing model, a clear picture with lights and a clear picture without lights are obtained; the clear picture with lights is subjected to noise processing to obtain a first noise picture; the clear picture without lights is subjected to darkening and noise processing to obtain a second noise picture. The first noise picture and / or the second noise picture can both be training sample pictures, and the initial denoising processing model is iteratively trained through the training pictures. In this way, the noisy pictures taken by the mobile device under low light conditions can be restored as much as possible, so that the training sample picture data of the training process is closer to the input image data to be processed in the inference process, and further the trained target denoising processing model can complete the image denoising processing more accurately, thereby improving the accuracy of the processing of the target denoising processing model.
[0031] Embodiments of the present application can be used to process the environment in which a mobile terminal device captures pictures, especially in low light conditions, such as night scenes, to improve the noise reduction effect of the pictures captured by the mobile terminal and restore the original features of the pictures as much as possible. Embodiments of the present application can also be applied to network-side devices, such as servers, to simulate the processing effect of mobile terminal devices and complete lightweight image noise reduction.
[0032] It should be understood that the discussions of the same noun or situation in the embodiments of the present application can be referred to before and after. That is to say, the discussion of a certain noun or situation in one embodiment can also be applied to the description of the same noun or situation in other embodiments, as long as there is no logical contradiction.
[0033] The following will describe in detail the technical solutions provided by the embodiments of the present application with reference to the accompanying drawings.
[0034] The image noise reduction method provided by the embodiments of the present application can be executed by a target device. Specifically, in implementation, the image noise reduction method provided by the embodiments of the present application can be executed by one target device or by multiple target devices cooperating with each other. Among them, the target device can be, for example, a mobile terminal device such as a mobile phone, a laptop computer, a tablet, etc., or a server, such as an independent physical server, a server cluster composed of multiple servers, and a cloud server capable of performing cloud computing.
[0035] Figure 1 is a flowchart of an image noise reduction method provided by the embodiments of the present application. As Figure 1 shown, the image noise reduction method provided by the embodiments of the present application can include:
[0036] Step 110, obtain an image to be processed, where the image to be processed is a noisy image.
[0037] In the embodiments of the present application, the image to be processed can be any noisy image. For example, it can be a directly captured noisy image, or a noisy image obtained by adding noise (such as Gaussian noise, Poisson noise, etc.) to a high-definition image. The format of the image to be processed can be the format of an original unprocessed image file (for example, RAW format). The content of the image to be processed can be any content under any lighting conditions, such as a certain building photographed during the day, or a photo of the lights of a certain city at night.
[0038] When the target device is a mobile terminal device, the target device can directly obtain a noisy image through the shooting function of its own camera. When the target device is a network-side device, the target device can obtain the image to be processed by receiving the noisy image captured by the mobile terminal device.
[0039] In an exemplary embodiment, after obtaining the image to be processed, the target device may perform image denoising processing on the image to be processed.
[0040] Step 120: Extract the features of the image to be processed and perform downsampling processing to obtain a first target feature map.
[0041] In the embodiments of the present application, the features of the image to be processed may be extracted by a target denoising processing model, and then the feature map of the image to be processed is downsampled to obtain a downsampled feature map (i.e., the first target feature map). During the downsampling process, the feature map of the image to be processed can be downsampled through multiple downsampling layers to reduce the resolution of the feature map of the image to be processed and obtain the first target feature map. The target denoising processing model is a deep learning model that can perform denoising processing on the image to be processed. For example, a deep learning model based on a convolutional neural network can extract and retain the image features of the image to be processed to achieve the purpose of denoising the image to be processed. The target denoising processing model can also be a deep learning model using an encoder-decoder architecture. The encoder and the decoder can form a symmetric structure and form skip connections. The decoder can not only obtain the output result from the previous decoder but also obtain the output result of the encoder with skip connections and process them together.
[0042] When the target device is a mobile terminal device, after obtaining the image to be processed, it can be processed by the target denoising processing model on the target device. When the target device includes multiple electronic devices, one electronic device can obtain the image to be processed, and the target denoising processing model on another electronic device can perform denoising processing.
[0043] After the resolution of the image to be processed is reduced through downsampling, the resolution can be restored through upsampling.
[0044] Step 130: Perform upsampling processing on the first target feature map to obtain a second target feature map.
[0045] In the embodiments of the present application, the second target feature map may be a feature map with the same resolution as the original resolution of the image to be processed after upsampling. During the upsampling process, the first target feature map can be upsampled by means of transposed convolution or bilinear interpolation.
[0046] Step 140: Use the decoder unit of the target denoising processing model to perform decoding operations on the second target feature map to obtain a target image; where the target image is a denoised image predicted by inputting the image to be processed into a target denoising processing model including an encoder unit and a decoder unit.
[0047] In the embodiments of the present application, the target image may be a high-definition display image obtained after denoising the image to be processed. For example, when the image to be processed is a photo of the lights of a city at night, the target image may be a high-definition photo of the lights of a city at night. The format of the target image may be the format of the original unprocessed image file (e.g., RAW format), or may also be the format of other high-definition display denoised images, such as the Portable Network Graphics (PNG) format.
[0048] In the embodiments of the present application, the input image to be processed may first be subjected to feature extraction processing to obtain a feature map of the image to be processed. Subsequently, the feature map of the image to be processed is subjected to downsampling processing to obtain a feature map after downsampling processing (i.e., the first target feature map). Then, the first target feature map is subjected to upsampling processing to obtain the second target feature map, and the resolution of the second target feature map may be the same as the original resolution of the image to be processed. Then, a decoding operation is performed on the second target feature map, and finally, a predicted denoised image (i.e., the target image) is obtained. In this way, the target image after denoising processed by the target denoising processing model can clearly represent the features of the image to be processed, achieving a good denoising effect and solving the problem of poor denoising effect in the related art.
[0049] In the embodiments of the present application, the target denoising processing model includes an encoder unit and a decoder unit; the encoder unit includes a convolutional layer and M downsampling layers; the decoder unit includes M upsampling layers and a convolutional layer. To better understand the architecture of the target denoising processing model provided in the embodiments of the present application, to better understand the concept of the architecture of the target denoising processing model provided in the embodiments of the present application, M is taken as 4 for example here. It should be understood that M is only an example here and not a limitation. In fact, M may also take other values other than 4. As Figure 2 shown, the target denoising processing model provided in the embodiments of the present application includes the encoder unit and the decoder unit. The encoder unit includes a convolutional layer and 4 downsampling layers, and the decoder unit includes 4 upsampling layers and a convolutional layer. The encoder and decoder units may use skip connections. The decoder unit can not only obtain the output result from the previous decoder, but also obtain the output result of the encoder with skip connections, making the processing effect of the target denoising processing model better.
[0050] When the target device is a mobile terminal device, the target noise reduction processing model can be set according to the characteristics of the mobile terminal device, so as to ensure that the target noise reduction processing model is more matched with the performance characteristics of the mobile terminal device, and give full play to the processing power of the mobile terminal. For example, when the mobile terminal uses a Qualcomm Tensor Processor (HTP for short), the design within the convolutional layer can be carried out according to the characteristics of the Qualcomm Tensor Processor. For example, in the Qualcomm Tensor Processor, multiples of 32 channels are preferred choices. Then, within the encoder unit and the decoder unit, the number of convolutional kernels can also be a multiple of 32. In this way, by making the setting of the target noise reduction processing model fit the characteristics of the processor of the mobile terminal, the processing speed can be greatly improved. In this way, electronic devices with relatively weak computing power such as mobile terminals can also apply the image noise reduction method provided by the embodiments of the present application, greatly expanding the application scope.
[0051] Figure 3 is a flowchart of an image noise reduction method provided by an embodiment of the present application. As Figure 3 shown, the image noise reduction method provided by the embodiments of the present application may include:
[0052] Step 310, obtain an image to be processed, where the image to be processed is a noisy image.
[0053] Step 320, input the image to be processed into a target noise reduction processing model, where the target noise reduction processing model is a deep learning model, and the target noise reduction processing model includes an encoder unit and a decoder unit; the encoder unit includes a convolutional layer and M downsampling layers; the decoder unit includes M upsampling layers and a convolutional layer; where M≥2.
[0054] Step 330, extract the features of the image to be processed through the convolutional layer of the encoder unit to obtain a first feature map.
[0055] In the embodiments of the present application, the first feature map may be a feature map of the image to be processed, which contains the basic features of the image to be processed.
[0056] When the target device is a mobile terminal device, the number of convolutional kernels in the convolutional layer of the encoder unit can be a multiple of 32. For example, within the convolutional layer of the encoder unit, there are 32 convolutional kernels. After the first feature map is obtained through the processing of 32 convolutional kernels, the first feature map can be input into the downsampling layer of the encoder unit.
[0057] Step 340, perform downsampling processing on the first feature map through the M downsampling layers to obtain a first target feature map.
[0058] In the embodiments of the present application, M can be any integer, for example, M is 4 or 5. During the downsampling process, the first feature map can be downsampled through multiple downsampling layers to reduce the resolution of the first feature map and obtain a first target feature map.
[0059] As Figure 2 shown, when M is 4, the first feature map is downsampled through 4 downsampling layers. After being downsampled by the fourth downsampling layer, a first target feature map is obtained.
[0060] After obtaining the first target feature map with a smaller resolution, the resolution can be restored through upsampling.
[0061] Step 350, upsample the first target feature map through the M upsampling layers to obtain a second target feature map;
[0062] In the embodiments of the present application, the number of upsampling layers corresponds to the number of downsampling layers. During the upsampling process, the first target feature map can be upsampled through transposed convolution or bilinear interpolation to restore the resolution of the first target feature map and obtain a second target feature map.
[0063] For example, as Figure 2 shown, when the first feature map is downsampled through 4 downsampling layers, the first target feature map can be upsampled through 4 upsampling layers. After being upsampled by the fourth upsampling layer, the second target feature map is obtained.
[0064] In the embodiments of the present application, after obtaining the second target feature map, a target image can be obtained according to the second target feature map.
[0065] Step 360, perform a decoding operation on the second target feature map through the convolutional layer of the decoder unit to obtain a target image.
[0066] In the embodiments of the present application, after obtaining the second target feature map, further feature extraction can be performed on the second target feature map through a convolutional layer to obtain a target image.
[0067] In the embodiment of the present application, after the feature extraction of the image to be processed to obtain the first feature map, downsampling processing is performed to obtain the first target feature map with a resolution lower than that of the image to be processed, thereby reducing the convolutional processing on the resolution scale of the image to be processed, improving the calculation speed, and enabling more rapid noise reduction processing of the image to be processed. And when the target device is a mobile terminal device, the embodiment of the present application can set the number of convolutional kernels according to the characteristics of the mobile terminal device, thereby more fully invoking the processing system of the mobile terminal device and improving the processing speed of the target noise reduction processing model.
[0068] In the embodiment of the present application, in step 340, a specific way to perform downsampling processing on the first feature map through the M downsampling layers to obtain the first target feature map may be:
[0069] Perform downsampling processing on the first feature map through the first downsampling layer to obtain a second feature map;
[0070] Perform downsampling processing on the i-th feature map through the i-th downsampling layer to obtain the (i + 1)-th feature map;
[0071] In the embodiment of the present application, the input of each downsampling layer is the output result of the previous upsampling. For example, after obtaining the second feature map, the second feature map is input into the second downsampling layer (at this time, i is 2) for downsampling processing to obtain the third feature map.
[0072] Perform downsampling processing on the M-th feature map through the M-th downsampling layer to obtain the (M + 1)-th feature map, and use the (M + 1)-th feature map as the first target feature map; where M > i > 1.
[0073] To better understand the method for obtaining the first target feature map provided by the embodiment of the present application, an example is given below. It should be noted that the example is not restrictive. As Figure 2 shown, when M is 4, i can be 2 or 3. After performing downsampling processing on the first feature map (the output result of the convolutional layer of the encoder unit for feature extraction of the image to be processed) through the first downsampling layer, a second feature map is obtained. Then, the second feature map is subjected to downsampling processing through the second downsampling layer to obtain the third feature map. And so on, the fourth feature map (that is, the output result of the third downsampling layer) is subjected to downsampling processing through the fourth downsampling layer to obtain the fifth feature map, and the fifth feature map is used as the first target feature map, that is, corresponding to step 240, when the M (as Figure 2 shown, M is 4) downsampling layers are processed, the obtained first target feature map.
[0074] In the embodiments of the present application, by performing downsampling processing on the first feature map through M downsampling layers, while reducing the resolution of the first feature map, the features of the first feature map can be retained as much as possible, improving the noise reduction effect of the target noise reduction processing model.
[0075] In an embodiment of the present application, the M-th downsampling layer includes: a feature enhancement convolutional layer. A specific process of performing downsampling processing on the M-th feature map through the M-th downsampling layer to obtain the (M + 1)-th feature map, that is, the first target feature map, includes: inputting the M-th feature map into the feature enhancement convolutional layer; performing blind spot denoising enhancement processing and pixel independence feature enhancement processing on the M-th feature map through the target convolutional kernel of the feature enhancement convolutional layer to obtain a first processed feature map; obtaining the (M + 1)-th feature map based on the first processed feature map; where the target convolutional kernel is of size A x A, where A is an odd number greater than or equal to 3, the convolution values of the even rows and even columns of the target convolutional kernel are all 0, and the weight value at the center position of the convolutional kernel is also 0; the stride of the target convolutional kernel is an integer greater than 1, and the resolution of the first processed feature map is smaller than the resolution of the M-th feature map.
[0076] To better understand the embodiments of the present application, an example is given here. It should be noted that the example is not restrictive. When M is 4, after completing the downsampling processing in the third downsampling layer to obtain the fourth feature map, the fourth feature map is input into the feature enhancement convolutional layer in the fourth downsampling layer. The target convolutional kernel of the feature enhancement convolutional layer performs blind spot denoising processing on the fourth feature map to obtain a first processed feature map.
[0077] In the process of performing pixel independence enhancement in the feature enhancement convolutional layer, the target convolutional kernel is of size AxA, where A is an odd number greater than or equal to 3, and the convolution values of the even rows and even columns of the target convolutional kernel are all 0; the stride of the target convolutional kernel is an integer greater than 1. As Figure 4 shown, when A is 5, the size of the target convolutional kernel is 5x5, and the stride is an integer greater than 1. When using a target convolutional kernel of size 5x5 and stride 2 for pixel independence enhancement, the convolution values of the even rows (the 2nd and 4th rows) and even columns (the 2nd and 4th columns) of the target convolutional kernel can be set to 0. In the process of performing blind spot denoising enhancement processing in the feature enhancement convolutional layer, in order to make full use of the blind spot denoising characteristics, the convolution value at the center point of the target convolutional kernel can also be set to 0. For example, the convolution value at the coordinate point (3, 3) is also set to 0, and then convolution and stacking are performed to obtain a first processed feature map with a resolution that is half of the resolution of the fourth feature map (the input result of the fourth downsampling layer).
[0078] In another embodiment of the present application, the M-th downsampling layer further includes: a first edge enhancement convolutional layer. The first edge enhancement convolutional layer can be an Angular Pixel Difference Convolution (APDC layer for short). A specific process of obtaining the (M + 1)-th feature map based on the first processed feature map can be: inputting the first processed feature map (the output result of the feature enhancement convolutional layer in the M-th sampling layer) into the edge enhancement convolutional layer for image gradient feature enhancement processing to obtain a second processed feature map; and obtaining the (M + 1)-th feature map based on the second processed feature map.
[0079] In the embodiments of the present application, the image gradient can be the directional change of image intensity or color, and the image gradient can be represented by the difference change of neighborhood pixel values. By performing image gradient feature enhancement processing on the first processed feature map through the first edge enhancement convolutional layer, during the process of performing image gradient feature enhancement processing on the first processed feature map through the first edge enhancement convolutional layer, the convolution of the differences of pixel values in the neighborhood (such as adjacent ways like up and down, left and right, diagonal, etc.) can be performed to make the pixel point features at the image edge more distinct, thereby enhancing the edge features of the first processed feature map and improving the noise reduction effect of the target noise reduction processing model.
[0080] In another embodiment of the present application, the M-th downsampling layer further includes: a first convolutional layer and a second convolutional layer. After the M-th downsampling layer obtains the M-th feature map, the M-th feature map can also be input into the first convolutional layer for feature extraction processing to obtain a third processed feature map; at the same time, the M-th feature map can be input into the feature enhancement convolutional layer and the first edge enhancement convolutional layer to obtain a second processed feature map. Then, the second processed feature map is input into the second convolutional layer for feature extraction processing to obtain a fourth processed feature map. Adding the third processed feature map and the fourth processed feature map together to obtain the (M + 1)-th feature map, that is, the first target feature map. It should be noted that the first convolutional layer can be a single convolutional layer or multiple convolutional layers, and the embodiments of the present application do not make specific limitations on this. Similarly, the second convolutional layer can also be a single or multiple convolutional layers.
[0081] To better understand the M-th downsampling layer provided in the embodiments of the present application, an example is given here. It should be noted that the example is not restrictive. The conceptual diagram of the M-th downsampling layer provided in the embodiments of the present application is as Figure 5As shown, the M-th downsampling layer first obtains the output result of the (M - 1)-th downsampling layer (i.e., the M-th feature map), and then the M-th feature map can be input into the branches where the feature enhancement convolutional layer and the first convolutional layer of the M-th downsampling layer are located respectively. When the M-th feature map passes through the feature enhancement convolutional layer, a first processed feature map is obtained. Subsequently, taking the first processed feature map as the input, it is input into the first edge enhancement convolutional layer. After the image gradient feature enhancement processing of the first edge enhancement convolutional layer, a second processed feature map is obtained. Then, the second processed feature map is convolved through the second convolutional layer to obtain a fourth processed feature map. While the M-th feature map is being processed by the feature enhancement convolutional layer, the first edge enhancement convolutional layer, and the second convolutional layer, the M-th feature map is subjected to feature extraction processing through the first convolutional layer to obtain a third processed feature map. The third processed feature map and the fourth processed feature map are added together to obtain the (M + 1)-th feature map, which is also the output result of the encoder unit of the target noise reduction processing model, the first target feature map.
[0082] In an embodiment of the present application, the (M - 1)-th downsampling layer includes: a third convolutional layer and a second edge enhancement convolutional layer. A specific process of downsampling the (M - 1)-th feature map through the (M - 1)-th sampling layer to obtain the M-th feature map includes: inputting the (M - 1)-th feature map into the third convolutional layer for feature extraction processing to obtain a first specified feature map; wherein, the stride of the convolutional kernel of the third convolutional layer is an integer greater than 1, and the resolution of the first specified feature map is smaller than that of the (M - 1)-th feature map; inputting the first specified feature map into the second edge enhancement convolutional layer for image gradient feature enhancement processing to obtain a second specified feature map. In the embodiment of the present application, the second edge enhancement convolutional layer and the first edge enhancement convolutional layer can both be APDC layers and adopt the same edge enhancement convolutional method, or they can adopt different edge enhancement convolutional methods. The embodiment of the present application does not make specific limitations in this regard. It should be understood that the third convolutional layer can be a single convolutional layer or multiple convolutional layers. The embodiment of the present application does not make specific limitations in this regard. Similarly, the fourth convolutional layer and the fifth convolutional layer can be a single layer or multiple layers.
[0083] In another embodiment of the present application, the (M - 1)-th downsampling layer further includes a fourth convolutional layer and a fifth convolutional layer. Inputting the (M - 1)-th feature map into the fourth convolutional layer for feature extraction processing to obtain a third specified feature map; inputting the second specified feature map into the fifth convolutional layer for feature extraction processing to obtain a fourth specified feature map. Adding the third specified feature map and the fourth specified feature map together to obtain the M-th feature map.
[0084] To better understand the (M - 1)-th downsampling layer provided in the embodiment of the present application, an example is given as follows Figure 6As shown in the figure, the (M-1)-th downsampling layer provided by the embodiment of the present application includes a third convolutional layer, a second edge enhancement convolutional layer, a fourth convolutional layer, and a fifth convolutional layer. After the (M-1)-th downsampling layer obtains the (M-1)-th feature map, the (M-1)-th feature map is input to both the third convolutional layer and the fourth convolutional layer. After the third convolutional layer obtains the (M-1)-th feature map, feature extraction processing is performed to obtain a first specified feature map. The first specified feature map is input to the second edge enhancement convolutional layer for image gradient feature enhancement processing to obtain a second specified feature map, and then the fifth convolutional layer performs convolutional processing on the second specified feature map to obtain a fourth specified feature map. While obtaining the fourth specified feature map, the (M-1)-th feature map is processed by the fourth convolutional layer to obtain a third specified feature map. The third specified feature map and the fourth specified feature map are added together to obtain the M-th feature map, which is the input image of the M-th downsampling layer.
[0085] Through the (M-1)-th downsampling layer provided by the embodiment of the present application, the M-th feature map can be further downsampled, the resolution of the M-th feature map can be compressed, and the pixel data volume of the M-th feature map can be reduced, thereby improving the processing speed of the target noise reduction processing model. And the edge features of the M-th feature map are enhanced through the second edge enhancement layer, making the processing effect of the target noise reduction processing model better.
[0086] In the embodiment of the present application, after the first feature map is downsampled by the M downsampling layers to obtain a first target feature map, the first target feature map can be upsampled by the M upsampling layers to obtain a second target feature map. A specific process of upsampling the first target feature map by the M upsampling layers to obtain a second target feature map includes: upsampling the first target feature map by the first upsampling layer to obtain a first intermediate feature map; upsampling the (j-1)-th intermediate feature map by the j-th upsampling layer to obtain the j-th intermediate feature map; where M≥j>1; when j = M, the M-th intermediate feature map is the second target feature map. For example Figure 2 As shown in the figure, when M is 4, the first upsampling layer upsamples the first target feature map (i.e., the (M+1)-th feature map, which is also the output image of the M-th downsampling layer) to obtain the first intermediate feature map. Then, the second upsampling layer (at this time j is 2) upsamples the first intermediate feature map (at this time, the (j-1)-th intermediate feature map is the first intermediate feature map, which is also the output image of the first upsampling layer) to obtain the second intermediate feature map.
[0087] In the embodiments of the present application, the format of the first intermediate feature map may be the same as the format of the first target feature map. By performing upsampling processing on the first target feature map through M upsampling layers, the resolution of the first target feature map can be enlarged, the image display content of the first target feature map can be restored, and the noise reduction effect of the target noise reduction processing model can be improved.
[0088] In the embodiments of the present application, the first upsampling layer includes a first branch and a second branch, and there is at least one convolutional layer in each branch; when the first upsampling layer obtains the output result of the M-th downsampling, the first target feature map, the first target feature map is input into the first branch and the second branch of the first upsampling layer at the same time, and the output result of the first branch and the output result of the second branch are added to obtain the first target output result; the first target output result is input into the first transposed convolutional layer for image upsampling processing to obtain the first intermediate feature map; wherein, the step size of the convolutional kernel of the first transposed convolutional layer is an integer greater than 1, and the resolution of the first intermediate feature map is greater than the resolution of the first target feature map.
[0089] For a better understanding of the first upsampling layer provided in the embodiments of the present application, an example is given below. The architecture of the first upsampling layer provided in the embodiments of the present application is as Figure 7 shown. After the first upsampling layer obtains the output result of the M-th downsampling layer, the first target feature map (i.e., the (M + 1)-th feature map), the first target feature map is input into the first branch and the second branch of the first upsampling layer respectively, and the output result of the first branch and the output result of the second branch are added to obtain the first target output result; the first target output result is input into the first transposed convolutional layer for image upsampling processing to obtain the first intermediate feature map. In the embodiments of the present application, the structure of the first branch of the first upsampling layer may be the same as the structure in the M-th downsampling layer as Figure 5 shown, for example, including a feature enhancement convolutional layer, a first edge enhancement convolutional layer and a second convolutional layer, and the embodiments of the present application do not make specific limitations on this.
[0090] In the embodiments of the present application, the first upsampling layer is set to a two-branch structure, making use of the symmetry of the encoder-decoder architecture to form a symmetric structure with the M-th downsampling layer. At the same time, the resolution of the first intermediate feature map is gradually restored through a convolutional kernel with a step size greater than 1, so that the resolution of the first intermediate feature map can be improved, and the noise reduction effect of the target noise reduction processing model is improved.
[0091] In an embodiment of the present application, after the first upsampling layer performs upsampling processing on the first target feature map to obtain a first intermediate feature map, the j-th upsampling layer may perform upsampling processing on the (j - 1)-th intermediate feature map to obtain a j-th intermediate feature map. A specific implementation may be: the j-th upsampling layer performs upsampling processing on the (j - 1)-th intermediate feature map to obtain an output feature map; the output feature map is added to the output result of the (M - j + 1)-th downsampling layer to obtain a target output feature map. For example, as Figure 2 shown, when M is 4 and j is 2, the second upsampling layer performs upsampling processing on the first intermediate feature map (the output result of the first upsampling layer) to obtain an output feature map, and the output feature map is added to the output result of the 3rd (4 - 2 + 1 = 3) downsampling layer (i.e., the fourth feature map) to obtain a target output feature map. The target output feature map is input into a second transposed convolutional layer for image upsampling processing to obtain a second intermediate feature map; wherein, the stride of the convolutional kernel of the second transposed convolutional layer is an integer greater than 1, and the resolution of the second intermediate feature map is greater than the resolution of the first intermediate feature map.
[0092] In an embodiment of the present application, through repeated upsampling processing by the j-th upsampling layer, the resolution of the j-th intermediate feature map is gradually enlarged to restore to the original resolution of the image to be processed, so that the features of the first target feature map can be displayed at as high a resolution as possible, achieving the effect of denoising the image to be processed and improving the effect of the target denoising processing model.
[0093] Figure 8 is a flowchart of an image denoising method provided by an embodiment of the present application. As Figure 8 shown, the image denoising method provided by an embodiment of the present application may include:
[0094] Step 810, obtain an image to be processed, where the image to be processed is a noisy image.
[0095] Step 815, extract the features of the image to be processed through the convolutional layer of the encoder unit to obtain a first feature map.
[0096] Step 820, perform downsampling processing on the first feature map through the first downsampling layer to obtain a second feature map.
[0097] Step 825, perform downsampling processing on the i-th feature map through the i-th downsampling layer to obtain an (i + 1)-th feature map.
[0098] Step 830, when i = M - 1, perform downsampling processing on the (M - 1)-th feature map through the (M - 1)-th downsampling layer to obtain an M-th feature map.
[0099] In step 830, a specific implementation process of downsampling the (M - 1)-th feature map through the (M - 1)-th downsampling layer to obtain the M-th feature map may be as follows:
[0100] Input the (M - 1)-th feature map into the fourth convolutional layer for feature extraction to obtain a third specified feature map
[0101] Input the (M - 1)-th feature map into the third convolutional layer for feature extraction to obtain a first specified feature map; wherein, the stride of the convolutional kernel of the third convolutional layer is an integer greater than 1, and the resolution of the first specified feature map is less than that of the (M - 1)-th feature map;
[0102] Input the first specified feature map into the second edge enhancement convolutional layer for image gradient feature enhancement to obtain a second specified feature map;
[0103] Input the second specified feature map into the fifth convolutional layer for feature extraction to obtain a fourth specified feature map;
[0104] Add the third specified feature map and the fourth specified feature map to obtain the M-th feature map.
[0105] Step 835: Downsample the M-th feature map through the M-th downsampling layer to obtain the (M + 1)-th feature map, and use the (M + 1)-th feature map as the first target feature map; where M > i > 1.
[0106] In step 835, a specific implementation process of downsampling the M-th feature map through the M-th downsampling layer to obtain the (M + 1)-th feature map, and using the (M + 1)-th feature map as the first target feature map may be as follows:
[0107] Input the M-th feature map into the feature enhancement convolutional layer;
[0108] Perform blind spot denoising on the M-th feature map through the target convolutional kernel of the feature enhancement convolutional layer to obtain a first processed feature map;
[0109] Input the first processed feature map into the first edge enhancement convolutional layer for image gradient feature enhancement to obtain a second processed feature map;
[0110] Input the second processed feature map into the second convolutional layer for feature extraction to obtain a fourth processed feature map;
[0111] Input the M-th feature map into the first convolutional layer for feature extraction to obtain a third processed feature map;
[0112] Add the third processed feature map and the fourth processed feature map to obtain the (M + 1)-th feature map, i.e., the first target feature map.
[0113] Step 840: Upsample the first target feature map through the first upsampling layer to obtain the first intermediate feature map.
[0114] In step 840, the first upsampling layer includes a first branch and a second branch, and there is at least one convolutional layer in each branch. A specific implementation of upsampling the first target feature map through the first upsampling layer to obtain the first intermediate feature map can be as follows:
[0115] Add the output result of the first branch and the output result of the second branch to obtain the first target output result;
[0116] Input the first target output result into the first transposed convolutional layer for image upsampling to obtain the first intermediate feature map;
[0117] Wherein, the stride of the convolutional kernel of the first transposed convolutional layer is an integer greater than 1, and the resolution of the first intermediate feature map is greater than the resolution of the first target feature map.
[0118] Step 845: Upsample the (j - 1)-th intermediate feature map through the j-th upsampling layer to obtain the j-th intermediate feature map; where M ≥ j > 1.
[0119] In step 845, a specific implementation can be as follows:
[0120] Upsample the (j - 1)-th intermediate feature map through the j-th upsampling layer to obtain the output feature map;
[0121] Add the output feature map and the output result of the (M - j + 1)-th downsampling layer to obtain the target output feature map;
[0122] Input the target output feature map into the i-th transposed convolutional layer for image upsampling to obtain the j-th intermediate feature map;
[0123] Wherein, the stride of the convolutional kernel of the j-th transposed convolutional layer is an integer greater than 1, and the resolution of the j-th intermediate feature map is greater than the resolution of the (j - 1)-th intermediate feature map.
[0124] Step 850: When j = M - 1, upsample the (M - 2)-th intermediate feature map through the (M - 1)-th upsampling layer to obtain the (M - 1)-th intermediate feature map.
[0125] Step 855: Upsample the (M-1)th intermediate feature map through the Mth upsampling layer to obtain the Mth feature map, and use the Mth feature map as the second target feature map.
[0126] Step 860: Decode the second target feature map through the convolutional layer of the decoder unit to obtain a target image, where the target image is the predicted denoised image.
[0127] In the embodiment of the present application, a to-be-processed image is obtained, where the to-be-processed image is a noisy image; the to-be-processed image is input into a target denoising processing model, and the target denoising processing model is a deep learning model; the to-be-processed image is subjected to image denoising processing through the target denoising processing model to obtain a target image, where the target image is the predicted denoised image; wherein, in terms of image denoising processing, the target denoising processing model extracts the features of the to-be-processed image for downsampling to obtain a first target feature map; upsamples the first target feature map to obtain a second target feature map; and decodes the second target feature map to obtain a target image. In this way, the target image after denoising processed by the target denoising processing model can clearly represent the features of the to-be-processed image, achieving a good denoising effect and solving the problem of poor denoising effect in the related art.
[0128] To better understand the image denoising method provided in the embodiment of the present application, the embodiment of the present application also provides a conceptual diagram of an image denoising method. The conceptual diagram of an image denoising method provided in the embodiment of the present application is as Figure 9 shown, where M can be 4 at this time.
[0129] As described in step 810, the target denoising processing model obtains a to-be-processed image.
[0130] After the target denoising processing model obtains the to-be-processed image, as described in step 815, the convolutional layer of the encoder unit extracts the features of the to-be-processed image to obtain a first feature map.
[0131] Then, as described in step 820, the first feature map is downsampled through the first downsampling layer to obtain a second feature map.
[0132] This cycle continues, and the ith feature map is downsampled through the ith downsampling layer to obtain the (i + 1)th feature map. When i = M - 1, as described in step 830, the (M - 1)th feature map is downsampled through the (M - 1)th downsampling layer to obtain the Mth feature map. Subsequently, as described in step 835, after obtaining the (M + 1)th feature map, that is, the first target feature map, it is input to the first upsampling layer.
[0133] After obtaining the first intermediate feature map through the upsampling process of the first upsampling layer as described in step 840, the j-th upsampling layer is used to perform upsampling processing on the (j - 1)-th intermediate feature map to obtain the j-th intermediate feature map; where M ≥ j > 1. At this time, j = 2, that is, the second upsampling layer is used to perform upsampling processing on the first intermediate feature map to obtain the second intermediate feature map. After obtaining the second intermediate feature map, j can be M - i + 1 and M in sequence, so as to obtain the second target feature map.
[0134] Finally, as described in step 860, the convolutional layer of the decoder unit is used to perform decoding operation on the second target feature map to obtain the target image, and the target image is the denoised predicted image.
[0135] Before using the target denoising processing model for inference to obtain the image to be processed, the image denoising method may further include the process of training the target denoising processing model. The training process of the target denoising processing model is explained below, and the training process of the target denoising processing model includes:
[0136] Obtain clear pictures with light and clear pictures without light;
[0137] Perform noise addition processing on the clear pictures with light to obtain the first noisy picture;
[0138] Perform darkening and noise addition processing on the clear pictures without light to obtain the second noisy picture;
[0139] Iteratively train the initial denoising processing model with the training sample pictures to obtain the trained target denoising processing model, where the training sample pictures include the first noisy picture and the second noisy picture.
[0140] In the embodiments of the present application, clear pictures can be classified into the well-lit clear pictures and the unlit clear pictures according to the lighting conditions on the pictures. The clear pictures (i.e., the well-lit clear pictures and the unlit clear pictures) are high-definition pictures containing any content. The clear pictures can be obtained by the target device receiving high-definition complete pictures taken by a device with a high-definition lens (such as a single-lens reflex camera), or after the target device receives high-definition complete pictures taken by a device with a high-definition lens (such as a single-lens reflex camera), the high-definition complete pictures are cut, and different parts of the high-definition pictures are used as separate complete high-definition pictures. To better restore the pixel and noise effects of the image to be processed during the inference process, when using clear pictures (i.e., the well-lit clear pictures and the unlit clear pictures) to train the initial noise reduction processing model, the clear pictures can be darkened (by performing division processing on the pixel data on the clear pictures) and then noise is added. To avoid the problem of color cast when darkening occurs after saturation truncation in a well-lit scene, the clear pictures are divided into well-lit clear pictures and unlit clear pictures. Noise is added to the well-lit clear pictures to obtain the first noise picture; the unlit clear pictures are darkened and noise is added to obtain the second noise picture.
[0141] To make the target noise reduction processing model perform better when processing low-light pictures, during the training process of the initial noise reduction processing model, more unlit clear pictures can be used. The picture data volume of the unlit clear pictures can be twice or more times that of the well-lit clear pictures, such as three times, four times, etc.
[0142] A specific process of processing well-lit clear pictures and unlit clear pictures into noisy training sample pictures can be as follows: When the clear picture is an unlit clear picture, first, the unlit clear picture is darkened by two times, and then noise is added, such as Gaussian and Poisson noise are added according to the parameter (k_sigma value) to obtain the second noise picture. When the clear picture is a well-lit clear picture, noise can be directly added, such as Gaussian and Poisson noise are added according to the parameter (k_sigma value) to obtain the first noise picture. After obtaining the first noise picture and / or the second noise picture, variance stabilization transformation (such as Freeman-Tukey variance stabilization transformation) can also be performed on the first noise picture and / or the second noise picture to stabilize the variance of the synthesized noise picture to the first threshold, and the first threshold can be any positive integer, such as 1, so as to obtain the training sample picture.
[0143] During each training process, a training sample image can be input into the initial noise reduction model. For example, a first noise image can be input into the initial noise reduction model, or a second noise image can be input into the initial noise reduction model. The initial noise reduction model performs prediction processing to obtain a predicted noise reduction result;
[0144] In the embodiment of the present application, the initial noise reduction model includes an encoder unit and a decoder unit; the encoder unit includes a convolutional layer and M downsampling layers; the decoder unit includes M upsampling layers and a convolutional layer; the convolutional layer, M downsampling layers, M upsampling layers of the encoder unit and the convolutional layer of the decoder unit are connected in sequence; the Mth downsampling layer includes a feature enhancement convolutional layer, a first edge enhancement convolutional layer, a guided filter layer, and a first target convolutional layer connected in sequence; the output of the first edge enhancement convolutional layer serves as the input of the guided filter layer and the first target convolutional layer. To better understand the Mth downsampling layer in the initial noise reduction model provided in the embodiment of the present application, an example is given below. It should be noted that the example is not restrictive. The Mth downsampling layer in the initial noise reduction model provided in the embodiment of the present application can be as Figure 10 shown. When the Mth downsampling layer obtains the input feature image from the (M - 1)th downsampling layer, the input feature image is input into the feature enhancement convolutional layer and the first convolutional layer simultaneously. Subsequently, the feature enhancement convolutional layer performs blind spot denoising. After the processing is completed, the output feature image is output to the first edge enhancement convolutional layer, and the first edge enhancement convolutional layer performs edge enhancement processing, and then the processing result is output to the guided filter layer and the first target convolutional layer. At this time, the guided filter layer can supervise and learn the output result of the first target convolutional layer to make the processing result of the first target convolutional layer more accurate. Finally, the output results of the first target convolutional layer and the first convolutional layer are added to obtain the output feature image of the Mth downsampling layer. In the embodiment of the present application, the first target convolutional layer and the first convolutional layer can be one convolutional layer or multiple convolutional layers, and the embodiment of the present application does not make specific limitations on this.
[0145] After obtaining the predicted noise reduction result through the initial noise reduction model, a loss function can be calculated based on the predicted noise reduction result for backpropagation to adjust the weights of each layer in the noise reduction model.
[0146] Based on the predicted noise reduction result and the annotation result of the training sample image, a first loss function is obtained;
[0147] In the embodiments of the present application, the annotation result of the training sample picture can be a high-definition picture of the training sample picture without adding noise (for example, if it is a picture in a picture set without light, the annotation result of the training sample picture can also be a high-definition picture without dimming processing). The first loss function can be the mean absolute error loss function (L1 loss). During the calculation of the mean absolute error loss function, the following function can be used:
[0148]
[0149] where n is the total number of samples, f(x i ) is the predicted value of the i-th sample, and y i is the true value of the i-th sample.
[0150] Obtain a second loss function based on the output result of the guided filter layer and the output result of the first target convolutional layer;
[0151] In the embodiments of the present application, the second loss function can be the same as the first loss function or different from the first loss function. The embodiments of the present application do not make specific limitations on this. For better understanding of the second loss function, an example is given below. As Figure 10 shown, the M-th downsampling layer of the initial noise reduction processing model used in the training process includes the guided filter layer and the first target convolutional layer. The result of the second loss function can be obtained based on the output result of the guided filter layer and the output result of the first target convolutional layer.
[0152] Perform weighted summation processing on the first loss function and the second loss function to obtain the target loss function.
[0153] In the embodiments of the present application, the training sample picture dataset is divided into a picture set with light and a picture set without light. Then, the pictures in the picture set without light are dimmed and then noise is added to try to restore the noise pictures generated during actual shooting (for example, night scene situations, etc.). Compared with the dataset processing method based on the clear-noise data pair in the prior art, the dataset processing method provided in the embodiments of the present application has a lower data acquisition cost. At the same time, the processing method of the training sample picture dataset provided in the embodiments of the present application can, to a certain extent, solve the problem of color cast that occurs after dimming when saturation truncation occurs in the light scene.
[0154] For better understanding of the training process of the initial noise reduction processing model provided in the embodiments of the present application, an example is given below. It should be noted that the example is not restrictive. The training process of the initial noise reduction processing model can be as Figure 11 shown, where M is 4 at this time. The initial noise reduction processing model can be referred to as Figure 2The encoder-decoder architecture shown.
[0155] First, obtain clear pictures with light and clear pictures without light; add Gaussian and Poisson noise to the clear pictures with light according to the parameter (k_sigma value) to obtain the first noisy picture; after darkening the clear pictures without light, then add Gaussian and Poisson noise according to the parameter (k_sigma value) to obtain the second noisy picture.
[0156] Perform variance stabilization transformation (such as Freeman-Tukey) on the first noisy picture and / or the second noisy picture to obtain the input of the initial noise reduction processing model network (i.e., the training sample picture). The variance stabilization transformation can stabilize the variance of the noisy image to a constant 1 to reduce the learning difficulty of the initial noise reduction processing model network. In each training process, a kind of training sample picture can be input into the initial noise reduction processing model network, and the training sample picture can come from the first noisy picture or the second noisy picture.
[0157] In the embodiment of the present application, in order to better adapt to mobile terminal devices, the convolution kernel design of the initial noise reduction processing network model can be carried out according to the design criteria of the Qualcomm Tensor Processor (HTP for short), for example, the number of channels of the feature image is preferably designed to be a multiple of 32, the convolution on the high-resolution scale is minimized, the number of branches is reduced, and the use of Tightly Coupled Memories (TCM) is reduced, etc.
[0158] The convolutional layer of the encoder unit can be a 3x3 convolution with a stride of 1, and the number of its convolution kernels is 32. The multiple of 32 channels in the Qualcomm Tensor Processor (HTP for short) is the optimal choice, and the convolution on the high resolution will increase the burden on the Tightly Coupled Memories (TCM), so the convolutional layer of the encoder unit adopts a 32-channel convolution. Input the training sample picture into the convolutional layer of the encoder unit
[0159] The first upsampling layer can be two 3x3 convolutions, and the number of convolution kernels is 32 for both. The stride of the first convolution is 2, and the stride of the second convolution is 1. The training sample picture is quickly downsampled to 1 / 2 of the original resolution, reducing the calculation time-consuming and extracting important features.
[0160] The second upsampling layer can be three 3x3 convolutions. The number of convolution kernels of the first two convolutions is 32, and the number of convolution kernels of the last convolution is 64. The stride of the first convolution is 2, and the strides of the last two convolutions are 1. The training sample picture is downsampled to 1 / 4 of the original resolution.
[0161] The third downsampling layer includes a detail enhancement module, which is mainly a convolutional kernel for edge feature extraction. The third downsampling layer can be a two-branch structure, and the two-branch structure is to ensure that the feature information is reduced after enhancement to reduce the loss of feature information. The first branch contains three convolutional layers. The first convolutional layer can be a 5x5 convolution with 64 convolutional kernels and a stride of 2. The second convolutional layer can be an Angular Pixel Difference Convolution (APDC layer) to enhance the image gradient features and thus enhance the edge features. The convolutional kernel size can be 3x3, the number of convolutional kernels is 64, and the stride is 1. The third convolutional layer can be a 3x3 convolution with 64 convolutional kernels and a stride of 1. The second branch is a single 3x3 convolution with 64 convolutional kernels and a stride of 2. Finally, the output results of the two-branch network are added to obtain the output features of the third downsampling layer, and the third downsampling layer downsamples the image to 1 / 8 of the original resolution.
[0162] The fourth downsampling layer includes a detail enhancement module, and the fourth downsampling layer is a two-branch structure. The first branch contains three convolutional layers. The first convolution is a feature enhancement convolutional module, which can make the pixel points further satisfy the assumption of independent and identical distribution and can further generate the central pixel value to reduce the influence of noise on the output feature map. The number of its convolutional kernels is 128, the stride is 2, and the convolutional kernel size is 5x5. The second convolution is an Angular Pixel Difference Convolution (APDC layer), which further enhances the features of the previous convolution. The convolutional kernel size is 3x3, the number of convolutional kernels is 128, and the stride is 1. And a guided filter branch is added after the second convolutional layer, and this branch is used as supervision information to alleviate the influence of the noise in the feature map on the final output result of the network (the guidance map is the input map, which can be used as an edge-preserving filter, and its window size value can be set to 5, and the parameter (e.g., ε) can be set to 0.1). The output result of this guided filter branch is used to calculate the first loss with the output of the third convolutional layer. This guided filter module can further supervise the production of important features, alleviate the influence of noise in the feature map, and thus further enhance the important details. The third convolution is a normal convolution with a convolutional kernel size of 3x3 and a stride of 1. The other branch is a single normal convolution with 128 convolutional kernels, a convolutional kernel size of 3x3, and a stride of 2. Finally, the output results of the two-branch network are added to obtain the output feature image of the fourth downsampling layer, and the fourth downsampling layer downsamples the image resolution to 1 / 16 of the original image resolution.
[0163] After the downsampling process of the encoding layer, after compressing the resolution of the training sample image, the training sample image can be input into the decoding layer for upsampling processing.
[0164] The first upsampling layer can be a two-branch network. The first branch can have 3 convolutional layers. The first layer is a 5x5 feature enhancement convolutional module with 128 convolutional kernels and a stride of 1. The second and third layers are 3x3 ordinary convolutional layers with 128 convolutional kernels and a stride of 1. The other branch is a single ordinary convolutional layer with 128 convolutional kernels, a convolutional kernel size of 3x3, and a stride of 1. Finally, the output results of the two-branch network are added and input into a transposed convolution with a stride of 2 to upsample the image, and finally output a feature image with a resolution of 1 / 8 of the training sample image.
[0165] The second upsampling layer can be a two-branch network. Its input is the sum of the output of the first upsampling layer and the output of the third downsampling layer. The first branch includes 3 ordinary convolutional layers. The first layer is a 5x5 ordinary convolutional layer with 64 convolutional kernels and a stride of 1. The second and third layers are 3x3 ordinary convolutional layers with 64 convolutional kernels each and a stride of 1. The other branch is a single ordinary convolutional layer with 64 convolutional kernels, a convolutional kernel size of 3x3, and a stride of 1. Finally, the output results of the two-branch network are added and input into a transposed convolution with a stride of 2 to upsample the image, and finally output a feature image with a resolution of 1 / 4 of the training sample image.
[0166] The third upsampling layer can be a residual structure. Its input is the sum of the output of the second upsampling layer and the output of the second downsampling layer, and it contains a total of three convolutions, all with a kernel_size equal to 3x3, 64 convolutional kernels, and a stride of 1. The output of this residual module is input into a transposed convolution with a stride of 2 to upsample the image, and finally output a feature image with a resolution of 1 / 2 of the training sample image.
[0167] The fourth downsampling layer can be two convolutional layers plus a transposed convolution. Its input is the sum of the output of the third upsampling layer and the output of the first downsampling layer. The convolutional kernel sizes of the two convolutional layers are both 3x3, with 32 convolutional kernels and a stride of 1. Finally, the output of the two convolutional layers is input into a transposed convolution with a stride of 2, and finally output a feature image with the original resolution of the training sample image.
[0168] The last layer of the decoder is a convolutional layer. Its input is the sum of the output of the fourth downsampling layer and the output of the first convolutional layer of the encoder unit. Its convolutional kernel size is 3x3, with 4 convolutional kernels and a stride of 1. Finally, it outputs the output feature image of the training sample image processed by the initial noise reduction processing model.
[0169] Since the training sample images undergo a variance stabilization transformation (e.g., Freeman-Tukey transformation) when input into the initial noise reduction processing model network, an inverse variance stabilization transformation (e.g., Freeman-Tukey inverse transformation) can be performed on the output feature map of the initial noise reduction processing model network after the initial noise reduction processing model network outputs the feature map to obtain the denoised training sample images.
[0170] The loss function of the initial noise reduction processing model network is the weighted sum of a first loss function and a second loss function. The weight value of the first loss function can be 0.1, and the weight value of the second loss function can be 1. The embodiments of the present application do not make specific limitations on this. The first loss function can be the mean absolute error loss function (L1 loss) between the output result of the guided filter in the M-th downsampling layer and the output of the third convolutional layer of the first branch in the M-th downsampling layer (i.e., Figure 10 the first target convolutional layer shown in). The second loss function can be the mean absolute error loss function (L1 loss) between the output feature image of the training sample image processed by the initial noise reduction processing model and the image without added noise in the training sample picture dataset.
[0171] Regarding the production problem of the training sample picture dataset for the initial noise reduction processing model provided by the embodiments of the present application, since the pixel values of the original unprocessed image file format (e.g., RAW format) are affected by factors such as the amount of incident light, gain, and exposure time, the exposure time is relatively long during the acquisition of the high-definition picture dataset in the training process, and its pixel values are generally larger than those of the images to be processed in the actual inference process. Therefore, the training sample picture dataset is prone to domain shift. In the training process of the initial noise reduction processing model provided by the embodiments of the present application, the training sample picture dataset is divided into a set of pictures with light and a set of pictures without light. The pictures in the set of pictures without light are darkened (e.g., darkened by a factor of two). And for better training effects, the number of pictures in the set of pictures without light in the training sample picture dataset provided by the embodiments of the present application is preferably more than twice the number of pictures in the set of pictures with light.
[0172] When the initial noise reduction processing model is trained and the target noise reduction processing model is obtained, it can enter as Figure 1 、 Figure 3 、 Figure 8 or Figure 9The actual reasoning process shown. The target noise reduction processing model can be used on the target device. The target noise reduction processing model provided by the embodiments of the present application can greatly improve the processing speed on mobile terminal devices while enhancing the details of the image to be processed. The backbone network part of the target noise reduction processing model provided by the embodiments of the present application can be based on a symmetric network structure (for example, the unet structure), and according to the design criteria of the Qualcomm Tensor Processor (HTP for short), for example, the number of channels of the feature image is preferably designed to be a multiple of 32, the convolution on the high-resolution scale is minimized, the number of branches is reduced, and the use of Tightly Coupled Memories (TCM for short) is reduced. A new network structure is proposed to better utilize the parallel processing mechanism of the Qualcomm Tensor Processor (HTP for short) and improve the running speed of the target noise reduction processing network on mobile terminal devices. After measurement, the running time of the target noise reduction processing model on mobile terminal devices can be about 16 ms, and the inference time on terminal devices can be 27 ms, which shows the effectiveness of the target noise reduction processing model designed for the Qualcomm Tensor Processor (HTP for short). Finally, for the detail enhancement part, the target noise reduction processing model provided by the embodiments of the present application proposes a new convolution kernel design method and a detail enhancement module to enhance network details and improve the network loss function, further enhancing the noise reduction result and details.
[0173] Figure 12 is a structural block diagram of an image noise reduction device provided by an embodiment of the present application, as Figure 12 shown, the image noise reduction device 1200 provided by the embodiment of the present application includes an acquisition module 1210 and a processing module 1220.
[0174] The acquisition module 1210 is configured to acquire an image to be processed, and the image to be processed is a noisy image;
[0175] The processing module 1220 is configured to extract features of the image to be processed and perform downsampling processing to obtain a first target feature map; perform upsampling processing on the first target feature map to obtain a second target feature map; and use the decoder unit of the target noise reduction processing model to perform decoding operations on the second target feature map to obtain a target image, where the target image is a noise reduction image predicted by inputting the image to be processed into a target noise reduction processing model including an encoder unit and a decoder unit.
[0176] The image denoising device provided in the embodiment of the present application can obtain an image to be processed, wherein the image to be processed is a noisy image; extract the features of the image to be processed and perform downsampling processing to obtain a first target feature map; perform upsampling processing on the first target feature map to obtain a second target feature map; use the decoder unit of the target denoising processing model to perform a decoding operation on the second target feature map to obtain a target image; wherein the target image is a denoised image predicted by inputting the image to be processed into a target denoising processing model including an encoder unit and a decoder unit. In this way, the target image after denoising processing by the target denoising processing model can clearly represent the features of the image to be processed, achieve a good denoising effect, and solve the problem of poor denoising effect in related technologies.
[0177] It should be noted that the embodiments of the image denoising device in this specification and the embodiments of the image denoising method in this specification are based on the same inventive concept, so the specific implementation of this embodiment can refer to the implementation of the corresponding image denoising method in the previous text, and the repeated parts will not be repeated.
[0178] In addition, if Figure 13 As shown, the embodiment of the present application further provides an electronic device 1300, which may be various types of computers, etc. The electronic device 1300 includes: a processor 1310 and a memory 1320, the memory 1320 stores programs or instructions, and when the programs or instructions are executed by the processor 610, the steps of any of the methods described above are implemented, such as Figure 1 , Figure 3 , Figure 8 , Figure 9 as well as Figure 11 The steps of the image denoising method are shown in . The target image after denoising by the target denoising processing model can clearly represent the features of the image to be processed, achieve a good denoising effect, and solve the problem of poor denoising effect in related technologies.
[0179] The embodiment of the present application also provides a readable storage medium, which stores a program or instruction, and when the program or instruction is executed by the processor 1310, the steps of any method described above are implemented. The target image after denoising by the target denoising processing model can clearly represent the features of the image to be processed, achieve a good denoising effect, and solve the problem of poor denoising effect in the related art.
[0180] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.
[0181] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0182] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0183] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0184] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0185] The memory may include non-permanent memory in the computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0186] A computer-readable medium includes permanent and non-permanent, removable and non-removable media that can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information accessible by a computing device. As defined herein, a computer-readable medium does not include transitory computer-readable media, such as modulated data signals and carrier waves.
[0187] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0188] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, system or computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.
[0189] The above are only the embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. An image denoising method, characterized in that, it includes: Obtain an image to be processed, where the image to be processed is a noisy image; Extract the features of the image to be processed and perform downsampling to obtain a first target feature map; Perform upsampling on the first target feature map to obtain a second target feature map; Use the decoder unit of the target denoising processing model to perform decoding operations on the second target feature map to obtain a target image; Wherein, the target image is a denoised image obtained by inputting the image to be processed into a target denoising processing model including an encoder unit and a decoder unit for prediction.
2. The method according to claim 1, characterized in that, The encoder unit includes a convolutional layer and M downsampling layers; the decoder unit includes M upsampling layers and a convolutional layer; where M≥2; The extracting the features of the image to be processed and performing downsampling to obtain a first target feature map includes: Extract the features of the image to be processed through the convolutional layer of the encoder unit to obtain a first feature map; Perform downsampling on the first feature map through the M downsampling layers to obtain a first target feature map; The performing upsampling on the first target feature map to obtain a second target feature map includes: Perform upsampling on the first target feature map through the M upsampling layers to obtain a second target feature map; The using the decoder unit of the target denoising processing model to perform decoding operations on the second target feature map to obtain a target image includes: Perform decoding operations on the second target feature map through the convolutional layer of the decoder unit to obtain a target image.
3. The method according to claim 2, characterized in that, The performing downsampling on the first feature map through the M downsampling layers to obtain a first target feature map includes: Perform downsampling on the first feature map through the first downsampling layer to obtain a second feature map; Perform downsampling on the i-th feature map through the i-th downsampling layer to obtain an (i + 1)-th feature map; Perform downsampling on the M-th feature map through the M-th downsampling layer to obtain an (M + 1)-th feature map, and use the (M + 1)-th feature map as the first target feature map; Where M>i>1.
4. The method according to claim 3, characterized in that, The M-th downsampling layer includes: a feature enhancement convolutional layer; the performing downsampling on the M-th feature map through the M-th downsampling layer to obtain an (M + 1)-th feature map includes: Input the M-th feature map into the feature enhancement convolutional layer; Perform blind spot denoising enhancement processing and pixel independence feature enhancement processing on the M-th feature map through the target convolutional kernel of the feature enhancement convolutional layer to obtain a first processed feature map; Based on the first processed feature map, obtain the (M + 1)-th feature map; Wherein, the target convolutional kernel is of size AxA, where A is an odd number greater than or equal to 3, the convolution values of the even rows and even columns of the target convolutional kernel are all 0, and the convolution value at the center position is also 0; the stride of the target convolutional kernel is an integer greater than 1, and the resolution of the first processed feature map is smaller than that of the M-th feature map.
5. The method according to claim 4, wherein, the M-th downsampling layer further includes: a first edge enhancement convolutional layer; obtaining the (M + 1)-th feature map based on the first processed feature map includes: inputting the first processed feature map into the first edge enhancement convolutional layer for image gradient feature enhancement processing to obtain a second processed feature map; obtaining the (M + 1)-th feature map based on the second processed feature map.
6. The method according to claim 3, wherein, when i = M - 1, the (M - 1)-th downsampling layer includes: a third convolutional layer and a second edge enhancement convolutional layer; performing downsampling processing on the (M - 1)-th feature map through the (M - 1)-th downsampling layer to obtain the M-th feature map, including: inputting the (M - 1)-th feature map into the third convolutional layer for feature extraction processing to obtain a first specified feature map; wherein, the stride of the convolutional kernel of the third convolutional layer is an integer greater than 1, and the resolution of the first specified feature map is less than the resolution of the (M - 1)-th feature map; inputting the first specified feature map into the second edge enhancement convolutional layer for image gradient feature enhancement processing to obtain a second specified feature map; obtaining the M-th feature map based on the second specified feature map.
7. The method according to any one of claims 2 - 6, wherein, performing upsampling processing on the first target feature map through the M upsampling layers to obtain a second target feature map includes: performing upsampling processing on the first target feature map through the first upsampling layer to obtain a first intermediate feature map; performing upsampling processing on the (j - 1)-th intermediate feature map through the j-th upsampling layer to obtain the j-th intermediate feature map; wherein, M ≥ j > 1; when j = M, the M-th intermediate feature map is the second target feature map.
8. The method according to claim 7, wherein, the first upsampling layer includes a first branch and a second branch, and there is at least one convolutional layer in each branch; performing upsampling processing on the first target feature map through the first upsampling layer to obtain a first intermediate feature map includes: adding the output results of the first branch and the second branch to obtain a first target output result; inputting the first target output result into a first transposed convolutional layer for image upsampling processing to obtain a first intermediate feature map; wherein, the stride of the convolutional kernel of the first transposed convolutional layer is an integer greater than 1, and the resolution of the first intermediate feature map is greater than the resolution of the first target feature map.
9. The method according to claim 7, wherein, performing upsampling processing on the (j - 1)-th intermediate feature map through the j-th upsampling layer to obtain the j-th intermediate feature map includes: performing upsampling processing on the (j - 1)-th intermediate feature map through the j-th upsampling layer to obtain an output feature map; adding the output feature map and the output result of the (M - j + 1)-th downsampling layer to obtain a target output feature map; inputting the target output feature map into the j-th transposed convolutional layer for image upsampling processing to obtain the j-th intermediate feature map; Among them, the stride of the convolution kernel of the j-th transposed convolution layer is an integer greater than 1, and the resolution of the j-th intermediate feature map is greater than that of the (j - 1)-th intermediate feature map.
10. The method according to claim 1, wherein, before obtaining the image to be processed, the method further includes: obtaining a clear picture with light and a clear picture without light; performing noise addition processing on the clear picture with light to obtain a first noise picture; performing darkening and noise addition processing on the clear picture without light to obtain a second noise picture; iteratively training an initial noise reduction processing model with training sample pictures to obtain a trained target noise reduction processing model, wherein the training sample pictures include the first noise picture and the second noise picture.
11. The method according to claim 10, wherein, the iteratively training the initial noise reduction processing model with the training sample pictures includes: inputting the training sample pictures into the initial noise reduction processing model for prediction processing to obtain a predicted noise reduction result; obtaining a target loss function based on the predicted noise reduction result and the annotation result of the training sample pictures; determining the loss value of the training sample pictures through the target loss function; iteratively training the initial noise reduction processing model based on the loss value.
12. The method according to claim 11, wherein, the initial noise reduction processing model includes an encoder unit and a decoder unit; the encoder unit includes a convolution layer and M downsampling layers; the decoder unit includes M upsampling layers and a convolution layer; the convolution layer, M downsampling layers, M upsampling layers of the encoder unit and the convolution layer of the decoder unit are connected in sequence; the M-th downsampling layer includes a feature enhancement convolution layer, a first edge enhancement convolution layer, a guided filter layer and a first target convolution layer connected in sequence; the output of the edge enhancement convolution layer is used as the input of the guided filter layer and the first target convolution layer; the obtaining a target loss function based on the predicted noise reduction result and the annotation result of the training sample pictures includes: obtaining a first loss function based on the predicted noise reduction result and the annotation result of the training sample pictures; obtaining a second loss function based on the output result of the guided filter layer and the output result of the first target convolution layer; performing weighted summation processing on the first loss function and the second loss function to obtain a target loss function.
13. An electronic device, wherein, it includes a processor and a memory, the memory stores a program or instruction running on the processor, and when the program or instruction is executed by the processor, the steps of the method according to any one of claims 1 - 12 are implemented.
14. A computer-readable storage medium, wherein, the computer-readable storage medium stores a program or instruction, and when the program or instruction is executed by a processor, the steps of the method according to any one of claims 1 - 12 are implemented.