Image Denoising Processing Method, Device, Equipment, Storage Medium and Program Product
A lightweight image denoising model for ISP chips processes all channels simultaneously, addressing the challenge of achieving effective denoising and real-time performance by optimizing the cascaded structure and reducing computational load.
Patent Information
- Application Number
- CN202211128666.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-16
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-09-16
AI Technical Summary
The existing image noise reduction algorithm cannot simultaneously have excellent noise reduction effect and meet the real-time requirements of ISP chips.
The image noise reduction model using a cascaded downsampling model and an upsampling model include n cascaded downsampling modules and n cascaded upsampling modules. Combined with the fusion module, the pixel values of each channel of the target image are directly processed, improving data processing efficiency and real-time noise reduction.
While meeting the real-time requirements of ISP chips, it significantly improves the image noise reduction effect, reduces the amount of data and calculation, and improves the image signal-to-noise ratio and processing efficiency.
Smart Images

Figure CN115471417B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular, to an image denoising processing method, apparatus, device, storage medium, and program product. Background Art
[0002] Image denoising technology is a key task in the field of image processing. The ISP (Image Signal Processor) chip is mainly used for image processing of real-time video images captured by a terminal. In terms of image denoising processing, it has high requirements for the real-time performance of the image denoising algorithm.
[0003] However, the current image denoising algorithms cannot simultaneously have excellent denoising effects and meet the real-time requirements of the ISP chip. Summary of the Invention
[0004] Based on this, it is necessary to provide an image denoising processing method, apparatus, device, storage medium, and program product that can meet the real-time requirements of the ISP chip and is used for denoising processing of video images in view of the above technical problems.
[0005] In a first aspect, this application provides an image denoising processing method. The method includes:
[0006] Inputting target image data into an image denoising model to obtain denoised image data output by the image denoising model, where the target image data includes pixel values of each channel of the target image;
[0007] Wherein, the image denoising model includes a cascaded downsampling model, an upsampling model, and an output layer. The downsampling model includes n cascaded downsampling modules, and the upsampling model includes n cascaded upsampling modules corresponding one-to-one to the n downsampling modules; the downsampling module includes a first downsampling module, a second downsampling module, and a fusion module cascaded with both the first downsampling module and the second downsampling module; the first downsampling module includes a cascaded first downsampling layer and a first convolutional layer, and the second downsampling module includes a second downsampling layer.
[0008] In one embodiment, the inputting target image data into an image denoising model to obtain denoised image data output by the image denoising model includes:
[0009] Input the target image data into the downsampling model, and each downsampling module in the downsampling model performs downsampling processing on the target image data to obtain downsampled feature data; input the downsampled feature data into the upsampling model, and each upsampling module in the upsampling model performs upsampling processing on the downsampled feature data to obtain upsampled feature data; the output layer obtains the denoised image data based on the upsampled feature data and the target image data.
[0010] In one embodiment, the image data resolutions of each channel of the target image are the same. The process of each downsampling module in the downsampling model performing downsampling processing on the target image data to obtain downsampled feature data includes:
[0011] For the i-th downsampling module, perform downsampling processing on the input data of the i-th downsampling module to obtain the intermediate downsampled feature data output by the i-th downsampling module; wherein, when i = 1, the input data of the i-th downsampling module is the target image data, and when i > 1, the input data of the i-th downsampling module is the intermediate downsampled feature data output by the (i - 1)-th downsampling module; use the intermediate downsampled feature data output by the last downsampling module as the downsampled feature data.
[0012] In one embodiment, the process of each upsampling module in the upsampling model performing upsampling processing on the downsampled feature data to obtain upsampled feature data includes:
[0013] For the i-th upsampling module, perform upsampling processing on the input data of the i-th upsampling module to obtain the intermediate upsampled feature data output by the i-th upsampling module; wherein, when i = 1, the input data of the i-th upsampling module is the downsampled feature data, and when i > 1, the input data of the i-th upsampling module is the aggregated feature data obtained by fusing the intermediate upsampled feature data output by the (i - 1)-th upsampling module and the intermediate downsampled feature data output by the corresponding downsampling module of the i-th upsampling module; use the intermediate upsampled feature data output by the last upsampling module as the upsampled feature data.
[0014] In one embodiment, the process of the output layer obtaining the denoised image data based on the upsampled feature data and the target image data includes:
[0015] Input the upsampled feature data and the target image data into the output layer for fusion processing to obtain the denoised image data output by the output layer.
[0016] In one embodiment, the resolution of the image data of each channel of the target image is different. The downsampling model further includes an additional downsampling module. The target image data is input into the downsampling model, and each downsampling module in the downsampling model performs downsampling processing on the target image data to obtain downsampling feature data, including:
[0017] The pixel values of the first channel of the target image included in the target image data are input into the additional downsampling module to obtain channel feature data output by the additional downsampling module; the channel feature data is fused with the pixel values of the second channel of the target image included in the target image data to obtain candidate target image data; for the i-th downsampling module, the input data of the i-th downsampling module is downsampled to obtain intermediate downsampling feature data output by the i-th downsampling module; wherein, when i = 1, the input data of the i-th downsampling module is the candidate target image data, and when i > 1, the input data of the i-th downsampling module is the intermediate downsampling feature data output by the (i - 1)-th downsampling module; the intermediate downsampling feature data output by the last downsampling module is used as the downsampling feature data.
[0018] In one embodiment, the upsampling model further includes an additional upsampling module. The downsampling feature data is input into the upsampling model, and each upsampling module in the upsampling model performs upsampling processing on the downsampling feature data to obtain upsampling feature data, including:
[0019] For the i-th upsampling module, the input data of the i-th upsampling module is upsampled to obtain intermediate upsampling feature data output by the i-th upsampling module; wherein, when i = 1, the input data of the i-th upsampling module is the downsampling feature data, and when i > 1, the input data of the i-th upsampling module is the aggregated feature data obtained by fusing the intermediate upsampling feature data output by the (i - 1)-th upsampling module and the intermediate downsampling feature data output by the downsampling module corresponding to the i-th upsampling module; the first intermediate channel feature data corresponding to the pixel values of the first channel included in the intermediate upsampling feature data output by the last upsampling module is input into the additional upsampling module to obtain the upsampling feature data output by the additional upsampling module.
[0020] In one embodiment, the output layer obtains the denoised image data based on the upsampling feature data and the target image data, including:
[0021] Input the upsampled feature data and the pixel values of the first channel in the target image data into the output layer for fusion processing to obtain candidate denoised image data output by the output layer; obtain the denoised image data according to the candidate denoised image data and the second intermediate channel feature data corresponding to the pixel values of the second channel included in the intermediate upsampled feature data output by the last upsampling module.
[0022] In one embodiment, performing downsampling processing on the input data of the i-th downsampling module to obtain intermediate downsampled feature data output by the i-th downsampling module includes:
[0023] Using the first downsampling layer to perform downsampling processing on the input data of the i-th downsampling module to obtain first downsampled feature data output by the first downsampling layer; using the first convolutional layer to perform convolutional processing on the first downsampled feature data to obtain first convolutional feature data output by the first convolutional layer; using the second downsampling layer to perform downsampling processing on the input data of the i-th downsampling module to obtain second downsampled feature data output by the second downsampling layer; using the fusion module to perform fusion processing on the first convolutional feature data and the second downsampled feature data to obtain the intermediate downsampled feature data output by the fusion module.
[0024] In one embodiment, the upsampling module includes a cascaded second convolutional layer and an upsampling layer; performing upsampling processing on the input data of the i-th upsampling module to obtain intermediate upsampled feature data output by the i-th upsampling module includes:
[0025] Using the second convolutional layer to perform convolutional processing on the input data of the i-th upsampling module to obtain second convolutional feature data output by the second convolutional layer; using the upsampling layer to perform upsampling processing on the second convolutional feature data to obtain the intermediate upsampled feature data output by the upsampling layer.
[0026] In one embodiment, the image denoising model is used in a RAW image denoising module, an RGB image denoising module, or a YUV image denoising module in an ISP chip; correspondingly, the format of the target image is RAW format, RGB format, or YUV format.
[0027] In one embodiment, the upsampling layer performs upsampling processing on the input data of the upsampling layer through convolutional processing, anti-pooling processing, or interpolation processing.
[0028] In a second aspect, the present application also provides an image denoising processing device. The device includes:
[0029] A noise reduction module, which is used to input target image data into an image denoising model to obtain denoised image data output by the image denoising model. The target image data includes pixel values of each channel of the target image;
[0030] Wherein, the image denoising model includes a cascaded downsampling model, an upsampling model and an output layer. The downsampling model includes n cascaded downsampling modules, and the upsampling model includes n cascaded upsampling modules corresponding one by one to the n downsampling modules; The downsampling module includes a first downsampling module, a second downsampling module and a fusion module cascaded with both the first downsampling module and the second downsampling module; The first downsampling module includes a cascaded first downsampling layer and a first convolutional layer, and the second downsampling module includes a second downsampling layer.
[0031] In one embodiment, the noise reduction module is specifically configured to:
[0032] Input the target image data into the downsampling model, and each downsampling module in the downsampling model performs downsampling processing on the target image data to obtain downsampled feature data; Input the downsampled feature data into the upsampling model, and each upsampling module in the upsampling model performs upsampling processing on the downsampled feature data to obtain upsampled feature data; The output layer obtains the denoised image data based on the upsampled feature data and the target image data.
[0033] In one embodiment, the image data resolutions of each channel of the target image are the same. The noise reduction module is specifically configured to:
[0034] For the i-th downsampling module, perform downsampling processing on the input data of the i-th downsampling module to obtain intermediate downsampled feature data output by the i-th downsampling module; Wherein, when i = 1, the input data of the i-th downsampling module is the target image data, and when i > 1, the input data of the i-th downsampling module is the intermediate downsampled feature data output by the (i - 1)-th downsampling module; Use the intermediate downsampled feature data output by the last downsampling module as the downsampled feature data.
[0035] In one embodiment, the noise reduction module is specifically configured to:
[0036] For the i-th upsampling module, perform upsampling processing on the input data of the i-th upsampling module to obtain the intermediate upsampling feature data output by the i-th upsampling module. Among them, when i = 1, the input data of the i-th upsampling module is the downsampling feature data. When i > 1, the input data of the i-th upsampling module is the aggregated feature data obtained by fusing the intermediate upsampling feature data output by the (i - 1)-th upsampling module and the intermediate downsampling feature data output by the downsampling module corresponding to the i-th upsampling module. Use the intermediate upsampling feature data output by the last upsampling module as the upsampling feature data.
[0037] In one embodiment, the noise reduction module is specifically configured to:
[0038] Input the upsampling feature data and the target image data into the output layer for fusion processing to obtain the denoised image data output by the output layer.
[0039] In one embodiment, the resolution of the image data of each channel of the target image is different. The downsampling model further includes an additional downsampling module. The noise reduction module is specifically configured to:
[0040] Input the pixel values of the first channel of the target image included in the target image data into the additional downsampling module to obtain the channel feature data output by the additional downsampling module. Perform fusion processing on the channel feature data and the pixel values of the second channel of the target image included in the target image data to obtain candidate target image data. For the i-th downsampling module, perform downsampling processing on the input data of the i-th downsampling module to obtain the intermediate downsampling feature data output by the i-th downsampling module. Among them, when i = 1, the input data of the i-th downsampling module is the candidate target image data. When i > 1, the input data of the i-th downsampling module is the intermediate downsampling feature data output by the (i - 1)-th downsampling module. Use the intermediate downsampling feature data output by the last downsampling module as the downsampling feature data.
[0041] In one embodiment, the upsampling model further includes an additional upsampling module. The noise reduction module is specifically configured to:
[0042] For the i-th upsampling module, perform upsampling processing on the input data of the i-th upsampling module to obtain the intermediate upsampling feature data output by the i-th upsampling module; wherein, when i = 1, the input data of the i-th upsampling module is the downsampling feature data, and when i > 1, the input data of the i-th upsampling module is the aggregated feature data obtained by fusing the intermediate upsampling feature data output by the (i - 1)-th upsampling module and the intermediate downsampling feature data output by the downsampling module corresponding to the i-th upsampling module; input the first intermediate channel feature data corresponding to the first channel pixel value included in the intermediate upsampling feature data output by the last upsampling module into the additional upsampling module to obtain the upsampling feature data output by the additional upsampling module.
[0043] In one embodiment, the noise reduction module is specifically configured to:
[0044] Input the upsampling feature data and the first channel pixel value in the target image data into the output layer for fusion processing to obtain the candidate noise reduction image data output by the output layer; obtain the noise reduction image data according to the candidate noise reduction image data and the second intermediate channel feature data corresponding to the second channel pixel value included in the intermediate upsampling feature data output by the last upsampling module.
[0045] In one embodiment, the noise reduction module is specifically configured to:
[0046] Use the first downsampling layer to perform downsampling processing on the input data of the i-th downsampling module to obtain the first downsampling feature data output by the first downsampling layer; use the first convolutional layer to perform convolutional processing on the first downsampling feature data to obtain the first convolutional feature data output by the first convolutional layer; use the second downsampling layer to perform downsampling processing on the input data of the i-th downsampling module to obtain the second downsampling feature data output by the second downsampling layer; use the fusion module to perform fusion processing on the first convolutional feature data and the second downsampling feature data to obtain the intermediate downsampling feature data output by the fusion module.
[0047] In one embodiment, the upsampling module includes a cascaded second convolutional layer and an upsampling layer; the noise reduction module is specifically configured to:
[0048] Use the second convolutional layer to perform convolutional processing on the input data of the i-th upsampling module to obtain the second convolutional feature data output by the second convolutional layer; use the upsampling layer to perform upsampling processing on the second convolutional feature data to obtain the intermediate upsampling feature data output by the upsampling layer.
[0049] In one embodiment, the image denoising model is used in a RAW image denoising module, an RGB image denoising module, or a YUV image denoising module in an ISP chip; correspondingly, the format of the target image is RAW format, RGB format, or YUV format.
[0050] In one embodiment, the upsampling layer performs upsampling processing on the input data of the upsampling layer through convolution processing, de-pooling processing, or interpolation processing.
[0051] In a third aspect, the present application further provides an electronic device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the method described in any one of the first aspects above are implemented.
[0052] In a fourth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method described in any one of the first aspects above are implemented.
[0053] In a fifth aspect, the present application further provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the method described in any one of the first aspects above are implemented.
[0054] The above image denoising method, device, equipment, storage medium and program product can directly input the target image data including the pixel values of each channel of the target image into the image denoising model, and then obtain the denoised image data output by the image denoising model, so as to realize the denoising processing of the target image. Among them, usually in the ISP chip, it is necessary to perform denoising processing based on the pixel values of the Y channel and the UV channel of the YUV image respectively to obtain the image data after denoising processing. Since the Y channel and the UV channel are processed simultaneously, it is necessary to repeatedly call the image data, and its processing efficiency is poor, which cannot meet the real-time requirements of the ISP chip. In this application, since the target image data including the pixel values of each channel of the target image can be directly input into the image denoising model for denoising processing, that is, the data of each channel of the target image is denoised simultaneously. Compared with denoising each channel separately, the data volume and calculation volume are greatly reduced, which can effectively improve the data processing efficiency and meet the real-time requirements; moreover, during the denoising processing, the information between each channel in the target image can be mutually referred to achieve a better denoising effect. Among them, the network structure of the image denoising model is streamlined, which includes a cascaded downsampling model, an upsampling model and an output layer. The downsampling model includes n cascaded downsampling modules, and the upsampling model includes n cascaded upsampling modules corresponding to the n downsampling modules one by one; the downsampling module includes a first downsampling module, a second downsampling module and a fusion module cascaded with both the first downsampling module and the second downsampling module; the first downsampling module includes a cascaded first downsampling layer and a first convolutional layer, and the second downsampling module includes a second downsampling layer. Through the streamlined image denoising model, the effective denoising processing of the target image data can be realized, so that the image denoising model can be fully adapted to the ISP chip for denoising real-time video images. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 is a schematic structural diagram of an image denoising model in an embodiment;
[0056] Figure 2 is a schematic structural diagram of a multi-convolution parallel module in an embodiment;
[0057] Figure 3 is a schematic diagram of the traditional ISP chip image processing flow in an embodiment;
[0058] Figure 4 is a schematic diagram of the first improved ISP chip image processing flow in an embodiment;
[0059] Figure 5 is a schematic diagram of the second improved ISP chip image processing flow in an embodiment;
[0060] Figure 6 Schematic diagram of the noise reduction process in an embodiment
[0061] Figure 7 Schematic diagram of the structure of a noise reduction neural network in an embodiment
[0062] Figure 8 Schematic diagram of the structure of another image noise reduction model in an embodiment
[0063] Figure 9 Schematic diagram of the structure of another noise reduction neural network in an embodiment
[0064] Figure 10 Schematic diagram of the element-wise addition fusion process in an embodiment
[0065] Figure 11 Schematic diagram of the channel-wise concatenation fusion process in an embodiment
[0066] Figure 12 Block diagram of the structure of an image noise reduction processing device in an embodiment
[0067] Figure 13 Internal structure diagram of a computer device in an embodiment Detailed implementation manners
[0068] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0069] In the field of image processing, image noise reduction has always been an image processing task that is difficult to perfectly handle. Image noise reduction is a type of image restoration technology, and its purpose is to accurately find the signal values or noise values in the image, or to separate the signal part and the noise part in the image.
[0070] Currently, image noise reduction algorithms are mainly divided into traditional algorithms and neural network-based algorithms. The current basic situation is that the noise reduction effect of traditional algorithms is poor and cannot meet the requirements of the noise reduction effect; while the neural network algorithm has a very large computational amount and is not friendly enough to existing chips, and cannot meet the real-time requirements of ISP chips.
[0071] Among them, traditional image denoising can be classified into spatial domain denoising, frequency domain denoising, and combined spatial-frequency domain denoising according to the feature space for separating signal and noise; and can be classified into local denoising and non-local denoising according to the image range used for denoising processing. Specific traditional denoising methods include mean filtering, median filtering, Gaussian filtering, bilateral filtering, non-local mean filtering, guided filtering, discrete cosine domain filtering, wavelet transform domain filtering, etc. Traditional denoising methods are all based on simple assumptions about the differences in signal and noise characteristics in statistics, and use a fixed set of methods to separate signals and noise. Due to the overly simple assumption of noise characteristics, part of the signal will be mixed during noise separation, or the noise separation is not thorough enough, resulting in noise residue. In actual scenarios, especially when the noise is relatively significant (such as imaging in low-illumination environments), the denoising effect is poor.
[0072] In recent years, in addition to traditional image denoising algorithms, various neural network-based image denoising algorithms have greatly improved the image denoising effect. Several representative networks of such neural networks include the straight-chain DnCNN (Denoising Convolutional Neural Network), CBDNet (Convolutional Blind Denoising Network) including a sub-network for evaluating the noise level, and RIDNet (Real Image Denoising Based on Feature Attention) based on the attention mechanism, etc. The advantage of neural network-based image denoising algorithms is that the effect is significantly improved compared with traditional algorithms, but the corresponding disadvantage is that the computational complexity far exceeds that of traditional algorithms, and it is difficult to implement on an ISP (Image Signal Processor) chip with high real-time requirements in practical applications.
[0073] In view of this, in the embodiments of the present application, an image denoising processing method that meets the real-time requirements of the ISP chip and can better perform denoising processing on video images is provided.
[0074] In one embodiment, an image denoising processing method is provided. In the embodiments of the present application, this method is exemplified by being applied to a terminal including an ISP chip. It can be understood that this method can also be applied to a server, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. Specifically, the execution subject of this method can be the ISP chip in the terminal. The terminal can be, but is not limited to, various computer devices or shooting devices, etc., and the server can be implemented by an independent server or a server cluster composed of multiple servers.
[0075] In the embodiments of the present application, this method includes: inputting target image data into an image denoising model to obtain denoised image data output by the image denoising model, where the target image data includes the pixel values of each channel of the target image. Among them, as Figure 1As shown in the figure, it shows a schematic structural diagram of an image denoising model provided by an embodiment of the present application. The image denoising model includes a cascaded downsampling model, an upsampling model, and an output layer. The downsampling model includes n cascaded downsampling modules, and the upsampling model includes n cascaded upsampling modules corresponding one by one to the n downsampling modules. The downsampling module includes a first downsampling module, a second downsampling module, and a fusion module cascaded with both the first downsampling module and the second downsampling module. The first downsampling module includes a cascaded first downsampling layer and a first convolutional layer, and the second downsampling module includes a second downsampling layer.
[0076] Among them, the ISP (Image Signal Processor) chip is used to obtain the image captured by the front-end image sensor in the terminal, perform a series of image processing, and output the processed image. Generally, the steps for the ISP chip to process the image transmitted in RAW format from the image sensor include: performing bad pixel correction, dark current correction, lens shading correction, RAW image denoising, white balance, color interpolation, etc. on the RAW image to obtain an RGB image, and then performing Gamma correction, color correction, RGB image to YUV image conversion, etc. to obtain a YUV image. After the YUV image is denoised, edge enhanced, and brightness / contrast / hue / saturation adjusted, etc., finally, the image data is encoded to obtain the final output video image. Optionally, the image processed by the ISP chip can be a single image or a video image composed of continuous frames. Various image processing algorithms can be integrated in the ISP chip to implement the above-mentioned image processing steps of the ISP chip. In the embodiment of the present application, the image denoising model is an algorithm applied to the ISP chip to implement the image denoising processing step.
[0077] It should be noted that Figure 1 only three downsampling modules and three upsampling modules are taken as examples, which are not used to limit the present application.
[0078] In the embodiment of the present application, the target image is the image that needs to be denoised in the ISP chip, and the target image data is the pixel values of each channel of the target image. Among them, if the front-end image sensor obtains a single image, the target image is a single image; if the front-end image processor obtains a real-time video image, the target image is a single frame image in the real-time video image.
[0079] Optionally, the target image can be in RAW format, RGB format, YUV format, etc., and the embodiment of the present application does not make specific limitations on this.
[0080] The image denoising model can perform different processes according to whether the resolutions of the image data of each channel of the target image to be processed are consistent. Specifically, for a target image with consistent resolutions of the image data of each channel, the image denoising model has single input and output, and can directly perform denoising processing on the input target image data; for a target image with inconsistent resolutions of the image data of each channel, the image denoising model has multiple inputs and outputs, and can input the image data of different channels in the target image data into the image denoising model respectively for denoising processing. Thus, the image denoising model can be adapted to perform denoising processing on various types of target images.
[0081] Among them, currently, image denoising is divided into types such as single-image denoising and multi-frame image joint denoising. Considering the computational load and data caching of the ISP chip, single-image denoising is more suitable for real-time processing scenarios than multi-frame image joint denoising. Combining the actual denoising effect, network computational complexity, data read-write volume, etc., the embodiment of the present application combines the U-Net network, optimizes the structure and computing units of the network, and proposes a very concise single-image denoising network structure, that is, the image denoising model. The main body of the image denoising model is the U-Net network. The downsampling model of the image denoising model is used to perform downsampling processing on the target image data, the upsampling model is used to perform upsampling processing on the feature data obtained after upsampling processing, and the output layer is used to output the denoised image data corresponding to the target image based on the data output by the upsampling model and the target image data. The input of each upsampling module is the output data of the previous upsampling module and the output data of the downsampling module corresponding to this upsampling module. Thus, the image features of deep and shallow layers can be fused to improve the denoising effect.
[0082] The downsampling model is composed of n cascaded downsampling modules, and each downsampling module is a multi-convolution parallel module for extracting more image features. For a clearer understanding of this multi-convolution parallel module, please refer to Figure 2, which shows a schematic structural diagram of a multi-convolution parallel module provided by an embodiment of the present application. Exemplarily, the multi-convolution parallel module includes two downsampling layers 201, a convolution layer 202, and a fusion layer 203. Optionally, the multi-convolution parallel module may further include other numbers of downsampling layers and convolution layers, which are not specifically limited in the embodiments of the present application. In other words, correspondingly, in addition to including a first convolution layer, a first downsampling layer, and a second downsampling layer, the downsampling module may further include other numbers of convolution layers and downsampling layers, which can be specifically determined based on parameters such as the computing power and bandwidth of the ISP chip itself, and are not specifically limited in the embodiments of the present application. It should be noted that the image denoising model can achieve a good downsampling effect as long as it includes the downsampling module with the structure provided in the embodiments of the present application, and can meet the real-time requirements of the ISP chip while ensuring the denoising effect.
[0083] Correspondingly, the numbers of the downsampling module and the upsampling module in the image denoising model can be determined based on parameters such as the computing power and bandwidth of the ISP chip itself, which are not specifically limited in the embodiments of the present application.
[0084] Optionally, the channel fusion method of the fusion module in the downsampling module can be selected as element-wise addition or channel splicing, etc., which are not specifically limited in the embodiments of the present application.
[0085] The above image denoising method can directly input the target image data including the pixel values of each channel of the target image into the image denoising model, and then obtain the denoised image data output by the image denoising model, so as to realize the denoising process of the target image. Among them, usually in the ISP chip, it is necessary to perform denoising processing based on the pixel values of the Y channel and the pixel values of the UV channel of the YUV image respectively to obtain the image data after denoising processing. Since the Y channel and the UV channel are processed simultaneously, it is necessary to repeatedly call the image data, and its processing efficiency is poor and cannot meet the real-time requirements of the ISP chip. In this application, since the target image data including the pixel values of each channel of the target image can be directly input into the image denoising model for denoising processing, that is, the data of each channel of the target image is denoised simultaneously. Compared with denoising each channel separately, the data volume and calculation volume are greatly reduced, which can effectively improve the data processing efficiency and meet the real-time requirements; moreover, during the denoising process, the information between the channels in the target image can be mutually referenced to achieve a better denoising effect. Among them, the network structure of the image denoising model is streamlined, which includes a cascaded downsampling model, an upsampling model and an output layer. The downsampling model includes n cascaded downsampling modules, and the upsampling model includes n cascaded upsampling modules corresponding one by one to the n downsampling modules; the downsampling module includes a first downsampling module, a second downsampling module and a fusion module cascaded with both the first downsampling module and the second downsampling module; the first downsampling module includes a first convolutional layer and a first downsampling layer, and the second downsampling module includes a second downsampling layer. Through the streamlined image denoising model, effective denoising processing of the target image data can be realized, so that the image denoising model can be fully adapted to the ISP chip for denoising real-time video images.
[0086] In one embodiment, the image denoising model is used in the RAW image denoising module, the RGB image denoising module or the YUV image denoising module in the ISP chip; correspondingly, the format of the target image is RAW format, RGB format or YUV format.
[0087] Please refer to Figure 3, which shows a schematic diagram of the image processing flow of a traditional ISP chip provided by an embodiment of the present application. The processing process of the traditional ISP chip includes: obtaining the RAW images corresponding to each frame in the video image transmitted by the front-end image sensor, and obtaining RGB images through dead pixel correction, dark current correction, lens shading correction, RAW image noise reduction, white balance, color interpolation, etc. Then, through Gamma correction, color correction, RGB to YUV conversion, etc., YUV images are obtained. The Y (luminance) channel data of the YUV image is subjected to noise reduction processing and edge enhancement and luminance / contrast adjustment, and the data of the UV (color) channel of the YUV image is subjected to hue / saturation adjustment to implement noise reduction processing on the YUV image. Then, the data code is performed on the noise-reduced YUV image to obtain the finally output video image.
[0088] The RAW image is the acquisition format of the image sensor, which is essentially a special RGB format. After a series of processing, the RAW image undergoes color interpolation to obtain a normal RGB image, and then is converted to a YUV image after a series of processing. The YUV image format is a format that separates the luminance and color of the image, where the luminance is represented by the Y channel and the color is represented by the UV two channels. In the processing of the YUV image, the Y channel and the UV channel are processed separately, and noise reduction algorithms are used for the Y (luminance) channel and the UV (color) channel respectively. Then, edge enhancement, luminance and contrast adjustment are performed on the luminance component, and hue and saturation adjustment are performed on the color component.
[0089] Since the traditional ISP chip processing requires separate channel processing of the YUV image, during this process, it is necessary to repeatedly read and write the image data of each channel for processing, resulting in poor noise reduction effect and low processing efficiency. Based on this, in the embodiment of the present application, the image noise reduction model is applied to the ISP chip, and can be directly used for noise reduction processing of RGB images, RAW images or YUV images. For example, if the image noise reduction model is applied to the RAW image noise reduction module in the ISP chip, the format of the corresponding target image is RAW format; if the image noise reduction model is applied to the RGB image noise reduction module in the ISP chip, the format of the corresponding target image is RGB format; if the image noise reduction model is applied to the YUV image noise reduction module in the ISP chip, the format of the corresponding target image is YUV format.
[0090] To replace the process of separate channel noise reduction processing of the YUV image in the traditional method, so as to improve the noise reduction effect and processing efficiency.
[0091] In one case, please refer to Figure 4, which shows a schematic diagram of the first improved ISP chip image processing flow provided by the embodiments of the present application. The image denoising model is applied to the YUV image denoising module in the ISP chip to directly perform denoising processing on the data of each channel of the YUV image. In this case, the target image data input to the image denoising model is the pixel values of each channel of the YUV image. Thus, the denoising processing effect and processing efficiency of the YUV image can be greatly improved. In the scenario of processing real-time video images, the achievement of the processing efficiency of this image denoising model is more significant.
[0092] In another case, please refer to Figure 5 , which shows a schematic diagram of the second improved ISP chip image processing flow provided by the embodiments of the present application. The image denoising model is applied to the RGB image denoising module in the ISP chip to perform denoising processing on the data of each channel of the RGB image. Then, the ISP chip can directly convert the RGB image obtained by the denoising processing into a YUV image and perform subsequent processing, without the need to perform denoising processing on each channel of the YUV image again, improving the efficiency of image denoising processing. In this case, the target image data input to the image denoising model is the pixel values of each channel of the RGB image.
[0093] Optionally, the image denoising model can also be applied to the RAW image denoising module in the ISP chip to directly perform denoising processing on the RAW image. Correspondingly, in this case, the target image data input to the image denoising model is the pixel values of each channel of the RAW image.
[0094] In the embodiments of the present application, the denoising module in the traditional ISP chip is slightly adjusted, and the luminance denoising and color denoising are combined to achieve overall denoising processing through the image denoising model, which can improve the ability to extract image signals in the denoising processing, improve the signal-to-noise ratio and subjective denoising effect. For YUV images, when denoising each channel (Y, U, V) of the image, the information between each channel can be referred to to achieve a better denoising effect; in addition, the actual calculation amount and data read / write amount of processing by combining channels are less than the sum of processing by separating channels. Therefore, combining the channels of the image for processing has positive benefits both in terms of performance and effect. Therefore, applying the image denoising model to the ISP chip for image denoising processing meets the requirements of the ISP chip both in terms of denoising effect and real-time performance.
[0095] Next, the specific processing process of the image denoising model for the target image data will be described.
[0096] In one embodiment, as Figure 6As shown, it shows a schematic flowchart of a noise reduction process provided by an embodiment of the present application; inputting target image data into an image denoising model to obtain denoised image data output by the image denoising model, including:
[0097] Step 601, input the target image data into a downsampling model, and each downsampling module in the downsampling model performs downsampling processing on the target image data to obtain downsampled feature data.
[0098] Step 602, input the downsampled feature data into an upsampling model, and each upsampling module in the upsampling model performs upsampling processing on the downsampled feature data to obtain upsampled feature data.
[0099] Step 603, the output layer obtains denoised image data based on the upsampled feature data and the target image data.
[0100] Among them, the downsampling model is composed of n cascaded downsampling modules. Optionally, the value of n can be determined based on the actual operation and storage conditions of the chip, etc., so as to determine the structures of the downsampling model and the upsampling model. The application embodiment does not specifically limit the number of upsampling modules and downsampling modules. Of course, the number of upsampling modules should be the same as the number of downsampling modules.
[0101] Each downsampling module is used to perform downsampling processing on the input data to extract more image features. After the downsampling processing, the image is reduced. Correspondingly, an upsampling module is used for upsampling processing to restore the image size.
[0102] The last module in the n cascaded downsampling modules outputs downsampled feature data, and the last module in the n cascaded upsampling modules outputs upsampled feature data.
[0103] Input the upsampled feature data and the target image data into the output layer for fusion processing, and the denoised image data can be obtained based on the data output by the output layer.
[0104] In the embodiment of the present application, currently almost all single image denoising algorithms are difficult to run in real time on an ISP chip. By streamlining the network structure and constructing a lightweight neural network model adapted to chip operation, an image denoising model is obtained, enabling the single image denoising algorithm to be applied to the ISP chip for real-time denoising, while meeting the requirements of image denoising effect and algorithm real-time performance, solving the problems of neural network deployment and real-time operation on the chip. At the same time, there is a significant improvement in the denoising effect compared to traditional chips.
[0105] Specifically, for the target image data corresponding to different types of target images, the process of inputting them into the image denoising model for processing is also different. The following separately describes the processes of processing the target image data corresponding to two types of target images.
[0106] In one case, the target image is an image in a format where the image data resolution of each channel is the same, such as a target image in RAW format, RGB format, or YUV444 format. At this time, each downsampling module in the downsampling model performs downsampling processing on the target image data to obtain downsampled feature data, including:
[0107] For the i-th downsampling module, perform downsampling processing on the input data of the i-th downsampling module to obtain the intermediate downsampled feature data output by the i-th downsampling module; use the intermediate downsampled feature data output by the last downsampling module as the downsampled feature data. Among them, when i = 1, the input data of the i-th downsampling module is the target image data, and when i > 1, the input data of the i-th downsampling module is the intermediate downsampled feature data output by the i - 1-th downsampling module.
[0108] In the ISP chip, images are stored in different data formats (such as RAW, RGB, YUV444, YUV420) in different modules. When the input and output of the image denoising model are in formats with the same resolution for each channel such as RAW, RGB, YUV444, etc., for the first downsampling module, the entire target image data is used as the input of the first downsampling module, and for other downsampling modules, the intermediate downsampled feature data output by the previous downsampling module is used as the input data of this downsampling module. Each downsampling module is used to extract image features to obtain intermediate downsampled feature data.
[0109] Correspondingly, in one embodiment, each upsampling module in the upsampling model performs upsampling processing on the downsampled feature data to obtain upsampled feature data, including:
[0110] For the i-th upsampling module, perform upsampling processing on the input data of the i-th upsampling module to obtain the intermediate upsampled feature data output by the i-th upsampling module; use the intermediate upsampled feature data output by the last upsampling module as the upsampled feature data. Among them, when i = 1, the input data of the i-th upsampling module is the downsampled feature data, and when i > 1, the input data of the i-th upsampling module is the aggregated feature data obtained by fusing the intermediate upsampled feature data output by the i - 1-th upsampling module and the intermediate downsampled feature data output by the downsampling module corresponding to the i-th upsampling module.
[0111] Among them, for the first upsampling module, the upsampled feature data is used as the input of the first upsampling module. For other upsampling modules, the intermediate upsampled feature data output by the previous upsampling module and the intermediate downsampled feature data output by the corresponding downsampling module of this upsampling module are fused to obtain aggregated feature data, and the aggregated feature data is used as the input data of this upsampling module. Thus, the deep features, shallow features, and features of different resolutions in the target image data can be fully fused, improving the effect of noise reduction processing.
[0112] Correspondingly, in one embodiment, the output layer obtains the denoised image data based on the upsampled feature data and the target image data, including: inputting the upsampled feature data and the target image data into the output layer for fusion processing to obtain the denoised image data output by the output layer.
[0113] Specifically, the output layer is mainly used for fusing the input data. In other words, both the output layer and the fusion module in the downsampling module are used for feature fusion processing. Optionally, the output layer can perform simple element-wise addition processing or channel-wise stacking processing on the input data to achieve fusion processing.
[0114] After a series of processes, the upsampled feature data output by the last upsampling module contains the noise feature data of the target image. By fusing the target image data with this upsampled feature data, the noise features can be removed from the target image data, and the denoised image data output by the output layer is the denoised image data after removing the noise feature data, realizing the denoising processing of the target image. That is, the entire image denoising model actually outputs the noise residual, and then in the output layer, the noise residual and the target image are superimposed to obtain the denoised image output by the output layer.
[0115] Optionally, the upsampling module can be composed of a cascaded convolutional layer and an upsampling layer.
[0116] Based on the above content, a single-input and single-output image denoising model provided by an embodiment of the present application can be obtained. Please refer to Figure 7 , which shows a schematic structural diagram of a denoising neural network provided by an embodiment of the present application. Exemplarily, Figure 7 the image denoising model shown in includes three downsampling modules, three upsampling modules, the output layer, the fusion module in the downsampling module, and other fusion processes take the element-wise addition fusion method as an example. The downsampling layer and the upsampling layer can achieve downsampling or upsampling through convolutional processing.
[0117] It can be seen that the image denoising model based on the U-net network structure provided by the embodiment of the present application has a simple structure, directly performs denoising processing on the data of each channel of the image as a whole, ensuring a good denoising effect and high processing efficiency.
[0118] When the neural network framework and the structure of the basic computing unit remain unchanged, the number of downsampling modules and upsampling modules can be modified according to the actual computing power and bandwidth limitations, which can appropriately improve the noise reduction effect. Figure 7 Taking three downsamplings and three upsamplings as an example, the number of times can actually be increased (such as 4 times, 5 times, etc.) or decreased (such as 2 times). The way of channel fusion can be selected as element-wise addition or channel concatenation, etc.
[0119] In another case, the target image is an image in a format where the image data resolutions of each channel are different, such as a target image in the YUV420 format, etc. At this time, the downsampling model in the image noise reduction model further includes an additional downsampling module. The upsampling model in the image noise reduction model further includes an additional upsampling module. Please refer to Figure 8 , which shows a schematic structural diagram of another image noise reduction model provided by an embodiment of the present application.
[0120] Correspondingly, input the target image data into the downsampling model, and each downsampling module in the downsampling model performs downsampling processing on the target image data to obtain downsampled feature data, including:
[0121] Input the pixel values of the first channel of the target image included in the target image data into the additional downsampling module to obtain the channel feature data output by the additional downsampling module. Perform fusion processing on the channel feature data and the pixel values of the second channel of the target image included in the target image data to obtain candidate target image data as the input data of the first downsampling module. For the i-th downsampling module, perform downsampling processing on the input data of the i-th downsampling module to obtain the intermediate downsampled feature data output by the i-th downsampling module. Take the intermediate downsampled feature data output by the last downsampling module as the downsampled feature data.
[0122] Among them, when i = 1, the input data of the i-th downsampling module is the candidate target image data, and when i > 1, the input data of the i-th downsampling module is the intermediate downsampled feature data output by the (i - 1)-th downsampling module.
[0123] In the ISP chip, images are stored in different data formats (such as RAW, RGB, YUV444, YUV420) in different modules. If the input and output image format of the image noise reduction model embedded in the chip is a position where the resolutions of each channel are inconsistent, such as YUV420, etc., then a multi-input multi-output network structure needs to be used, that is Figure 8 the structure in
[0124] Specifically, for the target image of each channel resolution, the target image data corresponding to the target image is composed of the first channel pixel values and the second channel pixel values. Taking the target image in YUV420 format as an example, the first channel pixel values are the Y channel pixel values, and the second channel pixel values are the UV channel pixel values. Since the resolutions are different, it is necessary to first perform downsampling processing on the first channel pixel values, that is, the Y channel pixel values, using an additional downsampling module. The channel feature data output by the additional downsampling module can then have the same resolution as the second channel pixel values, that is, the UV channel pixel values. At this time, the channel feature data and the second channel pixel values can be directly fused and the normal downsampling processing can be performed using each downsampling module.
[0125] Correspondingly, in one embodiment, the downsampled feature data is input into an upsampling model, and each upsampling module in the upsampling model performs upsampling processing on the downsampled feature data to obtain upsampled feature data, including:
[0126] For the i-th upsampling module, the input data of the i-th upsampling module is upsampled to obtain the intermediate upsampled feature data output by the i-th upsampling module. The first intermediate channel feature data corresponding to the first channel pixel values included in the intermediate upsampled feature data output by the last upsampling module is input into an additional upsampling module to obtain the upsampled feature data output by the additional upsampling module.
[0127] Among them, when i = 1, the input data of the i-th upsampling module is the downsampled feature data. When i > 1, the input data of the i-th upsampling module is the aggregated feature data obtained by fusing the intermediate upsampled feature data output by the (i - 1)-th upsampling module and the intermediate downsampled feature data output by the downsampling module corresponding to the i-th upsampling module. Thus, the deep features, shallow features, and features of different resolutions in the target image data can all be fully fused, improving the effect of noise reduction processing.
[0128] Corresponding to the target image data including the first channel pixel values and the second channel pixel values, the denoised image data output by the image denoising model should also include denoised image data of different channels with different resolutions.
[0129] For the upsampling module, similar to the image denoising model in the previous text, each upsampling module performs upsampling on the input data to obtain the intermediate upsampling feature data output by the last upsampling module. At this time, the intermediate upsampling feature data output by the last upsampling module includes the first intermediate channel feature data and the second intermediate channel feature data. The first intermediate channel feature data corresponds to the first channel pixel values, that is, the Y-channel image data. In other words, the first intermediate channel feature data is obtained by denoising the first channel pixel values; correspondingly, the second intermediate channel feature data is obtained by denoising the second channel pixel values.
[0130] Since the first channel pixel values are downsampled by the additional downsampling module, correspondingly, the first intermediate channel feature data also needs to be upsampled by the additional upsampling module, so that the output of the additional upsampling module is used as the upsampling feature data finally output by the upsampling model. This enables the output layer to perform further fusion processing based on the upsampling feature data and the first channel pixel values in the target image data.
[0131] Correspondingly, in one embodiment, the output layer obtains the denoised image data based on the upsampling feature data and the target image data, including: inputting the upsampling feature data and the first channel pixel values in the target image data into the output layer for fusion processing to obtain the candidate denoised image data output by the output layer. The denoised image data is obtained according to the candidate denoised image data and the second intermediate channel feature data corresponding to the second channel pixel values included in the intermediate upsampling feature data output by the last upsampling module.
[0132] Specifically, after a series of processing, the upsampling feature data contains the noise features corresponding to the first channel pixel values, that is, the noise residuals of the target image are extracted. By fusing the upsampling feature data with the first channel pixel values in the output layer, that is, superimposing the noise residuals and the target image in the output layer, the noise features in the first channel pixel values can be removed to obtain the candidate denoised image data output by the output layer.
[0133] Based on the candidate image denoising data and the second intermediate channel feature data included in the intermediate upsampling feature data output by the last upsampling module, the denoised image data can be obtained.
[0134] Thus, for images with different channel resolutions, the multi-input and multi-output image denoising model can also perform denoising processing, expanding the application of the ISP chip.
[0135] Optionally, the additional downsampling module is composed of a cascaded downsampling layer and a convolutional layer, and the additional upsampling module is composed of a cascaded convolutional layer and an upsampling layer.
[0136] Based on the above content, a schematic structural diagram of the multi-input and multi-output image denoising model provided in the embodiments of the present application can be obtained. Please refer to Figure 9 , which shows a schematic structural diagram of another denoising neural network provided in the embodiments of the present application. Exemplarily, Figure 9 the image denoising model shown in includes two downsampling modules, two upsampling modules, an output layer, and the fusion modules in the downsampling modules, and other fusion processes are exemplified by the element-wise addition fusion method. The downsampling layer and the upsampling layer can achieve downsampling through convolutional processing.
[0137] In the embodiments of the present application, the image denoising processing module, as an important module in the ISP chip, the input and output formats and data arrangement methods should be consistent with those of the denoising processing module in the traditional ISP chip to reduce the modification of the original layout of the ISP chip and speed up the application process. Therefore, a denoising neural network with multi-input and multi-output formats is provided in the embodiments of the present application to replace the original denoising module without significantly modifying the layout of the ISP chip. Compared with the traditional neural network, the multi-input and multi-output network structure is used in the embodiments of the present application and embedded in the ISP chip, which preferably solves the problems that the single-image denoising neural network cannot adapt to the layout and real-time performance of the ISP chip.
[0138] Compared with the image processing algorithms in the existing ISP chips, the image denoising model in the embodiments of the present application can fully fuse the information of the pixel values of each channel in the target image data, better mine the information in the image, and remove the noise in the image. Among them, the ISP denoising algorithm with channel fusion can effectively improve the signal-to-noise ratio of the image. Compared with the traditional denoising algorithm, using a lightweight denoising neural network can greatly improve the clarity of the image and reduce the noise. While reducing the noise, due to the improvement of the image quality, the trailing noise of the moving objects in the image can be improved at the same time, and the accuracy of the subsequent image tasks such as target detection or face recognition on the target image can be improved, and the application range of the target image after denoising processing can be expanded.
[0139] As mentioned above, each downsampling module includes a first downsampling module, a second downsampling module, and a fusion module cascaded with both the first downsampling module and the second downsampling module; the first downsampling module includes a first convolutional layer and a first downsampling layer, and the second downsampling module includes a second downsampling layer. The processing process of the downsampling module will be described below.
[0140] In one embodiment, downsampling processing is performed on the input data of the i-th downsampling module to obtain intermediate downsampling feature data output by the i-th downsampling module, including: performing downsampling processing on the input data of the i-th downsampling module using a first downsampling layer to obtain first downsampling feature data output by the first downsampling layer; performing convolution processing on the first downsampling feature data using a first convolutional layer to obtain first convolutional feature data output by the first convolutional layer; performing downsampling processing on the input data of the i-th downsampling module using a second downsampling layer to obtain second downsampling feature data output by the second downsampling layer; and performing fusion processing on the first convolutional feature data and the second downsampling feature data using a fusion module to obtain intermediate downsampling feature data output by the fusion module.
[0141] Optionally, the first downsampling layer and the second downsampling layer can implement downsampling processing through convolution processing.
[0142] Optionally, each downsampling module can select a suitable convolution structure. For example, the first downsampling layer and the first convolutional layer can be selected for convolution processing with a 5x5 convolution kernel, and the second convolutional layer can be selected for convolution processing with a 3x3 convolution kernel, etc. Of course, the convolution processing of each layer in the downsampling module can use any combination of direct connection, 1x1 convolution, 3x3 convolution, 5x5 convolution, 7x7 convolution, etc., and the embodiments of the present application do not make specific limitations on this.
[0143] The processing process of the upsampling module will be described below.
[0144] In one embodiment, the upsampling module includes a cascaded second convolutional layer and an upsampling layer; upsampling processing is performed on the input data of the i-th upsampling module to obtain intermediate upsampling feature data output by the i-th upsampling module, including: performing convolution processing on the input data of the i-th upsampling module using the second convolutional layer to obtain second convolutional feature data output by the second convolutional layer; and performing upsampling processing on the second convolutional feature data using the upsampling layer to obtain intermediate upsampling feature data output by the upsampling layer.
[0145] Optionally, the upsampling layer performs upsampling processing on the input data of the upsampling layer through convolution processing, anti-pooling processing, or interpolation processing.
[0146] Optionally, each upsampling module can also include other numbers of convolutional layers and upsampling layers, which can be specifically determined based on the computing power, bandwidth, or storage energy parameters of the ISP chip.
[0147] Among them, the second convolutional layer and the upsampling layer can select a suitable convolution structure, and the convolution processing of each layer can use any combination of direct connection, 1x1 convolution, 3x3 convolution, 5x5 convolution, 7x7 convolution, etc., and the embodiments of the present application do not make specific limitations on this.
[0148] In one embodiment, for the fusion processing of the fusion module, the fusion processing of the output layer, and other fusion processing in the image denoising model described above, the fusion processing can perform feature fusion on the deep and shallow layers of the image, and both can implement the fusion processing in the way of element-wise addition or channel concatenation. Please refer to Figure 10 , which shows a schematic diagram of an element-wise addition-based fusion processing provided by an embodiment of the present application; please refer to Figure 11 , which shows a schematic diagram of a channel concatenation-based fusion processing provided by an embodiment of the present application.
[0149] Among them, the method of fusing feature channels in the element-wise addition manner can significantly reduce the data reading volume and the calculation volume, but it may also lose some of the features that have been extracted. The channel concatenation method can be selected when the chip computing power and cache are sufficient. The element-wise addition method needs to ensure that the resolutions and the number of channels of the two sets of features to be fused are exactly the same, while the channel concatenation method does not require the number of channels of the two sets of features to be fused to be the same. Based on this, for different ISP chips, different channel fusion methods and upsampling methods are selected, and there are slight differences in the final denoising effect, but for some chips, the operation time and efficiency are very different. Therefore, the specific way of the fusion processing can be determined as element-wise addition or channel concatenation based on the image processing requirements pre-determined for the ISP chip, so as to fully improve the efficiency of the chip image processing.
[0150] In one embodiment, the present application provides a neural network denoising algorithm deployed on an ISP chip, and the main deployment process is as follows:
[0151] Step 1: Construct the basic structure of the U-Net neural network. Specifically, it includes determining the network framework, the input and output formats of the network (RAW, RGB, YUV444, etc.), the number of downsampling and upsampling layers, and the sampling rate of each layer.
[0152] Step 2: Determine the structure of the downsampling module in the network structure.
[0153] Step 3: Select the channel fusion method and the upsampling method. For example, the channel fusion adopts the element-wise addition processing method, and the upsampling is implemented by transposed convolution for upsampling.
[0154] Step 4: Determine the specific position where the network is embedded in the ISP chip image processing process, such as the RAW image denoising module, the RGB image denoising module, or the YUV image denoising module.
[0155] Step 5: Run and debug the complete neural network denoising algorithm.
[0156] Further, as described in step 1, exemplarily, the entire network may include 3 multi-convolution parallel sub-modules for downsampling processing, and correspondingly 3 ordinary convolution layers and 3 upsampling layers for upsampling processing. The features of the original resolution are retained before each downsampling and fused after the corresponding upsampling layer, so that the deep and shallow features in the image and the features of different resolutions can be fully fused. The entire network has a total of 15 convolution and upsampling layers, and an activation layer is added after each convolution. Taking a YUV format image with a size of 256x256x3 as the input data, the data sizes output by the 15-layer network are 128x128x16, 128x128x16, 128x128x16, 64x64x32, 64x64x32, 64x64x32, 32x32x64, 32x32x64, 32x32x64, 32x32x16, 64x64x16, 64x64x16, 128x128x16, 128x128x16, 256x256x3 (output in YUV format).
[0157] Optionally, the number of downsampling layers in step 1 can be N layers, where N is greater than or equal to 1. When the computing power and cache of the ISP chip are sufficient, N can be set to 4 or a larger integer.
[0158] Optionally, the downsampling rate in step 1 can be N:1, where N is greater than 1. N can be set to 3, 4, etc. Different downsampling layers can also use different sampling rates.
[0159] Optionally, the upsampling method in step 3 can be anti-pooling or interpolation algorithm, etc. Anti-pooling or interpolation methods can effectively reduce the number of parameters in the upsampling layer.
[0160] Compared with the image processing algorithms in existing ISP chips, in the embodiments of the present application, in the noise reduction process, the information of each channel of the image is fully integrated, which can better mine the information in the image and eliminate the noise in the image. The ISP noise reduction algorithm with channel fusion can effectively improve the signal-to-noise ratio of the image. Compared with the traditional noise reduction algorithm, using a lightweight noise reduction neural network can greatly improve the clarity of the image and reduce the noise. While reducing the noise, due to the improvement of the image quality, the trailing noise of moving objects in the image can be improved at the same time, and the accuracy of image tasks such as object detection and face recognition can be improved. Compared with ordinary noise reduction neural networks, the neural network used in the embodiments of the present application improves the network structure and basic operators, and can run in real time on the ISP chip better, thus solving the problem that ordinary neural networks cannot run in real time on mobile devices. Compared with ordinary neural networks, the embodiments of the present application use a multi-input multi-output network structure. While being embedded in the image processing flow of the ISP chip, it better solves the problem that the single-image noise reduction neural network cannot adapt to the image processing flow of the ISP chip.
[0161] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least some of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or steps or stages in other steps.
[0162] Based on the same inventive concept, the embodiments of the present application also provide an image noise reduction processing device for implementing the above-mentioned image noise reduction processing method. The solution provided by this device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more of the following embodiments of the image noise reduction processing device can refer to the limitations on the image noise reduction processing method in the above text, and will not be repeated here.
[0163] In one embodiment, as Figure 12 shown, an image noise reduction processing device is provided. The image noise reduction processing device 1200 includes: a noise reduction module 1201, where:
[0164] The noise reduction module 1201 is configured to input the target image data into an image noise reduction model to obtain the noise-reduced image data output by the image noise reduction model. The target image data includes the pixel values of each channel of the target image. The image noise reduction model includes a cascaded downsampling model, an upsampling model, and an output layer. The downsampling model includes n cascaded downsampling modules, and the upsampling model includes n cascaded upsampling modules that correspond one-to-one to the n downsampling modules. The downsampling module includes a first downsampling module, a second downsampling module, and a fusion module cascaded with both the first downsampling module and the second downsampling module. The first downsampling module includes a cascaded first downsampling layer and a first convolutional layer, and the second downsampling module includes a second downsampling layer.
[0165] In one embodiment, the noise reduction module 1201 is specifically configured to: input the target image data into the downsampling model, and each downsampling module in the downsampling model performs downsampling processing on the target image data to obtain downsampled feature data; input the downsampled feature data into the upsampling model, and each upsampling module in the upsampling model performs upsampling processing on the downsampled feature data to obtain upsampled feature data; and the output layer obtains the noise-reduced image data based on the upsampled feature data and the target image data.
[0166] In one embodiment, the image data resolution of each channel of the target image is the same. The noise reduction module 1201 is specifically configured to: for the i-th downsampling module, perform downsampling processing on the input data of the i-th downsampling module to obtain the intermediate downsampled feature data output by the i-th downsampling module. Wherein, when i = 1, the input data of the i-th downsampling module is the target image data, and when i > 1, the input data of the i-th downsampling module is the intermediate downsampled feature data output by the (i - 1)-th downsampling module; and use the intermediate downsampled feature data output by the last downsampling module as the downsampled feature data.
[0167] In one embodiment, the noise reduction module 1201 is specifically configured to: for the i-th upsampling module, perform upsampling processing on the input data of the i-th upsampling module to obtain the intermediate upsampled feature data output by the i-th upsampling module. Wherein, when i = 1, the input data of the i-th upsampling module is the downsampled feature data, and when i > 1, the input data of the i-th upsampling module is the aggregated feature data obtained by fusing the intermediate upsampled feature data output by the (i - 1)-th upsampling module and the intermediate downsampled feature data output by the downsampling module corresponding to the i-th upsampling module; and use the intermediate upsampled feature data output by the last upsampling module as the upsampled feature data.
[0168] In one embodiment, the noise reduction module 1201 is specifically configured to: input the upsampled feature data and the target image data into an output layer for fusion processing to obtain the noise-reduced image data output by the output layer.
[0169] In one embodiment, the image data resolutions of the channels of the target image are different, and the downsampling model further includes an additional downsampling module. The noise reduction module 1201 is specifically configured to: input the pixel values of the first channel of the target image included in the target image data into the additional downsampling module to obtain the channel feature data output by the additional downsampling module; perform fusion processing on the channel feature data and the pixel values of the second channel of the target image included in the target image data to obtain candidate target image data; for the i-th downsampling module, perform downsampling processing on the input data of the i-th downsampling module to obtain the intermediate downsampling feature data output by the i-th downsampling module; wherein, when i = 1, the input data of the i-th downsampling module is the candidate target image data, and when i > 1, the input data of the i-th downsampling module is the intermediate downsampling feature data output by the (i - 1)-th downsampling module; use the intermediate downsampling feature data output by the last downsampling module as the downsampling feature data.
[0170] In one embodiment, the upsampling model further includes an additional upsampling module. The noise reduction module 1201 is specifically configured to: for the i-th upsampling module, perform upsampling processing on the input data of the i-th upsampling module to obtain the intermediate upsampling feature data output by the i-th upsampling module; wherein, when i = 1, the input data of the i-th upsampling module is the downsampling feature data, and when i > 1, the input data of the i-th upsampling module is the aggregated feature data obtained by fusing the intermediate upsampling feature data output by the (i - 1)-th upsampling module and the intermediate downsampling feature data output by the corresponding downsampling module of the i-th upsampling module; input the first intermediate channel feature data corresponding to the pixel values of the first channel included in the intermediate upsampling feature data output by the last upsampling module into the additional upsampling module to obtain the upsampled feature data output by the additional upsampling module.
[0171] In one embodiment, the noise reduction module 1201 is specifically configured to: input the upsampled feature data and the pixel values of the first channel in the target image data into the output layer for fusion processing to obtain the candidate noise-reduced image data output by the output layer; obtain the noise-reduced image data based on the candidate noise-reduced image data and the second intermediate channel feature data corresponding to the pixel values of the second channel included in the intermediate upsampling feature data output by the last upsampling module.
[0172] In one embodiment, the noise reduction module 1201 is specifically configured to: perform downsampling processing on the input data of the i-th downsampling module by using a first downsampling layer to obtain first downsampled feature data output by the first downsampling layer; perform convolution processing on the first downsampled feature data by using a first convolutional layer to obtain first convolutional feature data output by the first convolutional layer; perform downsampling processing on the input data of the i-th downsampling module by using a second downsampling layer to obtain second downsampled feature data output by the second downsampling layer; and perform fusion processing on the first convolutional feature data and the second downsampled feature data by using a fusion module to obtain intermediate downsampled feature data output by the fusion module.
[0173] In one embodiment, the upsampling module includes a cascaded second convolutional layer and an upsampling layer; the noise reduction module 1201 is specifically configured to: perform convolution processing on the input data of the i-th upsampling module by using the second convolutional layer to obtain second convolutional feature data output by the second convolutional layer; and perform upsampling processing on the second convolutional feature data by using the upsampling layer to obtain intermediate upsampled feature data output by the upsampling layer.
[0174] In one embodiment, the image noise reduction model is used in the RAW image noise reduction module 1201, RGB image noise reduction module 1201, or YUV image noise reduction module 1201 in an ISP chip; correspondingly, the format of the target image is RAW format, RGB format, or YUV format.
[0175] In one embodiment, the upsampling layer performs upsampling processing on the input data of the upsampling layer through convolution processing, anti-pooling processing, or interpolation processing.
[0176] Each module in the above image noise reduction processing device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in or independent of the processor in the computer device in the form of hardware, or stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above respective modules.
[0177] In one embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 13As shown in the figure. The computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements an image noise reduction processing method. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0178] Those skilled in the art can understand that Figure 13 the structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0179] In one embodiment, an electronic device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.
[0180] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by the processor, the steps in the above method embodiments are implemented.
[0181] In one embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by the processor, the steps in the above method embodiments are implemented.
[0182] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.
[0183] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0184] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. An image noise reduction processing method, characterized in that, The method includes: Inputting target image data into an image denoising model to obtain denoised image data output by the image denoising model, where the target image data includes pixel values of each channel of the target image; Among them, the image denoising model includes a cascaded downsampling model, an upsampling model, and an output layer. The downsampling model includes n cascaded downsampling modules, and the upsampling model includes n cascaded upsampling modules corresponding one-to-one to the n downsampling modules. The downsampling module includes a first downsampling module, a second downsampling module, and a fusion module cascaded with both the first downsampling module and the second downsampling module. The first downsampling module includes a cascaded first downsampling layer and a first convolutional layer, and the second downsampling module includes a second downsampling layer; The step of inputting the target image data into the image denoising model to obtain the denoised image data output by the image denoising model includes: inputting the target image data into the downsampling model, and performing downsampling processing on the target image data by each of the downsampling modules in the downsampling model to obtain downsampled feature data; inputting the downsampled feature data into the upsampling model, and performing upsampling processing on the downsampled feature data by each of the upsampling modules in the upsampling model to obtain upsampled feature data; and obtaining the denoised image data by the output layer based on the upsampled feature data and the target image data.
2. The method according to claim 1, wherein The target image is an image that needs to be denoised in an ISP chip.
3. The method according to claim 1, wherein The image data resolutions of each channel of the target image are the same. The step of performing downsampling processing on the target image data by each of the downsampling modules in the downsampling model to obtain downsampled feature data includes: For the i-th downsampling module, performing downsampling processing on the input data of the i-th downsampling module to obtain intermediate downsampled feature data output by the i-th downsampling module. Wherein, when i = 1, the input data of the i-th downsampling module is the target image data; when i > 1, the input data of the i-th downsampling module is the intermediate downsampled feature data output by the (i - 1)-th downsampling module; Taking the intermediate downsampled feature data output by the last downsampling module as the downsampled feature data.
4. The method according to claim 3, wherein The step of performing upsampling processing on the downsampled feature data by each of the upsampling modules in the upsampling model to obtain upsampled feature data includes: For the i-th upsampling module, performing upsampling processing on the input data of the i-th upsampling module to obtain intermediate upsampled feature data output by the i-th upsampling module. Wherein, when i = 1, the input data of the i-th upsampling module is the downsampled feature data; when i > 1, the input data of the i-th upsampling module is the aggregated feature data obtained by fusing the intermediate upsampled feature data output by the (i - 1)-th upsampling module and the intermediate downsampled feature data output by the downsampling module corresponding to the i-th upsampling module; Use the intermediate upsampling feature data output by the last one of the upsampling modules as the upsampling feature data.
5. The method according to claim 4, wherein The output layer obtaining the noise reduction image data based on the upsampling feature data and the target image data includes: Input the upsampling feature data and the target image data into the output layer for fusion processing to obtain the noise reduction image data output by the output layer.
6. The method according to claim 1, wherein The resolution of the image data of each channel of the target image is different, and the downsampling model further includes an additional downsampling module. Inputting the target image data into the downsampling model, and each of the downsampling modules in the downsampling model performing downsampling processing on the target image data to obtain downsampling feature data includes: Input the pixel values of the first channel of the target image included in the target image data into the additional downsampling module to obtain the channel feature data output by the additional downsampling module; Perform fusion processing on the channel feature data and the pixel values of the second channel of the target image included in the target image data to obtain candidate target image data; For the i-th downsampling module, perform downsampling processing on the input data of the i-th downsampling module to obtain the intermediate downsampling feature data output by the i-th downsampling module; wherein, when i = 1, the input data of the i-th downsampling module is the candidate target image data, and when i > 1, the input data of the i-th downsampling module is the intermediate downsampling feature data output by the (i - 1)-th downsampling module; Use the intermediate downsampling feature data output by the last one of the downsampling modules as the downsampling feature data.
7. The method according to claim 6, wherein The upsampling model further includes an additional upsampling module. Inputting the downsampling feature data into the upsampling model, and each of the upsampling modules in the upsampling model performing upsampling processing on the downsampling feature data to obtain upsampling feature data includes: For the i-th upsampling module, perform upsampling processing on the input data of the i-th upsampling module to obtain the intermediate upsampling feature data output by the i-th upsampling module; wherein, when i = 1, the input data of the i-th upsampling module is the downsampling feature data, and when i > 1, the input data of the i-th upsampling module is the aggregated feature data obtained by fusing the intermediate upsampling feature data output by the (i - 1)-th upsampling module and the intermediate downsampling feature data output by the corresponding downsampling module of the i-th upsampling module; Input the first intermediate channel feature data corresponding to the pixel values of the first channel included in the intermediate upsampling feature data output by the last upsampling module into the additional upsampling module to obtain the upsampling feature data output by the additional upsampling module.
8. The method according to claim 7, wherein The output layer obtaining the noise reduction image data based on the upsampling feature data and the target image data includes: Input the first-channel pixel values in the upsampled feature data and the target image data into the output layer for fusion processing to obtain candidate denoised image data output by the output layer; Obtain the denoised image data according to the candidate denoised image data and the second intermediate channel feature data corresponding to the second-channel pixel values included in the intermediate upsampled feature data output by the last upsampling module.
9. The method according to any one of claims 3 or 6, characterized in that, The downsampling the input data of the i-th downsampling module to obtain intermediate downsampled feature data output by the i-th downsampling module includes: Using the first downsampling layer to downsample the input data of the i-th downsampling module to obtain first downsampled feature data output by the first downsampling layer; Using the first convolutional layer to perform convolutional processing on the first downsampled feature data to obtain first convolutional feature data output by the first convolutional layer; Using the second downsampling layer to downsample the input data of the i-th downsampling module to obtain second downsampled feature data output by the second downsampling layer; Using the fusion module to perform fusion processing on the first convolutional feature data and the second downsampled feature data to obtain the intermediate downsampled feature data output by the fusion module.
10. The method according to claim 4 or 7, characterized in that, The upsampling module includes a cascaded second convolutional layer and an upsampling layer; the upsampling the input data of the i-th upsampling module to obtain intermediate upsampled feature data output by the i-th upsampling module includes: Using the second convolutional layer to perform convolutional processing on the input data of the i-th upsampling module to obtain second convolutional feature data output by the second convolutional layer; Using the upsampling layer to upsample the second convolutional feature data to obtain the intermediate upsampled feature data output by the upsampling layer.
11. The method according to claim 1, wherein The image denoising model is used in a RAW image denoising module, an RGB image denoising module, or a YUV image denoising module in an ISP chip; correspondingly, the format of the target image is RAW format, RGB format, or YUV format.
12. The method according to claim 10, characterized in that, The upsampling layer upsamples the input data of the upsampling layer through convolutional processing, anti-pooling processing, or interpolation processing.
13. An image noise reduction processing device, characterized in that, The device includes: A denoising module for inputting target image data into an image denoising model to obtain denoised image data output by the image denoising model, where the target image data includes pixel values of each channel of a target image; Wherein, the image denoising model includes a cascaded downsampling model, an upsampling model, and an output layer, the downsampling model includes n cascaded downsampling modules, and the upsampling model includes n cascaded upsampling modules corresponding one-to-one to the n downsampling modules; the downsampling module includes a first downsampling module, a second downsampling module, and a fusion module cascaded with both the first downsampling module and the second downsampling module; the first downsampling module includes a cascaded first downsampling layer and a first convolutional layer, and the second downsampling module includes a second downsampling layer; The noise reduction module is specifically configured to input the target image data into the downsampling model, and each of the downsampling modules in the downsampling model performs downsampling processing on the target image data to obtain downsampled feature data; input the downsampled feature data into the upsampling model, and each of the upsampling modules in the upsampling model performs upsampling processing on the downsampled feature data to obtain upsampled feature data; and the output layer obtains the denoised image data based on the upsampled feature data and the target image data.
14. An electronic device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 12.
15. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 12.
16. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Image processing method and device, equipment and readable storage medium
CN111192215A
Image restoration method and device, computer equipment, storage medium and program product
CN114913094A
Cited By
Image noise reduction processing method and apparatus, device, storage medium, and program product
WO2024055458A1