RAW domain dark light image noise reduction method based on improved U-net
By improving the U-net model, designing a lightweight network structure, and using global and local information modeling, the problem of high complexity of image denoising calculation under low light conditions is solved, and efficient and real-time image denoising capabilities are achieved, which is suitable for device deployment with limited resources.
Patent Information
- Application Number
- CN202311815934.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-27
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art when denoising images under low light conditions has high computational complexity, making it difficult to achieve real-time processing, and the amount of parameters is large, and the storage and computing bandwidth requirements are high.
Improve the U-net model, design lightweight network structures through global and local information modeling, use 5*5 convolution and attention modules to reduce the amount of parameters and calculation complexity, and is suitable for deployment on end-side devices with limited resources.
The ability to efficiently denoise images under low light conditions is realized, the calculation complexity and parameter amount is reduced, suitable for real-time image processing, and suitable for resource-constrained device deployment.
Smart Images

Figure CN120219211A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image denoising, and particularly relates to a method for denoising RAW domain low-light images by improving U-net. Background Art
[0002] With the development of technology, especially the advent of the era of artificial intelligence, image processing technology has become increasingly important. Among them, image noise is an unnecessary random signal introduced during the image acquisition and processing process, which will distort the image, reduce the visual quality and information content of the image. The process of image denoising can be achieved through technologies such as filters. The denoised image can display details more clearly and reflect the information expressed by the original image more accurately, which is very important for many applications, such as medical images, satellite images, security images, etc. In addition, image denoising is also an important issue in the field of computer vision, which can provide better image input for tasks such as image segmentation, object detection, and image recognition. Therefore, image denoising has important significance and application value in various fields.
[0003] In low-light environments, the signals contained in the image are very weak, while the noise is very obvious, resulting in a very low signal-to-noise ratio of the image, which makes it difficult for the method to distinguish which details are noise and which are useful image information. In order to remove noise under low-light conditions, complex calculations and processing are required. This makes the time and space complexity of the method very high and needs to be optimized and accelerated to achieve real-time denoising. Therefore, image denoising under low-light conditions is a challenging problem that requires advanced methods and technologies to overcome various difficulties and challenges.
[0004] With the continuous development of deep learning technology, image denoising methods based on deep learning have also been greatly developed and improved, such as denoising autoencoders and denoising convolutional neural networks based on deep learning. These methods can automatically learn the relationship between noise and images by learning a large amount of training data, so as to perform more accurate and adaptive denoising processing on the images. The image denoising method based on convolutional neural network (CNN) trains a neural network to learn the probability distribution of the noise distribution from the noisy image, and then denoises the image. The image denoising method based on recurrent neural network (RNN) learns the long-term dependence relationship of the image through the recurrent neural network structure and denoises on this basis. This method has better denoising effect than the method based on CNN, but the computational cost is relatively large. The image denoising method based on generative adversarial network (GAN) realizes denoising by training the adversarial generator and discriminator to make the generator generate more realistic images. This method can remove more complex noise while retaining the details of the image.
[0005] However, in order to obtain extreme noise reduction effects, most existing image noise reduction methods adopt deep neural networks with complex structures, large numbers of parameters, and high computational complexities, resulting in large storage overheads and bandwidth requirements, slow computational speeds, and difficulty in achieving the frame rates required for real-time image processing. Summary of the Invention
[0006] To solve the above problems, the purpose of this application is to provide an improved U-net RAW domain low-light image noise reduction method. The improved U-net model can model complex noise maps according to the global and local information of low-light images, thereby removing the noise in low-light images, and has the characteristics of few parameters and low computational complexity, which is conducive to deployment on resource-limited edge devices.
[0007] Specifically, the present invention provides an improved U-net RAW domain low-light image noise reduction method, and the method includes the following steps:
[0008] S1. Collect noisy images and corresponding noise-free reference images, and perform preprocessing and image enhancement operations to expand the scale of the training data set;
[0009] S2. Refer to the U-net form to design a lightweight network structure. The input of the network is a noisy image, and the output is a predicted noise map;
[0010] S3. Train the network: Input the preprocessed and data-augmented noisy image X input into the improved U-net, and obtain the predicted noise map through forward propagation Calculate the loss between the predicted image and the reference image using the MAE function, update the network parameters using the Adam optimizer, use the cosine annealing learning rate decay strategy during training, and set the half-period of the cosine function to the maximum number of iterations;
[0011] S4. Load the preprocessed uncropped noisy image into the trained network for denoising to obtain the denoised image.
[0012] The step S1 further includes:
[0013] S1.1, Image acquisition: Select indoor and outdoor scenes. For the same scene, adjust the illuminance through a fill light, collect K noisy images at each illuminance, take the average as the reference image, and randomly select one from the K images as the paired noisy image; where the value of K depends on the illuminance of the scene;
[0014] S1.2, Data preprocessing: Perform black level correction and normalization on the Bayer format RAW image, as shown in the following formula (1):
[0015]
[0016] In formula (1), bit is the number of bits of the RAW image. After black level correction and normalization, the single-channel output image is converted into a 4-channel image, and the image spatial resolution is reduced to 1 / 4 of the original;
[0017] S1.3, Data augmentation: Randomly crop the image after data preprocessing, and randomly rotate and flip the cropped image patches. The enhanced image patches are used as the input of the network.
[0018] In step S1.1, the value of K depends on the illuminance of the scene. Specifically, K is set to 500 at an illuminance of 0.1 lux and K is set to 200 at an illuminance of 1 lux.
[0019] The network structure described in step S2 includes:
[0020] S2.1, To reduce the network depth, only two downsamplings are performed in the encoding stage. For an image with an input spatial resolution of H×W, after passing through the encoder, the spatial resolution is only reduced to The entire encoding stage consists of two convolutional layers and two encoding modules. The encoding module consists of a downsampling module and N basic modules, where N is used to control the network depth and computational complexity. To ensure that the encoder has a large enough receptive field to capture context information, 5*5 convolutions are used instead of 3*3 convolutions; S2.2, The decoding stage consists of two decoding modules, an attention module, and two convolutional layers. The decoding module consists of a basic residual module and a transposed convolution; To reduce the number of parameters, 3*3 convolutions are used in the convolutional layers of the decoder; An attention module is inserted after the two decoding modules, which is composed of a channel attention module and a spatial attention module in series; Among them, the channel attention obtains global information through global pooling to guide the learning of the decoder; The spatial attention implicitly obtains the noise level map through backpropagation, which can distinguish the noise intensity of the input image at different positions and generate a more accurate noise map.
[0021] In step S2.1, N is set to 3.
[0022] In step S2.1, in order to further reduce the number of parameters, depthwise separable convolutions are used to reduce the number of parameters.
[0023] In step S2.2, the process of applying the channel attention to the input is shown in formula (2):
[0024]
[0025] The generation of a more accurate noise map is shown in formula (3):
[0026]
[0027] Among them, in Formulas (2) and (3), FC represents the fully connected layer, gap and gmp respectively represent global average pooling and global max pooling, Conv represents 1*1 convolution, and σ represents the sigmoid activation function. represents element-wise multiplication.
[0028] In step S3, the loss between the predicted image and the reference image is calculated using the MAE function, as shown in Formula (4):
[0029]
[0030] Thus, the advantages of this application are as follows:
[0031] (1) Using 5*5 convolution in the encoder part increases the receptive field and can capture more context information. Inserting channel attention and spatial attention modules in the decoder part introduces global information and noise intensity information at different positions, enhancing the denoising ability and detail recovery ability of the network.
[0032] (2) The improved U-net has fewer parameters and lower computational complexity, making it more suitable for deployment on resource-constrained edge devices. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] The drawings described herein are used to provide a further understanding of the present invention, form a part of this application, and do not limit the present invention.
[0034] Figure 1 is a schematic flowchart of the method of this application.
[0035] Figure 2 is a schematic diagram of the improved U-net network structure involved in this application.
[0036] Figure 3 is a schematic diagram showing that the encoding module involved in this application is composed of a downsampling module and N basic modules.
[0037] Figure 4 is a schematic diagram showing that the decoding module involved in this application is composed of a basic residual module and a transposed convolution.
[0038] Figure 5 is a schematic diagram showing that the attention module involved in this application is composed of a channel attention module and a spatial attention module connected in series. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0039] In order to more clearly understand the technical content and advantages of the present invention, the present invention will be further described in detail below with reference to the drawings.
[0040] This application provides a method for denoising RAW domain low-light images by improving U-net. As Figure 1 shown, this method includes:
[0041] Step S1. Collect a RAW domain low-light image dataset and preprocess the image data: Collect noisy images and corresponding noise-free reference images, and perform preprocessing and image enhancement operations to further expand the scale of the training dataset. The specific steps are as follows:
[0042] Step S1.1, Image acquisition. Select indoor and outdoor scenes. For the same scene, adjust the illuminance through a fill light, and collect K noisy images at each illuminance. After averaging, use it as the reference image, and randomly select one from the K images as the paired noisy image.
[0043] The value of K depends on the illuminance of the scene. For example, set K = 500 at an illuminance of 0.1 lux and set K = 200 at an illuminance of 1 lux.
[0044] Step S1.2, Data preprocessing. Perform black level correction and normalization on the RAW image in Bayer format, as shown in the following formula:
[0045]
[0046] In formula (1), bit is the number of bits of the RAW image. After black level correction and normalization, convert the single-channel output image to a 4-channel image, and reduce the image spatial resolution to 1 / 4 of the original. On the one hand, reducing the resolution can reduce the computational overhead of the network. On the other hand, it can avoid destroying the Bayer format when performing data augmentation operations.
[0047] Step S1.3, Data augmentation. Randomly crop the image after data preprocessing, and randomly rotate and flip the cropped image block. Use the enhanced image block as the input of the network.
[0048] Step S2. Load it into the improved U-net model: Design a lightweight network structure with reference to the U-net form. The input of the network is the noisy image, and the output is the predicted noise map. The network structure is as Figure 2 shown. The specific improvements are as follows:
[0049] Step S2.1, In order to reduce the network depth, only perform downsampling twice in the encoding stage. For an image with an input spatial resolution of H×W, after passing through the encoder, the spatial resolution is only reduced to The entire encoding stage consists of two convolutional layers and two encoding modules. Among them, the encoding module consists of a downsampling module and N basic modules, as Figure 3As shown, where N is used to control the depth and computational complexity of the network, and is set to 3 in this embodiment. To ensure that the encoder has a large enough receptive field to capture context information, 5*5 convolutions are used instead of 3*3 convolutions. To further reduce the number of parameters, depthwise separable convolutions are used. The receptive field is the size of the area on the input image that the pixels on the feature map output by each layer of the convolutional neural network are mapped back to.
[0050] Step S2.2: The decoding stage consists of two decoding modules, an attention module, and two convolutional layers. The decoding module consists of a basic residual module and a transposed convolution, as Figure 4 shown. To reduce the number of parameters, 3*3 convolutions are used in the convolutional layers of the decoder. An attention module is inserted after the two decoding modules, which is composed of a channel attention module and a spatial attention module in series, as Figure 5 shown. Among them, channel attention obtains global information through global pooling to guide the learning of the decoder. The process of applying channel attention to the input is shown in the following formula:
[0051]
[0052] Spatial attention implicitly obtains the noise level map through backpropagation, can distinguish the noise intensity of the input image at different positions, and generates a more accurate noise map, as shown in the following formula:
[0053]
[0054] In formulas (2) and (3), FC represents the fully connected layer, gap and gmp respectively represent global average pooling and global max pooling, Conv represents 1*1 convolution, σ represents the sigmoid activation function, represents element-wise multiplication.
[0055] Step S3. Train the network: The model is trained using the MAE loss function and the Adam optimizer; the preprocessed and data-augmented noisy image X input is input into the improved U-net, and the predicted noise map is obtained through forward propagation The MAE function is used to calculate the loss between the predicted image and the reference image, as shown in the following formula:
[0056]
[0057] The Adam optimizer is used to update the network parameters. During the training process, the cosine annealing learning rate decay strategy is used, and the half-period of the cosine function is set to the maximum number of iterations.
[0058] Step S4. Load the complete image into the trained model and output the denoised image: Load the uncropped noisy image after preprocessing into the trained network for denoising to obtain the denoised image.
[0059] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, various changes and modifications can be made to the embodiments of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. An improved RAW domain low-light image denoising method for U-net, characterized in that, The method includes the following steps: S1. Collect noisy images and corresponding noise-free reference images, and perform preprocessing and image enhancement operations to expand the scale of the training data set; S2. Refer to the U-net form and design a lightweight network structure. The input of the network is the noisy image, and the output is the predicted noise map; S3. Train the network: Input the preprocessed and data-augmented noisy image X input into the improved U-net, and obtain the predicted noise map through forward propagation Calculate the loss between the predicted image and the reference image using the MAE function, update the network parameters using the Adam optimizer, use the cosine annealing learning rate decay strategy during training, and set the half period of the cosine function to the maximum number of iterations; S4. Load the uncropped noisy image after preprocessing into the trained network for denoising to obtain the denoised image.
2. An improved U-net RAW domain low-light image denoising method according to claim 1, characterized in that, Step S1 further includes: S1.1, Image acquisition: Select indoor and outdoor scenes. For the same scene, adjust the illuminance through a fill light, collect K noisy images at each illuminance, take the average as the reference image, and randomly select one from the K images as the paired noisy image; where the value of K depends on the illuminance of the scene; S1.2, Data preprocessing: Perform black level correction and normalization on the RAW image in Bayer format as shown in the following formula (1): In formula (1), bit is the number of bits of the RAW image. After black level correction and normalization, convert the single-channel output image to 4 channels, and reduce the image spatial resolution to 1 / 4 of the original; S1.3, Data augmentation: Randomly crop the image after data preprocessing, and randomly rotate and flip the cropped image patches. Use the enhanced image patches as the input of the network.
3. An improved U-net RAW domain low-light image denoising method according to claim 2, characterized in that In step S1.1, the value of K depends on the illuminance of the scene, including setting K = 500 at an illuminance of 0.1 lux and setting K = 200 at an illuminance of 1 lux.
4. An improved U-net RAW domain low-light image denoising method according to claim 1, characterized in that, The network structure described in step S2 includes: S2.
1. To reduce the network depth, only two downsamplings are performed during the encoding stage. For an image with an input spatial resolution of H×W, after passing through the encoder, the spatial resolution is only reduced to The entire encoding stage consists of two convolutional layers and two encoding modules. Each encoding module is composed of a downsampling module and N basic modules, where N is used to control the network depth and computational complexity. To ensure that the encoder has a large enough receptive field to capture context information, 5*5 convolutions are used instead of 3*3 convolutions; S2.
2. The decoding stage consists of two decoding modules, an attention module, and two convolutional layers. Each decoding module is composed of a basic residual module and a transposed convolution; To reduce the number of parameters, 3*3 convolutions are used in the convolutional layers of the decoder; An attention module is inserted after the two decoding modules, which is composed of a channel attention module and a spatial attention module in series; Among them, the channel attention obtains global information through global pooling to guide the learning of the decoder; The spatial attention implicitly obtains the noise level map through backpropagation, which can distinguish the noise intensity of the input image at different positions and generate a more accurate noise map.
5. An improved U-net RAW domain low-light image denoising method according to claim 4, characterized in that, In step S2.1, N is set to 3.
6. An improved U-net RAW domain low-light image denoising method according to claim 4, characterized in that In step S2.1, in order to further reduce the number of parameters, depthwise separable convolutions are used to reduce the number of parameters.
7. An improved U-net RAW domain low-light image denoising method according to claim 4, characterized in that, In step S2.2, the process of applying the channel attention to the input is shown in formula (2): The generation of a more accurate noise map is shown in formula (3): Among them, in Formulas (2) and (3), FC represents the fully connected layer, gap and gmp respectively represent global average pooling and global max pooling, Conv represents the 1*1 convolution, and σ represents the sigmoid activation function. represents element-wise multiplication.
8. An improved U-net RAW domain low-light image denoising method according to claim 1, characterized in that, In step S3, the MAE function is used to calculate the loss between the predicted image and the reference image, as shown in formula (4):