A low-illumination color image enhancement method based on image decomposition
Patent Information
- Application Number
- CN202310160230.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-24
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2043-02-24
AI Technical Summary
[0005]针对现有技术的不足,本发明提供了一种基于图像分解的低照度彩色图像增强方法,解决现有的低照度彩色图像增强方法得到的图像对比度差和细节丢失的问题
[0022]与现有技术相比,本发明提供了一种基于图像分解的低照度彩色图像增强方法,具备以下有益效果:
Smart Images

Figure CN116258645B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing methods, specifically to a low-light color image enhancement method based on image decomposition. Background Technology
[0002] Images are the most intuitive way for humans to understand the world and are the most important channel for acquiring and exchanging information. However, image quality is closely related to the shooting environment. For example, images captured in adverse weather, at night, or under satellite remote control often suffer from low contrast, severe color distortion, and significant noise. Low-light color image enhancement primarily studies how to quickly and effectively improve image quality and enhance the visual experience of images acquired in low-light conditions, which often exhibit low contrast and high noise levels, thus enabling more timely responses to adverse weather and nighttime events. However, existing technologies suffer from high computational costs, weak image enhancement effects, and an inability to adequately reveal details in dark areas.
[0003] Chinese patent publication number "CN115018711A" is titled "An Improved Retinex-Net Low-Light and Dark-Vision Image Enhancement Method." This method first preprocesses the low-light or dark-visual image to obtain the reflectance and illumination components. Then, the processed image is input into an improved network. Next, two copies of the adjusted illumination map are copied and stitched together with the original image to obtain a stitched illumination map. Then, two copies of the denoised reflectance map are copied and stitched together with the original image to obtain a stitched reflectance map. Finally, the obtained images are fused to form an enhanced image. However, the enhanced image obtained by this method does not conform to human visual perception and suffers from artifacts at the stitching points, poor image contrast, and loss of detail. Summary of the Invention
[0004] (a) Technical problems to be solved
[0005] To address the shortcomings of existing technologies, this invention provides a low-light color image enhancement method based on image decomposition, which solves the problems of poor image contrast and loss of detail obtained by existing low-light color image enhancement methods.
[0006] (II) Technical Solution
[0007] To achieve the above objectives, the present invention specifically adopts the following technical solution:
[0008] A low-light color image enhancement method based on image decomposition includes the following steps:
[0009] Step 1: Construct the network model: The entire network model consists of three parts: a decomposition module, an enhancement module, and a reconstruction module. These three modules comprise ten convolutional blocks, each containing skip connections, concatenation operations, convolutional layers, and activation functions. Convolutional blocks 1, 2, 3, and 4 belong to the decomposition module; convolutional blocks 5, 6, 7, 8, and 9 belong to the enhancement module; and convolutional block 10 is the reconstruction module.
[0010] Step 2: Preprocess the data: Convert the low-light color image from RGB to HSV. The HSV color space is represented by an inverted cone, where H represents hue, S represents saturation, and V represents brightness. Here, the H and S components of the image are preserved, and the V component of the image is extracted.
[0011] Step 3: Input the V component image obtained in Step 2 into the network model obtained in Step 1 for training;
[0012] Step 4: Select the loss function and optimal evaluation metric: Train the relative total variation model by defining the loss function until the number of training iterations reaches a set threshold or the value of the loss function reaches a set range. The model parameters are then considered to have been pre-trained and saved. At the same time, select the optimal evaluation metric to measure the accuracy of the algorithm and evaluate the performance of the system.
[0013] Step 5: Fine-tune the model: Train and fine-tune the model using low-light color images to obtain stable and usable model parameters;
[0014] Step 6: Save the model: Fix the finalized model parameters. When low-light color image enhancement is needed, simply input the image into the network to obtain the final enhanced V component image. Then, combine it with the H and S components retained in Step 2 to synthesize the complete enhanced low-light color image.
[0015] Furthermore, the network model structure in step 1 consists of three parts: a decomposition module, an enhancement module, and a reconstruction module. The three modules together consist of ten convolutional blocks, each of which comprises skip connections, concatenation operations, convolutional layers, and activation functions. The decomposition module consists of convolutional block 1, convolutional block 2, convolutional block 3, and convolutional block 4; the enhancement module consists of convolutional block 5, convolutional block 6, convolutional block 7, convolutional block 8, and convolutional block 9; and the reconstruction module consists of convolutional block 10.
[0016] Furthermore, the decomposition module in step 1 mainly decomposes the V component image obtained in step 2 to obtain the structural layer and detail layer; it consists of convolutional block one, convolutional block two, convolutional block three and convolutional block four. The system consists of convolutional layers, upsampling layers, downsampling layers, linear activation functions, and sigmoid activation functions. Except for convolutional block four, which uses a sigmoid activation function to form the output block, producing the structure layer and detail layer, the other three convolutional blocks use linear activation functions. Convolutional block one has 3 input feature channels and 64 output feature channels; convolutional blocks two and three both have 64 input and output feature channels; convolutional block four has 64 input feature channels and 3 output feature channels. The skip connections between convolutional blocks one and two, and between convolutional blocks one and three, are mainly for extracting and fusing feature maps during the sampling process, stacking them according to the number of feature map channels. Convolutional blocks one, two, and three are mainly for extracting and fusing feature maps during the sampling process, then inputting them into convolutional block four. Convolutional block four serves as the output block, outputting the structure layer and detail layer to the enhancement module to complete image decomposition.
[0017] Further, in step 1, the enhancement module mainly performs brightness enhancement, denoising, and detail enhancement on the structural layer image and detail layer image obtained from power requirement 3, respectively; it consists of convolutional blocks five, six, seven, eight, and nine; each convolutional block includes a convolutional layer, a linear activation function, a concatenation operation, and bilinear interpolation, with kernel sizes of 3×3 and 1×1; convolutional blocks five, six, and seven enhance the brightness of the structural layer obtained from the decomposition module, while convolutional blocks eight and nine denoise and enhance the detail layer obtained from the decomposition module; finally, the image is input into the reconstruction module to complete the image enhancement.
[0018] Furthermore, in step 1, the reconstruction module fuses the structural layer and detail layer obtained according to the power requirement 4 to complete image reconstruction and obtain the final enhanced V image component; and it is composed of convolutional blocks.
[0019] Furthermore, a relative total variation loss function is used during training.
[0020] Furthermore, the low-light color image dataset is the low dataset of our485 in LOL.
[0021] (III) Beneficial Effects
[0022] Compared with existing technologies, this invention provides a low-light color image enhancement method based on image decomposition, which has the following beneficial effects:
[0023] This invention proposes using a relative total variation model as the loss function to amplify the difference between structural edges and texture details, and effectively remove texture information from intersecting regions in the image, resulting in a smoother structural layer image. The network model structure of the decomposition module also reduces the learning pressure on the network, accelerates its convergence speed, and improves its stability. The resulting enhanced image has less noise, better contrast, and more detail, effectively improving the image quality of low-light color images. Attached Figure Description
[0024] Figure 1 This is a flowchart of a low-light color image enhancement method based on image decomposition according to the present invention.
[0025] Figure 2 This is a network structure diagram of a low-light color image enhancement method based on image decomposition according to the present invention.
[0026] Figure 3 This is a diagram showing the specific components of the decomposition module in the network model.
[0027] Figure 4 This is a diagram showing the specific components of the enhancement module in the network model;
[0028] Figure 5 A comparison chart of relevant indicators of existing technologies and the method of this invention. Detailed Implementation
[0029] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0030] Example
[0031] like Figure 1 As shown, a low-light color image enhancement method based on image decomposition is presented, which specifically includes the following steps:
[0032] Step 1: Construct the network model: The entire network model consists of three parts: a decomposition module, an enhancement module, and a reconstruction module. These three modules comprise ten convolutional blocks, each containing skip connections, concatenation operations, convolutional layers, and activation functions. Convolutional blocks 1, 2, 3, and 4 belong to the decomposition module; convolutional blocks 5, 6, 7, 8, and 9 belong to the enhancement module; and convolutional block 10 is the reconstruction module.
[0033] Step 2: Preprocess the data: Convert the low-light color image from RGB to HSV. The HSV color space is represented by an inverted cone, where H represents hue, S represents saturation, and V represents brightness. Here, the H and S components of the image are preserved, and the V component of the image is extracted.
[0034] Step 3: Input the V component image obtained in Step 2 into the network model obtained in Step 1 for training;
[0035] Step 4: Select the loss function and optimal evaluation metric: The relative total variation model is trained by defining the loss function until the number of training iterations reaches a set threshold or the loss function value falls within a set range. The model parameters are then considered pre-trained and saved. Simultaneously, the optimal evaluation metric is selected to measure the algorithm's accuracy and evaluate the system's performance. Suitable evaluation metrics include structural similarity (SSIM), information entropy (H(x)), and peak signal-to-noise ratio (PSNR), which can effectively evaluate the algorithm's accuracy and efficiency and measure the role of the image enhancement network.
[0036] Step 5: Fine-tune the model: Train and fine-tune the model using low-light color images to obtain stable and usable model parameters;
[0037] Step 6: Save the model: Fix the finalized model parameters. When low-light color image enhancement is needed, simply input the image into the network to obtain the final enhanced V component image. Then, combine it with the H and S components retained in Step 2 to synthesize the complete enhanced low-light color image.
[0038] The network model structure in step 1 is as follows: Figure 2 As shown.
[0039] The entire network model structure is mainly divided into three parts: decomposition module, enhancement module, and reconstruction module. The decomposition module mainly decomposes the V component image obtained in step 2 to obtain the structural layer and detail layer; the enhancement module mainly performs detail enhancement and denoising operations on the detail layer obtained by the decomposition module, and brightness enhancement on the obtained structural layer to obtain the enhanced detail layer and structural layer; the reconstruction module mainly fuses the enhanced structural layer and detail layer images obtained by the enhancement module to obtain the final enhanced V image component.
[0040] The decomposition module mainly consists of convolutional block 1, convolutional block 2, convolutional block 3, and convolutional block 4. The specific structure of the decomposition module is shown in the diagram below. Figure 3As shown, the convolutional blocks are composed of the following: Convolutional Block 1 consists of a convolutional layer, a linear activation function, and a downsampling layer; the downsampling layer is a deconvolutional layer with a stride of 2 and a kernel size of 3×3; it has 3 input feature channels and 64 output feature channels. Convolutional Block 2 consists of a convolutional layer, a linear activation function, a downsampling layer, and an upsampling layer; the upsampling layer is a convolutional layer with a stride of 2, the downsampling layer is a deconvolutional layer with a stride of 2, and the kernel size of 3×3. Convolutional Block 3 consists of a convolutional layer, a linear activation function, and an upsampling layer; the upsampling layer is a convolutional layer with a stride of 2 and a kernel size of 3×3. Both Convolutional Block 2 and Convolutional Block 3 have 64 input and 64 output feature channels. Convolutional block 4 consists of convolutional layers and a sigmoid activation function; the kernel size is 3×3; the input feature channels are 64, and the output feature channels are 3; the sigmoid activation function controls the output range within [0,1], making the data less prone to divergence during transmission. The skip connections between convolutional blocks 1 and 2, and between convolutional blocks 1 and 3, are mainly for extracting and fusing feature maps during the sampling process, stacking them according to the number of feature map channels, and inputting the fused feature map into convolutional block 4; convolutional block 4 then serves as the output block, outputting the decomposed structural and detail layers. In summary, convolutional blocks 1, 2, and 3 mainly complete the extraction and fusion of feature maps, while convolutional block 4 completes the output. This network structure can effectively extract image structure, retain more texture information, reduce the learning pressure on the network, and accelerate network convergence.
[0041] The enhancement module mainly consists of convolutional block five, convolutional block six, convolutional block seven, convolutional block eight, and convolutional block nine. The specific structure of the enhancement module is shown in the diagram below. Figure 4 As shown, the network consists of the following convolutional blocks: Convolutional block 5 consists of cubic convolution and a linear activation function; the kernel size is 3×3. Convolutional block 6 consists of cubic convolution and bilinear interpolation; the kernel size is 3×3. Convolutional block 7 consists of cubic convolution and a linear activation function; the kernel size is 1×1, and the concatenation operation uses concat. Convolutional block 8 consists of convolution and a linear activation function; the kernel size is 3×3. Convolutional block 9 consists of residual connections and convolutional layers; residual connections help the network preserve the signal, and the 1×1 convolution learned end-to-end throughout the network adjusts the trade-off between denoising and signal preservation. This removes a large amount of noise from the detail layer image while preserving important information. In summary, the functions of convolutional blocks 5, 6, and 7 are to enhance the brightness of low-light images at different scales, complete the brightness restoration of the structural layer image, and make the enhanced structural layer image brightness close to that of the natural light image. Convolutional blocks 8 and 9 are mainly responsible for restoring details and removing noise from the detail layer image, thereby reducing noise and making the details clearer.
[0042] Convolutional block 10 belongs to the reconstruction module, which mainly fuses the structural layer and detail layer images generated by the enhancement module to obtain the final enhanced V image component.
[0043] The linear activation function and the sigmoid activation function are defined as follows:
[0044]
[0045]
[0046] In step 2, the dataset preprocessing involves converting the low-light color images in the dataset from RGB to HSV images. The HSV color space is represented by an inverted cone, where H represents hue, S represents saturation, and V represents brightness. Here, the H and S components of the image are preserved, and the V component image is extracted for subsequent network training.
[0047] In step 4, a loss function and optimal evaluation metric are selected. To make the output structure layer as similar as possible to the structure of the input image and to amplify the difference between structure edges and texture details, resulting in better edge smoothing of the output structure layer, a relative total variation model is chosen as the loss function. This model can effectively remove texture information from intersecting regions in the image. The formula for the relative total variation loss function is as follows:
[0048]
[0049]
[0050]
[0051] S represents the decomposed structural layer, and I represents the original input image. This refers to the difference between the input image and the output structural layer image, D x (p) and D y (p) refers to the structural and texture measurements in the x and y directions at point p. It represents the gradient of the structural component in the x and y directions at point p. It is an Lp norm. and The values of M are both in the range [0,1], representing the structure confidence and texture confidence of point p, respectively; x Structural edges are filtered out; when the structural confidence of point p is high... Approaching 1, When the gradient confidence level approaches 0, the gradient regularization term also approaches 0, thus preserving the structural edges of point p; conversely, when the texture confidence level of point p is high, When the gradient regularization term approaches 1, it is relatively large and can separate texture edges located at point p in the structural components. and yes and The balance coefficient (the same applies to the x and y directions, so it will not be repeated).
[0052] In step 4, appropriate evaluation metrics are selected during the training process: structural similarity (SSIM), information entropy (H(x)), and peak signal-to-noise ratio (PSNR).
[0053] Structural similarity is used to compare the similarity between two images. The value is fixed in the range of [0,1]. The closer the value is to 0, the lower the similarity between the two images. Conversely, the closer the value is to 0, the higher the similarity between the two images.
[0054] Information entropy is used to measure the amount of information in an image; the higher the information entropy, the richer the information contained in the image.
[0055] Peak signal-to-noise ratio (PSNR) is a commonly used objective method for evaluating image quality. It is usually simply represented by mean square error, and a larger value indicates better image quality.
[0056] The training iterations are set to 100, with each iteration inputting approximately 8-16 images. The upper limit for the number of images input depends on the computer's graphics processing unit (GPU) performance; generally, a larger number of images input per iteration is better for network stability. The learning rate is set to 0.0001 to ensure fast network fitting without overfitting. The adaptive moment estimation algorithm is chosen as the network parameter optimizer. Its advantage lies in the fact that after bias correction, the learning rate has a defined range for each iteration, resulting in relatively stable parameters. The loss function threshold is set to approximately 0.0003; a value less than 0.0003 or more iterations than a preset value indicates that the network training is essentially complete.
[0057] In step 5, the LOL low-light image dataset is used during the fine-tuning of model parameters. Specifically, a subset of this dataset, our485, is used, which provides 485 low-light color images. We use 400 images for training and 200 images for testing.
[0058] After the network training is completed in step 6, all parameters of the network need to be saved. Then, the V component image to be enhanced is input into the network to obtain the enhanced V component image. This enhanced V component image is then synthesized with the H and S components retained in step 2 to form a complete enhanced low-light color image. This network has no requirements on the size of the input image; any size is acceptable.
[0059] Among them, the implementation of convolution, activation function, skip connection, splicing operation and Gaussian filtering are algorithms known to those skilled in the art, and the specific process and method can be found in the relevant textbooks or technical documents.
[0060] This invention constructs a low-light color image enhancement method based on image decomposition, which can directly obtain the enhanced image without intermediate steps, thus avoiding manually designed low-light image enhancement rules. Under the same conditions, the feasibility and superiority of this method are further verified by calculating the relevant indicators of the images obtained by existing methods. A comparison of the relevant indicators of existing technologies and the method proposed in this invention is provided below. Figure 5 As shown.
[0061] from Figure 5 As can be seen from the above, the method proposed in this invention has higher structural similarity (SSIM), information entropy (H(x)), and peak signal-to-noise ratio (PSNR), which further demonstrates that the method proposed in this invention has better image quality.
[0062] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A low-light color image enhancement method based on image decomposition, characterized in that: Includes the following steps, Step 1: Construct the network model: The entire network model consists of three parts: a decomposition module, an enhancement module, and a reconstruction module. These three modules comprise ten convolutional blocks, each containing skip connections, concatenation operations, convolutional layers, and activation functions. Convolutional blocks 1, 2, 3, and 4 belong to the decomposition module; convolutional blocks 5, 6, 7, 8, and 9 belong to the enhancement module; and convolutional block 10 is the reconstruction module. Step 2: Preprocess the data: Convert the low-light color image from RGB to HSV. The HSV color space is represented by an inverted cone, where H represents hue, S represents saturation, and V represents brightness. Here, the H and S components of the image are kept unchanged, and the V component of the image is extracted. Step 3: Input the V component image obtained in Step 2 into the network model obtained in Step 1 for training; Step 4: Select the loss function and optimal evaluation metric: Train the relative total variation model by defining the loss function until the number of training iterations reaches a set threshold or the value of the loss function reaches a set range. The model parameters are then considered to have been pre-trained and saved. At the same time, select the optimal evaluation metric to measure the accuracy of the algorithm and evaluate the performance of the system. The formula for the relative total variation loss function is as follows: ; ; ; S represents the decomposed structural layer, and I represents the original input image. This refers to the difference between the input image and the output structural layer image. and This refers to the structural and texture measurements in the x and y directions at point p. It represents the gradient of the structural component in the x and y directions at point p. It is L p Norm; and The values of are all in the range of [0, 1], representing the structure confidence and texture confidence of point p, respectively; and yes and The balance coefficients are calculated similarly for the x and y directions; Step 5: Fine-tune the model: Train and fine-tune the model using low-light color images to obtain stable and usable model parameters; Step 6: Save the model: Fix the final model parameters. When low-light color image enhancement is needed, simply input the image into the network to obtain the final enhanced V component image. Then, combine it with the H and S components retained in Step 2 to synthesize the complete enhanced low-light color image.
2. The low-light color image enhancement method based on image decomposition according to claim 1, characterized in that: In step 1, the network model structure consists of three parts: a decomposition module, an enhancement module, and a reconstruction module. The three modules together consist of ten convolutional blocks. Each convolutional block is composed of skip connections, splicing operations, convolutional layers, and activation functions. The decomposition module consists of convolutional block 1, convolutional block 2, convolutional block 3, and convolutional block 4. The enhancement module consists of convolutional block 5, convolutional block 6, convolutional block 7, convolutional block 8, and convolutional block 9. The reconstruction module consists of convolutional block 10.
3. The low-light color image enhancement method based on image decomposition according to claim 1, characterized in that: In step 1, the decomposition module decomposes the V component image obtained in step 2 to obtain the structural layer and detail layer. It consists of convolutional blocks 1, 2, 3, and 4, each comprising convolution, upsampling, downsampling, linear activation functions, and sigmoid activation functions. Except for convolutional block 4, which uses a sigmoid activation function to form the output block, outputting the structural and detail layers, the other three convolutional blocks use linear activation functions. Convolutional block 1 has 3 input feature channels and 64 output feature channels. Convolutional blocks 2 and 3... Both input and output feature channels are 64; Convolutional Block 4 has 64 input feature channels and 3 output feature channels; the skip connections between Convolutional Block 1 and Convolutional Block 2, and between Convolutional Block 1 and Convolutional Block 3 are for extracting and fusing feature maps during the sampling process, stacking them according to the number of feature map channels; Convolutional Block 1, Convolutional Block 2, and Convolutional Block 3 are for completing the feature map extraction and fusing of feature maps during the sampling process, and then inputting them into Convolutional Block 4; Convolutional Block 4 serves as the output block, outputting the structural layer and detail layer to the enhancement module to complete image decomposition.
4. The low-light color image enhancement method based on image decomposition according to claim 1, characterized in that: In step 1, the enhancement module performs brightness enhancement, noise reduction, and detail enhancement on the structural layer image and detail layer image obtained according to claim 3, respectively. It consists of convolutional blocks five, six, seven, eight, and nine. Each convolutional block comprises a convolutional layer, a linear activation function, a concatenation operation, and bilinear interpolation, with kernel sizes of 3×3 and 1×1. Convolutional blocks five, six, and seven enhance the brightness of the structural layer obtained from the decomposition module, while convolutional blocks eight and nine perform noise reduction and detail enhancement on the detail layer obtained from the decomposition module. Finally, the image is input into the reconstruction module to complete the image enhancement.
5. The low-light color image enhancement method based on image decomposition according to claim 1, characterized in that: In step 1, the reconstruction module fuses the structural layer and detail layer obtained according to claim 4 to complete image reconstruction and obtain the final enhanced V image component; and it is composed of convolutional blocks.
6. The low-light color image enhancement method based on image decomposition according to claim 1, characterized in that: The relative total variation loss function is used during training.
7. The low-light color image enhancement method based on image decomposition according to claim 1, characterized in that: The low-light color image dataset is the low dataset from our485 in LOL.
Citation Information
Patent Citations
Low-illumination color image enhancement method based on Retinex and convolutional neural network
CN110232661A