Low-light image enhancement method based on HVI color space and multi-scale cavity convolution

By converting low-light images into the HVI color space and combining it with a multi-scale dilated convolution module, the problem of impaired image quality under low-light conditions is solved, the brightness and color are effectively enhanced, and the visual effect and applicability of the image are improved.

CN120672630APending Publication Date: 2025-09-19QINGHAI NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510579065.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Image quality is impaired under low-light conditions, and existing technologies find it difficult to effectively restore the brightness, contrast, and details of images, resulting in limited information acquisition and utilization for high-level visual tasks.

Method used

A low-light image enhancement method based on the HVI color space and multi-scale dilated convolution is used. By converting the low-light image to the HVI color space, it is split into three independent components: hue (H), brightness (V), and illumination (I). Brightness and color enhancement are then performed separately. Multiple dilated convolution modules are then used to extract multi-scale features, enhance the image's brightness, and ultimately restore the image to normal illumination.

Benefits of technology

While enhancing brightness, it maintains relative color stability, avoids color cast problems, and simulates the human eye's perception of brightness and color, making the enhanced image visual effect more natural and realistic. It is suitable for image processing tasks in a variety of complex lighting environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120672630A_ABST
    Figure CN120672630A_ABST
Patent Text Reader

Abstract

The invention discloses a low-light image enhancement method based on an HVI color space and multi-scale cavity convolution, and the method comprises the steps: S1, converting an inputted low-light image from an RGB color space to the HVI color space, and carrying out the enhancement of the color features of the low-light image, and obtaining a color feature enhanced image; s2, grouping the low-light image and the color feature enhanced image in the S1 through channel dimensions, and respectively extracting space attention information of each group; s3, fusing the color feature enhanced image in the S1 and each group of space attention information extracted in the S2; and S4, taking the fused feature information in the S3 as input information, performing image brightness enhancement through a brightness enhancement network, and finally outputting a recovered normal illumination image. According to the invention, the brightness enhancement and the color adjustment are respectively carried out, the separation mode is beneficial to keeping the relative stability of the color while enhancing the brightness, and the problem of color cast possibly introduced during overall adjustment is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention mainly relates to the technical field of image enhancement methods, and specifically to a low-light image enhancement method based on HVI color space and multi-scale dilated convolution. Background Art

[0002] Images are a vital source of information for humans, and vision is one of the primary ways we receive external information in daily life. Object detection, however, leverages multidisciplinary knowledge, including image processing, pattern recognition, and machine learning, to computationally locate and identify objects of interest in images or videos. Combining object detection and recognition is more challenging than image classification, but the performance of detection and recognition algorithms often depends on the quality of the captured images. However, the image capture process often suffers from a variety of uncontrollable physical factors, which can degrade image quality and, in turn, hinder the acquisition and utilization of information in high-level visual tasks such as face recognition, semantic segmentation, and feature extraction. Among the various factors that affect image quality, low light, such as in cloudy or nighttime scenes, is particularly common and unavoidable. While deep learning has made significant progress in low-light image enhancement in recent years, restoring brightness, contrast, and detail remains a significant challenge. Summary of the Invention

[0003] The purpose of this invention is to provide a low-light image enhancement method based on HVI color space and multi-scale dilated convolution. The specific technical solution is as follows:

[0004] A low-light image enhancement method based on HVI color space and multi-scale dilated convolution includes: S1, converting the input low-light image from RGB color space to HVI color space, and enhancing the color features of the low-light image to obtain a color feature enhanced image; S2, grouping the low-light image and the color feature enhanced image in S1 according to the channel dimension, and extracting the spatial attention information of each group respectively; S3, fusing the color feature enhanced image in S1 with the spatial attention information of each group extracted in S2; S4, using the fused feature information in S3 as input information to enhance the brightness of the image through a brightness enhancement network, and finally outputting a restored normal light image.

[0005] When converted to the HVI color space in S1, the color information of the low-light image is split into three independent components: hue H, brightness V, and illumination component I. A weight-aware adaptive factor adjustment mechanism is introduced to adjust and / or enhance the brightness and color of the low-light image respectively.

[0006] The spatial attention information of each group is extracted in S2, including: S2.1, performing global average pooling and global maximum pooling in the height H direction and width W direction respectively to obtain global spatial information; S2.2, using two 1×1 convolution layers to perform feature dimensionality reduction and recovery to enhance feature representation capabilities; S2.3, calculating the attention weights in the H direction and W direction through the Sigmoid activation function, and the calculated attention weights in the H direction and W direction.

[0007] The brightness enhancement network in S4 adopts a U-Net-based encoding-decoding structure and combines the image with a multi-dilated convolutional module (MDFA) for feature extraction.

[0008] The multi-dilation convolution module (MDFA) includes: the first branch, which adopts 1×1 convolution operation without setting the dilation rate, keeps the channel dimension of the input image unchanged, and is used to extract local detail features and process detail information without changing the size of the receptive field; the second branch, which adopts 3×3 convolution with a dilation rate of 9, increases the size of the receptive field and is used to capture medium-scale contextual information, while retaining local details while obtaining more contextual information; the third branch, which adopts 3×3 convolution with a dilation rate of 18, further expands the receptive field and extracts larger-scale image features to capture a wider range of information; the fourth branch, which adopts 3×3 convolution with a dilation rate of 27, maximizes the receptive field and is used to extract global information of the image and restore global brightness and color.

[0009] The multiple dilated convolution module (MDFA) also includes a global feature extraction branch, which extracts global information of the image through global average pooling and 1×1 convolution to help the model understand the brightness and color distribution of the entire image and fuse it with the feature maps of other branches.

[0010] The above-mentioned technical solution of the present invention has the following beneficial technical effects: the present invention performs brightness enhancement and color adjustment separately. This separation method helps to maintain the relative stability of color while enhancing brightness, avoids the color cast problem that may be introduced during overall adjustment, and can more effectively simulate the human eye's perception characteristics of brightness and color, making the enhanced image visual effect more natural and realistic, and can be applied to image processing tasks in a variety of complex lighting environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 It is a schematic flow diagram of the present invention;

[0012] Figure 2 This is a schematic diagram of the GGCA module design in the present invention;

[0013] Figure 3 It is a schematic diagram of the MDFA module design in the present invention. DETAILED DESCRIPTION

[0014] To make the objectives, technical solutions, and advantages of the present invention more clearly understood, the present invention will be further described in detail below in conjunction with specific embodiments and with reference to the accompanying drawings. It should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessary confusion of the concepts of the present invention.

[0015] like Figure 1-3 As shown, a low-light image enhancement method based on HVI color space and multi-scale dilated convolution includes: S1, converting the input low-light image from RGB color space to HVI color space, and enhancing the color features of the low-light image to obtain a color feature enhanced image; S2, grouping the low-light image and the color feature enhanced image in S1 by channel dimension, and extracting spatial attention information of each group respectively; S3, fusing the color feature enhanced image in S1 with the spatial attention information of each group extracted in S2; S4, using the fused feature information in S3 as input information to enhance the brightness of the image through a brightness enhancement network, and finally outputting a restored normal light image. The present invention performs brightness enhancement and color adjustment separately. This separation method helps to maintain the relative stability of color while enhancing brightness, avoids the color cast problem that may be introduced during overall adjustment, and can more effectively simulate the human eye's perception characteristics of brightness and color, making the enhanced image visual effect more natural and realistic, and can be applied to image processing tasks in a variety of complex lighting environments.

[0016] When converted to the HVI color space in S1, the color information of the low-light image is split into three independent components: hue H, brightness V, and illumination component I. A weight-aware adaptive factor adjustment mechanism is introduced to adjust and / or enhance the brightness and color of the low-light image respectively.

[0017] Extracting spatial attention information for each group in S2 involves: S2.1. Global average pooling and global max pooling are performed in the height H and width W directions, respectively, to obtain global spatial information; S2.2. Two 1×1 convolutional layers are used to reduce and restore features to enhance feature representation; and S2.3. Attention weights in the H and W directions are calculated using a sigmoid activation function. This is specifically accomplished using the Global Grouped Coordinate Attention (GGCA) module.

[0018] The brightness enhancement network in S4 adopts a U-Net-based encoding-decoding structure and combines the image with a multi-dilated convolutional module (MDFA) for feature extraction.

[0019] The multi-dilation convolution module (MDFA) includes: the first branch, which adopts 1×1 convolution operation without setting the dilation rate, keeps the channel dimension of the input image unchanged, and is used to extract local detail features and process detail information without changing the size of the receptive field; the second branch, which adopts 3×3 convolution with a dilation rate of 9, increases the size of the receptive field and is used to capture medium-scale contextual information, while retaining local details while obtaining more contextual information; the third branch, which adopts 3×3 convolution with a dilation rate of 18, further expands the receptive field and extracts larger-scale image features to capture a wider range of information; the fourth branch, which adopts 3×3 convolution with a dilation rate of 27, maximizes the receptive field and is used to extract global information of the image and restore global brightness and color.

[0020] The multiple dilated convolution module (MDFA) also includes a global feature extraction branch, which extracts global information of the image through global average pooling and 1×1 convolution to help the model understand the brightness and color distribution of the entire image and fuse it with the feature maps of other branches.

[0021] In order to make the present invention easier to understand, it is described in more detail below:

[0022] Weight-aware HVI color space: The core goal of the HVI color space is to transform the input image's color space and then perform brightness enhancement and color restoration in the transformed space. The conversion process first separates the image's color information into three independent components: hue (H), brightness (V), and illumination (I), enabling separate brightness enhancement and color adjustment. This separation helps maintain relative color stability while enhancing brightness, avoiding color casts that may be introduced during overall adjustments. It also more effectively simulates the human eye's perception of brightness and color, resulting in a more natural and realistic enhanced image. This approach is suitable for image processing tasks in a variety of complex lighting environments. To address color casts, uneven color enhancement, and the potential for uneven or unnatural colors in extreme low light conditions, this paper designs an adaptive HVI color space transformation method that intelligently adjusts color enhancement to varying lighting conditions, improving color fidelity in low-light areas and avoiding over-enhancement in highlight areas. A dynamic weight adjustment strategy based on visual perception optimizes the HVI color space transformation and proposes a color adaptation factor for adaptive adjustment, thereby improving the color enhancement effect. In low-light areas, the color adaptation factor is increased to enhance color enhancement, thereby improving color fidelity and reducing color loss in low-light environments. Furthermore, in bright areas, the color adaptation factor can be lowered to reduce the amount of color enhancement, preventing color distortion caused by over-enhancement and maintaining a natural color distribution. This strategy dynamically adjusts based on the local brightness of the image, resulting in a more balanced color enhancement across the entire image, avoiding over- or under-enhancement and improving the overall visual quality.

[0023] Global Grouped Coordinate Attention (GGCA): is a lightweight attention mechanism that enhances feature representation capabilities through channel grouping and global coordinate attention. It combines global pooling and shared feature transformation to extract spatial information and calculate independent attention weights in the height (H) and width (W) directions. GGCA mainly consists of the following parts: the grouping mechanism divides the input feature map into multiple groups according to the number of channels, so that features in different groups can independently extract spatial attention information; the global pooling module performs global average pooling and global maximum pooling in the height H direction and width W direction respectively to obtain global spatial information; the shared feature transformation module uses two 1×1 convolutional layers to perform feature dimensionality reduction and recovery to enhance feature representation capabilities; finally, the attention weights in the H and W directions are calculated through the Sigmoid activation function, and the calculated attention weights in the H and W directions are applied to the input feature map to achieve adaptive enhancement.

[0024] The Multi-Dilated Convolutional Module (MDFA) aims to capture image features at different scales and enhance the representation of important regions using an attention mechanism. This not only requires restoring brightness details but also removing noise. The MDFA module achieves this goal through multiple branches and dilated convolutions of different scales. Its core components are multiple convolutional branches, configured as follows: The first branch uses only 1×1 convolutions with no dilation ratio, preserving the channel dimension of the input image. This branch primarily extracts local detail features while maintaining the receptive field size and effectively processing detail information. The second branch uses 3×3 convolutions with a dilation ratio of 9, increasing the receptive field size to capture medium-scale contextual information while preserving local details. The third branch uses 3×3 convolutions with a dilation ratio of 18, further expanding the receptive field to extract larger-scale image features and capture a wider range of information. The fourth branch uses 3×3 convolutions with a dilation ratio of 27 to maximize the receptive field and extract global information, helping to restore global brightness and color. In addition, it also contains a global feature extraction branch, which extracts the global information of the image through global average pooling and 1×1 convolution, which helps the model understand the brightness and color distribution of the entire image and fuse it with the feature maps of other branches.

[0025] Below is a comparison between this application and other models:

[0026] Table 1 shows the quantitative comparison results of different methods on the League of Legends v1 dataset using the PSNR, SSIM, LPIPS, and NIQE metrics. The models involved in the comparison include 15 algorithms, including RetinexNet, DRBN, RUAS, RetinexDIP, URetinex-Net, SCRnet, and Kind++.

[0027] Table 1 Quantitative comparison results of LOLv1 dataset

[0028]

[0029] Table 2 shows the verification results on the two sub-datasets LOLv2-real and LOLv2-Syn. It performs well in multiple indicators, especially in structure preservation, brightness restoration and detail restoration, demonstrating its strong adaptability and high-quality image enhancement capabilities in complex scenes and low-light scenes.

[0030] Table 2 Quantitative comparison results on LOLv2-real and LOLv2-Syn datasets

[0031]

[0032] Table 3 shows the quantitative comparison on the SICE_Mix dataset. The results show that the proposed method not only has a strong objective quality improvement in extreme low-light image restoration, but also performs well in subjective visual experience.

[0033] Table 3 Quantitative comparison on two extreme low-light datasets (SICEMix, Grad)

[0034]

[0035] Table 4 shows the performance of our models trained using LOLv1 and LOLv2-Syn on five unpaired datasets: DICM, LIME, MEF, NPE, and VV. These models were evaluated using various methods and measured using the BRISQUE and NIQE metrics. As can be seen in Table 4, our method outperforms previous state-of-the-art methods on all unpaired datasets. Our method significantly surpasses existing methods on both the BRISQUE and NIQE metrics, with overall improvements ranging from 20% to 60%.

[0036] Table 4 Quantitative comparison on five unpaired datasets

[0037]

[0038] To explore the impact of different dilation rate settings in dilated convolution on low-light image enhancement performance, Table 5 shows four different dilation rate combinations: (3, 6, 9), (6, 12, 18), (9, 18, 27), and (12, 24, 36). Each configuration consists of three increasing dilation rates, corresponding to three parallel dilated convolution branches in the model. It aims to enhance feature extraction capabilities through multi-scale receptive fields, thereby improving the modeling effect of structure and details in the image enhancement process. On the LOLv1 dataset, ablation experiments were conducted on the above four configurations, and the four evaluation indicators of PSNR, SSIM, LPIPS, and NIQE were used for testing. From the experimental results in Table 5, it can be seen that when the dilation rate combination of (9, 18, 27) is used, the performance is the best in terms of PSNR and LPIPS indicators, indicating that this configuration can achieve the best in terms of detail recovery and perceptual quality.

[0039] Table 5 Void rate ablation experiment

[0040]

Claims

1. A low-light image enhancement method based on HVI color space and multi-scale dilated convolution, characterized in that: include: S1. Converting an input low-light image from an RGB color space to an HVI color space, and enhancing color features of the low-light image to obtain a color feature enhanced image; S2, grouping the low-light image and the color feature enhanced image in S1 by channel dimension, and extracting spatial attention information of each group respectively; S3, fusing the color feature enhanced image in S1 with each group of spatial attention information extracted in S2; S4. Using the feature information fused in S3 as input information, the brightness of the image is enhanced through a brightness enhancement network, and finally a restored normal illumination image is output.

2. The low-light image enhancement method based on HVI color space and multi-scale dilated convolution as claimed in claim 1, characterized in that: When converting to the HVI color space in S1, the color information of the low-light image is split into three independent components: hue H, brightness V, and illumination component I, so as to introduce a weight-aware adaptive factor adjustment mechanism to adjust and / or enhance the brightness and color of the low-light image respectively.

3. The low-light image enhancement method based on HVI color space and multi-scale dilated convolution as claimed in claim 1, characterized in that: Extracting the spatial attention information of each group in S2 includes: S2.

1. Perform global average pooling and global maximum pooling in the height direction H and width direction W respectively to obtain global spatial information; S2.2, using two 1×1 convolutional layers to perform feature dimensionality reduction and recovery to enhance feature representation capabilities; S2.

3. Calculate the attention weights in the H and W directions through the Sigmoid activation function, and calculate the attention weights in the H and W directions.

4. The low-light image enhancement method based on HVI color space and multi-scale dilated convolution as claimed in claim 1, characterized in that: The brightness enhancement network in S4 adopts a U-Net-based encoding-decoding structure and combines the image with a multi-dilated convolutional module (MDFA) for feature extraction.

5. The low-light image enhancement method based on HVI color space and multi-scale dilated convolution as claimed in claim 4, characterized in that: The multiple dilated convolution module (MDFA) includes: The first branch uses a 1×1 convolution operation without setting a dilation rate, keeping the channel dimension of the input image unchanged, and is used to extract local detail features and process detail information without changing the size of the receptive field; The second branch uses 3×3 convolution with a dilation rate of 9 to increase the size of the receptive field and capture medium-scale context information, thereby obtaining more context information while retaining local details. The third branch uses 3×3 convolution with a dilation rate of 18 to further expand the receptive field and extract larger-scale image features to capture a wider range of information. The fourth branch uses 3×3 convolution with a void ratio of 27 to maximize the receptive field, which is used to extract the global information of the image and restore the global brightness and color.

6. The low-light image enhancement method based on HVI color space and multi-scale dilated convolution as claimed in claim 5, characterized in that: The multiple dilated convolution module (MDFA) also includes a global feature extraction branch, which extracts global information of the image through global average pooling and 1×1 convolution to help the model understand the brightness and color distribution of the entire image and fuse it with the feature maps of other branches.