Low-light image enhancement method based on combination of Retinex and wavelet transform
By combining Retinex theory with wavelet transform Retinex WT architecture, the problem of noise and detail blur in low-light image enhancement is solved, and higher quality image enhancement and detail recovery is achieved.
Patent Information
- Application Number
- CN202411939652.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2044-12-26
AI Technical Summary
When existing low-light image enhancement methods deal with problems such as noise, uneven lighting and blurred details, they are prone to excessive enhancement and image distortion, and it is difficult to effectively separate lighting information and reflectivity.
Combining Retinex theory and wavelet transformation, the RetinexWT architecture is proposed, and the lighting information and image degradation are processed through the illumination estimator and the degradation recovery device, the noise and detail information are decomposed using the wavelet transformation, and the characteristics are selectively fusion through the gating mechanism, and the self-attention mechanism of Transformer is used for global enhancement.
It improves the visual quality and detail clarity of low-light images, effectively suppresses noise, preserves the integrity and clarity of image structure, and is better than traditional Retinex and deep learning methods.
Smart Images

Figure CN119963465A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing and computer vision, and in particular to a low-light image enhancement method based on combining Retinex with wavelet transform. Background Art
[0002] As technology advances, the need to extract clear and accurate information from images continues to expand. In the field of image processing and computer vision, low-light image enhancement is an important research direction. However, environmental factors often hinder the capture of high-quality images. Night or low-light images often suffer from noise, uneven lighting, and blurred details, all of which reduce visual quality and lead to incomplete information. These limitations not only impair visual perception, but also have a chain effect on advanced computer vision tasks such as object recognition, autonomous driving, and image classification, ultimately reducing the accuracy of downstream processing tasks. To address these challenges, various low-light image enhancement algorithms have been proposed, including basic methods such as gamma correction and histogram equalization. However, these methods often lead to over-enhancement and image distortion because their success depends largely on the accuracy of manually set priors. In real-world scenarios, the complexity of lighting conditions complicates the determination of low-light factors. Traditional cognitive methods, such as Retinex theory, inspired by the human visual system, decompose images into two components - illumination and reflectance - to simulate how human perception interprets color and brightness under different lighting conditions. According to Retinex theory, the purpose of low-light enhancement is to mitigate the effects of low-light illumination, amplify the reflective component, and restore details and colors. However, traditional Retinex methods require complex parameter tuning and often introduce significant noise and artifacts, especially in low-light conditions.
[0003] With the advancement of deep learning, convolutional neural networks (CNNs) and Transformer-based models have set new benchmarks for low-light image enhancement. CNNs effectively capture local image features, such as edges and textures, which are critical for detail recovery and noise suppression in low-light environments. However, in low-light conditions, image details and textures are often lost, while accurately separating illumination information remains critical. Traditional convolution operations have difficulty balancing detail preservation with accurate illumination estimation. In addition, traditional convolution upsampling and downsampling mechanisms may degrade details because texture and edge occlusion or blurring are common in low-light images. Direct convolution downsampling may exacerbate the problem by confusing noise with details, especially in dim and degraded areas where noise is often mistakenly amplified.
[0004] In feature fusion during upsampling and downsampling, standard channel concatenation is often used. This technique concatenates features along the channel dimension, lacks the ability to selectively emphasize relevant features, and often leads to the accumulation of redundant information, which dilutes key details and hinders the model's ability to effectively capture key features. In addition, since CNNs mainly capture local features, relying on these features alone may not be sufficient to address global illumination defects in low-light images. Although Transformer models, with their self-attention mechanism, can achieve a global perspective and model long-range dependencies more effectively, they can enhance detail and structure recovery in low-light images. However, applying the original Transformer architecture is computationally intensive and involves a complex training process, making it challenging to adopt in real-time low-light image enhancement tasks.
[0005] Therefore, it is necessary to provide a low-light image enhancement method based on the combination of Retinex and wavelet transform to solve the above technical problems. Summary of the invention
[0006] Since images taken under low-light conditions are usually disturbed by noise and the illumination distribution is uneven, the image quality is poor and the details are blurred. Existing methods often amplify noise and produce artifacts during the enhancement process, making it difficult to ensure the integrity and clarity of the image structure. In order to make up for the shortcomings of the existing methods, the present invention provides a low-light image enhancement method based on the combination of Retinex and wavelet transform, which aims to improve the visual quality and detail clarity of images in low-light environments.
[0007] To achieve the purpose of the present invention, the present invention proposes a RetinexWT architecture, which is a novel method that combines Retinex theory with wavelet transform for robust low-light image enhancement. RetinexWT mainly includes an illumination estimator and a degradation restorer. On the basis of the traditional Retinex model, the present invention introduces perturbation terms to the reflectance and illumination components to accurately simulate the degradation that usually occurs under low-light conditions. The illumination estimator is combined with the wavelet transform to generate enhanced illumination information for low-light images, while the degradation restorer repairs and suppresses various forms of degradation, including noise, artifacts, underexposure / overexposure, and color distortion. By integrating the wavelet feature decomposer into the downsampling module of the degradation restorer, the model can enhance brightness and suppress noise of different frequency components respectively. In addition, a gating mechanism is used to selectively fuse the downsampled and upsampled features, enabling the model to suppress noise while retaining edge and detail information. The self-attention mechanism of the Transformer is further exploited to capture long-range dependencies within the image, thereby facilitating accurate global enhancement and restoration.
[0008] The low-light image enhancement method based on combining Retinex and wavelet transform provided by the present invention comprises the following steps:
[0009] S1. Decompose the input low-light image into reflection component and illumination component using Retinex theory;
[0010] S2, introducing a disturbance term to simulate the degradation under low light conditions and generating degraded reflection and illumination components;
[0011] S3, performing frequency domain analysis on the illumination component through wavelet transform to obtain illumination features, and enhancing the low-light image by combining the illumination features and the result of wavelet transform;
[0012] S4, using a gating mechanism to selectively fuse downsampled and upsampled features to suppress noise and preserve edge and detail information;
[0013] S5. Apply Transformer’s self-attention mechanism to capture long-range dependencies within the image and achieve global enhancement and restoration.
[0014] Preferably, the wavelet transform is a Haar wavelet transform, which is used to decompose the input features into high-frequency and low-frequency components.
[0015] Preferably, the gating mechanism includes a GatedFusion Module for adaptively determining which features are more important and controlling the contribution of features from different channels.
[0016] Preferably, the Transformer's self-attention mechanism is implemented through a Transformer-based hybrid attention network to ensure effective modeling of long-range dependencies during the enhancement process.
[0017] Preferably, the upsampling and downsampling processes are divided into two levels, each level of downsampling consists of Enhanced Illumination-Guided Attention Model (EIGAM) and Wavelet Transform Feature Decomposer Downsampling (WTFDown); each level of upsampling consists of a deconv 2×2 (stride=2) and a conv 1×1 and an EIGAM.
[0018] Preferably, the EIGAM comprises the following steps:
[0019] The input features first pass through the illumination guided attention (IGA), and IGA also uses the illumination features generated in IE as input to guide the calculation of attention;
[0020] Reduce computational complexity through nonlinear activation freedom (NAF), help simplify model architecture and reduce computing resource consumption;
[0021] After IGA and NAF, residual connections are performed to alleviate the gradient disappearance and retain the original output features through a layer normalization (LN) and a feed-forward network (FFN).
[0022] Preferably, the NAF comprises the following steps:
[0023] Use a point-by-point convolution (1x1 conv) to adjust the number of channels of the input feature and achieve linear transformation of the channel dimension;
[0024] Depth-wise Convolution performs independent convolution on each channel to better extract spatial features such as edges and details;
[0025] After simplified channel attention (SCA) and SimpleGate, the number of channels is restored by point-by-point convolution (1x1 conv) and connected with the input feature residual to obtain the first part of the output;
[0026] The first part of the output is convolved and gated to obtain the second part of the output, which is added to the first part of the output through a residual connection to form the final output.
[0027] Preferably, the WTFDown comprises the following steps:
[0028] The input features are first subjected to a 1*1conv operation to improve nonlinearity, and then the spatial features are converted into four frequency domain components through Haar wavelet transform, which are a low-frequency component A and three high-frequency components. The high-frequency components can be divided into horizontal high-frequency components H, vertical high-frequency components V and diagonal high-frequency components D.
[0029] Low-frequency features are represented by convolution to learn features and extract global structural information. High-frequency features are concatenated by concatenating the horizontal, vertical and diagonal high-frequency components, and are subjected to dimensionality reduction and nonlinear operations through point-by-point convolution (1×1conv) and BatchNorm.
[0030] The two parts are added together to obtain global and local composite features, which retains the overall information while enhancing the edges and details and balancing the influence of different frequency components in the image.
[0031] Compared with the related art, the low-light image enhancement method based on the combination of Retinex and wavelet transform provided by the present invention has the following beneficial effects:
[0032] 1. This paper proposes a Transformer-based hybrid attention network for low-light image enhancement to ensure effective modeling of long-range dependencies during the enhancement process. The method of the present invention uses the frequency domain information obtained by wavelet transform combined with the Retinex model to obtain more accurate lighting information.
[0033] 2. The present invention avoids traditional downsampling and instead introduces a Haar wavelet decomposer to preserve information. The input features are decomposed into high-frequency and low-frequency components through wavelet transform, so that the structural information and detail information of the image can be preserved and separated during downsampling, thereby improving the enhancement performance. In addition, by providing frequency domain features, the model can use this rich information in the reconstruction process and perform more targeted noise suppression, edge refinement, noise removal and other operations based on the frequency domain features.
[0034] 3. This paper introduces a gating mechanism to selectively fuse upsampled and downsampled features. Compared with direct channel concatenation, this mechanism can adaptively determine which features are more important. By controlling the contribution of different channel features, noise suppression and detail preservation are better balanced. During the feature fusion process, convolution and activation operations are applied to further refine and enhance the information flow, reducing the risk of information loss during upsampling and downsampling.
[0035] 4. Qualitative and quantitative experiments show that the RetinexWT of the present invention outperforms all previous Retinex-based deep learning methods and achieves results that are superior to the most advanced (SOTA) methods on multiple datasets. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 A schematic diagram of using Haar wavelet transform to perform feature decomposition in the present invention;
[0037] Figure 2 An architecture diagram of a low-light image enhancement method based on combining Retinex and wavelet transform provided by the present invention;
[0038] Figure 3 It is a detailed diagram of IGTW in the present invention;
[0039] Figure 4 for (a) Gated Fusion Module (GFM) and (b) Nonlinear Activation Freedom (NAF);
[0040] Figure 5 This is a schematic diagram of the WTFDown module of the present invention introducing frequency domain feature mapping into the network for down sampling;
[0041] Figure 6 The figure shows the qualitative experimental results of the SOTA method on the LOL-v1 (52) and LOL-v2-synthetic (45) datasets;
[0042] Figure 7 This is a qualitative experimental result diagram of the SOTA method on the LOL-v2-real (45) dataset;
[0043] Figure 8 The visual result diagrams of the present invention on LIME (11), ExDark (19), VV (34), MEF (25), and DICM (19);
[0044] Fig. 9 Graph showing a visual comparison of the impact of the proposed method on low-light object detection. DETAILED DESCRIPTION
[0045] The present invention will be further described below in conjunction with the accompanying drawings and implementation modes.
[0046] Figure 2 The overall architecture of the method of the present invention is described. Figure 2 As shown in (a), the RetinexWT proposed in the present invention mainly consists of two parts: Illumination Estimator and Corruption Restorer. Among them, the Illumination Estimator of the present invention is influenced by the traditional Retinex model. On this basis, the disturbance term is introduced to combine the frequency domain information. In the estimation of illumination features, wavelet transform is also combined. The design of Corruption Restorer is based on Illumination Guided Transformer with Wavelet (IGTW), as shown in Figure 2 As shown in (b), the basic unit of IGTW is the Enhanced Illumination-Guided Attention Model (EIGAM), which consists of Illumination-Guided Attention (IGA), Nonlinear Activation Free Block (NAF), normalization (LN) and feed-forward network (FFN). Figure 3 (a) shows the details of IGTW.
[0047] Retinex-based Framework
[0048] In the field of low-light image enhancement, Retinex theory is used to simulate the human visual system's perception of brightness and color. The traditional Retinex algorithm decomposes the input image I into a reflection component R and an illumination component L, and the formula is expressed as
[0049]
[0050] Represents the multiplication of elements. This method can effectively handle illumination changes and color distortion in images, but in low-light conditions, since this method does not consider the noise and artifacts caused by factors such as insufficient light, it may be amplified and damaged during the enhancement process. In order to overcome these shortcomings, the present invention adopts the perturbation modeling proposed in [2]. Different from the original formula that considers I as undamaged, this study introduces perturbation terms through the illumination component L and the reflection component R to simulate the loss caused under low-light conditions.
[0051]
[0052] and represents the loss, where R is regarded as a well-exposed image. Next, feature extraction is performed through convolution to obtain the illumination map use Brighten the low-light image I by element-wise multiplication, where The formula is expressed as
[0053]
[0054] can be simplified to
[0055]
[0056] Among them I lu To light up the image, C represents the overall loss term, which is Amplified noise and artifacts Exposure problems and color distortion caused by the enhancement process Therefore, the RetinexWT of the present invention can be expressed as
[0057] (I lu ,F lu )=IE(I,L p ), I en =CR(I lu ,F lu ) (5)
[0058] IE stands for IlluminationEstimator, and CR stands for CorruptionRestorer. p =mean c (I), mean c Represents the operation of averaging each pixel along the channel dimension, which is used to evaluate the overall illumination level of the image. IE is expressed in terms of I and L p As input, the output result is the light image I lu and lighting characteristics F lu CR is based on Ilu and F lu As input, the noise and distortion introduced in the image are processed to finally generate the repaired image I en In addition, in order to improve the performance of the model, the model uses a convolutional neural network (CNN) to extract lighting features and process the lighting map, combining image information to achieve effective enhancement of images under complex lighting conditions.
[0059] IlluminationEstimator
[0060] like Figure 2 As shown in (a), the IlluminationEstimator (IE) combines the original low-light image I with the illumination prior L obtained by calculating the pixel average of the I channel dimension. p Merge and increase the channel dimension as input. First, fuse I and L through 1×1 convolution p , that is, applying the lighting prior to the low-light image. Then, a depth-separable 5×5 convolution is used to upsample the input to further extract features and generate preliminary lighting features. It is known that the low-frequency information of the wavelet transform mainly contains the overall illumination and structural information of the image, and L p Ignoring the details of color components and focusing on the overall brightness and lighting information of the image, the lighting prior L p The low-frequency information of L is more focused on the nature of illumination and focuses on illumination information. Therefore, in order to generate more accurate illumination features and make them more robust under low-light and complex illumination conditions, the present invention uses the illumination prior L p Perform wavelet transform and combine the obtained low-frequency components with the preliminary obtained lighting features to generate the final lighting features F lu , where the feature dimension n feat Set to 40. Finally, another 1×1 convolutional layer is used for downsampling to restore the 3-channel illumination map Then multiply the original low-light image I by element to obtain the bright image I lu .
[0061] Illumination GuidedTransformerwithWavelet
[0062] In the RetinexWT framework, CorruptionRestorer (IGTW) consists of an encoder based on IlluminationGuidedTransformerwithWavelet and a decoder that introduces a gated fusion mechanism. The encoder represents the downsampling process, while the decoder represents the upsampling process. Figure 2(b) is shown. Both the upsampling and downsampling processes are divided into two levels. First, the illuminated image I obtained by IlluminationEstimator (IE) lu Downsampling is performed through conv 3×3 (stride=2) in order to match the lighting feature F lu , to facilitate subsequent operations. Next, two levels of downsampling are performed to extract deep features. Each level of downsampling consists of EnhancedIllumination-Guided AttentionModel and WaveletTransform Feature DecomposerDownsampling (WTFDown). Using WTFDown for downsampling can effectively retain the high-frequency information of the image, improve the detail recovery ability of IGTW, and effectively suppress noise amplification during the enhancement process. After each WTFDown, the width and height of the image are halved, and the feature dimension is doubled, that is, the input feature dimension is C, the feature dimension after one level of downsampling is 2C, and after two levels of downsampling, the deepest feature dimension is 4C. After extracting features through downsampling, IGTW needs to continue upsampling to restore the image. Like downsampling, upsampling is also performed in two levels, each level consisting of a deconv2×2 (stride=2) and a conv 1×1 and an EIGAM. After each deconv, the shape size of the image doubles, and the feature dimension is halved. Then, the output of deconv is fused with the output of EIGAM in the corresponding level of downsampling through the Gated Fusion Module (GFM) to reduce the image information loss during downsampling and flexibly and efficiently connect the up and down sampling features to achieve the effects of noise suppression, detail enhancement and information fusion. Finally, the image is convolved through a conv3×3 (stride=2) convolution to reduce the feature dimension and restore it to a three-channel RGB format. The restored image and the illuminated image I lu Make residual connection to get the final enhanced image I en .
[0063] The structure of EIGAM.EnhancedIllumination-Guided Attention Model (EIGAM) is as follows Figure 3 (a) As shown in EIGAM, input feature F in First, it passes through Illumination-Guided Attention (IGA), and IGA also uses the illumination feature F generated in IE luAs input to guide the calculation of attention, NonlinearActivationFree(NAF) is then used to reduce the computational complexity, which helps simplify the model architecture and reduce the consumption of computing resources. After IGA and NAF, residual connections are performed to alleviate the gradient disappearance and retain the original detail information. Finally, a LayerNormalization(LN) and Feed-Forward Network(FFN) are used to obtain the output feature F. out Among them, IGA is used to process the illumination feature F lu And guide the calculation of multi-head self-attention, such as Figure 3 (b). In order to solve the huge computational cost problem of global multi-head self-attention in Transformer, IGA rescales the input features into k-head tokenX(HW*c), and further splits it into k-head X i (HW*dk), for each head, is calculated through three fully connected layers and
[0064]
[0065] Where W Q,i , W K,i and W V,i , i is the learnable parameter of the fc layer, and T represents the matrix transpose. Next, use the illumination feature F lu Encode lighting information, provide global lighting conditions, and adjust them into k-head tokens Y (HW*c) and also split into k heads Y i (HW*dk) guides the calculation of self-attention of each head:
[0066]
[0067] where α i It is the scaling parameter of the science department, which is used to adjust the calculation results. Finally, the features of the k heads are passed through the fully connected layer and position encoding to obtain the output features.
[0068] GFM.GFM (Gated Fusion Module) is used in Illumination Guided Transformer with Wavelet (IGTW) to fuse the features of up and down sampling. Its structure is as follows Figure 4 (a). GFM (first, upsample the feature F u (H*W*C) and downsample F dThe (H*W*C) features are concatenated in the channel dimension to integrate information, and then a Point-wise Convolution is used to mix the channel information, integrate the feature information, and use Depth-wise Convolution to process the local spatial information, retaining the spatial structure, especially the edges and details. After convolution, GFM divides the mixed features into two parts along the channel dimension, F gate (H*W*C) and F content (H*W*C). Among them, F gate After being processed by the GELU activation function, it is used as a gating feature to control the information flow. content and F gate HadamardProduct is performed to achieve adaptive filtering of feature information. Finally, the gated features are combined with the original input features ResidualConnection to obtain the output features. Compared with using direct channel connection to preserve details, GFM uses convolution and activation operations when fusing features in upsampling and downsampling to further extract and enhance information during the flow process. In addition to reducing the loss of information during downsampling and upsampling, it also maintains feature consistency and improves feature representativeness, which can help the model better learn and express the brightness and color distribution in low-light images, thereby improving the enhancement quality.
[0069] NAF. This invention uses a module NAF (NonlinearActivation Free) that removes the traditional nonlinear activation function, aiming to reduce the complexity of calculation and improve performance. Figure 4As shown in (b), NAF first uses a point-wise convolution (1x1 conv) to adjust the number of channels of the input feature and realize the linear transformation of the channel dimension. Then, each channel is independently convolved by Depth-wise Convolution to better extract spatial features such as edges and details. Then, after simplified channel attention (SCA) and SimpleGate, the number of channels is restored by point-wise convolution (1x1 conv), and the first part of the output is residually connected with the input feature to obtain the first part of the output. Finally, the first part of the output is convolved and gated to obtain the second part of the output, and added to the first part of the output through the residual connection to form the final output. The core feature of the NAF module is that it does not contain traditional nonlinear activation functions (such as ReLU, Sigmoid, etc.), but replaces channel attention / GELU with simplified channel attention (SCA) and SimpleGate. By removing nonlinear activation and adopting simple element-by-element operations, the NAF module reduces the computational complexity and retains the feature information. It is an efficient network structure design. In addition, NAFBlock's deep convolution and SCA can enhance the areas with significant differences in brightness and contrast in the feature map, that is, the features at the edges and contours are given extra attention and enhancement. Through a series of efficient feature processing operations, the model can better restore the edge and boundary information of the image in a low-light environment.
[0070] Wavelet Transform Feature Decomposer Downsampling
[0071] Wavelet Transform is a signal processing technology that can decompose a signal into sub-signals of different frequencies for analysis in the time domain and frequency domain. In image processing, Wavelet Transform is used to decompose an image into multiple sub-bands, including low-frequency components and high-frequency components. Among them, the low-frequency component retains the overall brightness and structure of the image, which is convenient for enhancing the global brightness; the high-frequency component contains detail information such as edges and textures, and also contains noise. Therefore, in the task of low-light image enhancement, the image can be enhanced in brightness while retaining delicate texture and edge information by processing the low-frequency and high-frequency parts separately. However, previous studies often directly apply Wavelet Transform in the downsampling layer of the network, replacing spatial domain features with frequency domain features, but this may cause the loss of spatial information, making the enhanced image blurred or distorted. In order to solve this problem, the present invention introduces a method that combines frequency domain and spatial domain information, namely Wavelet Transform Feature Decomposer Downsampling (WTFDown), which is used to introduce frequency domain feature mapping in the low-light enhancement network for downsampling. Its structure is as follows Figure 5 . In WTFDown, the input features are first subjected to a 1*1conv operation to improve nonlinearity, and then the spatial features are converted into four frequency domain components through Haar wavelet transform, namely a low-frequency component A and three high-frequency components. The high-frequency components can be divided into horizontal high-frequency components H, vertical high-frequency components V and diagonal high-frequency components D. Low-frequency features are learned through convolution to extract global structural information. High-frequency features concatenate the three high-frequency components of horizontal, vertical and diagonal, and perform dimensionality reduction and nonlinear operations through point-by-point convolution (1×1conv) and BatchNorm. This processing method enhances the edge and local detail information in the high-frequency components, which helps to preserve details in low-light images. Finally, the two parts are added to obtain global and local composite features, which retains the overall information while enhancing the edges and details, and balancing the influence of different frequency components in the image. WTFDown uses Haar wavelet transform to simply and efficiently decompose the image signal, introduces frequency information into the network, and achieves a more comprehensive feature representation. Figure 1 The process of Haar wavelet transform is shown. A(X) represents a low-pass filter, which performs low-pass filtering on the original data, and D(X) represents a high-pass filter, which performs high-pass filtering on the original data. cA represents the low-frequency component, while cH, cV, and cD represent the horizontal high-frequency component, the vertical high-frequency component, and the diagonal high-frequency component, respectively. Each time the filter is filtered, the feature size is reduced by half.
[0072] Specifically, the input feature X(H*W*C) first undergoes a 1*1conv operation to improve nonlinearity, and then the spatial feature is converted into four frequency domain components through Haar wavelet transform, which are a low-frequency component A and three high-frequency components. The high-frequency components can be divided into horizontal high-frequency components H, vertical high-frequency components V and diagonal high-frequency components D. For each channel X c (H*W), the wavelet transform is calculated as follows:
[0073]
[0074] Among them A c (i,j) and D c (i, j) represent the low-frequency approximation coefficient and high-frequency approximation coefficient of channel c respectively, i is the row index, ranging from 1 to H. j is the column index, ranging from 1 to Next, the Haar wavelet transform is applied to the approximation coefficients and detail coefficients of each column, expressed as:
[0075]
[0076] Among them, A c is the low-frequency approximation coefficient of a single channel, Hc 、V c , D c They are the high-frequency detail coefficients in the horizontal, vertical and diagonal directions respectively. Subsequently, the three high-frequency components are concatenated and point-by-point convolution is used to reduce the dimension, remove unimportant information (noise, etc.), and retain key features (edge and detail information), thereby obtaining the final high-frequency feature low-frequency feature F l and F h , the formula is:
[0077]
[0078] in, Represents a batch normalization operation. Finally, the obtained high-frequency and low-frequency features are added to provide a downsampled feature map containing frequency domain features.
[0079] Compared with the related art, the low-light image enhancement method based on the combination of Retinex and wavelet transform provided by the present invention has the following technical effects:
[0080] The proposed method is compared with various deep learning-based SOTA methods listed in Table 1, including SID, 3DLUT, DeepUPE, SCI, RetinecNet, etc. The datasets used for comparison include synthetic data of LOL-v1 and real data and synthetic data of LOL-v2. For fair comparison, the official pre-trained model of each method and its public code are used to obtain quantitative results.
[0081] Quantitative analysis: The evaluation indicators used for comparison are PSNR and SSIM, where PSNR reflects the overall enhancement quality, and the higher the value, the better the performance. SSIM measures the retention of high-frequency details and structural information, and the higher the value, the better the retention of image content. The RetinexWT method proposed in this paper shows significant performance advantages over the above SOTA methods on the LOL dataset.
[0082] Specifically, compared with other Retinex-based deep learning SOTA methods, including SID, DeepUPE, SCI, LIME, RetinexNet, RUAS, FIDE, KinD, and Retinexformer, the method of the present invention achieves significant improvements in PSNR and SSIM on LOL-v1 and LOL-v2 datasets. In terms of PSNR, RetinexWT achieves 0.27dB, 0.07dB, and 0.35dB enhancements on LOL-v1, LOLv2-real, and LOL-v2-synthetic datasets, respectively. Similarly, SSIM improvements of 0.007dB, 0.02dB, and 0.001dB were observed on the same datasets, which emphasizes the superior ability of the method of the present invention in balancing enhancement quality and detail preservation. As shown in Table 4.1, the effectiveness of the method of the present invention is further highlighted.
[0083] Quantitative analysis: In order to provide a more comprehensive and intuitive comparison, the present invention is to RetinexWT
[0084] Our method is visually evaluated against other state-of-the-art (SOTA) methods. Figure 6 and Figure 7 Taken from the LOL dataset (LOL-v1 and LOL-v2), where the input consists of severely degraded low-light images. The results reveal several limitations of existing methods: noise amplification (e.g. Figure 7 RetinexNet in ), underexposure or overexposure (e.g. Figure 6 LEDNet and RUAS in ), color distortion (e.g. Figure 7 Restormer in ) and the introduction of black spots and artifacts (such as Figure 7 EnGAN and Kind in ). By Figure 6 , Figure 7 The comparison of the results shows the obvious advantages of the method. After careful examination by magnification, the method of the present invention achieves excellent visual effects.
[0085]
[0086] Table 1:Performance comparison of various methods on LOL(v1(52)and v2(45))datasets in terms of PSNR and SSIM.Retinex indicates whether the method is based on Retinex theory.The bold-red font indicates the best result, and the second-best result is formatted with bold-blue font.
[0087] In contrast, the method of the present invention shows significant improvements in both global and local enhancement. It can effectively manage exposure levels, significantly improve visibility and contrast, and minimize noise, thereby providing excellent visual quality. Figure 8 Unpaired benchmark results are presented in , where our approach demonstrates precise exposure control and vivid color restoration. This further highlights its excellent generalization ability across a variety of low-light image scenarios.
[0088] Low-light Object Detection:
[0089] In this section, we study the effects of various preprocessing methods on the efficiency of object detection under low-light conditions. Specifically, we conduct experiments on the ExDark dataset, which contains 7,363 real-world nighttime images covering 12 object categories. To evaluate the impact of preprocessing, we first apply our proposed RetinexWT method and several comparison methods, and then use YOLO-v3 as the object detection model to evaluate the detection performance.
[0090] Quantitative Analysis: In Table 2, we present the average precision (AP) scores obtained by different methods used as a preprocessing step for object detection. RetinexWT stands out, not only achieving the highest overall average AP score, but also the highest AP in five specific categories: bicycle, cat, dog, person, and table. In addition, it achieves the second highest AP in the ship, bus, and car categories, demonstrating strong performance. These results highlight the effectiveness and robustness of RetinexWT in enhancing the performance of object detection in different categories.
[0091]
[0092] Table2: Low-light object detection results on ExDark(2)dataset using different preprocessing methods. The bold-red font indicates the best result, and the second-best result is formatted with bold-blue font.
[0093] Quantitative analysis: The present invention visually compares the target detection results of the original low-light image and the image enhanced by the present invention, such as Fig. 9 As shown in Figure 2, some categories are missed in the detection results of the original low-light image, and the overall detection result accuracy is greatly limited. In contrast, the enhanced image processed by the technology of the present invention shows comprehensive category detection and significantly improved detection accuracy.
[0094] The above descriptions are merely embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A low-light image enhancement method based on the combination of Retinex and wavelet transform, characterized in that: The following steps are involved: S1. Decompose the input low-light image into reflection component and illumination component using Retinex theory; S2, introducing a disturbance term to simulate the degradation under low light conditions and generating degraded reflection and illumination components; S3, performing frequency domain analysis on the illumination component through wavelet transform to obtain illumination features, and enhancing the low-light image by combining the illumination features and the result of wavelet transform; S4, using a gating mechanism to selectively fuse downsampled and upsampled features to suppress noise and preserve edge and detail information; S5. Apply the Transformer’s self-attention mechanism to capture long-range dependencies within the image and achieve global enhancement and restoration.
2. The low-light image enhancement method based on combining Retinex and wavelet transform as claimed in claim 1, characterized in that: The wavelet transform is a Haar wavelet transform, which is used to decompose input features into high-frequency and low-frequency components.
3. The low-light image enhancement method based on combining Retinex and wavelet transform as claimed in claim 1, characterized in that: The gating mechanism includes a Gated Fusion Module for adaptively determining which features are more important and controlling the contribution of features from different channels.
4. The low-light image enhancement method based on combining Retinex and wavelet transform as claimed in claim 1, characterized in that: The Transformer’s self-attention mechanism is implemented via a Transformer-based hybrid attention network to ensure effective modeling of long-range dependencies during the reinforcement process.
5. The low-light image enhancement method based on combining Retinex and wavelet transform as claimed in claim 1, characterized in that: The upsampling and downsampling processes are divided into two levels. Each level of downsampling consists of Enhanced Illumination-GuidedAttentionModel (EIGAM) and Wavelet Transform Feature DecomposerDownsampling (WTFDown); each level of upsampling consists of a deconv 2×2 (stride=2) and a conv 1×1 and an EIGAM.
6. The low-light image enhancement method based on combining Retinex and wavelet transform as claimed in claim 5, characterized in that: The EIGAM comprises the following steps: The input features first pass through the illumination guided attention (IGA), and IGA also uses the illumination features generated in IE as input to guide the calculation of attention; Reduce computational complexity through nonlinear activation freedom (NAF), help simplify model architecture and reduce computing resource consumption; After IGA and NAF, residual connections are performed to alleviate the gradient disappearance and retain the original output features through a layer normalization (LN) and a feed-forward network (FFN).
7. The low-light image enhancement method based on combining Retinex and wavelet transform as claimed in claim 5, characterized in that: The NAF comprises the following steps: Use a point-by-point convolution (1x1 conv) to adjust the number of channels of the input feature and achieve linear transformation of the channel dimension; Depth-wise Convolution performs independent convolution on each channel to better extract spatial features such as edges and details; After simplified channel attention (SCA) and SimpleGate, the number of channels is restored by point-by-point convolution (1x1 conv) and connected with the input feature residual to obtain the first part of the output; The first part of the output is convolved and gated to obtain the second part of the output, which is added to the first part of the output through a residual connection to form the final output.
8. The low-light image enhancement method based on combining Retinex and wavelet transform as claimed in claim 5, characterized in that: The WTFDown includes the following steps: The input features are first subjected to a 1*1conv operation to improve nonlinearity, and then the spatial features are converted into four frequency domain components through Haar wavelet transform, which are a low-frequency component A and three high-frequency components. The high-frequency components can be divided into horizontal high-frequency components H, vertical high-frequency components V and diagonal high-frequency components D. Low-frequency features are represented by convolution to learn features and extract global structural information. High-frequency features are concatenated by concatenating the horizontal, vertical and diagonal high-frequency components, and are subjected to dimensionality reduction and nonlinear operations through point-by-point convolution (1×1conv) and BatchNorm. The two parts are added together to obtain global and local composite features, which retains the overall information while enhancing the edges and details and balancing the influence of different frequency components in the image.
Citation Information
Patent Citations
Low-illumination image enhancement method based on Retinex-Net and wavelet transform fusion
CN116228595A
Low-illumination image enhancement method based on feature fusion and attention embedding
CN116797488A
Low-illumination image enhancement method and device based on wavelet transform and Retinex-Net
CN117094907A
Weak light image enhancement method based on soft gating fusion mechanism and adaptive frequency domain perception
CN118982472A