Learable illumination prior residual network for enhancing low-light image
By designing a learning lighting prior residual network (LIPResNet), the combination of the light awakening module and the denoising and enhancement module is used to solve the problem of insufficient utilization of lighting information in low-illumination image enhancement, and better image enhancement effect and visual quality are achieved.
Patent Information
- Application Number
- CN202510267813.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-24
AI Technical Summary
The prior art fails to fully utilize the lighting information in low-illumination image enhancement, resulting in unsatisfactory enhancement effect and prone to problems such as noise amplification and color distortion.
A learning lighting prior residual network (LIPResNet) is designed, adopting a dual residual structure, including a light awakening module and a denoising and enhancement module. The light awakening module extracts lighting information through the learning lighting prior module. The denoising and enhancement module adopts a U-shaped structure with jump connection, and combines the light-guided convolution attention module for feature fusion and denoising.
By making full use of lighting prior information, LIPResNet significantly improves the visual quality of the image in low-illumination image enhancement, avoids noise amplification and color distortion problems, and has smaller parameters and calculations, and has better performance than the comparison method.
Smart Images

Figure CN120198336A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and particularly relates to a learnable illumination prior residual network for enhancing low-light images. Background Art
[0002] Low-light image enhancement is an important challenge in image processing and computer vision, aiming to improve the visual quality of images in low-light environments. Low-light images usually have problems such as insufficient brightness, high noise, low contrast, and detail loss, which affect human visual perception and the performance of downstream visual tasks (such as object detection, image segmentation, and object recognition, etc.). These problems are particularly common in scenarios such as night monitoring and autonomous driving. Therefore, how to effectively enhance low-light images and avoid problems such as noise amplification and color distortion not only improves human visual perception of images but also has important significance for downstream tasks such as object detection, image segmentation, and object recognition.
[0003] For low-light image enhancement, early methods had no learning ability, could not effectively utilize illumination and semantic information, had poor adaptability to complex scenes, and were prone to color distortion; methods based on the Retinex theory reduce the influence of illumination by decomposing the image into illumination and reflection components. Although such methods have theoretical support, due to the failure to consider the particularity of low light, they often amplify noise and produce color deviations in low light.
[0004] With the development of neural network technology, convolutional neural networks (CNNs) have been widely used in low-light image enhancement due to their excellent ability to automatically extract local features. Methods based on CNNs are mainly divided into two categories: end-to-end mapping and deep Retinex decomposition. The end-to-end mapping method learns the mapping function from low-light images to normal images through convolutional blocks to enhance low-light images; the deep Retinex decomposition method uses convolutional blocks to decompose a color image into a reflection component and an illumination component, denoise the reflection component and adjust the illumination of the illumination component, and then reconstruct the enhanced image using the two processed components. However, these methods do not fully utilize illumination information, resulting in unsatisfactory enhancement effects.
[0005] CNN-based methods are limited in capturing long-range dependencies and non-local information modeling. The Transformer model effectively extracts global information through the self-attention mechanism and performs well, but its computational complexity grows quadratically with the image size. SNR-Net uses the Transformer at the bottom of the U-shaped network to reduce the computational amount of the Transformer module while obtaining the ability of the Transformer to extract global information, resulting in good image enhancement effects. Retinexformer further explores the attention mechanism with linear complexity, reducing the computational complexity; at the same time, it rethinks the Retinex decomposition of low-light images and achieves good results by modeling the illumination component loss and the reflection component loss. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to provide a Learnable Illumination Priorbased Residual Network (LIPResNet) for enhancing low-light images. By making full use of the advantages of illumination prior in image enhancement, an efficient learnable illumination prior extractor is designed. For two different types of features, illumination and content, a lightweight convolutional attention module is constructed in the feature fusion stage to achieve effective feature fusion and ensure the enhancement of low-illumination images.
[0007] To solve the above technical problems, the present invention is implemented as follows:
[0008] A learnable illumination prior residual network for enhancing low-light images. The residual network adopts a double-residual structure, which includes an illumination awakening module and a denoising and enhancement module. The illumination awakening module is provided with a learnable illumination prior module for extracting illumination information and generating an illumination map. The denoising and enhancement module adopts a U-shaped structure with skip connections for enhancing brightness and suppressing noise. In the downsampling branch of the U-shaped structure, an illumination-guided convolutional attention module is provided, which includes an illumination channel prior attention composed of channel convolutional attention and spatial convolutional attention. In the upsampling branch of the U-shaped structure, an illumination-guided attention module is provided for restoring image details from deep features.
[0009] Further, the specific process of the illumination awakening module is as follows:
[0010] Extract the illumination prior map L from the input image I through the learnable illumination prior module p , use three-layer convolution to fuse the concatenation result of I and L p During the fusion process, for the illumination-enhanced feature F that guides later denoising lu, a 5×5 depth convolutional layer is used to model different regions in space, and a 1×1 convolutional layer is used to aggregate the scattered luminance features. The aggregated result is multiplied by the original input image I to generate the first residual image I res1 , and then added to the original image to obtain the preliminary illumination enhancement image The expression of the illumination awakening module process is as follows:
[0011] (I res1 ,F lu )=H1(I)
[0012] I1=I + I res1
[0013] Among them, H1(I) represents the overall operation process of the illumination awakening module.
[0014] Furthermore, the specific process of the denoising and enhancement module is as follows:
[0015] The preliminary illumination enhancement image I1 generates a feature map of C channels through a 3×3 convolutional layer, and then successively passes through an Illumination-Guided Convolutional Attention Block (IGCAB) in the downsampling branch, a 4×4 convolutional layer with a stride of 2 for downsampling, two Illumination-Guided Convolutional Attention Blocks, and a strided convolutional layer to extract three scales of feature maps;
[0016] In the bottleneck layer, two Illumination-Guided Attention Blocks (IGAB) are used to recover image information from the features. Then, in the upsampling branch, it is upsampled through a 2×2 transposed convolutional layer with a stride of 2 and concatenated with the features of the same scale in the downsampling branch. After fusing the features through a 1×1 convolutional layer, it is input into two Illumination-Guided Attention Modules. After another upsampling and feature fusion, it is input into an Illumination-Guided Attention Module. Finally, the number of channels is adjusted to 3 through a 3×3 convolutional layer to obtain the second residual image I res2 ;
[0017] In the process of generating the second residual image I res2 , the illumination enhancement feature F obtained by the illumination awakening module lu is used to guide the image luminance enhancement. Finally, I res2 is added to I1 to retain the common features between the two and fade the differences, obtaining the final enhanced image I2. The expression of the denoising and enhancement module process is as follows:
[0018] I res2 =H2(I1,F lu)
[0019] I2 = I1 + I res2
[0020] where H2(I1, F lu ) represents the overall operation process of the denoising and enhancement module.
[0021] Preferably, the learnable illumination prior module (LIP module) includes a convolutional layer and an affine transformation layer. The learnable illumination prior module extracts the initial features of the image through the convolutional layer, and then uses the affine transformation layer to adjust the brightness and contrast of the initial feature map, enhancing the details and contrast of the image. For the input image the process of obtaining the illumination prior map through the LIP module is expressed as follows:
[0022] L p = Conv 3×3 (Aff(LeakReLu(Conv 3×3 (I))))
[0023] where Conv 3×3 (·) represents the 3×3 convolutional layer operation, Aff(x) represents the affine transformation layer operation, and LeakReLu represents the activation function;
[0024] The expression of the affine transformation layer operation is as follows:
[0025]
[0026] where C represents the number of channels, H and W represent the height and width of the feature map respectively, and represent the learnable scaling factor and offset, which are vectors with initial values of 1 and 0 respectively, represents the color transformation matrix, with an initial value of the identity matrix and continuously optimized through training.
[0027] Preferably, the illumination-guided convolutional attention module (IGCAB module) uses layer normalization to improve the generalization ability of the model, and uses residual connections to alleviate gradient disappearance and retain the original information. Finally, it uses a feedforward neural network (FFN) to extract features. The operation expression of the IGCAB module is as follows:
[0028] f′(fea, ill) = ICPA(LN(fea), ill) + fea
[0029] IGCAB(fea, ill) = FFN(LN(f′(fea, ill))) + f′(fea, ill)
[0030] Among them, fea represents the content feature map of the input module, ill represents the illumination feature map of the input module, and LN() represents layer normalization processing.
[0031] Furthermore, the illumination channel prior attention is composed of channel convolutional attention (CA) and spatial convolutional attention (SA). First, the content feature map and the illumination feature map are input into the channel convolutional attention to obtain the channel fusion weight with illumination channel prior information, and then it is multiplied element-wise with the input content feature map to generate the channel attention feature F enhanced by the illumination channel prior. c ; The spatial convolutional attention uses multi-scale depth convolution to obtain the spatial weight, and then multiplies it with the channel attention feature F c to obtain the spatial attention feature The expression of its specific process is as follows:
[0032]
[0033] The specific process expression of the channel convolutional attention is as follows:
[0034]
[0035] Among them, σ represents the sigmoid activation function, MLP represents the fully connected layer, ⊙ represents element-wise multiplication, AvgPool and MaxPool respectively represent spatial average pooling and spatial max pooling, and respectively represent the content intermediate feature and the illumination intermediate feature after pooling, multi-layer perceptron, and activation function processing, represents the multiplication of the two intermediate features;
[0036] The specific process expression of the spatial convolutional attention is as follows:
[0037] x init = DW 5×5 (F c )
[0038]
[0039] a s = σ(Conv 1×1 (S x ))
[0040]
[0041] Among them, DW(·) represents depthwise separable convolution, x initDenotes the initial state feature map obtained by passing the channel convolution attention feature map through a 5×5 depthwise separable convolution, S x Denotes the multi-scale spatial feature map captured from the image using depth convolution kernels of different scales, a s Denotes the spatial attention weight after 1×1 convolution and activation function.
[0042] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0043] The learnable illumination prior residual network of this application adopts a double residual structure. Given the low overall brightness value of low-light images, a learnable illumination prior module is set up to accurately estimate the illumination information. The inherent illumination information of the original image is activated through several convolutions and activation functions and superimposed on the original image to achieve preliminary illumination enhancement. To address problems such as noise amplification, artifacts, and color deviation that may occur during the enhancement process, a multi-level U-Net network is used as the residual part in the second stage to enhance the brightness again and denoise. Through experiments and analysis on the dataset, it is proved that while maintaining low parameters and computational complexity, the objective indicators and subjective feelings of the model are overall better than the comparative methods. Brief Description of the Drawings
[0044] Figure 1 Schematic diagram of the residual network structure with learnable illumination prior of the present invention;
[0045] Figure 2 Schematic diagram of the learnable illumination prior module of the present invention;
[0046] Figure 3 Schematic diagram of the illumination-guided convolutional attention module of the present invention;
[0047] Figure 4 Schematic diagram of the illumination channel prior attention of the present invention;
[0048] Figure 5 Schematic diagram of the comparison of enhancement effects of the present invention on the LOL dataset;
[0049] Figure 6 Schematic diagram of the comparison of enhancement effects of the present invention on 5 datasets without ground truth maps. Detailed Description of the Invention
[0050] The following further elaborates on the specific implementation manners of the present invention in detail in conjunction with the accompanying drawings and specific embodiments.
[0051] As Figure 1As shown in the figure, a Learnable Illumination Prior based Residual Network (LIPResNet) for enhancing low-light images. The residual network adopts a double-residual structure, which includes an illumination awakening module and a denoising and enhancement module. The illumination awakening module is provided with a learnable illumination prior module for extracting illumination information and generating an illumination map to make up for the insufficient illumination under low-light conditions and achieve preliminary enhancement. The denoising and enhancement module adopts a U-shaped structure with skip connections for enhancing the brightness and suppressing noise. An illumination-guided convolutional attention module is arranged in the downsampling branch of the U-shaped structure. This module includes an illumination-channel prior attention composed of channel convolutional attention and spatial convolutional attention, which can efficiently combine illumination information with image detail information. An illumination-guided attention module is arranged in the upsampling branch of the U-shaped structure for recovering image details from deep features.
[0052] The specific process of the illumination awakening module is as follows:
[0053] Extract the illumination prior map L from the input image I through the learnable illumination prior module (LIP) p , use three-layer convolution to fuse the concatenation result of I and L p During the fusion process, for the illumination enhancement feature F that guides the subsequent denoising lu , use a 5×5 depth convolutional layer to model different regions in space, and use a 1×1 convolutional layer to aggregate the scattered brightness features to obtain an intuitive illumination enhancement map; multiply the aggregation result by the original input image I to generate the first residual image I res1 , and then add it to the original image to obtain the preliminary illumination enhancement image The expression of the illumination awakening module process is as follows:
[0054] (I res1 , F lu ) = h1(I)
[0055] I1 = I + I res1
[0056] where H1(I) represents the overall operation process of the illumination awakening module, which not only obtains I res1 , but also saves the intermediate illumination enhancement feature F lu .
[0057] The specific process of the denoising and enhancement module is as follows:
[0058] The preliminary illumination enhancement image I1 generates a feature map of the C channel through a 3×3 convolutional layer, and then passes through an illumination-guided convolutional attention block (IGCAB) in the downsampling branch, a 4×4 convolutional layer with a stride of 2 for downsampling, two illumination-guided convolutional attention modules and a strided convolutional layer to extract H×W×C, Feature maps of three scales;
[0059] In the bottleneck layer, two illumination-guided attention blocks (IGABs) are used to recover image information from the features. Then, in the upsampling branch, a 2×2 transposed convolution layer with a step size of 2 is used for upsampling and concatenated with the features of the same scale in the downsampling branch. After a 1×1 convolution layer, the features are fused and input into two illumination-guided attention blocks. After another upsampling and feature fusion, the feature map is restored to its original size and input into an illumination-guided attention block. Finally, a 3×3 convolution layer is used to adjust the number of channels to 3 to obtain the second residual image I res2 ;
[0060] In generating the second residual image I res2 In the process of using the illumination awakening module, the illumination enhancement feature F lu To guide the image brightness improvement, finally I res2 Add it to I1, retain the common features between the two and dilute the differences, and get the final enhanced image I2. The expression of the denoising and enhancement module process is as follows:
[0061] I res2 =H2(I1,F lu )
[0062] i2=I1+I res2
[0063] Among them, H2(I1,F lu ) represents the overall operation process of the denoising and enhancement module.
[0064] In traditional channel prior processing methods, prior information of an image is usually obtained by calculating the average or maximum value of color channels. This method is widely used because of its low computational cost and ability to quickly estimate the overall illumination level of the image, and it does not require complex parameter adjustment. However, when dealing with low-light images, due to the low overall brightness, this method results in a narrow brightness dynamic range and is difficult to effectively map high-range pixel values. In addition, this method only considers the information of individual pixels and ignores the spatial correlation between pixels. Therefore, in the enhanced image, details and textures are often lost, and the guiding effect of high-brightness regions on low-brightness regions is not fully utilized.
[0065] As Figure 2 shown, the learnable illumination prior module (LIP module) proposed in this application includes a convolutional layer and an affine transformation layer, which can optimize the illumination prior information of the image during the training process and better meet the requirements of different scenarios. The learnable illumination prior module extracts the initial features of the image through the convolutional layer, and then uses the affine transformation layer to adjust the brightness and contrast of the initial feature map, enhancing the details and contrast of the image. For the input image after passing through the LIP module, the process of obtaining the illumination prior map is expressed as follows:
[0066] L p = Conv 3×3 (Aff(LeakReLu(Conv 3×3 (I))))
[0067] where Conv 3×3 (·) represents the 3×3 convolutional layer operation, Aff(x) represents the affine transformation layer operation, and LeakReLu represents the activation function;
[0068] The expression of the affine transformation layer operation is as follows:
[0069]
[0070] where, C represents the number of channels, H and W represent the height and width of the feature map respectively, and represent the learnable scaling factor and offset, which are vectors with initial values of 1 and 0 respectively, represents the color transformation matrix, with an initial value of the identity matrix and continuously optimized through training.
[0071] The LIP module introduces a learnable affine transformation mechanism. This mechanism finely adjusts the brightness and contrast of the image by learning the scaling factor and offset channel by channel, thereby enhancing the detail performance of the image; the affine transformation adjusts the feature map through element-wise scaling and offset, and optimizes the color transformation matrix C during the training processm To enhance the correlation between the channels of the feature map, the LIP module can optimize the illumination prior information of the image and improve the adaptability in different scenarios.
[0072] For the U-shaped network, it mainly focuses on extracting multi-level features and performing effective information fusion in the encoder stage, while in the decoder stage, it uses the fused information to restore the image. Considering the insufficient fusion of illumination information and image information in the encoder part, in order to enhance the fusion, this application specifically designs an illumination-guided convolutional attention module in the encoder part. As Figure 3 shown, the illumination-guided convolutional attention module (IGCAB module) uses layer normalization to improve the generalization ability of the model, and uses residual connections to alleviate the vanishing gradient and retain the original information. Finally, a feedforward neural network (FFN) is used to extract features. The operation expression of the IGCAB module is as follows:
[0073] f′(fea, ill) = ICPA(LN(fea), ill) + fea
[0074] IGCAB(fea, ill) = FFN(LN(f′(fea, ill))) + f′(fea, ill)
[0075] Among them, fea represents the content feature map input to the module, ill represents the illumination feature map input to the module, and LN() represents the layer normalization process.
[0076] The core of the IGCAB module is the illumination channel prior attention (ICPA). As Figure 4 shown, the illumination channel prior attention is composed of channel attention (CA) and spatial attention (SA). First, the content feature map and the illumination feature map are input into the channel attention to obtain the channel fusion weight with illumination channel prior information, and then multiplied element-wise with the input content feature map to generate the channel attention feature F enhanced by the illumination channel prior c ; the spatial attention uses multi-scale depth convolution to obtain the spatial weight, and then multiplies it with the channel attention feature F c to obtain the spatial attention feature The expression of its specific process is as follows:
[0077]
[0078] Different from the traditional CBAM, ICPA fuses the channel weights of content features and illumination features in the channel attention part, obtains the prior information of the illumination channel, and thus improves the feature representation ability.
[0079] The expression of the specific process of the channel convolution attention is as follows:
[0080]
[0081] Among them, σ represents the sigmoid activation function, MLP represents the fully connected layer, ⊙ represents element-wise multiplication, AvgPool and MaxPool represent spatial average pooling and spatial max pooling respectively, and represent the content intermediate feature and the illumination intermediate feature after being processed by pooling, multi-layer perceptron and activation function respectively, represents the multiplication of the two intermediate features;
[0082] CA guides the model to focus on the corresponding content feature channels through the prior information of the illumination feature channels, thus realizing that the illumination prior guides LIPResNet to denoise and enhance low-light images.
[0083] The spatial convolution attention focuses on strengthening the regions with significant spatial features in the image and also plays a crucial role in the ICPA module. The expression of the specific process is as follows:
[0084] x init = DW 5×5 (F c )
[0085]
[0086] a s = σ(Conv 1×1 (S x ))
[0087]
[0088] Among them, DW(·) represents depthwise separable convolution, x init represents the initial state feature map obtained by passing the channel convolution attention feature map through a 5×5 depthwise separable convolution, S x represents the multi-scale spatial feature map capturing the image using depth convolution kernels of different scales, and a s represents the spatial attention weight after passing through a 1×1 convolution and activation function.
[0089] Low-light images usually lack details and contrast spatially. SA helps LIPResNet better enhance low-light images by emphasizing the important regions in the image.
[0090] Examples:
[0091] 1) Datasets
[0092] The experimental data includes LOL-v1, LOL-v2-real, LOL-v2-synthesis, SID, FiveK, LSRW-Huawei, LSRW-Nikon, SMID, SDSD-indoor, and SDSD-outdoor datasets with ground truth images, and DICM, LIME, MEF, NPE, and VV datasets without ground truth images. The datasets with ground truth images mainly test the performance of the quantitative metrics of the model, and the datasets without ground truth images mainly test the generalization performance of the model.
[0093] 2) Implementation details
[0094] The LIPResNet model is implemented in Pytorch. During the training process, it is optimized by the Adam optimizer with default parameters β1 = 0.9 and β2 = 0.999. The Batch size is set to 8, and the total number of iterations is 1.5×10 5 times. The learning rate is initially set to 2×10 -4 and steadily increases to 3×10 4 during the first 4.6×10 -4 iterations through cosine annealing, and then steadily decreases to 1×10 -6 in the subsequent iterations. The training samples are randomly cropped from paired low / normal illumination images with a patch size of 256×256 as training samples, and the training data is randomly rotated and flipped to enhance the variability and robustness of the model. The training objective is to minimize the mean absolute error (MAE) between the enhanced image and its corresponding ground truth image. The peak signal-to-noise ratio (PSNR) and structural similarity (SSIM) are used as evaluation metrics. The experiments are conducted on a processor of Intel(R) Xeon(R) Platinum 8358P CPU@2.60GHz and an NVIDIA GeForce RTX 3090 GPU computing platform.
[0095] 3) Comparison and evaluation
[0096] Quantitative results:
[0097] Table 1 Comparison results on LOL, SID, SMID, and SDSD datasets
[0098]
[0099] Table 1 shows the comparison results of the method of this application and other proposed SOTA methods on a total of 7 datasets, namely LOL, SID, SMID, and SDSD. The method of this application has achieved the best values in terms of both PSNR and SSIM on most datasets. It is only slightly inferior to the best method in 4 comparison items on 3 datasets, but still achieved excellent results in the top 3. Compared with the other two relatively good methods (SNR-Net and Retinexformer), our method has fewer parameters and computational costs, and the overall performance exceeds the compared methods, achieving SOTA results.
[0100] Comparison results on the LOL-v1 dataset in Table 2
[0101]
[0102] Since the LOL-v1 dataset is a commonly used dataset in recent years, some of the latest methods only conduct quantitative index tests on this dataset. Therefore, a comparison was made with these latest methods on the LOL-v1 dataset, and the comparison results are shown in Table 2. The method of this application achieved the best results in terms of PSNR and SSIM metrics.
[0103] Comparison results on the LSRW dataset in Table 3
[0104]
[0105] Some methods also conducted experiments on the LSRW dataset, and a comparison was made with these methods on the LSRW dataset. The results are shown in Table 3. LIPResNet achieved the optimal overall performance on two sub-datasets of LSRW and also achieved SOTA results.
[0106] Comparison results on the MIT-FiveK dataset in Table 4
[0107]
[0108] In addition, experiments were conducted on the FiveK dataset in the standard RGB format. Correspondingly, a comparison was also made with the methods in these literatures on the FiveK dataset, and the comparison results are shown in Table 4. The method of this application demonstrated SOTA results and achieved the optimal overall performance.
[0109] Qualitative results:
[0110] A visual comparison was made with some SOTA methods to compare the effects after low-light image enhancement, as Figure 5 shown. The 3 original images from top to bottom are from LOL-v1, LOL-v2-re, and LOL-v2-syn respectively; as Figure 6As shown, problems such as image blurring or color deviation (such as KinD and MIRNet in the first row), overexposure (such as Restormer and SNR-Net in the second row), or insufficient enhancement (such as Retinexformer in the third row) will occur when comparing methods. LIPResNet better avoids these problems. At the same time, 1 image is selected from each of the 5 datasets without ground-truth reference images for enhancement effect comparison. LIPResNet has a more excellent enhancement effect on 5 low-light images. Some of the other methods being compared sometimes have obvious noise (such as in the second column of the first row), overexposure (such as in the third and fourth columns of the second row), insufficient enhancement (such as in the fifth and sixth columns of the third row), or loss of detail texture (such as in the seventh column of the fourth row) and other problems, while LIPResNet better avoids these problems.
[0111] It can be found through visual qualitative display that LIPResNet can effectively improve brightness, increase contrast, and at the same time achieve a balance between low-light and normal-light areas. The enhanced images obtained have more natural colors and fine details and textures are retained.
[0112] 4) Ablation study
[0113] The LOL-v1 dataset is the most commonly used dataset in the field of low-light image enhancement. However, the number of test set pictures in this dataset is too small. Therefore, the extended version (LOL-v2-real) with more pictures is selected for ablation study to ensure more stable results. Ablation experiment analysis and comparison are carried out for the proposed LIP and IGCAB modules, as well as the improvements within the two modules; at the same time, the selection of attention modules in the two branches of the U-shaped network in the denoising and enhancement parts is also explored; finally, ablation is also done on whether residual connection should be used for the light-awakening part.
[0114] Table 5 Ablation for LIP and IGCAB
[0115]
[0116] Table 5 shows the influence of the LIP module and the IGCAB module on network gain. The baseline model (No. A) extracts the illumination prior using the average value of color channels and uses IGAB fusion in the downsampling stage of the U-shaped network; No. B uses the LIP module to extract the illumination prior and uses IGAB fusion in the downsampling stage; No. C uses Mean to extract the illumination prior and uses the IGCAB module in the downsampling stage; No. D is the model we proposed. The results show that the LIP module is more effective than extracting the illumination prior using the average value of color channels, and using the IGCAB module in the downsampling stage has better effects than IGAB. The use of these two modules not only improves the enhancement effect but also reduces the number of parameters.
[0117] Ablation of the LIP module
[0118]
[0119] To deeply explore the performance of the illumination prior module, we compared the average value, maximum value, channel weighted value, and simple convolution value of each pixel color channel in Table 6. The results show that using only the maximum value within the channel is lower than other value-taking methods. The main reason is that the maximum value cannot combine the information of the other two channels. Although the average value and the channel weighted value are obtained based on the values of the three channels and achieve better results than the maximum value, since they do not consider the spatial correlation between pixel points, the effect is lower than that of the simple convolution value and LIP. The reason why the LIP module achieves better results than simple convolution is that the low-light image is appropriately enhanced through affine transformation to obtain a better illumination prior value.
[0120] Table 7. Ablation of channel attention in the ICPA module
[0121]
[0122] The ICPA module consists of two main parts: channel attention and spatial attention. In the channel attention part, in addition to retaining the content features, illumination features are added. Since the illumination features are newly incorporated features, the impact of using two pooling methods for the illumination features is explored. The results are shown in Table 7. It can be seen that the best results are achieved when using both max pooling and average pooling for the illumination features simultaneously.
[0123] Table 8. Comparison of separable convolution groups in ICPA
[0124]
[0125] The impact of separable convolutions of different sizes on spatial attention is explored. The results are shown in Table 8. The results indicate that the best effect is achieved when using the convolution kernels separately. Therefore, we choose this combination of convolution kernels in the spatial attention part.
[0126] Table 9. Comparison of the selection of illumination guidance modules for the two branches of the U-shaped network
[0127]
[0128] The issue of how to match the two modules, IGAB and IGCAB, in the downsampling and upsampling branches of the U-shaped network is further explored. The experimental results are shown in Table 9. It can be seen that better performance can be achieved by using IGCAB in the downsampling and IGAB in the upsampling.
[0129] Table 10. Ablation of residual connections in Light Awakening
[0130]
[0131] Finally, it is also discussed whether residual connections should be performed on the light illumination awakening part. LIPResNet adopts residual connections. From the results in Table 10, adopting residual connections has better performance than not adopting residual connections.
[0132] The above are only the implementation manners of the present invention. Once again, it is stated that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements can still be made to the present invention, and these improvements are also included in the protection scope of the claims of the present invention.
Claims
1. A learnable illumination prior residual network for enhancing low-light images, characterized by: The residual network adopts a double residual structure, which includes an illumination awakening module and a denoising and enhancement module. The illumination awakening module is provided with a learnable illumination prior module for extracting illumination information and generating an illumination map; the denoising and enhancement module adopts a U-shaped structure with a jump connection for improving brightness and suppressing noise; An illumination-guided convolutional attention module is set in the downsampling branch of the U-shaped structure, which includes an illumination channel prior attention composed of channel convolutional attention and spatial convolutional attention. An illumination-guided attention module is set in the upsampling branch of the U-shaped structure to recover image details from deep features.
2. The learnable illumination prior residual network for enhancing low-light images according to claim 1, characterized in that: The specific process of the light awakening module is as follows: Extract the illumination prior map L from the input image I through a learnable illumination prior module p , using three layers of convolution to fuse I and L p The splicing result, in the fusion process, is used to guide the illumination enhancement feature F in the later denoising process. lu , a 5×5 deep convolutional layer is used to model different regions in the space, a 1×1 convolutional layer is used to aggregate the scattered brightness features, and the aggregation result is multiplied with the original input image I to generate the first residual image I res1 , and then add it to the original image to get the initial lighting enhancement image The expression of the light awakening module process is as follows: (AND res1 ,F lu )=H1(I) I1=I+I res1 Among them, H1(I) represents the overall operation process of the light awakening module.
3. The learnable illumination prior residual network for enhancing low-light images according to claim 1, characterized in that: The specific process of the denoising and enhancement module is as follows: The preliminary illumination-enhanced image I1 generates a feature map of the C channel through a 3×3 convolutional layer, and then passes through a lighting-guided convolutional attention module in the downsampling branch, a 4×4 convolutional layer with a stride of 2 for downsampling, two lighting-guided convolutional attention modules and a strided convolutional layer to extract feature maps of three scales; In the bottleneck layer, two illumination-guided attention modules are used to recover image information from the features. Then, in the upsampling branch, upsampling is performed through a 2×2 transposed convolution layer with a step size of 2, and concatenated with the features of the same scale in the downsampling branch. After a 1×1 convolution layer, the features are fused and input into two illumination-guided attention modules. After the same upsampling and feature fusion, they are input into an illumination-guided attention module. Finally, a 3×3 convolution layer is used to adjust the number of channels to 3, and the second residual image I is obtained. res2 ; In generating the second residual image I res2 In the process of using the illumination awakening module, the illumination enhancement feature F lu To guide the image brightness improvement, finally I res2 Add it to I1, retain the common features between the two and dilute the differences, and get the final enhanced image I2. The expression of the denoising and enhancement module process is as follows: I res2 =H2(I1,F lu ) I2=I1+I res2 Among them, H2(I1,F lu ) represents the overall operation process of the denoising and enhancement module.
4. The learnable illumination prior residual network for enhancing low-light images according to claim 2, characterized in that: The learnable illumination prior module includes a convolution layer and an affine transformation layer. The learnable illumination prior module extracts the preliminary features of the image through the convolution layer, and then uses the affine transformation layer to adjust the brightness and contrast of the preliminary feature map to enhance the details and contrast of the image. The illumination prior map is obtained through the LIP module The process expression is as follows: L p =Conv 3×3 (Aff(LeakReLu(Conv 3×3 (I)))) Among them, Conv 3×3 (·) represents a 3×3 convolutional layer operation, Aff(x) represents an affine transformation layer operation, and LeakReLu represents an activation function; The expression of the affine transformation layer operation is as follows: in, C represents the number of channels, H and W represent the height and width of the feature map respectively. and represents the learnable scaling factor and offset, Represents the color transformation matrix, which is initially the unit matrix and is continuously optimized through training.
5. The learnable illumination prior residual network for enhancing low-light images according to claim 3, characterized in that: The illumination-guided convolutional attention module uses layer normalization to improve the generalization ability of the model, and uses residual connections to alleviate gradient disappearance and retain the original information. Finally, a feedforward neural network is used to extract features. The operation expression of the IGCAB module is as follows: f′(fea,ill)=ICPA(LN(fea),ill)+fea IGCAB(fea,ill)=FFN(LN(f′(fea,ill)))+f′(fea,ill Among them, fea represents the content feature map of the input module, ill represents the illumination feature map of the input module, and LN() represents layer normalization processing.
6. The learnable illumination prior residual network for enhancing low-light images according to claim 1, characterized in that: The illumination channel prior attention is composed of channel convolution attention and spatial convolution attention. The content feature map and the illumination feature map are first input into the channel convolution attention to obtain the channel fusion weight with illumination channel prior information, and then multiplied element by element with the input content feature map to generate the channel attention feature F enhanced by illumination channel prior. c ; The spatial convolution attention uses multi-scale deep convolution to obtain spatial weights, which are then combined with the channel attention feature F c Multiply to get the spatial attention feature The specific process is expressed as follows: The expression of the specific process of channel convolution attention is as follows: Among them, σ represents the sigmoid activation function, MLP represents the fully connected layer, ⊙ represents element-wise multiplication, AvgPool and MaxPool represent spatial average pooling and spatial maximum pooling respectively. and They represent the content intermediate features and illumination intermediate features after processing by pooling, multi-layer perceptron and activation function, respectively. Represents the multiplication of two intermediate features; The expression of the specific process of spatial convolution attention is as follows: x init =DW 5×5 (F c ) a s =σ(Conv 1×1 (S x )) Where DW(·) represents depth-wise separable convolution, x init represents the initial state feature map obtained by a 5×5 depth-separable convolution of the channel convolution attention feature map, S x It represents the use of deep convolution kernels of different scales to capture the multi-scale spatial feature map in the image, a s Represents the spatial attention weight after 1×1 convolution and activation function.
Citation Information
Patent Citations
Low-illumination image enhancement method based on Retinex theory and guided by color prior
CN119273602A
Cited By
Low-light medical image enhancement method, system and equipment based on brightness guidance and detail recovery, and medium
CN120876314A