Progressive low-illumination image enhancement system, computer program product and apparatus implementing same
By adopting the combination of global brightness enhancement, coarse-grained denoising and FGEM modules in the low-light image enhancement system, the hidden features are refined using the U-shaped MetaFormer network, the problem of poor quality of low-light image is solved and efficient and accurate image enhancement effect is achieved.
Patent Information
- Application Number
- CN202510183054.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-06-06
AI Technical Summary
Images captured under low light conditions have problems such as reduced visibility, improved noise levels and color distortion, which affect the visual quality of the image and the performance of subsequent computer vision tasks.
A progressive low-light image enhancement system is adopted, which includes a global brightness enhancement module, a coarse-grained noise denoising module and a FGEM module. The hidden features are gradually refined using the U-shaped MetaFormer network, fine-grained enhancement features are generated and added to the brightening image.
While significantly reducing computing overhead, it provides high enhancement accuracy, significantly improving the visual quality of images and performance performance of subsequent computer vision tasks.
Smart Images

Figure CN120107104A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of GPU branch image processors in artificial intelligence hardware platforms, and in particular to a progressive low-light image enhancement system, and a computer program product and device for implementing the system. Background Art
[0002] Low-light image enhancement (LLIE) is a key challenge in the field of computer vision, which has a significant impact on downstream tasks such as surveillance, security, night photography, and object detection. Images taken under low-light conditions often suffer from reduced visibility, increased noise levels, and color distortion. These problems not only degrade the visual quality of the image, but also impair the performance of subsequent computer vision tasks that rely on clear and detailed visual input. Summary of the invention
[0003] 1. Technical issues to be solved
[0004] The present invention is expected to at least partially solve one of the above technical problems.
[0005] 2. Technical Solution
[0006] The first aspect of the present invention provides a progressive low-light image enhancement system. The progressive low-light image enhancement system comprises: a global brightness enhancement module for improving the original image I input Perform global brightness enhancement to obtain the brightened image I g ; Coarse-grained denoising module, used to brighten image I g Perform coarse-grained denoising to obtain the denoised shallow feature F 0 ; For denoising shallow features F 0 Refine the features to obtain content features F C ; FGEM module, which uses a U-shaped MetaFormer network to denoise shallow features F 0 and content feature F C As input, the hidden features are gradually refined to obtain the fine-grained enhanced features E; the image synthesis module is used to add the fine-grained enhanced features E to the brightened image I g , and get the enhanced image I output .
[0007] A second aspect of the present invention provides a computer program product, which includes a computer program, which implements the above progressive low-light image enhancement system when executed by a processor.
[0008] A third aspect of the present invention provides a computer device, which includes: a processor; and a memory on which a computer program is stored; wherein the processor executes the computer program to implement the above progressive low-light image enhancement system.
[0009] 3. Beneficial Effects
[0010] It can be seen from the above technical solution that the present invention has at least one of the following beneficial effects compared with the prior art:
[0011] 1. U-shaped MetaFormer network
[0012] In the present invention, the progressive low-light image enhancement system is a progressive enhancement framework that uses a U-shaped MetaFormer (UMFormer) network to perform LLIE. Through the above scheme, the present invention can produce the following effects:
[0013] ①UMFormer implements diverse token mixing techniques at different feature levels, thereby minimizing computational complexity without sacrificing performance.
[0014] ②The FGEM module adopts the UMFormer network to provide fine-grained image restoration capabilities. The other two modules operate at a broader level, ensuring that the training of the fine-grained module is more efficient.
[0015] ③ By combining the progressive enhancement framework with UMFormer, the present invention significantly reduces the computational overhead while providing high enhancement accuracy, becoming an effective solution for the LLIE task.
[0016] 2. Cross-attention Token mixer in the outer layer; Deep convolutional Token mixer in the inner layer
[0017] In the present invention, the U-shaped MetaFormer network includes multiple levels: a cross-attention submodule of the outer N1 layer; a deep convolution submodule of the inner N2 layer, the cross-attention submodule includes: a cross-attention Token mixer and a feedforward network; the deep convolution submodule includes: a deep convolution Token mixer and a feedforward network.
[0018] Such a setting can produce the following beneficial effects:
[0019] ① By integrating different Token mixers at different levels, each mixer is customized to capture features at different scales. Therefore, the UMFormer network can gradually refine the hidden features.
[0020] ②UMFormer uses MetaFormer modules with different Token mixers (especially cross attention and deep convolution) at different feature map levels to replace the conventional self-attention mechanism. This strategy can effectively reduce redundant calculations and reduce the overall computational cost.
[0021] ③ At the initial full-resolution level, features retain fine-grained spatial details, where low-level features mainly contain local patterns and limited semantic information. This makes inter-channel interactions more important than spatial dependencies, and accurate channel mixing becomes crucial to preserve detailed structural information. As features enter deeper levels, they exhibit increased semantic abstraction and reduced spatial resolution, which makes them suitable for efficient processing by deep convolutional Token mixers.
[0022] 3. Excellent actual performance
[0023] Through testing, the beneficial effects of the progressive low-light image enhancement system of the present invention are mainly reflected in the following aspects:
[0024] ①The computational complexity is greatly reduced
[0025] This paper conducts a comprehensive analysis of UMFormer and proves that its time complexity is O(K 2 HWC), where K is the kernel size (usually set to 3), H and W represent the image height and width, and C represents the number of channels. This is faster than the time complexity of the conventional U-shaped transformer network, O(H 2 W 2 C) has been substantially improved. The method of the present invention achieves this efficiency by minimizing redundant computations and using streamlined token mixing operations. As a result, compared with the current state-of-the-art method RetinexFormer, the floating point operations (FLOPs) and model parameters of the present G-UMFormer are reduced by 26% and 21%, respectively.
[0026] ②Advanced LLIE performance
[0027] Extensive benchmark experiments show that the proposed G-UMFormer achieves advanced LLIE performance while significantly reducing computational and parameter overhead, as verified by peak signal-to-noise ratio (PSNR) and structural similarity index (SSIM). These results are achieved on LOL-v2-syn, SID, and FIVEK datasets, while maintaining high accuracy on LOL-v1 and LOL-v2-real. This highlights the ability of G-UMFormer to provide high-quality enhancement with higher efficiency.
[0028] ③Higher comprehensive performance
[0029] We propose G-UMFormer as a novel and effective progressive low-light image enhancement system. By integrating content-aware insights of spatial and channel image features, it is able to strike a good balance between computational efficiency and high performance, ensuring a detailed and precise enhancement process. Comprehensive experiments (including ablation studies) verify that our method has competitive performance in comparison with the state-of-the-art. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 and Figure 2 They are respectively a design principle diagram and an overall architecture diagram of a progressive low-light image enhancement system according to an embodiment of the present invention.
[0031] Figure 3 for Figure 1 Schematic diagram of the cross-attention Token mixer in the FGEM module in the progressive low-light image enhancement system.
[0032] Figure 4 for Figure 1 Schematic diagram of the deep convolutional Token mixer in the FGEM module in the progressive low-light image enhancement system.
[0033] Figure 5 The following is a visual comparison of low-light images processed by the progressive low-light image enhancement system of this embodiment and a representative LLIE method in the prior art. DETAILED DESCRIPTION
[0034] The present invention proposes a progressive low-light image enhancement system, which adopts a progressive U-shaped MetaFormer network (Gradual U-Shaped MetaFormer, referred to as "G-UMFormer") framework. The progressive low-light image enhancement system includes a global brightness enhancement module and a coarse denoising module for overall image quality improvement, and a fine-grained enhancement module (Fine-Grained Enhancement Module, referred to as "FGEM module") using a U-shaped MetaFormer network (UMFormer). The U-shaped MetaFormer network integrates MetaFormer blocks with different Token mixing techniques, such as cross attention and depthwise separable convolution, effectively reducing the computational complexity from quadratic to linear in the image size while keeping the performance unaffected.
[0035] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with specific implementation methods and with reference to the accompanying drawings.
[0036] A first aspect of the present invention provides a progressive low-light image enhancement system. Figure 1 and Figure 2 The following are the design principle diagram and overall architecture diagram of the progressive low-light image enhancement system of the embodiment of the present invention. As shown in the figure, the progressive low-light image enhancement system of the embodiment includes:
[0037] The global brightness enhancement module is used to improve the original image I input Perform global brightness enhancement to obtain the brightened image I g , expressed in formula: I g =GlobalBrightening(I input );
[0038] Coarse-grained denoising module, used to brighten image I g Perform coarse-grained denoising to obtain the denoised shallow feature F 0 ; For denoising shallow features F 0 Refine the features to obtain content features F C , expressed in formula: F 0 ,F C =CoarseGrainedDenoising(I g );
[0039] The FGEM module uses a U-shaped MetaFormer network to denoise the shallow features F 0 and content feature F C As input, the hidden features are gradually refined to obtain the enhanced features E, which can be expressed by the formula: E = FineGrainedEnhancement (F 0 ,F C );
[0040] Image synthesis module, used to add the enhanced features E to the brightened image I g , and get the enhanced image I output , expressed in formula: I output =I g +E.
[0041] It can be seen that the progressive low-light image enhancement system of this embodiment is a progressive enhancement framework that uses the U-shaped MetaFormer (UMFormer) network for LLIE. It first uses the global brightness enhancement module and the coarse denoising module to improve the image quality on a broader level, and then uses the FGEM module to refine the image details of each area. FGEM uses the UMFormer network and replaces the conventional transformer module in the U-shaped transformer with the MetaFormer module. Through the above scheme, this embodiment can produce the following effects:
[0042] ①UMFormer implements diverse token mixing techniques at different feature levels, thereby minimizing computational complexity without sacrificing performance.
[0043] ②The FGEM module adopts the UMFormer network to provide fine-grained image restoration capabilities. The other two modules operate at a broader level, ensuring that the training of the fine-grained module is more efficient.
[0044] ③ By combining the progressive enhancement framework with UMFormer, this embodiment provides high enhancement accuracy while significantly reducing the computational overhead, becoming an effective solution for the LLIE task.
[0045] The various parts of the progressive low-light image enhancement system of this embodiment are described in detail below.
[0046] In this embodiment, the global brightness enhancement module uses data-driven gamma correction to achieve initial adjustment of the image so that its brightness is aligned with the target domain. The initial enhanced learnable gamma correction is defined as: γ is a learnable parameter that is adjusted during training;
[0047] Those skilled in the art should understand that, in addition to gamma correction, the global brightness enhancement stage may also be performed by histogram equalization or other well-known methods in the art, which may also implement the present invention and is also within the protection scope of the present invention.
[0048] In the prior art, the brightened image from the global brightness enhancement module is used to extract basic content information related to the pixel spatial position and intensity, which is usually affected by noise.
[0049] In order to collect information from a larger receptive field without losing spatial resolution, in this embodiment, average pooling with a 7×7 kernel is used in the coarse-grained denoising (CGD) module. This step aggregates the responses of neighboring pixels and performs coarse-grained denoising, capturing a comprehensive representation.
[0050] Specifically, the resulting feature representation F C Merging local details with contextual information from the extended neighborhood is crucial for LLIE. The denoised content features extracted by the CGD module are as follows:
[0051]
[0052] F C =Conv1(F 0 )
[0053] Among them, Conv1 is a convolution operation with a kernel size of 1, which maintains spatial resolution and focuses on local details. Pool7 is an average pooling with a kernel size of 7.0 The final convolution refines the features to produce F C . Feature F C and F 0 is the input of FGEM. 0 captures a wider context, while F C Detailed content information is provided to guide FGEM to generate fine-grained enhanced features.
[0054] Those skilled in the art should understand that the types and kernel sizes of the above-mentioned convolution operations and pooling operations can be adjusted according to the computing power of the actual scenario. In particular, the above Pool7 can also be maximum pooling and minimum pooling, and the kernel K of the pooling operation satisfies: 5≤K≤9.
[0055] like Figure 1 and Figure 2 As shown, the FGEM module uses a U-shaped and MetaFormer-based network. It is represented by a preliminary feature and denoising content feature F C As input, generate fine-grained enhanced features as output.
[0056] Furthermore, the U-shaped MetaFormer network includes: a cross-attention submodule of the outer N1 layer; a deep convolution submodule of the inner N2 layer, 1≤N1,N2≤3; the cross-attention submodule includes: a cross-attention Token mixer and a feedforward network.
[0057] The cross-attention submodule of each layer includes: symmetrical cross-attention submodules on the encoder side and the decoder side; the depth convolution submodule of each layer includes: symmetrical depth convolution submodules on the encoder side and the decoder side; the submodules on the encoder side and the decoder side of the same layer are jump-connected.
[0058] The cross-attention submodule includes: a cross-attention Token mixer and a feedforward network; the deep convolution submodule includes: a deep convolution Token mixer and a feedforward network; wherein the feedforward network function FFN is one of the following: a 3×3 convolution operation; a combination of a 1×1 convolution followed by a deep 3×3 convolution; a combination of several layers of MLP fully connected operations after Reshape.
[0059] On the encoder side, for the outermost cross-attention submodule, it denoises the shallow features F 0 and content feature F C As input, the output feature is F out ; For other Token mixers except the outermost cross attention submodule, the output of the Token mixer outside the current layer and the content feature F CAs input, the output feature is F out .
[0060] The decoder side is set symmetrically with the encoder side. For the outermost cross-attention Token mixer, its output feature F out The enhanced features E are output from the FGEM module.
[0061] In this embodiment, N1=1; N2=2, that is, the U-shaped MetaFormer network in this embodiment includes three levels: the cross attention submodule at level 1 (initial resolution level, level 1); the deep convolution submodule at levels 2 and 3 (deeper levels, level 2 and level 3). This setting can produce the following beneficial effects:
[0062] ① By integrating different Token mixers at different levels, each mixer is customized to capture features at different scales. Therefore, the UMFormer network can gradually refine the hidden features.
[0063] ②UMFormer uses MetaFormer modules with different Token mixers (especially cross attention and deep convolution) at different feature map levels to replace the conventional self-attention mechanism. This strategy can effectively reduce redundant calculations and reduce the overall computational cost.
[0064] ③ At the initial full-resolution level (level 1), features retain fine-grained spatial details, where low-level features mainly contain local patterns and limited semantic information. This makes inter-channel interactions more important than spatial dependencies, and accurate channel mixing becomes crucial to preserve detailed structural information. As features enter deeper levels, they exhibit increased semantic abstraction and reduced spatial resolution, which makes them suitable for efficient processing by deep convolutional Token mixers.
[0065] The following is a detailed description of the cross-attention Token mixer and the deep convolutional Token mixer.
[0066] (1) Content-aware cross-attention submodule
[0067] For the full-resolution level, this embodiment uses content-aware cross-attention (cross-attention submodule, referred to as "CCA submodule") to effectively process full-resolution features. The CCA submodule enhances feature representation by achieving dynamic channel-to-channel interaction while maintaining spatial information integrity. Different from the conventional cross-attention mechanism that directly uses external features as key-value pairs, this embodiment introduces a content-modulated attention mechanism, in which the content feature F C Modulate the key matrix and the value matrix by element-wise multiplication.
[0068] In this embodiment, the cross attention submodule includes: a cross attention token mixer and a feedforward network, which are described in detail as follows.
[0069] Figure 3 for Figure 1 Schematic diagram of the cross-attention Token mixer in the FGEM module in the progressive low-light image enhancement system. As shown in the figure, the content-aware cross-attention of the feature level 1, where F is the input feature, F C is the content feature of the modulation key value matrix, F N is a normalized feature. The Token mixer uses a content-modulated attention mechanism to process channel tokens.
[0070] like Figure 3 As shown, given the input features Where H×W represents the spatial dimension, and C represents the number of channels. In this embodiment, each channel is regarded as a token in the cross-attention token mixer, generating C tokens, each of which has H×W dimensions.
[0071] 1.1 Layer Normalization
[0072] F N =LayerNorm(F)
[0073] In this step, LayerNorm() is the layer normalization operation; F N is the normalized input feature;
[0074] If the current module is the outermost cross-attention token mixer, then F = F 0 ; Otherwise, F is the output of X from the upper-layer Token mixer out .
[0075] 1.2 Feature Projection
[0076] Q=W Q F N ,K=W K F N ,V=W V F N ,
[0077] In this step, is the feature projection operation, It means pixel-level 1×1 convolution followed by channel-level 3×3 convolution; The three are the query, key, and value matrices obtained by the feature projection operation.
[0078] Those skilled in the art should understand that the size of the convolution kernel, etc., can be adjusted based on the computing power of the actual scenario under the concept of this embodiment.
[0079] 1.3 Modulation and Reshaping
[0080]
[0081] It should be noted that in Figure 2 In the content feature F C After the "e" operation, it enters the Token mixer. Here, "e" is the embedding operation, that is, Fc = embedding (Fc). This operation is a common operation in this field, so Figure 3 Not given in .
[0082] In this step, ⊙ is the element-by-element multiplication and rs is the reshaping operation. C The key matrix K and the value matrix V are modulated by element-by-element multiplication (⊙), and the result is rs (re-shaped) to obtain C tokens, each of which has a dimension of H×W. Q does not need to be F C Instead of performing element-by-element multiplication, we directly perform the rs operation to obtain C Tokens, each of which has a dimension of H×W.
[0083] In this step, It is the matrix of queries, keys, and values that have been modulated and reshaped to achieve attention calculation between channel tokens.
[0084] 1.4 Attention Calculation
[0085]
[0086] In this step, is the attention result; α is a learnable scaling factor; softmax() is a normalized exponential function.
[0087] 1.5 Attention Projection and Connection
[0088]
[0089] In this step, the attention result is projected to the 1×1 convolution W 1x1 , and then perform a residual connection with the input feature F to obtain the output result X of the cross attention Token mixer out Through the above steps, content-aware cross-attention Token mixing is completed.
[0090] After completing the cross-attention Token mixing, enter the feedforward network:
[0091] F out =FFN(X out )
[0092] In this step, FFN is a feed-forward network function, which is a 3×3 convolution operation in this embodiment.
[0093] For the cross-attention submodule of the encoding layer, its output F out Input to the deep convolution submodule of the second layer. For the cross attention submodule of the decoding layer, its output F out Enhanced feature E as the output of the FGEM module.
[0094] (2) Deep convolution submodule at the deep level
[0095] For deeper layers, this embodiment introduces a content-aware deep convolution submodule that combines deep convolution with content-guided feature modulation. This design choice enables the network to effectively process features with larger receptive fields while maintaining computational efficiency. The content-aware design helps the mixer capture and integrate semantic information during feature refinement.
[0096] In this embodiment, the deep convolution submodule includes: a deep convolution Token mixer and a feedforward network, which are described in detail as follows.
[0097] Figure 4 for Figure 1 Schematic diagram of the deep convolutional Token mixer in the FGEM module in the progressive low-light image enhancement system. In this embodiment, the content-aware convolution of feature levels level 2 and level 3, where the input feature F is combined with the content feature F through a learnable scaling parameter β C Modulation is performed followed by depthwise convolution.
[0098] Specifically, in this embodiment, for the input feature The token mixing process is as follows:
[0099] 2.1 Element-by-element multiplication
[0100] F mix =F⊙(β·F C )
[0101] In this step, the input feature F is combined with the scaled content feature β·F C Perform element-by-element multiplication to obtain the guided features β is a learnable scaling parameter.
[0102] 2.2 Convolution
[0103] X out =DwConv3(Fmix )
[0104] Among them, the guide feature is subjected to a deep 3×3 convolution operation - DwConv3 to generate the output result of the deep convolutional Token mixer:
[0105] In this embodiment, the deep convolutional token mixer can effectively capture deep semantic information while maintaining computational efficiency. Through the above steps, deep convolutional token mixing is completed.
[0106] After completing the deep convolutional Token mixing, X out Input feedforward network:
[0107] F out =FFN(X out )
[0108] In this step, FFN is a feed-forward network function, which is a 3×3 convolution operation in this embodiment.
[0109] Those skilled in the art should understand that in the U-type MetaFormer network, for each deep convolution submodule, the obtained F out Input to the next submodule.
[0110] Regarding the image synthesis module, it is used to add the enhancement feature E to the brightened image I g , and get the enhanced image I output , specifically: I output =I g +E.
[0111] So far, the overall architecture of the progressive low-light image enhancement system of this embodiment has been introduced.
[0112] Before the progressive low-light image enhancement system of this embodiment is actually used, it needs to be trained. The following is an explanation of the relevant content of the training.
[0113] 1. Loss Function
[0114] The LLIE task is regarded as a supervised image restoration problem. Considering L 1 In order to achieve the goal of losing robustness in low-level image processing and obtaining a clear and sharp output image, this embodiment uses L 1 As part of the loss function:
[0115]
[0116] Among them, I out is the enhanced output image, I GT is the true image, and N is the total number of pixels.
[0117] To emphasize structural accuracy, this embodiment introduces SSIM into the loss function. SSIM captures the dependencies between pixels by evaluating brightness, contrast, and structural differences, especially for spatially adjacent pixels. A higher SSIM value indicates a higher similarity. 1 or L 2 Norm, SSIM is more suitable for measuring spatial pixel differences. SSIM loss is defined as:
[0118]
[0119] Among them, μ our and μ GT They are and The mean of out and σ GT They are and The variance of yes and The covariance of l and C 2 is a small constant used to stabilize the weak denominator, both of which are less than 10 -3 .
[0120] The total loss function combined with L 1 and L S :
[0121]
[0122] Where λ is the weight factor of SSIM loss, 0.05≤λ≤0.15.
[0123] 2. Training Data
[0124] The experiments in this embodiment used multiple low-light datasets, including LOL-V1, LOL-V2, SID, and MIT-AdobeFiveK, for training and evaluation. LOL-v1 contains 500 pairs of low-light and normal-light images. LOL-v2 is divided into two subsets, real and synthetic, containing 789 pairs of real low-light images and 1000 pairs of synthetic low-light images. These datasets are divided into training sets and test sets, with the ratios of 485:15 for LOL-v1, 689:100 for LOL-v2-real, and 900:100 for LOL-v2-synthetic. SID consists of 2697 pairs of short / long exposure RAW image pairs, which are converted from RAW format to RGB format. MIT-Adobe FiveK contains 5000 pairs of low-light / normal-light image pairs adjusted by five experts AE. This embodiment uses images adjusted by expert C. The division ratio of the training set and the test set is: 2099:598 for SID and 4500:500 for MIT-Adobe FiveK.
[0125] 3. Training Details and Experimental Setup
[0126] For LOLV1 and LOLV2, independent models are trained. The training parameters are kept the same for all experiments. The learned brightness factor in the global brightness module is initially set to γ = 2.1, because values around 2.1 are usually used for image brightening. FGEM adopts a 3-layer U-shaped architecture, and the distribution of UMFormer blocks in each layer is [2, 2, 1]. The total loss L total =L 1 +λL S The weight factor λ in is set to 0.1, which achieves a good balance between pixel-level accuracy and structural similarity in our experiments. The training adopts the AdamW optimizer with parameter β 1 =0.9,β 2 =0.999, weight decay is 1×10 -4 , 250K iterations. The initial learning rate is 2.5×10 -4 , reduced to 1×10 by cosine annealing scheduling -6 . Training uses randomly cropped 128×128 pixel patches from paired low-light and normal-light images. Data augmentation includes horizontal and vertical flipping. The experiments are performed on an Ubuntu 18.04 platform, using CUDA 11.1 and cuDNN 11.0. Training is performed on a single NVIDIA Tesla V100 GPU (32GB of video memory), and the code is implemented in PyTorch.
[0127] Regarding bias, the UMFormer block does not include bias in convolutional layers, fully connected layers, or normalization operations, as it is observed that this has a negligible impact on model performance. This configuration is used as the default setting for the model. The impact of bias in the restoration task via UMFormer is left for future research.
[0128] The computational complexity of the progressive low-light image enhancement system of this embodiment is analyzed below.
[0129] Table 1 Comparison of computational complexity (K=3<<H or W, and C<<H or W)
[0130]
[0131] In this example, the U-shaped MetaFormer network (UMFormer network) achieves linear computational complexity relative to the image size by strategically combining different token mixers, which is a significant advantage over the conventional U-shaped Transformer with quadratic complexity. This efficiency is achieved by using deep convolution and channel-cross attention token mixers at different feature levels, as shown in Table 1.
[0132] Conventional U-type Transformer uses spatial self-attention as a token mixing mechanism. For an image of size H×W×C, it is divided into H×W tokens (each dimension is C), resulting in a computational complexity of O(H2W2C). The quadratic dependency of H and W makes the computational cost high. To address this limitation, this embodiment adopts two alternative mechanisms: one is channel cross attention, which treats each channel as a token, obtains C tokens (each dimension is H×W), and calculates the inter-channel attention, with a complexity of O(HWC 2 ), because C < < H or W, the computational cost is significantly reduced; the second is the depth convolution, which uses a separate convolution filter to independently process each input channel. For the convolution with a kernel size of K, its complexity is O(K 2 HWC), because K (usually 3) <<H or W, the computational cost is also significantly reduced.
[0133] Table 2 Quantitative comparison results on LOLV1, LOLV2-real, LOLV2-synthetic and SID datasets.
[0134]
[0135] This embodiment uses PSNR and SSIM to evaluate performance. A high SSIM value indicates that the high-frequency details and structural integrity of the enhanced image are maintained. High PSNR and SSIM values reflect better image quality. The computational complexity is quantified by FLOPs, and the model size is represented by the number of parameters (Params, in millions). G-UMFormer is compared with existing LLIE methods, including SID, RF, UFormer, EnGAN, DRBN, KinD, Restormer, SNR-Net, Bread, LLFormer, PyDiff and RetinexFormer. To the best of the applicant's knowledge, RetinexFormer represents the leading level among current LLIE methods.
[0136] As shown in Table 2, the G-UMFormer proposed in this embodiment achieves a balance between computational efficiency and image quality. With 1.20M parameters and a relatively low computational cost of 11.63GFLOPs, G-UMFormer achieves a PSNR of 23.71dB and an SSIM of 0.832 on the LOL-v1 dataset, 21.74dB and 0.848 on the LOL-v2-real dataset, and 25.86dB and 0.937 on the LOL-v2-synthetic dataset. Compared directly with Restormer, the method of this embodiment achieves PSNR improvements of 1.28dB, 1.79dB, and 4.45dB, and SSIM improvements of 0.009, 0.02, and 0.107 on the LOL-v1, LOL-v2-real, and LOL-v2-synthetic datasets, respectively. This performance improvement is achieved with significantly reduced computational resources - a 92% reduction in FLOPs and a 95% reduction in parameters.
[0137] Table 3 Comparison of the FIVEK dataset in sRGB color space
[0138]
[0139] Compared with existing technologies such as SNR-Net, LLFormer, and RetinexFormer, G-UMFormer provides competitive performance on the LOL dataset while significantly reducing the computational cost. As can be seen from Tables 2 and 3, the performance on the large datasets SID and FiveK also achieves a comparable efficiency level. These results show that the method of this embodiment is competitive with the state-of-the-art.
[0140] Overall, the proposed method stands out for its ability to achieve competitive image quality while maintaining low computational burden and model complexity. This feature is particularly beneficial for scenarios that require high-quality LLIE and efficient processing of low-light images.
[0141] Table 4 Ablation study of the Token mixer module proposed by UMFormer.
[0142]
[0143] Table 4 presents an ablation study of the Token mixer module of UMFormer, where I, A, and C represent Identity Mapping, Attention, and Convolutional Token mixers, respectively, to evaluate the impact of different strategies on LLIE performance, and compare attention-based and convolution-based mixers with Identity Mapping as a benchmark. The individual impacts on model performance (including FLOPs, number of parameters, PSNR, and SSIM) are determined by modifying or omitting components. In the (A, C, C) configuration, A uses the attention mechanism as the initial mixer in the first layer, and C uses convolutional layers for mixing in the second and third layers. This embodiment removes or changes the attention and convolution components, retrains the evaluation model, and observes the performance changes. Although the (A,A,A) configuration has the highest PSNR, the (A,C,C) configuration achieves a better balance between performance (23.71dBPSNR and 0.832SSIM) and resource consumption (24.09% reduction in FLOPs and 25.47% reduction in parameters compared to (A,A,A)), and significantly improves PSNR and SSIM by 23.30% and 11.53% respectively compared to the baseline (I,I,I), with only a moderate increase in computing resources.
[0144] Table 5 Ablation study of Gamma correction in global brightness module and coarse-grained denoising module
[0145]
[0146] Table 5 shows an ablation study to examine the impact of the learnable gamma correction and coarse-grained denoising mechanism in G-UMFormer. Among them, S1-S4, S4 is the setting of this embodiment. The study outlines five different methods, from the basic Identity (no gamma correction) method to a comprehensive method that includes learnable gamma correction and denoising features. The results show that the introduction of learnable gamma correction and coarse-grained denoising mechanisms significantly improves the model performance. Specifically, compared with the baseline configuration (S1), the full model (S4) of this embodiment achieves a significant improvement of 3.18% in PSNR (from 22.98dB to 23.71dB) and a 5.58% improvement in SSIM (from 0.788 to 0.832). The transition from traditional to learnable gamma correction (S2 to S3) alone brings a 0.30% increase in PSNR and a 1.62% improvement in SSIM, highlighting the advantages of the adaptive method of this embodiment. In addition, adding coarse-grained denoising (S3 to S4) on top of learnable gamma correction further brings 0.72% PSNR and 1.96% SSIM improvements, demonstrating the synergistic effect of these components in optimizing the quality of images generated by the UMFormer model.
[0147] Figure 5 The progressive low-light image enhancement system of this embodiment is compared with a representative method of the prior art. The results show that the method performs well in improving brightness, visibility and signal-to-noise ratio, while maintaining the natural scene appearance and reducing noise and color distortion.
[0148] Through the above tests, the beneficial effects of the progressive low-light image enhancement system of this embodiment are mainly reflected in the following aspects:
[0149] ①The computational complexity is greatly reduced
[0150] This example conducts a comprehensive analysis of UMFormer and proves that its time complexity is O(K 2 HWC), where K is the kernel size (usually set to 3), H and W represent the image height and width, and C represents the number of channels. This is faster than the time complexity of the conventional U-shaped transformer network, O(H 2 W 2 C) has been substantially improved. The method of this embodiment achieves this efficiency by minimizing redundant computations and using streamlined token mixing operations. As a result, compared to the current state-of-the-art method RetinexFormer, G-UMFormer reduces the number of floating point operations (FLOPs) and model parameters by 26% and 21%, respectively.
[0151] ②Advanced LLIE performance
[0152] Extensive benchmark experiments show that the G-UMFormer of this embodiment achieves advanced LLIE performance while significantly reducing computational and parameter overhead, as verified by peak signal-to-noise ratio (PSNR) and structural similarity index (SSIM). These results are achieved on the LOL-v2-syn, SID, and FIVEK datasets, while maintaining high accuracy on LOL-v1 and LOL-v2-real. This highlights the ability of G-UMFormer to provide high-quality enhancement with higher efficiency.
[0153] ③Higher comprehensive performance
[0154] This example proposes G-UMFormer as a novel and effective progressive low-light image enhancement system. By integrating content-aware insights of spatial and channel image features, it is able to strike a good balance between computational efficiency and high performance, ensuring a detailed and precise enhancement process. Comprehensive experiments (including ablation studies) verify that the method of this example has competitive performance in comparison with the state-of-the-art.
[0155] So far, the introduction of the progressive low-light image enhancement system of this embodiment is completed.
[0156] Based on the above progressive low-light image enhancement system, the second aspect of the present invention provides a computer program product. The computer program product includes: a computer program, which implements the progressive low-light image enhancement system of the above embodiment when executed by a processor.
[0157] Based on the above progressive low-light image enhancement system, the third aspect of the present invention provides a computer device. The computer device includes: a processor; a memory on which a computer program is stored; wherein the processor executes the computer program to implement the progressive low-light image enhancement system of the above embodiment.
[0158] For specific contents of the progressive low-light image enhancement system in the above embodiment, reference may be made to the detailed description of the previous embodiment, which will not be repeated here.
[0159] At this point, the various embodiments of the present invention have been introduced. According to the above description, those skilled in the art should have a clear understanding of the present invention.
[0160] For certain implementations, if they are not the key content of the present invention and are well known to ordinary technicians in the relevant technical field, they are not described in detail in the drawings or text of the specification due to space limitations. In this case, reference can be made to relevant existing technologies for understanding.
[0161] It should be noted that, unless explicitly indicated to the contrary, the numerical parameters in the specification and claims of the present invention may be approximate values and may be changed according to the content of the present invention. Specifically, all the numbers indicating the content of the composition, reaction conditions, etc. recorded in the specification and claims should be understood to be modified by the term "about" in all cases, and the meaning of the expression is to include a change of ±10% from a specific number in some embodiments.
[0162] The present invention may also be implemented as a device or apparatus program (e.g., a computer program and a computer program product) for performing a portion or all of the methods described herein. Such a program implementing the present invention may be stored on a computer-readable medium, or may be in the form of one or more signals. Such a signal may be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.
[0163] The present invention can be implemented by means of hardware including several different elements and by means of a suitably programmed computer. The various component embodiments of the present invention can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. The physical implementation of the hardware structure includes but is not limited to physical devices, which include but are not limited to transistors, memristors, DNA computers, single-chip microcomputers, microprocessors or digital signal processors (DSPs). In addition, the present invention is not directed to any specific programming language. It should be understood that the content of the present invention can be implemented using various programming languages, and the description of specific languages herein is to disclose the best implementation of the present invention.
[0164] Those skilled in the art should understand that in the claims and description of the present invention, the word "comprising" does not exclude the existence of elements (or steps) not listed in the claims. The word "a" or "an" preceding an element (or step) does not exclude the existence of multiple such elements (or steps).
[0165] Furthermore, the above embodiments are provided merely to enable the present invention to satisfy legal requirements, and the present invention may be embodied in many different forms and should not be construed as limited to the embodiments set forth herein.
[0166] Similarly, it should be understood that in order to simplify the present invention, in the above description of the exemplary embodiments of the present invention, the various features of the present invention are sometimes grouped together into a single embodiment, figure, or description thereof. However, the method of the invention should not be interpreted as reflecting the following intention: the claimed invention requires more features than the features explicitly stated in each claim. More specifically, as reflected in the claims, each inventive aspect lies in less than all the features of the previous single embodiment. Moreover, the embodiments can be mixed and matched with each other or with other embodiments based on design and reliability considerations, that is, the technical features in different embodiments can be freely combined to form more embodiments. Therefore, the claims following the specific embodiment are hereby explicitly incorporated into the specific embodiment, wherein each claim itself serves as a separate embodiment of the present invention.
[0167] The above specific embodiments provide a detailed description of the objectives, technical means and beneficial effects of the present invention. It should be understood that the purpose of the detailed description is to enable those skilled in the art to understand the present invention more clearly, and it is not used to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A progressive low-light image enhancement system, characterized in that: include: The global brightness enhancement module is used to improve the original image I input Perform global brightness enhancement to obtain the brightened image I g ; Coarse-grained denoising module, used to brighten image I g Perform coarse-grained denoising to obtain denoised shallow features F0; refine the denoised shallow features F0 to obtain content features F C ; The FGEM module uses a U-shaped MetaFormer network to denoise the shallow features F0 and content features F C As input, the hidden features are gradually refined to obtain fine-grained enhanced features E; Image synthesis module, which is used to add fine-grained enhancement features E to the brightened image I g , and get the enhanced image I output .
2. The progressive low-light image enhancement system according to claim 1, characterized in that: In the FGEM module, The U-shaped MetaFormer network includes: an outer N1 layer cross attention submodule; an inner N2 layer deep convolution submodule, 1≤N1,N2≤3; The cross attention submodule of each layer includes: symmetrical cross attention submodules on the encoder side and the decoder side; the depth convolution submodule of each layer includes: symmetrical depth convolution submodules on the encoder side and the decoder side; the submodules on the encoder side and the decoder side of the same layer are skipped; The cross attention submodule includes: a cross attention Token mixer and a feedforward network; the deep convolution submodule includes: a deep convolution Token mixer and a feedforward network; On the encoder side, for the outermost cross attention submodule, it uses the denoised shallow feature F0 and the content feature F C As input, the output feature is F out ; For other submodules except the outermost cross attention submodule, the output and content features F of the submodule outside the current layer are used. C As input, the output feature is F out ; The decoder side is set symmetrically with the encoder side. For the outermost cross attention submodule, its output feature F out The enhanced features E are output from the FGEM module.
3. The progressive low-light image enhancement system according to claim 2, characterized in that: In the cross-attention submodule: For the Cross-Attention Token Mixer: Input features Where H×W represents the spatial dimension and C represents the number of channels. The token mixing process includes: F N =LayerNorm(F) In this step, LayerNorm() is the layer normalization operation; F N is the normalized input feature; If the current module is the outermost cross-attention Token mixer, then F = F0; otherwise, F is the output of the upper-layer Token mixer X out ; Q=W Q F N ,K=W K F N ,V=W V F N , In this step, is the feature projection operation, represents a pixel-level 1×1 convolution followed by a channel-level 3×3 convolution; Q, K, The three are the matrices of query, key, and value obtained by the feature projection operation; In this step, ⊙ is the element-by-element multiplication and rs is the reshaping operation; In this step, is the attention result; α is a learnable scaling factor; softmax() is a normalized exponential function; In this step, the attention result is projected to the 1×1 convolution W 1x1 , and then perform a residual connection with the input feature F to obtain the output result X of the cross attention Token mixer out ; For the feed-forward network: F out =FFN(X out ) In this step, FFN is a feed-forward network function.
4. The progressive low-light image enhancement system according to claim 2, characterized in that: In the depth convolution submodule: For the deep convolutional token mixer: Input features The token mixing process includes: F mix =F⊙(β·F C ) In this step, the input feature F is combined with the scaled content feature β·F C Perform element-by-element multiplication to obtain the guided features β is a learnable scaling parameter; X out =DwConv3(F mix ) Among them, the guide feature is subjected to a deep 3×3 convolution operation - DwConv3 to generate the output result of the deep convolutional Token mixer: For the feed-forward network: F out =FFN(X out ) In this step, FFN is a feed-forward network function.
5. The progressive low-light image enhancement system according to claim 2, characterized in that: N1=1; N2=2; The U-shaped MetaFormer network includes three levels: a cross-attention submodule at level 1; deep convolution submodules at levels 2 and 3; And / or, in the cross-attention submodule and the deep convolution submodule, the feedforward network function FFN is one of the following: a 3×3 convolution operation; a combination operation of a 1×1 convolution followed by a deep 3×3 convolution; a combination operation of several layers of MLP full connection after Reshaping.
6. The progressive low-light image enhancement system according to any one of claims 1 to 5, characterized in that: In the global brightness enhancement module, histogram equalization or gamma correction is used to enhance the global brightness; And / or, in the coarse-grained denoising module, a first convolution operation and a pooling operation are performed on the enhanced image to obtain a denoised shallow feature F0; a second convolution operation is performed on the denoised shallow feature F0 to obtain a content feature F C ; Wherein, the kernel of the first convolution operation and the second convolution operation is 1: the kernel K of the pooling operation satisfies: 5≤K≤9; And / or, in the image synthesis module: I output =I g +E.
7. The progressive low-light image enhancement system according to claim 6, characterized in that: In the global brightness enhancement module, gamma correction is used to enhance the global brightness: Among them, γ is a learnable parameter adjusted during the training process; And / or, in the coarse-grained denoising module: F C =Conv1(F0) Among them, Conv1 is a convolution operation with a kernel size of 1; Pool7 is an average pooling with a kernel size of 7.
8. The low-light image enhancement system according to any one of claims 1 to 7, characterized in that: During the training process, the loss function L total for: Where λ is the weight factor of SSIM loss, 0.05≤λ≤0.15; Among them, I out is the enhanced output image, I GT is the true image, N is the total number of pixels; Among them, μ our and μ GT They are and The mean of out and σ GT They are and The variance of yes and The covariance of l and C2 are small constants used to stabilize the weak denominator, both of which are less than 10 -3 .
9. A computer program product, characterized in that include: A computer program, which, when executed by a processor, implements the progressive low-light image enhancement system according to any one of claims 1 to 8.
10. A computer device, characterized in that: include: processor; a memory having a computer program stored thereon; The processor executes the computer program to implement the progressive low-light image enhancement system according to any one of claims 1 to 8.