Low-light image enhancement method based on deep perception multi-scale attention network
Patent Information
- Application Number
- CN202610893889.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-22
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2046-06-22
AI Technical Summary
[0009]针对现有低照度图像增强方法在引入深度先验时存在的模态差异大、信息融合困难以及低光条件下深度估计不稳定等问题,本文提出了一种用于低光图像增强的深度感知多尺度注意力网络(DMSA-Net) 的低照度图像增强方法,用于在低光环境下提取鲁棒的深度信息的同时,保证深度信息与RGB特征有效协同融合,并实现低光图像增强过程中结构一致性保持与细节信息恢复的优化
[0065]1、本发明引入光照不变的深度几何先验,显著提升低照度增强的结构一致性与稳定性:本发明通过基于 Retinex 理论对低照度图像进行分解,提取对光照变化不敏感的反射分量,并在此基础上进行深度估计,避免了传统方法直接在低照度图像上估计深度所带来的噪声放大与失真问题,从而获得更加鲁棒、稳定的场景几何结构先验信息。相较于仅依赖二维图像特征的现有技术,本发明能够有效保持物体边缘、轮廓及空间层次关系。频域-空间域协同:通过傅里叶处理模块,本发明不仅在空间域恢复图像,还在频域对幅度和相位进行精细调节,有效平衡了亮度提升与纹理细节保持,减少了伪影和模糊。
Smart Images

Figure CN122434800B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of image processing and computer vision technology, and in particular to a low-light image enhancement method based on a depth-sensing multi-scale attention network. Background Technology
[0002] Low-Light Image Enhancement (LLIE) is an important task in computer vision and image processing. Its main goal is to improve the quality of images captured in low-light conditions through algorithms, enhancing brightness, contrast, and overall visual effect. In real-world scenarios, challenges such as complex and variable lighting conditions and inherent limitations of capture equipment often result in directly acquired low-light images suffering from severe brightness deficiencies, color distortion, low contrast, and significant sensor noise and color deviation. These degradation phenomena not only severely impact subjective visual experience but also significantly restrict the performance and reliability of downstream advanced vision tasks, such as image classification, object detection, and semantic segmentation.
[0003] Early low-light enhancement methods primarily relied on histogram equalization or Retinex-based models. These traditional methods enhance contrast by adjusting the global or local pixel distribution of an image, but they often struggle to adapt to complex lighting conditions, easily leading to side effects such as over-enhancement, loss of detail, or noise amplification. Furthermore, inaccurate lighting priors can cause visual distortion and color inconsistencies.
[0004] In recent years, deep learning-based methods, particularly those utilizing convolutional neural networks (CNNs) and Transformer architectures, have become the mainstream paradigm for low-light image enhancement. These end-to-end models have achieved remarkable results on multiple benchmarks by learning complex mappings from large amounts of data. However, existing methods still exhibit insufficient generalization ability and robustness when dealing with more challenging real-world scenarios such as extreme lighting conditions and complex noise. A prominent limitation is that while improving overall brightness, these models often struggle to effectively preserve details such as edges and textures.
[0005] To address the challenges of detail preservation and model generalization, introducing external prior information has become an effective solution. However, existing methods rely on semantic or edge priors that operate at the two-dimensional image level, making it difficult to fully reflect the three-dimensional geometry of real-world scenes. Considering that depth can be viewed as a geometric complement to RGB images, exploring and fusing three-dimensional geometric priors such as depth can provide more physically consistent constraints for low-light enhancement.
[0006] However, significant modal differences exist between RGB images and depth maps. This difference leads to complementarity in information representation, but also presents fundamental challenges for their collaborative utilization: First, illumination significantly impacts depth estimation of the current image. The difficulty in low-light depth estimation lies in the model's need to simultaneously infer the scene's inherent reflectivity and geometric structure attributes from severely degraded visual signals. These two tasks are coupled and interfere with each other in low light, making end-to-end direct depth estimation highly susceptible to bias. Second, after obtaining the depth map, how to effectively fuse the geometric prior of depth information with the color and texture information of the RGB image at different network layers is crucial. Simple concatenation or addition operations are insufficient to adapt to the differences in modal contributions under different scenes and feature scales. Inappropriate fusion strategies can lead to mutual interference or even conflict between the two modalities. For example, noise in the depth map may contaminate RGB features, or low-light artifacts in RGB may distort the geometric structure, thus limiting model performance improvement. Therefore, effectively fusing depth priors into the low-light enhancement process is a challenging and critical task.
[0007] Therefore, how to effectively explore and integrate 3D geometric priors such as depth to guide low-light image enhancement while ensuring enhancement effect and model generalization ability remains a key problem that needs to be solved in current research. Summary of the Invention
[0008] The purpose of this invention is to provide a low-light image enhancement method based on a depth-sensing multi-scale attention network, which is applicable to image quality improvement, scene depth estimation and low-light reconstruction, low-light imaging quality improvement in computer vision perception systems and edge computing scenarios.
[0009] To address the problems of large modal differences, difficult information fusion, and unstable depth estimation under low light conditions in existing low-light image enhancement methods when introducing depth priors, this paper proposes a low-light image enhancement method based on a depth-aware multi-scale attention network (DMSA-Net). This method extracts robust depth information in low-light environments while ensuring effective synergistic fusion of depth information and RGB features, and optimizes the preservation of structural consistency and restoration of detail information during the low-light image enhancement process.
[0010] This invention constructs a depth-aware multi-scale attention network architecture for low-light image enhancement. By explicitly introducing depth information as a geometric prior and utilizing a multi-scale attention mechanism to fuse features, it achieves more accurate and natural enhancement results. The network consists of two core modules: a pre-deep estimation module for robust depth extraction in low-light environments and a multi-scale attention encoding / decoding fusion module for effectively fusing RGB and depth features and achieving efficient image restoration.
[0011] In low-light depth estimation, this invention proposes a pre-deep estimation module. By performing Retinex decomposition and feature encoding on the low-light image, it is decomposed into illumination and reflection components. The reflection component learning branch focuses on characterizing the inherent texture and structural information of the image, suppressing the interference of illumination changes on depth estimation. Therefore, the reflection component is used as input to an advanced monocular depth model to achieve stable inference of the scene's 3D structure, thus providing physical constraints for low-light enhancement.
[0012] Furthermore, this invention implements a multi-scale depth fusion mechanism in the encoder, adaptively integrating depth geometric information into the image feature space to achieve efficient interaction between depth and RGB features. Specifically, the multi-scale depth fusion module incorporates a multi-scale context extraction module to capture robust feature representations with rich receptive fields, introduces a depth-aware attention fusion module to modulate depth priors into image features through a cross-attention mechanism, and embeds residual dense blocks to recover texture details lost in low-light environments. Through the synergistic effect of these sub-modules, the multi-scale depth fusion module achieves fine-grained control over structure preservation and detail reconstruction in low-light images.
[0013] Compared to existing methods that simply stitch together or directly weighted fuse, the deep fusion strategy of this invention makes full use of the complementary information of depth and RGB, avoids modal conflicts and noise interference, and improves the visual quality and model generalization ability of low-light image restoration while ensuring the enhancement effect and stability.
[0014] Simultaneously, on various paired and unpaired datasets, it can generate enhanced images with more natural visual effects, more balanced brightness, and richer details. Furthermore, because the proposed pre-depth estimation and multi-scale depth fusion modules implement feature adaptive modulation in the encoder part, the objective evaluation index of the low-light image enhancement method proposed in this invention is effectively improved on paired datasets through the above technical solutions. It possesses efficient geometric and image feature coordination capabilities and stable low-light enhancement performance, achieving structural consistency and detail restoration, and has good engineering application potential and broad practical promotion value. To achieve the above objectives, the technical solution adopted by this invention is as follows:
[0015] S10. Obtain the low-light image to be processed;
[0016] S20. Construct an image decomposition module based on Retinex theory to decompose the low-illuminance image and obtain the reflection component and illuminance component;
[0017] S30. Based on the reflection component, the depth estimation module generates a depth map corresponding to the low-light image to serve as prior information on scene geometry. The image decomposition module and the depth estimation module maintain fixed parameters during subsequent network training to provide stable prior information on geometry.
[0018] S40. Construct a low-light image enhancement network with an encoder-decoder structure, and perform multi-scale feature extraction and context modeling processing on the low-light image during the encoding stage to obtain a multi-scale feature representation that integrates local detail information and global context information.
[0019] S50. Construct a depth-aware attention module to fuse the depth map with the multi-scale feature representation, and perform spatial adaptive modulation of the multi-scale feature representation through an attention mechanism; the depth-aware attention module maps the depth features only to query components and key components for generating attention weights, without constructing value components for feature value transmission, so that the depth information modulates the RGB features in a guided manner, avoiding the direct propagation of depth estimation noise to the RGB feature space;
[0020] S60. The features modulated by depth perception are passed through residual dense blocks and then input into the decoder for feature reconstruction to obtain the enhanced image.
[0021] S70. Construct a hybrid loss function containing multiple types and use paired datasets to train the constructed augmented network model end-to-end.
[0022] Specifically, in step S20, an image decomposition network module based on Retinex theory is constructed. It consists of two structurally symmetrical subnetworks that do not share weights. and Composition, used to extract from input image Extracting the reflection component and light component The decomposition process can be described as follows:
[0023]
[0024] Among them, the reflection component Defined as an inherent physical property of the scene, it describes the reflectivity of an object's surface to light of different wavelengths and is independent of the lighting conditions during imaging; while the illumination component... This is used to characterize the distribution of external light intensity. Through the above decomposition process, information related to changes in light intensity can be explicitly extracted from the image.
[0025] In step S30, the reflection component extracted in step S20 is... Input depth estimation module Generate a depth map corresponding to the low-light image. This serves as prior information about the scene's geometric structure, and can be specifically described as follows:
[0026]
[0027] Among them, the image decomposition module based on Retinex theory and depth estimation module During low-light image enhancement training, the parameters are kept unchanged and used only as a deep inference engine. This is due to the reflection component. Unaffected by lighting conditions, this depth estimation process can effectively avoid the interference of low-light noise and brightness distortion on depth prediction, thereby obtaining robust prior information on geometric structure.
[0028] In step S40, specifically, the encoder includes a multi-scale context extraction module that fully utilizes low-light RGB input to extract multi-scale appearance features rich in contextual information. It then uses a cascaded multi-scale feature modulation attention module to achieve cross-level feature selection and fusion, enabling collaborative modeling of local details, structural information, and global illumination priors. This module consists of five branches, the first branch... The first branch is a 1×1 convolution, mainly used for channel mapping and preservation of local texture information; the second, third, and fourth branches... Three with different void ratios The 3×3 dilated convolution branch aims to effectively expand the receptive field without reducing feature resolution by progressively increasing the dilation rate, thereby capturing structural and illumination changes at different scales; the fifth branch... It is a global average pooling branch used to model the overall illumination distribution and contextual prior information of an image.
[0029] We are , and Three scale branches introduce a multi-scale feature modulation attention module. For the global pooling branch... Output features It has already aggregated complete scene-level context information through global pooling, so there is no need to perform additional cross-scale modulation on it.
[0030] This invention will Features extracted from branches This is considered the shallowest scale feature, containing rich local detail information. Since the shallowest feature does not require guidance from other scale features, it is directly used as the initial state of the recursive process: Based on this, the multi-scale feature modulation attention module adopts a modulation strategy of low-level features guiding high-level features. Let... and Representing features at two adjacent scales, using low-level features As a guide, the characteristics of adjacent high-rise buildings To perform attention modulation, taking the first-level multi-scale feature modulation attention module as an example, the first step is to... Output As the initial feature, it is used as the branch at this level. The query input to the multi-scale feature modulation attention module, and the branch at that level. feature Interaction is then performed. Subsequently, the output of each level of the multi-scale feature modulation attention module serves as the query input for the next level of the multi-scale feature modulation attention module, achieving recursive feature modulation across scales. This process can be represented as:
[0031] .
[0032] Here, MFMA stands for Multi-Scale Feature Modulation Attention Module. In this module, firstly... Projected into query feature Q, and simultaneously After segmentation, each segment is projected into key features. Sum value characteristics Subsequently, the query features are calculated using a scaled dot product. Key features The correlation between them is determined, and the initial attention map is obtained by Softmax normalization. :
[0033]
[0034] in, From input features The query feature representation obtained through linear mapping, From input features The key feature representation obtained by linear mapping, The embedding dimension of the feature is represented by this layer. This is used to generate the initial correlation distribution. Based on this, MFMA introduces a D-TopK sparsity constraint to filter the initial attention mapping for saliency, resulting in a sparse attention matrix. It can be represented as:
[0035]
[0036] in, For the initial attention mapping; , TopK is used to filter multi-scale significant responses across spatial dimensions; this layer Used to renormalize sparse attention. In obtaining the sparse attention matrix... Then, MFMA uses this attention to value features Modulation is performed to achieve selective interaction across feature information, and its output is defined as:
[0037]
[0038] in The first multi-scale feature modulation attention module represents the... The output feature representation of the branch, This represents the query feature representation passed to the next level. Indicates input features The sparse attention weight matrix generated in the induced manner Represents matrix multiplication. Indicates input features The value feature representation obtained through linear mapping.
[0039] By adopting the sparse cross-attention modeling strategy described above, the MFMA module can effectively suppress interference responses from irrelevant regions while retaining the core modeling capabilities of cross-attention. This guides the model to focus on spatial regions with significant structures and consistent semantics, thereby providing contextual information support with both stability and high discriminative power for subsequent feature reconstruction and interaction.
[0040] Finally, the modulation output of all levels of the multi-scale context extraction module The final output of the module is formed after splicing and convolution fusion:
[0041]
[0042] in, Indicates feature splicing, This is a fusion convolutional layer. It effectively integrates multi-scale contextual feature information from RGB images.
[0043] In step S50, a depth-sensing attention module is constructed, using multi-scale RGB contextual feature information as the main information carrier and depth features as the attention guidance signal. Depth-sensing weights are constructed through a cross-modal attention mechanism, and spatial adaptive modulation is performed on the multi-scale feature representation to match the enhancement strategy with the scene geometry.
[0044] Let the RGB features extracted by the multi-scale context extraction module be represented as: Feature maps obtained from previous depth estimation The lightweight feature extraction submodule consists of several convolutional layers. The deep features obtained through processing are represented as follows: ,in, It provides rich texture and semantic information, and This primarily describes the geometric structure and spatial relationships of the scene. The depth-sensing attention module first uses a 1×1 convolution to... Map and decompose into , and Three components. Considering that different attention branches have different focuses in their effects on features, and that different subspaces within the RGB features also exhibit different functional divisions in structure enhancement and detail restoration, It is further decomposed into two complementary subspaces through two parallel linear mappings based on 1×1 convolutions. . The decomposition operation enables decoupled bidirectional cross-modal modulation, allowing appearance-guided attention and geometry-guided attention to function independently in complementary RGB subspaces. If this decomposition step is omitted, the two types of modulation signals with fundamentally different semantics will directly act on the same value space, leading to mutual interference between geometric constraints and appearance enhancement, ultimately reducing the effectiveness of feature optimization. Therefore, Designed as a structurally consistent RGB subspace, it modulates RGB features through depth-aware queries, thereby enhancing a spatially consistent appearance response under explicit structural constraints; accordingly... As an appearance-aware RGB subspace, it injects luminosity and texture cues into depth-guided attention to alleviate the oversmoothing problem caused by relying solely on structural constraints.
[0045] The difference is that deep features Mapped only to queries AND key Two components were used, but no corresponding Value component was constructed. This is because depth information in the model primarily serves as structural constraint and spatial guidance, rather than providing discriminative visual semantic information related to the task objective, such as object categories and texture patterns. By using only... and The two components emphasize describing the geometric properties of pixels or regions in 3D space, rather than directly injecting feature values as values, thus avoiding the explicit propagation of potential noise or errors in depth estimation into the RGB representation space. This strategy enhances structural consistency while effectively maintaining the semantic integrity and stability of RGB features, making the cross-modal fusion process more robust.
[0046] In the RGB branch, the query component of depth features With RGB button Dot product operation is performed on the components to construct cross-modal attention relationships. This encoding of the correlation between RGB features from a depth-sensing perspective highlights structurally consistent regions guided by geometry. Subsequently, this attention weight is applied to the subspace of the RGB features. This generated RGB features enhanced by depth perception. This branch focuses on maintaining the semantic consistency of RGB features and modeling global dependencies under the guidance of deep structures. The above process can be represented as:
[0047]
[0048]
[0049] In the deep branch, the query component of RGB Key to deep features Component-based attention mapping It captures the geometric correlation of appearance guidance, enabling RGB cues to modulate the response process of depth-sensing attention, and further acts on another subspace of RGB features. This generates RGB features that integrate depth information and possess both structure-aware enhancement and boundary constraint characteristics. The above process can be represented as:
[0050]
[0051]
[0052] After obtaining the output features from the two branches, feature integration is achieved through 1×1 convolution and normalization. To avoid excessive perturbation of the original representation by the attention enhancement features and to improve training stability, this module introduces two learnable scaling parameters. and This is used to make an adaptive trade-off between input features and augmented features. The above process can be expressed as:
[0053]
[0054]
[0055] in, Represents the characteristic transformation function.
[0056] In step S60, the features modulated by the depth-sensing attention module are processed through residual dense blocks and then input into the decoder for feature reconstruction. Specifically, the decoder maps the features obtained in the encoding stage back to the image space through stepwise feature fusion and reconstruction operations, and finally outputs the enhanced image.
[0057] In step S70, the low-light image enhancement method based on a deep-perception multi-scale attention network employs an image reconstruction loss mechanism. Perceived loss Structural similarity loss and gradient consistency loss The image reconstruction loss is trained using a joint loss function. Used to measure the difference between an enhanced image and a true, normally lit image in pixel space; perceptual loss. This is used to impose constraints in the feature space rather than the pixel space, ensuring consistency between the enhanced image and the target image in high-level semantic features; structural similarity loss. It focuses on preserving the structural information of the image, evaluating image quality by comparing the brightness, contrast, and structural information of local regions; gradient consistency loss. This is used to maintain edge consistency between the enhanced image and the original low-light image. The blending loss function is... The above process can be represented as:
[0058]
[0059]
[0060]
[0061]
[0062]
[0063] in, This represents the number of samples in the batch. and These represent the enhanced image and the normal illumination image, respectively, with superscript indicating the result. Indicates the first Training samples. Indicates the first in the VGG-16 network Features extracted from the layers, For the selected feature layer index, These represent the weight coefficients for the corresponding layers. Through multi-layer feature constraints, this loss function can simultaneously capture local texture and global structural information. During training, the VGG-16 network parameters remain frozen and do not participate in backpropagation. and These are the mean values of the local windows, and Standard deviation For covariance, these statistics are all calculated directly from the pixels of the corresponding local regions of the image. This represents the gradient magnitude calculated by the Sobel operator. This represents the input low-light image. , It is the stability constant. The loss weights are set to 1.0, 0.5, 0.1, and 0.3, respectively.
[0064] The beneficial effects of this invention are:
[0065] 1. This invention introduces illumination-invariant depth geometry priors, significantly improving the structural consistency and stability of low-light enhancement: This invention decomposes low-light images based on Retinex theory, extracting reflection components that are insensitive to illumination changes, and then performs depth estimation based on these components. This avoids the noise amplification and distortion problems caused by directly estimating depth on low-light images in traditional methods, thus obtaining more robust and stable scene geometry prior information. Compared to existing technologies that rely solely on two-dimensional image features, this invention effectively preserves object edges, contours, and spatial hierarchy. Frequency-spatial domain collaboration: Through a Fourier processing module, this invention not only recovers the image in the spatial domain but also finely adjusts the amplitude and phase in the frequency domain, effectively balancing brightness enhancement and texture detail preservation, reducing artifacts and blurring.
[0066] 2. Achieving depth-guided spatial adaptive enhancement to avoid over-enhancement and structural damage: This invention uses the depth map as an attention guidance signal and introduces a depth-aware attention mechanism to perform spatial adaptive modulation on multi-scale features. This enables the network to dynamically adjust the enhancement strategy according to the geometric differences in different regions, suppressing noise amplification in regions with gentle depth changes and strengthening the ability to maintain structure in structural edge regions with abrupt depth changes. This effectively overcomes the problems of over-enhancement, loss of details, and structural blurring that are common in existing technologies.
[0067] 3. Multi-scale feature and geometric prior collaborative modeling enhances augmentation effects in complex scenes: During the encoding stage, this invention obtains local texture information and global semantic information through multi-scale feature extraction and context modeling, and combines this with a depth-aware attention module to achieve collaborative fusion of multi-scale RGB features and depth features. Compared to existing methods that only use single-scale or simple feature concatenation, this invention can more fully utilize contextual information at different scales, achieving stable and natural enhancement effects even in complex lighting and structural scenes. Attached Figure Description
[0068] Figure 1 This is a flowchart of the present invention;
[0069] Figure 2 This is a diagram illustrating the overall architecture of the method according to an embodiment of the present invention.
[0070] Figure 3 This is a diagram showing the internal module structure of the method of the present invention;
[0071] Figure 4 A visualization comparing our method with representative methods on the paired dataset LOL;
[0072] Figure 5 A comparative visualization of our method and representative methods on the paired datasets SICEv2 and LOLv2 synthetic dataset;
[0073] Figure 6 The visualization results show how a model trained on the LOL dataset is directly applied to an unpaired dataset. Detailed Implementation
[0074] The present invention will be further described below with reference to the accompanying drawings and specific embodiments. The illustrative embodiments and descriptions herein are used to explain the present invention, but are not intended to limit the present invention.
[0075] We implemented this framework in PyTorch and experimented on an NVIDIA A100 GPU. The model uses the Adam optimizer. Perform parameter updates, initial learning rate set to And supplemented by Weight decay enhances its generalization ability. Training consists of 200 epochs with a fixed batch size of 4. Input images are uniformly adjusted to 256×256 pixels. Simultaneously, a cosine annealing scheduler dynamically adjusts the learning rate (minimum learning rate reduced to...). ), and apply gradient clipping (threshold 0.1) to ensure training stability.
[0076] We trained our network using low-light image pairs collected from the SICEv2 and LOL datasets. The LOL dataset has two versions: v1 and v2. LOL-v2 is split into real-world and synthetic subsets, unlike LOL-v1 which only contains real-world data. The SICEv2 dataset contains 512 low-light images. We selected 30 image pairs from SICEv2 and 15 from the official LOL evaluation set for performance evaluation. Since both datasets provide reference images, we used PSNR, SSIM, and LPIPS as evaluation metrics. Higher PSNR and SSIM values indicate higher similarity to the reference image, while lower LPIPS values indicate better enhancement quality.
[0077] Furthermore, the performance of our proposed model has been extensively validated on multiple unpaired datasets, including MEF, LIME, and DICM. These unpaired datasets contain only low-light images without corresponding normal-light images.
[0078] like Figure 1 As shown, the low-light image enhancement method based on a depth-aware multi-scale attention network (DMSA-Net) proposed in this invention mainly includes a pre-deep estimation module and a main encoder-decoder module. The implementation steps are as follows:
[0079] Step 1: Low-light depth estimation. First, the input low-light image is processed using an image decomposition module based on Retinex theory. Decomposed into reflection components and light component Reflection component Defined as an inherent physical property of a scene, it describes the reflectivity of an object's surface to light of different wavelengths. The extracted reflection component... Input depth estimation module Generate a depth map corresponding to the low-light image. This serves as prior information about the scene's geometric structure.
[0080] Step 2: RGB multi-scale feature extraction and modeling. This involves extracting and modeling the low-light input image. The input codec main network architecture utilizes a multi-scale context module to obtain multi-scale appearance features rich in contextual information from multiple branches, and uses a cascaded multi-scale feature modulation attention module to achieve cross-level feature filtering and fusion of features from different branches, thereby realizing collaborative modeling of local details, structural information, and global illumination priors.
[0081] Step 3: Depth-Aware Attention Modulation. The RGB features processed by the multi-scale context module in Step 2, along with the depth information extracted in Step 1, enter the depth-aware attention module. This module uses the multi-scale RGB context features as the main information carrier, the depth features as the attention guidance signal, constructs depth-aware weights through a cross-modal attention mechanism, and performs spatial adaptive modulation on the multi-scale feature representation to match the enhancement strategy with the scene geometry.
[0082] Step 4: The model training network is optimized using a hybrid loss function.
[0083] To verify the effectiveness of this invention, the following representative low-light image enhancement methods were selected as benchmarks for comparison. These methods cover different technical approaches: RetinexNet, KinD, and KinD++ based on Retinex decomposition modeling; RUAS and SCI based on unsupervised optimization and model unfolding; Zero-DCE based on zero-reference learning; EnlightenGAN based on generative adversarial learning; LLFlow based on normalized flow generative modeling; and PairLIE based on paired supervised learning. These methods represent the mainstream technical paradigms in the current low-light enhancement field, including physical model-driven, unsupervised learning, curve mapping learning, generative modeling, and supervised learning. Furthermore, Ground Truth (GT) serves as a reference benchmark for the performance upper limit of real-world, normally lit images. The method of this invention (Ours) achieves comprehensive optimization of the above technical paths by introducing a depth-aware multi-scale attention mechanism and a cross-modal feature fusion strategy. Experiments will comprehensively compare the proposed method with the aforementioned existing technologies from two dimensions: objective image quality indicators and subjective visual effects.
[0084] This method was tested on different datasets. Experiments on the real low-light dataset LOL-V2-real show that the proposed method achieves state-of-the-art performance in both PSNR and SSIM, and also achieves an excellent performance of 0.137 in the perceptual quality metric LPIPS, with a small gap from the best method. This verifies that the proposed method has stable enhancement capabilities in complex real-world scenarios.
[0085] On the synthetic dataset LOL-V2-Synthetic, our proposed method slightly lags behind the state-of-the-art in structural similarity metrics. Analysis indicates this is primarily due to the fact that depth information focuses more on the macroscopic geometry of the scene, while its ability to represent subtle internal textures of objects is relatively limited. However, it is noteworthy that on both the synthetic dataset LOL-V2-Synthetic and the multi-exposure dataset SICE, our method outperforms the comparative methods in the perceptual quality metric LPIPS. This result confirms the effectiveness of the spatial priors inherent in depth information in improving the visual naturalness of images, guiding the generation of augmented results that better align with human visual perception.
[0086] Based on the experimental results across four datasets, our proposed method achieves optimal or near-optimal performance on most metrics, particularly demonstrating stable and excellent performance in perceptual quality metrics. This fully demonstrates the effectiveness of incorporating depth information as a geometric prior, as the spatial structure guidance it provides significantly enhances the visual realism and naturalness of the augmented results. Table 1 compares our method with other methods on paired datasets.
[0087] Figure 2 This is an architecture diagram of the method of the present invention. The network includes a Retinex-based front-end decomposition module for decomposing the input image into reflectance and illumination components, and then using the illumination-independent reflectance for robust depth estimation. The extracted depth map is fused with RGB features generated by a depth-aware attention feature fusion module within the encoder-decoder framework and a multi-branch attention module, thereby achieving structure-aware enhancement. Residual dense blocks are embedded in the network to preserve fine texture details. The overall design ensures adaptive utilization of geometric prior information under challenging low-light conditions, thereby improving the brightness, contrast, and geometric consistency of the image.
[0088] Figure 3 This is a diagram of the internal module structure of the method of the present invention. (a) is a multi-scale context extraction module, which captures multi-scale context information through parallel convolution and embeds a multi-scale feature modulation attention module. Under low-light conditions, it enhances structure-aware features through sparse cross-attention and dual Top-K selection. (b) is a depth-aware attention module, which jointly models RGB appearance features and depth features through parallel attention branches and adaptively fuses them through attention-guided feature decomposition and modulation to enhance structural consistency and illumination robustness in low-light image enhancement.
[0089] Figure 4This paper compares the visual performance of DMSA-Net with representative contrasting methods on the paired dataset LOL. The dataset includes typical low-light samples such as textured indoor bookshelves, colorful toys, unevenly lit office scenes, and complex outdoor scenes. Under real low-light conditions, RetinexNet and KinD++, based on Retinex decomposition, suffer from detail loss and color distortion in dark areas; Zero-DCE tends to produce halo artifacts at edges when improving contrast; EnlightenGAN and LLFlow remain unstable in color reproduction and structure preservation in complex textured areas; RUAS and SCI, while showing good structural integrity, do not provide sufficient overall brightness improvement; PairLIE still retains some blurred details in extremely dark areas. In contrast, this invention, through multi-scale feature extraction and depth-aware attention modulation, maintains more natural tonal transitions and edge sharpness while improving overall brightness, and effectively suppresses background noise, outputting clean and clear enhancement results in various indoor and outdoor scenes.
[0090] Figure 5 This study compares the visual performance of DMSA-Net with representative contrasting methods on the paired datasets SICEv2 and LOLv2 synthetic datasets to evaluate the generalization ability of each method on multi-exposure data and synthetic low-light data. On the LOLv2 synthetic dataset, RetinexNet exhibits significant noise amplification and blocky artifacts. EnlightenGAN and LLFlow are relatively sensitive to synthetic noise, resulting in color shifts or local overexposure. The enhancement results of SCI and Zero-DCE are relatively conservative, with limited brightness improvement. On the SICEv2 multi-exposure data, KinD++ and SCI show different degrees of color shift, and LLFlow has the risk of overexposure in highlight areas. This invention demonstrates superior noise suppression, structural edge preservation, and white balance stability in the above scenarios. The enhanced results have natural tones and clear details, verifying its good generalization ability under complex lighting and different data distributions.
[0091] Figure 6 shows the visualization effect of directly applying the model trained on the LOL dataset to an unpaired dataset. We conducted systematic visualization experiments on multiple unpaired low-light datasets, including DICM, LIME, and MEF, to verify the model's generalization ability and robustness in real-world scenes and complex lighting conditions. The experimental results cover three significantly different unpaired dataset scenes: outdoor natural landscapes, indoor complex texture scenes, and scenes with strong backlighting. Since there is no correspondence between low-light and normal-light images in these datasets, the model's performance in real-world applications can be more realistically reflected. From the visualization results, it can be observed that while some contrast enhancement methods improve overall brightness during the enhancement process, they also introduce obvious color shifts, overexposure, or texture destruction. In addition, in complex scenes, problems such as insufficient contrast or structural blurring in local areas can also be observed. In contrast, the method of this invention can still achieve stable and natural low-light enhancement effects even without pairing. In terms of brightness recovery, it effectively improves the global and local brightness distribution of the image, allowing details in dark areas to be fully revealed, while effectively avoiding overexposure in highlight areas. Regarding structure preservation, it maintains clear structural consistency in building outlines, object edges, and complex geometric structures, without producing obvious artifacts or edge blurring. In terms of color consistency, the model can effectively suppress color casts while restoring the true colors of the image, making the overall image tone more natural and conforming to the visual perception characteristics of the human eye. In terms of noise suppression, it can effectively suppress noise amplification in extremely low-light input images, exhibiting excellent smoothing performance in flat areas and dark backgrounds. Furthermore, Figure 6 The sixth column shows the depth map extracted by the pre-task in this paper. The results show that the proposed pre-task model can stably extract depth prior information from low-light images, providing effective prior guidance for low-light enhancement tasks, thereby achieving more effective low-light enhancement.
[0092] Table 1: Quantitative comparison of different methods on the LOL-v1, LOL-v2-real, LOL-v2-syn, and SICEv2 datasets. Bold indicates the best result, and underline indicates the second-best result.
[0093]
[0094] The technical solutions of the present invention are not limited to the specific embodiments described above. Any technical modifications made in accordance with the technical solutions of the present invention fall within the protection scope of the present invention.
Claims
1. A low-light image enhancement method based on a deep-perception multi-scale attention network, characterized in that, Includes the following steps: S10. Obtain the low-light image to be processed; S20. Construct an image decomposition module based on Retinex theory to decompose the low-illuminance image and obtain the reflectance component and illuminance component; S30. Based on the reflection component, input the depth estimation module to generate a depth map corresponding to the low-light image, as prior information of scene geometry; The image decomposition module and the depth estimation module maintain fixed parameters during subsequent network training to provide stable geometric priors; S40. Construct a low-light image enhancement network with an encoder-decoder structure, and perform multi-scale feature extraction and context modeling processing on the low-light image during the encoding stage to obtain a multi-scale feature representation that integrates local detail information and global context information. S50. Construct a depth-sensing attention module, fuse the depth map with the multi-scale feature representation, and perform spatial adaptive modulation of the multi-scale feature representation through an attention mechanism; The depth-aware attention module maps depth features only to query and key components for generating attention weights, without constructing value components for feature value propagation. This allows depth information to modulate RGB features in a guided manner, preventing depth estimation noise from directly propagating into the RGB feature space. S60. The features modulated by depth perception are passed through residual dense blocks and then input into the decoder for feature reconstruction to obtain the enhanced image. S70. Construct a hybrid loss function containing multiple types, and use paired datasets to train the constructed network model end-to-end.
2. The low-light image enhancement method based on a deep-perception multi-scale attention network according to claim 1, characterized in that, In step S20, the image decomposition module based on Retinex theory includes a module for extracting reflection components. Reflection subnetwork and used to extract illuminance components Illumination subnetwork The decomposition process of the image decomposition module is represented as follows: and Used to characterize low-light images The scene's inherent attribute information and lighting distribution information.
3. The low-light image enhancement method based on a deep-perception multi-scale attention network according to claim 2, characterized in that, In step S30, the depth estimation module With the reflection component As input, a depth map is generated. To reduce the impact of illumination changes on depth estimation results and improve the stability and robustness of depth information, the process is as follows: The image decomposition module and depth estimation module based on Retinex theory keep their parameters unchanged during the low-light image enhancement training process, thereby obtaining stable depth information of the low-light image as a stable geometric prior to guide the image enhancement process.
4. The low-light image enhancement method based on a deep-perception multi-scale attention network according to claim 1, characterized in that, The multi-scale feature extraction in step S40 includes parallel or cascaded processing of input features at different spatial scales to simultaneously extract local texture features and global structural features.
5. The low-light image enhancement method based on a deep-perception multi-scale attention network according to claim 4, characterized in that, The context modeling process in step S40 is performed by a multi-scale context extraction module, which is used to aggregate the contextual relationships between features at different scales to enhance the global semantic expressive power of the features.
6. The low-light image enhancement method based on a deep-perception multi-scale attention network according to claim 5, characterized in that, The multi-scale context extraction module generates context-enhanced features to guide the subsequent deep perception attention module by interacting and fusing contextual information of features at different scales.
7. The low-light image enhancement method based on a deep-perception multi-scale attention network according to claim 1, characterized in that, The depth-sensing attention module in step S50 includes a depth feature encoding unit and an image feature encoding unit, and guides and modulates the multi-scale feature representation using the depth map through a cross-attention mechanism or an equivalent attention mechanism.
8. The low-light image enhancement method based on a deep-perception multi-scale attention network according to claim 7, characterized in that, The depth-sensing attention module adaptively adjusts the response weights of corresponding image features based on the geometric structure information at different spatial locations in the depth map, in order to enhance structural regions and suppress unstructured noise regions.
9. The low-light image enhancement method based on a deep-perception multi-scale attention network according to claim 1, characterized in that, In step S60, the decoder reconstructs the encoded features by progressive upsampling and feature fusion to generate an enhanced image with natural brightness distribution, good structural consistency, and clear details.
10. The low-light image enhancement method based on a deep-perception multi-scale attention network according to claim 1, characterized in that, In step S70, the low-light image enhancement network employs an image reconstruction loss mechanism. Perceived loss Structural similarity loss and gradient consistency loss The image reconstruction loss is trained using a joint loss function; wherein the image reconstruction loss... Used to measure the difference between an enhanced image and a true, normally lit image in pixel space; perceptual loss. This is used to impose constraints in the feature space rather than the pixel space, ensuring consistency between the enhanced image and the target image in high-level semantic features; structural similarity loss. It focuses on preserving the structural information of the image, evaluating image quality by comparing the brightness, contrast, and structural information of local regions; gradient consistency loss. To maintain edge consistency between the enhanced image and the original low-light image; the overall loss function is expressed as: in The loss weights are set to 1.0, 0.5, 0.1, and 0.3, respectively.
Citation Information
Patent Citations
Mobile terminal user identity authentication device and method based on multiple biological feature modals
CN104933344A
Low-illumination image enhancement method based on Retinex theory and guided by color prior
CN119273602A