Low-illumination image enhancement method based on double-domain prior collaboration

By constructing a luminance statistical prior (LSP) and a chrominance structure prior (CSP) in the YCrCb domain, and combining them with the fusion-then-encoding and bidirectional channel cross-attention modules of the end-to-end network BIP-CENet, the problem of luminance and chrominance coordination in low-light image enhancement is solved, achieving more accurate luminance representation and chrominance stability, and improving the image enhancement effect.

CN121961952APending Publication Date: 2026-05-01SICHUAN UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SICHUAN UNIV
Filing Date
2026-01-19
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing low-light image enhancement methods struggle to accurately describe overall brightness statistics, dynamic range, and local illumination. Furthermore, under extremely low signal-to-noise conditions, hue shift and edge weakening may occur, affecting the image enhancement effect.

Method used

A dual-domain prior collaboration approach is adopted to map the RGB color space to the HVI color space, and construct a luminance statistical prior (LSP) and a chromaticity structure prior (CSP) in the YCrCb domain. Image enhancement is performed through an end-to-end low-light enhancement network (BIP-CENet), which includes a prior-gated fusion module (MCPF), a bidirectional channel cross-attention module (GB-CCA), and a guided enhancement module (GSE-IRB) to achieve collaborative interaction and enhancement of luminance and chromaticity.

Benefits of technology

It improves the stability of low-light images, reduces noise and color shift caused by brightness enhancement, stabilizes chromaticity response and enhances texture edge details, and improves the fidelity and consistency of visual structure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121961952A_ABST
    Figure CN121961952A_ABST
Patent Text Reader

Abstract

The invention discloses a low-illumination image enhancement method based on double-domain prior collaboration. The method comprises the following steps: constructing a brightness statistics prior LSP and a chroma structure prior CSP for a low-illumination image in a YCrCb domain; in an end-to-end low illumination enhancement network BIP-CENet, a conditional injection mode of first fusion and then coding is adopted, and an LSP and a CSP are respectively injected into a brightness branch I and a chrominance branch HV through an MCPF at a shallow scale; in a multi-scale coding process, bidirectional cross-branch collaborative attention interaction of I and HV branches is realized through GB-CCA, and GSE-IRB is introduced into an LSP / CSP prior branch to carry out lightweight enhancement and noise reduction on prior features so as to provide more robust guidance information; and in a decoding stage, residual reconstruction is carried out on the enhanced HVI representation, and the HVI representation is approximately reversibly mapped back to RGB through PHVIT, so that an enhanced low-illumination image is obtained. According to the invention, through a cooperation normal form of hierarchical representation coding-prior-driven cross-domain cooperation-consistency reconstruction, a low-illumination enhancement effect with continuous colors, clear details and natural vision is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of image processing technology, specifically relating to a low-light image enhancement method based on dual-domain prior collaboration. Background Technology

[0002] Due to the rapid development of digital image processing, computer vision systems have been widely applied in fields such as smart construction, industrial inspection, and life sciences. However, in these application scenarios, the imaging process is inevitably affected by adverse lighting conditions such as low light and non-uniform illumination. This often leads to a significant decrease in the performance of computer vision algorithms, thus having a cascading impact on key vision tasks such as segmentation, detection, and classification. Therefore, to meet the quality requirements of related upstream tasks, low-light image enhancement is needed, the purpose of which is to improve image brightness while reducing the impact of noise and color deviation.

[0003] To address these needs, most existing low-light image enhancement methods directly learn the mapping from low light to normal light in the sRGB domain. While the pipeline is simple, brightness and chromaticity are strongly coupled, often resulting in color shift, noise amplification, and texture blurring during brightening. To alleviate this coupling, some works first convert the image from sRGB to a color decoupling space, and then process brightness and chromaticity separately; for example, HSV uses the maximum value of the three channels as the brightness / intensity components, but it has two typical problems: First, phase breaks occur at the red boundary of hue, separating similar reds at both ends of the hue ring, which easily produces red discontinuity artifacts after enhancement; second, in extremely dark regions, saturation and chromaticity radii are close to zero and mixed with sensor noise. During brightening, these noises are amplified simultaneously, forming black plane noise and color noise block artifacts.

[0004] To improve the above problems, HVI made two modifications based on HSV: (1) the ring representation of hue was rewritten as continuous planar components H and V, so that red no longer breaks at the hue boundary, reducing the Euclidean distance between similar reds and improving color continuity; (2) while retaining the chromaticity radius determined by saturation, an adaptive gating coefficient related to brightness was introduced, and the chromaticity radius in the dark area automatically shrinks to suppress black plane noise, and the color intensity is gradually restored as the brightness increases. Nevertheless, this decoupling route still has two key shortcomings: (1) in HSV and HVI, the intensity is usually taken as the maximum value of the three channels, which can only reflect the brightest component and is difficult to accurately describe the overall brightness statistics, dynamic range and local illumination; (2) H and V are still geometric reparameterizations and do not contain gradient or structural priors. Hue drift and edge weakening may still occur under extremely low signal-to-noise conditions. Summary of the Invention

[0005] To address the aforementioned shortcomings in existing technologies, this invention provides a low-light image enhancement method based on dual-domain prior collaboration. This method solves the problem that existing methods struggle to accurately describe overall brightness statistics, dynamic range, and local illumination. Furthermore, since H and V are geometrically reparameterized, hue shift and edge weakening may still occur under extremely low signal-to-noise conditions, thus affecting the low-light image enhancement effect.

[0006] To achieve the aforementioned objectives, the present invention employs the following technical solution: a low-light image enhancement method based on dual-domain prior collaboration, comprising the following steps: For low-light images, the RGB color space is mapped to the HVI color space, and two types of prior information are constructed in the YCrCb domain: the luminance statistics prior LSP and the chromaticity structure prior CSP. Construct and train an end-to-end low-light enhancement network BIP-CENet; the end-to-end low-light enhancement network BIP-CENet is an HV-I dual-branch encoding and decoding structure, including a priori gated fusion module MCPF, a bidirectional channel cross attention module GB-CCA, and a guided enhancement module GSE-IRB; In the end-to-end low-light enhancement network BIP-CENet: Based on the conditional injection method of fusion before encoding, at the shallow scale, the luminance statistical prior LSP and chrominance structure prior CSP are injected into the luminance branch I and chrominance branch HV respectively through the prior gated fusion module MCPF. During the multi-scale encoding process, the bidirectional cross-branch collaborative interaction between the luma branch I and the chroma branch HV is realized through the bidirectional channel cross-attention module GB-CCA to stabilize chroma and enhance texture edges. At the same time, the guided enhancement module GSE-IRB is introduced into the LSP / CSP prior branch to perform lightweight enhancement and noise reduction on the prior features, generating robust guided features for use by the bidirectional channel cross-attention module GB-CCA. In the decoding stage, the luminance branch I and chrominance branch HV are reconstructed, multi-scale upsampling is performed, and the enhanced HVI is fused with the initial HVI through skip-connection fusion. The enhanced low-light image is then approximately reversibly mapped back to RGB via PHVIT.

[0007] Furthermore, the luminance statistical prior LSP is a multi-view luminance statistical stack constructed from the Y channel of a low-light image in the YCrCb domain, including a normalized linear luminance Y. lin Logarithmic compression Y log Weber comparison, single-scale Retinex and HVI-gated isomorphism of the luminance sensitivity term Y sin_k ; The chromaticity structure prior (CSP) is constructed from the Cr / Cb channels in the YCrCb domain for low-light images, including: The Cr and Cb channels of low-light images were residualized with a neutral value of 0.5 and then standardized and normalized according to the sample standard deviation. The chromaticity magnitude and its gradient energy obtained by the Sobel operator are calculated on the normalized chromaticity plane. The gradient energies corresponding to the Cr and Cb channels are weighted and fused into the chromaticity structure energy, which is used as the chromaticity structure prior structure term. By adaptively gating and coupling the luminance sensitivity term with the chromaticity structure energy, the chromaticity radius is automatically reduced and structural noise is suppressed in the dark area, while color and edges are preserved in the medium and high luminance areas, thus constructing the chromaticity structure prior CSP.

[0008] Furthermore, the luminance statistical prior LSP is expressed as: The chromaticity structure prior CSP is represented as follows: In the formula, Indicates the brightness channel. Represents the local mean. This represents the low-pass approximation, where α>0, ε>0. , They represent C respectively b C r The chromaticity residuals after being residualized with a neutral value of 0.5 and normalized to the sample standard deviation are... Let k represent the Sobel operator. hvi An exponential parameter that is identical to the HVI brightness gating.

[0009] Furthermore, the prior gating fusion module MCPF injects the luminance statistical prior (LSP) and chrominance structure prior (CSP) into the luminance branch I and chrominance branch HV of the low-light image, respectively, including: The luminance statistical prior (LSP) and chrominance structure prior (CSP) are encoded and aligned to obtain the corresponding luminance guiding features and chrominance guiding features. The luminance features of the luminance branch I of the low-light image are used as the main path features, and the corresponding luminance guiding features are used as the auxiliary path features to form a luminance feature pair; the chrominance features of the chrominance branch HV of the low-light image are used as the main path features, and the corresponding chrominance guiding features are used as the auxiliary path features to form a chrominance feature pair. The luminance feature pair and the chrominance feature pair are respectively input into two prior gated fusion modules MCPF with identical structures for prior injection; wherein, one prior gated fusion module MCPF is used to inject the luminance guiding feature corresponding to the luminance statistical prior LSP into the luminance branch I, and the other prior gated fusion module MCPF is used to inject the chrominance guiding feature corresponding to the chrominance structure prior CSP into the chrominance branch HV. In the prior gating fusion module MCPF, the processing of luminance / chrominance feature pairs includes: The main road features and auxiliary road features are processed by 1×1 convolution for channel projection and 3×3 depth convolution respectively to obtain the local response features of the main road and the local response features of the auxiliary road. Extract value features from the auxiliary path features through a branch containing a 1×1 convolution; The local response features of the main road and the local response features of the auxiliary road are added element by element in spatial location, and then activated by Sigmoid to form a spatial gate; By using spatial gate and value features to perform element-wise multiplication to achieve position-wise modulation, intermediate gating results are obtained. The intermediate gating results are fused through a linear bottleneck consisting of two 1×1 convolutions, and ReLU activation is introduced between the two 1×1 convolutions before being fed back to the original number of channels to obtain the output features after spatial gating and bottleneck fusion. The channel gate is extracted from the main path features using an SE-style channel gate extraction branch. The SE-style channel gate extraction branch includes: performing global average pooling on the main path features, then sequentially passing them through two 1×1 convolution layers with ReLU activation introduced in the middle, and finally obtaining the channel gate through a Sigmoid function. The output features are channel scaled using a channel gate and added to the main features to form a residual output, thus obtaining the updated main features, thereby completing the injection of the luminance statistical prior (LSP) / chrominance structure prior (CSP).

[0010] Furthermore, through cross-branch collaborative attention via the bidirectional channel cross-attention module GB-CCA, bidirectional collaborative interaction between the luma branch I and the chroma branch HV is achieved at multiple scales to stabilize chroma and enhance texture edges, including: The input luminance feature and its corresponding first guide map are set, as well as the chrominance feature and its corresponding second guide map; wherein, the first guide map is used to guide the collaborative interaction direction from chrominance to luminance, and the second guide map is used to guide the collaborative interaction direction from luminance to chrominance; Channel-independent weight coefficient pairs and spatial guidance maps are generated based on the first guidance map and the second guidance map, respectively; wherein, the channel-independent weight coefficient pairs are generated by global average pooling of the first guidance map and the second guidance map and then obtained by convolutional mapping, and then split; the spatial guidance maps are generated by convolutional mapping of the first guidance map and the second guidance map, respectively. Residual element-wise gating is performed on luminance features and chrominance features using the weighted coefficient pairs and spatial guided mapping respectively; wherein, in the "chrominance to luminance" direction, the weighted coefficient pairs generated by the first guided map and spatial guided mapping are used, and in the "luminance to chrominance" direction, the weighted coefficient pairs generated by the second guided map and spatial guided mapping are used. After gating, bidirectional channel cross-attention is performed: using the luminance branch as the query and the chrominance branch as the key and value, the channel cross-attention from chrominance to luminance is calculated to obtain the channel attention luminance feature; using the chrominance branch as the query and the luminance branch as the key and value, the channel cross-attention from luminance to chrominance is calculated to obtain the channel attention chrominance feature; wherein, the channel cross-attention includes: normalizing the gating features; generating queries, keys, and values ​​through convolutional projection; rearranging the queries, keys, and values ​​in a multi-head manner based on the generated queries, keys, and values, normalizing them in the spatial dimension, calculating channel correlation, introducing a learnable scaling factor to stabilize channel energy, performing weight normalization in the channel dimension to aggregate value features; and setting residual connections to stabilize feature responses. Lightweight refinement and enhancement are performed on the channel attention luminance features and channel attention chrominance features respectively, including: normalizing and convolutional transforming the respective channel attention luminance / chrominance features and splitting them into two sub-features along the channel dimension; performing lightweight convolution and nonlinear refinement on the two sub-features respectively to obtain two refined sub-features; and then multiplying the two refined sub-features element by element to obtain the fusion enhancement result. The fusion enhancement results of the luminance branch and the chrominance branch are respectively subjected to channel back projection through 1×1 convolution, and fused with the corresponding channel attention features to obtain the refined output features of the luminance branch I and the chrominance branch HV, so as to stabilize the chrominance and enhance the texture edge.

[0011] Furthermore, the expressions for generating channel-independent weight vectors and spatial guided mappings are as follows: In the formula, and Indicates The two scalar gates cut off This represents the function of splitting in half. This represents the Sigmoid activation function. Represents the ReLU activation function. express convolution, express convolution, Indicates global average pooling. This represents a guide mask with fine spatial granularity. The superscript X indicates the guide map type, X=A indicates a luminance guide map, and X=B indicates a chrominance guide map.

[0012] Furthermore, channel attention is performed on the luminance / chrominance features, including: The gated luminance / chrominance features are pre-normalized using LayerNorm in the channel dimension; local projection is performed on the luminance branch I / chrominance branch HV, and then passed through 1×1 convolution and 3×3 depth convolution to obtain the query Q; local projection is performed on the chrominance branch HV / luminance branch I, and then passed through 1×1 convolution and 3×3 depth convolution to obtain the features, which are then split into two in the channel dimension to obtain the key K and value V; Rearrange Q, K, and V into a bullish pattern, and perform L2 normalization on Q and K in the last dimension N; The channel-dimensional attention weights are calculated based on the normalized Q and K, and V is weighted and aggregated. The heads are then merged and projected by 1×1 to obtain the channel attention output. The channel attention output is added to the corresponding gated luminance / chrominance feature by taking the residual, and the channel attention luminance / chrominance feature is obtained.

[0013] Furthermore, the guiding features of the prior branch LSP / CSP are enhanced with steady-state details and contrast through the guiding enhancement module GSE-IRB, including: The input features are first expanded by 1×1 point convolution to obtain intermediate features, and then 3×3 depth convolution and element-wise SiLU activation are performed to obtain local nonlinear representations. The local nonlinear representation is halved along the channel dimension into two sub-features. The second sub-feature is gated by Sigmoid and multiplied element-wise with the first sub-feature to form a position-dependent channel-gated feature. The channel-gated features are projected onto the target number of channels through a 1×1 convolution, and recalibrated features are obtained by performing channel recalibration through an SE structure. The recalibrated features and input features are scaled and residually fused to obtain the output features; wherein, when the number of input / output channels is inconsistent, the residual branches are aligned by 1×1 convolution, and the scaling factor is a preset constant.

[0014] Furthermore, channel recalibration is performed through the SE structure, as follows: In the formula, Indicates recalibration features, This indicates the feature after being fed back to the target number of channels, and ⊙ indicates element-wise multiplication.

[0015] Introducing parameter integration with coefficients to perform scaled residual fusion of recalibrated features and input features, expressed as: In the formula, Indicates the output of the residual branch. Let χ represent the input features of the i-th scale and branch. Indicates the number of input channels. Indicates the number of output channels. This represents a 1×1 convolution used for channel alignment. This represents the residual scaling factor.

[0016] Furthermore, the training loss function of the end-to-end low-light enhancement network BIP-CENet is: In the formula, ℓ(·,·) represents the composite reconstruction loss calculated in the image domain, and λ is the weight hyperparameter. This represents the enhanced HVI feature map output by the end-to-end low-light enhancement network BIP-CENet. This represents the HVI truth mapping constructed under the guidance of the YCrCb–HVI dual-domain prior. express The reconstructed image obtained by inverse transformation Φ_RGB mapping back to the sRGB domain This represents a reference, normally exposed sRGB image.

[0017] The beneficial effects of this invention are as follows: (1) Based on dual-domain prior collaboration (constructing a luminance statistical prior LSP and a chromaticity structure prior CSP in the YCrCb domain and collaborating with the HVI domain enhancement process), it can more accurately characterize the overall luminance statistics, dynamic range and local illumination changes of low-light images, thereby improving the enhancement stability in weak light and non-uniform lighting scenarios and reducing the noise and color shift amplification problems caused by luminance enhancement.

[0018] (2) By using the prior-gated fusion module MCPF to achieve the conditional injection of "fusion before encoding", the prior and backbone features can be aligned and adaptively gated at the shallow scale, thereby improving the efficiency of prior utilization and suppressing the interference of invalid / noisy priors on backbone features, and enhancing the robustness of the enhancement results.

[0019] (3) By using the bidirectional cross-attention module GB-CCA to achieve bidirectional collaborative interaction between the luminance branch I and the chrominance branch HV at multiple scales, the chrominance response can be stabilized and texture edge details can be enhanced during the enhancement process, thereby reducing hue drift and edge weakening under extremely low signal-to-noise conditions and improving visual structure fidelity. (4) By using the Guided Enhancement Module GSE-IRB to perform lightweight enhancement and noise reduction on prior / guided features, more robust guided information can be generated for cross-branch interaction, thereby further improving dark detail and contrast performance, and improving the consistency and generalization ability of the enhancement results. Attached Figure Description

[0020] Figure 1 The flowchart of the low-light image enhancement method based on dual-domain prior collaboration provided by the present invention is shown.

[0021] Figure 2 The diagram shows the structure of the end-to-end low-light enhancement network BIP-CENet provided by this invention.

[0022] Figure 3 The structural diagram of the prior gating fusion module MCPF provided by the present invention.

[0023] Figure 4 The structure diagram of the bidirectional channel cross-attention module GB-CCA provided by the present invention.

[0024] Figure 5 The diagram shows the structure of the boot enhancement module GSE-IRB provided by this invention.

[0025] Figure 6 The figure shows the comparison results of BIP-CENet provided by this invention and the comparison method on objective indicators PSNR, SSIM, and LPIPS.

[0026] Figure 7 The BIP-CENet method provided in this invention is visually compared with the LOLv2-Real dataset.

[0027] Figure 8 The BIP-CENet method provided in this invention is visually compared with the LOLv2-Syn dataset.

[0028] Figure 9 The figures show the NIQE comparison results between BIP-CENet and the comparison method provided by this invention. (a) NIQE score radar chart; (b) NIQE score bubble chart.

[0029] Figure 10 The BIP-CENet provided by this invention is visually compared with the comparison method on the LIME, NPE, and MEF datasets.

[0030] Figure 11 The image shows a comparison of BIP-CENet ablation effects and a statistical chart of local RGB color distribution provided by this invention. Detailed Implementation

[0031] The specific embodiments of the present invention are described below to enable those skilled in the art to understand the present invention. However, it should be understood that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the present invention as defined and determined by the appended claims. All inventions utilizing the concept of the present invention are protected.

[0032] This invention provides a low-light image enhancement method based on dual-domain prior collaboration. It introduces a dual-domain prior based on the YCrCb domain into the HVI color space, performing complementary modeling of intensity and chromaticity. Based on the HVI representation and dual-domain prior, an end-to-end low-light enhancement network, BIP-CENet, is further proposed. This network jointly models luminance and chromaticity in both the HVI and YCrCb color domains, following a collaborative paradigm of "hierarchical representation encoding—prior-driven cross-domain collaboration—consistent reconstruction," resulting in low-light enhancement effects that are color-continuous, detail-clear, and visually natural.

[0033] refer to Figure 1 A low-light image enhancement method based on dual-domain prior collaboration includes the following steps: For low-light images, the RGB color space is mapped to the HVI color space, and two types of prior information are constructed in the YCrCb domain: the luminance statistics prior LSP and the chromaticity structure prior CSP. The end-to-end low-light enhancement network BIP-CENet was constructed and trained. The end-to-end low-light enhancement network BIP-CENet is an HV-I dual-branch encoding and decoding structure, including the prior gated fusion module MCPF, the bidirectional channel cross attention module GB-CCA, and the guided enhancement module GSE-IRB. In the end-to-end low-light enhancement network BIP-CENet: Based on the conditional injection method of fusion before encoding, at the shallow scale, the luminance statistical prior LSP and chrominance structure prior CSP are injected into the luminance branch I and chrominance branch HV respectively through the prior gated fusion module MCPF. During the multi-scale encoding process, the bidirectional cross-branch collaborative interaction between the luma branch I and the chroma branch HV is realized through the bidirectional channel cross-attention module GB-CCA to stabilize chroma and enhance texture edges. At the same time, the guided enhancement module GSE-IRB is introduced into the LSP / CSP prior branch to perform lightweight enhancement and noise reduction on the prior features, generating robust guided features for use by the bidirectional channel cross-attention module GB-CCA. In the decoding stage, the luminance branch I and chrominance branch HV are reconstructed, multi-scale upsampling is performed, and the enhanced HVI is fused with the initial HVI through skip-connection fusion. The enhanced low-light image is then approximately reversibly mapped back to RGB via PHVIT.

[0034] The main degradation of low-light images manifests as underexposure and increased noise, while the chromaticity structure usually retains some usability in the local neighborhood. Traditional methods only perform convolution or attention modeling in the RGB domain. Due to the high coupling between brightness and chromaticity, strong brightening often leads to color cast, noise amplification, and loss of detail / texture. Therefore, color space decoupling has become a common approach, but relying solely on HVI representation still has two key limitations: First, the intensity component is usually defined as the maximum value of the three channels I = max(R,G,B), which only reflects the brightest channel and is difficult to stably characterize global brightness statistics, dynamic range, and local illumination and contrast changes; Second, H and V are still geometrically reparameterized and do not have explicitly embedded structures or gradient priors. Therefore, in dark areas with extremely low signal-to-noise ratios, phenomena such as planarization (black plane) chromaticity noise, edge weakening, and hue shift may still occur.

[0035] Based on this, the present invention introduces a joint Y+CrCb prior and works in conjunction with HVI.

[0036] Among them, the luminance statistical prior (LSP) is a multi-view luminance statistical stack constructed from the Y channel of the low-light image in the YCrCb domain, including normalized linear luminance Y. lin Logarithmic compression Y log Weber comparison, single-scale Retinex and HVI-gated isomorphism of the luminance sensitivity term Y sin_k ; The chromaticity structure prior (CSP) is constructed from the Cr / Cb channels in the YCrCb domain for low-light images, including: The Cr and Cb channels of low-light images were residualized with a neutral value of 0.5 and then standardized and normalized according to the sample standard deviation. The chromaticity magnitude and its gradient energy obtained by the Sobel operator are calculated on the normalized chromaticity plane. The gradient energies corresponding to the Cr and Cb channels are weighted and fused into the chromaticity structure energy, which is used as the chromaticity structure prior structure term. By adaptively gating and coupling the luminance sensitivity term with the chromaticity structure energy, the chromaticity radius is automatically reduced and structural noise is suppressed in the dark area, while color and edges are preserved in the medium and high luminance areas, thus constructing the chromaticity structure prior CSP.

[0037] The chromatic structure prior (CSP) obtained using the above method both perceives chromaticity amplitude and encodes edge and texture structures. After introducing the above LSP / CSP, the black plane color noise is significantly reduced and the BPCR is significantly decreased under the same dark threshold, verifying the noise suppression and brightening effect of the dual-domain prior.

[0038] In the process of constructing the aforementioned dual-domain prior, low-light images are used. Converting from RGB to HVI: Where I = max(R,G,B) is the intensity component, (H,V) is the orthogonal chromaticity component related to hue, and Ψ HVI For explicit color conversion from RGB to HVI, Φ RGB Analyze its inverse transform; implement the expression containing intensity-sensitive terms. .

[0039] Therefore, the luminance statistical prior LSP is constructed as follows: The chromatic structure prior CSP is represented as follows: In the formula, Indicates the brightness channel. Represents the local mean. This represents the low-pass approximation, where α>0, ε>0. , They represent C respectively b C r The chromaticity residuals after being residualized with a neutral value of 0.5 and normalized to the sample standard deviation are... Let k represent the Sobel operator. hvi An exponential parameter that is identical to the HVI brightness gating.

[0040] In this embodiment, the luminance statistical prior LSP provided by Y and the chromaticity structure prior CSP provided by Cr / Cb complement each other in the HVI expression: the former provides steady-state constraints on global luminance and local contrast, while the latter suppresses color shift and color noise and improves edge consistency under luminance gating, providing a more physically and statistically reasonable input representation for the subsequent joint enhancement of the network.

[0041] In this invention, to more effectively achieve consistent enhancement of brightness and chromaticity under low-light conditions, suppress false color and noise, and maintain the continuity of edge structures, a low-light image enhancement network, BIP-CENet, is constructed, such as... Figure 2As shown, the network is based on dual-domain priors of YCrCb and HVI: the conditional injection of fusion before encoding is achieved through the prior-gated fusion module MCPF, and LSP and CSP are modulated into the I / HV backbone respectively; in the multi-scale encoding-decoding process, the guided enhancement module GSE-IRB is introduced into the LSP / CSP prior branch to enhance the steady-state details and contrast of the prior features to generate guided features, and the bidirectional channel cross-attention module GB-CCA completes the collaborative interaction of the luminance branch I and the chrominance branch HV at each scale; finally, after the enhanced representation is obtained in the HVI domain, it is residually fused with the initial HVI representation, and mapped back to sRGB through the approximately reversible PHVIT inverse transform, so as to achieve a joint improvement in luminance stability, pseudo-color suppression and edge consistency.

[0042] In order to avoid explicitly introducing O((HW) 2 Under the premise of global attention overhead at the level of 0, this invention introduces a priori gating fusion module MCPF into BIP-CENet by robustly injecting luminance and chrominance priors into the HVI backbone. For example... Figure 3 As shown, the prior gating fusion module MCPF treats a pair of primary and secondary features as or By constructing a spatial gate using a lightweight Q–K–V form and combining it with SE channel recalibration, prior information is gated and injected at both the spatial and channel levels, thereby achieving "prior-driven local cross-fusion": on the one hand, it maintains the statistical consistency of the intensity / chroma backbone itself, and on the other hand, it effectively enhances the guiding ability of luminance statistical prior and chroma structure prior on backbone features; among them, the "fusion before encoding" injection occurs in the shallow stage, enabling prior guidance to propagate step by step to deeper multi-scale features along with the backbone encoding process.

[0043] Specifically, based on Figure 3 The structure shown illustrates that the prior-gated fusion module MCPF injects the luminance statistical prior (LSP) and chrominance structural prior (CSP) into the luminance branch I and chrominance branch HV of the low-light image, respectively, including: The luminance statistical prior (LSP) and chrominance structure prior (CSP) are encoded and aligned to obtain the corresponding luminance guiding features. and chromaticity guiding features ; Brightness features of low-light image brightness branch I As the main path feature, the corresponding brightness guidance feature will be used. As auxiliary features, they form a brightness feature pair; the chromaticity features of the chromaticity branch HV of low-light images are used. As the main path feature, the corresponding chroma guiding feature will be used. As auxiliary path features, they constitute chromaticity feature pairs; Luminance feature pairs and chrominance feature pairs are respectively input into two identical prior-gated fusion modules (MCPF) for prior injection; one of the MCPF modules is used to inject the luminance guiding features corresponding to the luminance statistical prior LSP. Injecting luminance branch I, another prior gating fusion module MCPF is used to inject the chroma guiding features corresponding to the chroma structure prior CSP. Inject chromaticity branch HV; In the prior gated fusion module MCPF, the processing of luminance / chrominance feature pairs includes: The main road features and auxiliary road features are processed by 1×1 convolution for channel projection and 3×3 depth convolution respectively to obtain the local response features of the main road and the local response features of the auxiliary road. Extract value features from the auxiliary path features through a branch containing a 1×1 convolution; The local response features of the main road and the local response features of the auxiliary road are added element by element in spatial location, and then activated by Sigmoid to form a spatial gate; Position-by-position modulation is performed using spatial gate pair value features to obtain intermediate gating results; The intermediate gating results are fused through a linear bottleneck consisting of two 1×1 convolutions, and ReLU activation is introduced between the two 1×1 convolutions before being fed back to the original number of channels to obtain the output features after spatial gating and bottleneck fusion. The channel gate is extracted from the main path features using the SE-style channel gate extraction branch. The SE-style channel gate extraction branch includes: performing global average pooling on the main path features, then passing them through two 1×1 convolution layers with ReLU activation in between, and finally obtaining the channel gate through Sigmoid. Channel scaling is applied to the output features using channel gates, and the residual output is added to the main path features to obtain the updated main path features, thereby completing the injection of the luminance statistical prior (LSP) / chrominance structure prior (CSP).

[0044] For example, combined Figure 3 With a pair of general inputs To illustrate the above prior injection process: The main road comes from First, use 1×1 convolution to perform channel projection to obtain the central features of the main path. Then, a 3×3 depthwise convolution is applied to obtain the main path local response features containing local statistics. The auxiliary road comes from... The same process is used to construct the keys, that is, first use 1×1 convolution to obtain... The local response features of the auxiliary path are obtained by performing a 3×3 depthwise convolution. Meanwhile, a branch containing a 1×1 convolution is used from... Extracted value features Local response characteristics of primary and secondary components and The elements are added one by one in spatial order and fed into a Sigmoid to form a spatial gate. The intermediate gating result is obtained by adjusting the value features of the spatial gate position by position. The modulation results are fused through a linear bottleneck header consisting of two 1×1 convolutional layers (with ReLU in the middle) and then fed back to the original number of channels to obtain the spatially gated result. From the characteristics of the main road Channel gates are extracted using the SE structure. Based on this, right Perform channel scaling and write back to the main path in residual form to obtain the updated output (injecting luminance / chrominance features from the luminance statistical prior LSP / chrominance structure prior CSP). .

[0045] To establish robust, low-overhead, and interpretable information exchange between the intensity branch I and the chromaticity branch HV, this embodiment constructs the following... Figure 4 The diagram shows the bidirectional channel cross-attention module GB-CCA. In many low-light and non-uniform lighting scenarios, there is a significant misalignment between intensity and color statistical distributions. Directly performing global fusion on both paths often spreads strong path noise or color bias to the other branch, causing artifacts and oversaturation. The design of the bidirectional channel cross-attention module GB-CCA follows the principle of "screening before cross-attention, local before global": first, a learnable gating is generated at both the spatial and channel levels by the guiding graph to suppress invalid responses and align the regions of interest; then, bidirectional cross-attention is performed in the channel dimension, with I and HV acting as the "query source" and "key / value source" respectively, explicitly modeling their conditional dependencies in the channel subspace; finally, extremely lightweight dual-branch nonlinear kernels are used to refine the two outputs, compensating for the nonlinear expression after attention with minimal parameters / computing power. The intuition behind this approach is that gating before attention confines the attention alignment problem to the "cleaned" subspace; choosing to perform attention in the channel dimension rather than the spatial dimension reduces complexity from... Reduced to At the same time, it preserves the discriminative nature of the channel subspace.

[0046] based on Figure 4 The structure shown, through cross-branch collaborative attention of the bidirectional channel cross-attention module GB-CCA, performs collaborative interaction between the luma branch I and the chroma branch HV at multiple scales, thereby stabilizing chroma and enhancing texture edges, including: The input luminance feature and its corresponding first guide map are set, as well as the chrominance feature and its corresponding second guide map; wherein, the first guide map is used to guide the collaborative interaction direction from chrominance to luminance, and the second guide map is used to guide the collaborative interaction direction from luminance to chrominance; Among them, the input luminance features and chrominance features Corresponding to the first guide image and the second guide image and set Number of heads is h, number of channels per head is .

[0047] Channel-independent weight coefficient pairs and spatial guidance maps are generated based on the first and second guidance maps, respectively. The channel-independent weight coefficient pairs are generated by global average pooling and convolution mapping of the first and second guidance maps, and then split. Preferably, they are two scalar coefficients, which are used to modulate the gating strength of the luminance branch and the chrominance branch, respectively. The spatial guidance map is generated by convolution mapping of the guidance map, and preferably forms a fine-grained gating mask per pixel and per channel through nonlinear constraints to provide element-wise guidance. Residual element-wise gating is performed on luminance and chrominance features using weighted coefficient pairs and spatial guided mapping, respectively. That is, the original features are weighted and modulated element-wise and then residually fused with the original features to obtain gated luminance features and gated chrominance features. Specifically, in the "chrominance to luminance" direction, weighted coefficient pairs generated by the first guided map and spatial guided mapping are used, while in the "luminance to chrominance" direction, weighted coefficient pairs generated by the second guided map and spatial guided mapping are used. After gating, bidirectional channel cross-attention is performed: Using the luminance branch as the query and the chrominance branch as the key and value, channel cross-attention from chrominance to luminance is calculated to obtain channel attention luminance features; using the chrominance branch as the query and the luminance branch as the key and value, preferably, the query projection includes pointwise convolution and depthwise convolution to extract local statistical information, and the key and value are obtained by splitting the projection results of the corresponding branch along the channel dimension; channel cross-attention from luminance to chrominance is calculated to obtain channel attention chrominance features; wherein, channel cross-attention includes: normalizing the gating features; generating queries, keys, and values ​​through convolutional projection; based on the generated queries, keys, and values, rearranging the queries, keys, and values ​​in a multi-head manner, calculating channel correlation based on the normalized representation of the channel vector in the spatial dimension, and introducing a learnable scaling coefficient to stabilize channel energy, performing weight normalization in the channel dimension to aggregate value features; and setting residual connections to stabilize feature responses. Lightweight refinement enhancement is performed on the channel attention luminance features and channel attention chrominance features respectively, including: normalizing and convolutionally transforming the respective channel attention luminance / chrominance features and splitting them into two sub-features along the channel dimension; performing lightweight convolution and nonlinear refinement on the two sub-features respectively to obtain two refined sub-features; and then multiplying the two refined sub-features element-wise to obtain the fused enhancement result; wherein, the nonlinear refinement preferably adopts an activation method that suppresses extreme responses, and optionally, residual fusion is performed on the two refined sub-features and their corresponding inputs respectively; furthermore, the luminance branch is used to more fully enhance texture details and suppress unstable responses, and the chrominance branch adopts the same or lighter refinement structure as the luminance branch to maintain output amplitude stability; The fusion enhancement results of the luminance branch and the chrominance branch are respectively subjected to channel back projection through 1×1 convolution, and fused with the corresponding channel attention features (preferably residual fusion) to obtain the refined output features of the luminance branch I and the chrominance branch HV, so as to stabilize the chrominance and enhance the texture edge.

[0048] The expressions for generating the channel-independent weight vector and the spatial guided mapping are as follows: In the formula, and Indicates The two scalar gates cut off This represents the function of splitting in half. This represents the Sigmoid activation function. Represents the ReLU activation function. express convolution, express convolution, Indicates global average pooling. This represents a guide mask with fine spatial granularity. The superscript X indicates the guide map type, X=A indicates a luminance guide map, and X=B indicates a chrominance guide map.

[0049] In the above process, channel attention is performed on the luminance / chrominance features, including: The gated luminance / chrominance features are pre-normalized using LayerNorm in the channel dimension; local projection is performed on the luminance branch I / chrominance branch HV, and then passed through 1×1 convolution and 3×3 depth convolution to obtain the query Q; local projection is performed on the chrominance branch HV / luminance branch I, and then passed through 1×1 convolution and 3×3 depth convolution to obtain the features, which are then split into two in the channel dimension to obtain the key K and value V; Rearrange Q, K, and V into a bullish pattern, and perform L2 normalization on Q and K in the last dimension N; The channel attention weights are calculated based on the normalized Q and K, and V is weighted and aggregated. The heads are merged and then projected by 1×1 to obtain the channel attention output. The channel attention output is added to the residual of the corresponding gated luminance / chrominance feature to obtain the channel attention luminance / chrominance feature; For example, combined Figure 4 Taking the chroma-to-luminance channel attention as an example, the above process is described as follows: Depend on Generate channel-independent weight vectors and spatially guided mappings, used as... Pre-selection of branches: In the formula, For global average pooling; and These are the ReLU and Sigmoid activation functions, respectively. through Cut into two scalar gates , It is used to modulate the intensity of the two channels separately.

[0050] in, As a spatially fine-grained guiding mask.

[0051] Based on the two formulas above, element-wise gating is performed on the two feature paths to align the region of interest and suppress noise / color cast before entering the attention path: The role of gating is to perform "coarse alignment" and "weak inhibition" first, providing a cleaner representational base for subsequent attention.

[0052] To execute For channel attention, first perform LayerNorm pre-normalization on the two gated features obtained from the above formula in the channel dimension to obtain... and Then, a local projection is performed on the I-path, followed by a 1×1 convolution and a 3×3 depthwise convolution to obtain Q. The same local projection is performed on the HV-path, and it is split in two along the channel dimension to obtain K and V. Finally, Q, K, and V are rearranged into a multi-head configuration. For Q, K is performed on the last dimension N. Normalization yields and Then, attention is calculated in the channel dimension, first using... and The correlation and multiplied by the head-by-head temperature The weights are obtained by performing a softmax operation along the channel dimension. Aggregate based on this get Then, after merging the heads (optional), the channel attention output is obtained by 1×1 projection. Features after gating Adding them together gives Completely symmetrically, channel attention is performed in the other direction using HV as the query and I as the key, resulting in... This creates a two-way relationship where each condition is another.

[0053] To compensate for the non-linear expression after cross-attention and enhance contrast / texture details, extremely lightweight dual-branch refinement kernels are connected to both paths: the intensity path (I-path) uses a residual intensity refinement unit, and the color path (HV-path) uses a color refinement unit. First, channel attention output is applied to the I-path. LayerNorm, 1×1 convolution, and 3×3 depthwise convolution are performed sequentially, and the channel dimension is halved to obtain two branches. , As shown in the following formula: Subsequently, a light enhancement was applied to each branch. For each branch, a 3×3 depthwise convolution was first applied followed by a tanh nonlinearity, and then the residuals from the input of that branch were added to obtain the enhanced result. and : in, It emphasizes mid-to-high frequency details while suppressing extreme values.

[0054] Next, element-wise multiplication of the two branches is performed to explicitly model complementarity and consistency. Then, 1×1 convolution is used for channel back-projection to fuse interactive information. Finally, the channel attention output is combined with the gated channel attention output. The external residuals are summed to obtain the I-path refined output. : The external residual of the I-path ensures that the enhancement does not excessively deviate from the attention output. The HV-path structure is the same as the above formula, but there is no external residual addition to maintain the steady-state amplitude of the color branches.

[0055] Based on the above process, the output of GB-CCA is: These two values ​​correspond to the results of intensity and color paths after a bidirectional cross-attention and lightweight refinement, respectively, and can be directly used for subsequent upsampling / downsampling and cross-layer fusion. Since attention is performed in the channel dimension, its correlation matrix size is c. h ×c h The overall complexity is approximately O(N) for spatial attention2 The noise level is significantly reduced; and the order of "gating first, then attention; local first, then global" helps to control noise propagation and color bias amplification without sacrificing expressiveness, making the interconversion between intensity and color both stable and interpretable.

[0056] To fully extract detailed information from multi-scale luminance and chrominance priors while maintaining a lightweight overall network, and to maintain a unified and stable feature enhancement unit across branches such as LSP, CSP, I, and HV, this invention designs and introduces a guided enhancement module GSE-IRB in BIP-CENet, such as... Figure 5 As shown, this module uses pointwise convolution and depthwise convolution as basic operators, and jointly models local texture and contrast through gating and channel attention. On the one hand, it serves as a general feature enhancement block within each branch, and on the other hand, it provides a more robust and prior-consistent intermediate representation for subsequent GB-CCA collaborative attention.

[0057] Given any branch Input A stats B represents the feature flow generated by the statistical prior branches of LSP. struct This represents the feature flow generated by the prior branches of the CSP structure. GSE-IRB, without changing the spatial resolution, uses the following steps: "pointwise dimensionality increase—depth convolution—SiLU—GLU position-related channel gating—pointwise re-projection—SE channel recalibration—residual with coefficients (the residual uses a scaling factor β, when C..."). in ≠C out A lightweight stacked structure (using 1×1 convolutional projection to match dimensions) selectively enhances effective textures and contrast, outputting... Take the middle channel In the implementation, e=2, and ensure that C m It is an even number.

[0058] Specifically, in combination Figure 5 The structure shown enhances steady-state details and contrast in the guiding features of the prior branch LSP / CSP through the guiding enhancement module GSE-IRB, including: The input features are first expanded by 1×1 point convolution to obtain intermediate features, and then 3×3 depthwise convolution and element-wise SiLU activation are performed to obtain local nonlinear representations. ; The local nonlinear representation is halved along the channel dimension into two sub-features; the second sub-feature is gated by Sigmoid to obtain a gating mask, and then multiplied element-wise with the first sub-feature to form a position-dependent channel-gated feature. The channel-gated features are projected onto the target number of channels via a 1×1 convolution, and then projected back onto the target number of channels via a 1×1 convolution to obtain... The recalibrated features are obtained by performing channel recalibration through the SE structure branch, and then recalibrated through the SE structure branch to obtain recalibrated features with steady-state detail and contrast enhancement. ; Recalibration features The output features are obtained by scaling and fusing the input features with the residuals; when the number of input / output channels is inconsistent, the residual branches are aligned by 1×1 convolution, and the scaling factor is a preset constant.

[0059] Among them, channel gating features are formed based on local nonlinear characteristics. The process is as follows: In the channel maintenance Divided into two branches and Use one last one The gate value is obtained by element-wise Sigmoid. Then use element-wise multiplication to... Acting on the previous one The channel gating characteristics are obtained. .

[0060] The process of channel recalibration via SE structure branches is as follows: First to Global average pooling is performed to obtain channel statistics, which are then sequentially processed through "1×1 convolution - ReLU - 1×1 convolution" to form the SE bottleneck mapping. Finally, the channel weights are obtained through Sigmoid, and then broadcast in the spatial dimension and compared with... Element-wise multiplication yields the recalibrated features. .

[0061] The above process can be represented as follows: In the formula, Indicates recalibration features, This indicates the characteristics after being recast to the target number of channels. express convolution, Indicates global average pooling. This represents the Sigmoid activation function; To maintain consistency with the input and stabilize training, a parameter fusion with coefficients is introduced to perform scaled residual fusion of the recalibrated features and the input features, which is expressed as: In the formula, Indicates the output of the residual branch. Let χ represent the input features of the i-th scale and branch. Indicates the number of input channels. Indicates the number of output channels. This represents a 1×1 convolution used for channel alignment. This represents the residual scaling factor.

[0062] When the number of channels is mismatched, the residual branches are projected using a 1×1 convolution; β is the residual scaling factor. The computation and memory usage of GSE-IRB vary with H. i W i C i Linear growth, for It is suitable for stacking at multiple scales of pyramids.

[0063] In this embodiment, after the above encoding process, the decoding stage employs symmetrical upsampling units, i.e., first performing a 3×3 convolution, then a Bilinear×2 upsampling, followed by a skip connection with the encoding end at the same resolution, concatenating them along the channel dimension, and then fusing them with PReLU via a 1×1 process. This yields the mesoscale decoding features, which are then paired and interacted with, as shown below: At full resolution, chroma and intensity are predicted separately using a linear head: Finally, residual write-back is performed in the HVI domain and back-projected to RGB to obtain the enhanced low-light image: In this embodiment of the invention, to provide sufficient constraints for the training of BIP-CENet, the network output is supervised simultaneously in both the HVI and sRGB spaces. Given a standard exposure sRGB image X, the corresponding HVI ground truth mapping is first constructed through RGB→HVI transformation, combined with dual-domain prior guidance based on YCrCb–HVI. During the forward inference process, BIP-CENet outputs an enhanced HVI feature map. The reconstructed sRGB image is obtained by inverse transformation Φ_RGB. The goal of training is to simultaneously reduce the difference between the network output and its respective reference target in both spaces.

[0064] Based on this, the training loss function of the end-to-end low-light enhancement network BIP-CENet is: In the formula, ℓ(·,·) represents the composite reconstruction loss calculated in the image domain, and λ is the weight hyperparameter. This represents the enhanced HVI feature map output by the end-to-end low-light enhancement network BIP-CENet. This represents the HVI truth mapping constructed under the guidance of the YCrCb–HVI dual-domain prior, indicating... The reconstructed image obtained by inverse transformation Φ_RGB mapping back to the sRGB domain This represents a reference, normally exposed sRGB image.

[0065] This embodiment uses a joint loss function of HVI and sRGB domains for end-to-end optimization. The HVI domain loss is used to constrain the consistency of the network output in the HVI representation, so as to introduce and strengthen the constraints of YCrCb-HVI dual-domain prior on luminance statistics and chromaticity structure. The sRGB domain loss is used to directly constrain the consistency of the reconstructed image in structure, texture and color, thereby obtaining better subjective perception and objective evaluation indicators in the visible field.

[0066] In this embodiment of the invention, an experimental example is provided to verify the effect of the above-mentioned low-light image enhancement method.

[0067] In this embodiment, BIP-CENet is included in the evaluation along with 19 other representative low-light image enhancement methods, for a total of 20 methods. These baselines cover a variety of technical approaches, from traditional Retinex models to frequency domain modeling, and then to Transformers and state-space networks. Specifically: classic algorithms based on priors and Retinex theory include NPE, LIME, and SRIE; deep methods that combine Retinex decomposition or convolutional networks for end-to-end learning in the spatial domain include KinD, KinD++, DHURE, and MIRNet; networks that explicitly model brightness and structural information in the frequency domain or wavelet / Fourier domain include FECNet, FourLLIE, UHDFour, DMFourLLIE, and CWFAS-Net; models that explicitly introduce signal-to-noise ratio information as a priori into the network include SNR-Aware and SNR-SKF; and advanced methods that further introduce Transformers, state-space modeling, or new color space design at the network structure level include Retinexformer, RetinexMamba, WaveMamba, CIDNet, and XGFu. The methods described above, together with BIP-CENet, form a comparative set that covers both traditional prior-class models without learning and various mainstream deep network architectures in the current low-light image enhancement field. To minimize the impact of implementation details on the results, all deep learning methods were reproduced based on the authors' publicly available code and trained or fine-tuned under a unified data partitioning, image preprocessing, and training strategy; traditional methods were inferred strictly according to the parameter configurations recommended in the original paper.

[0068] Quantitative and qualitative analysis on paired datasets: A systematic comparison of the quantitative results of the aforementioned 20 methods was conducted on four standard datasets: LOLv1, LOLv2-Real, LOLv2-Syn, and LSRW-Huawei. The relevant numerical statistics are summarized in Tables 1 and 2. It can be seen that BIP-CENet's overall performance on all four datasets is among the top tier: in the three real-world paired datasets (LOLv1, LOLv2-Real, and LSRW-Huawei), its PSNR and SSIM metrics reach the current best or near-best results; on the synthetic LOLv2-Syn dataset, BIP-CENet's PSNR is only slightly lower than the best method, while achieving the best performance on SSIM and LPIPS. Overall, the LPIPS error is consistently the lowest or near-lowest among all methods on each dataset, indicating strong comprehensive capabilities in terms of structure preservation and perceived quality. Furthermore, the Param and FLOPs metrics show that BIP-CENet has a significantly smaller parameter size and lower computational cost compared to large models such as MIRNet, SNR-Aware, and SNR-SKF, demonstrating a more balanced performance-efficiency trade-off compared to these complex baseline methods. Visualization of these results is shown below. Figure 6 As shown.

[0069] In terms of subjective visual quality, six baseline models with superior performance and representative structure were selected from all comparison methods: SNR-Aware, FourLLIE, SNR-SKF, Retinexformer, DMFourLLIE, and CIDNet. These models were then compared with BIP-CENet in a visual manner. Figure 7 Visual comparisons of these methods on four representative scenes in the LOLv2 real dataset are presented: the first row shows an indoor lighting and billboard scene, used to examine overexposure suppression and halo artifacts in bright areas; the second row shows a pool clock scene, focusing on the structural fidelity and noise control of small-sized bright numbers; the third row shows a large-scale stadium grandstand scene, used to evaluate detail recovery and shadow noise of repetitive structures such as grandstand seats; and the fourth row shows a complex bicycle pile scene, mainly reflecting foreground details and overall color reproduction. From the magnified green and red boxes, it can be observed that some baseline methods exhibit halos, black edges, or blurred numbers in bright areas such as overhead lights and clocks, and either insufficient brightening and blurred textures in dark areas such as grandstands and bicycle piles, or introduce obvious blocky noise. In contrast, BIP-CENet better suppresses overexposure and artifacts in all four scenes, clearly distinguishes structural details such as seats and handlebars, maintains clean shadows, and has colors and contrast closer to the ground truth (GT), resulting in more stable overall subjective quality.

[0070] In addition, such as Figure 8 As shown, further visual comparisons were performed on the LOLv2-Syn dataset to evaluate the model's robustness under synthetic degradation conditions. The figure also presents four representative synthetic scenes: the first row shows a human statue scene, used to examine the color shift and lighting levels of the statue's surface material and its shadow areas; the second row shows a large-area high-frequency texture scene, primarily examining whether fine textures are over-smoothed or produce artifacts; the third row shows a stadium panorama and stands scene, used to evaluate mid-to-far-field structural details and global exposure balance; the fourth row shows a high-contrast scene composed of city buildings and clouds (containing both strong highlights and deep shadows), focusing on examining shadow details, highlight levels, and overall stylistic consistency. As can be seen, some baseline methods produce slightly distorted or grayish-white colors on the surface of the statue in the synthetic data, and the textures of areas such as vegetation and audience seats are smoothed out or appear rather messy. The details of distant buildings are blurred, and the edges at the junction of the sky and buildings are slightly overflowing. In contrast, BIP-CENet can better restore the material and shadow levels of the statue in all four scenes, maintain the integrity of high-frequency textures such as wheat ears and audience seats, and take into account the details of dark areas and the layers of clouds in bright areas in urban scenes, presenting a more natural visual effect that is consistent with the GT style. This shows that it also has better overall visual quality than the comparison methods on LOLv2-Syn.

[0071] Table 1 Comparison of experimental metrics on the LOLv1 and LOLv2-Real datasets. Table 2 Comparison of experimental results on the LOLv2-Syn and LSRW-Huawei datasets. Quantitative and qualitative analysis on unpaired datasets: On five unpaired datasets—LIME, VV, DICM, NPE, and MEF—that only provide low-light images, the NIQE metric was used to evaluate the no-reference quality of each method (a lower value indicates that the visual statistical properties are closer to those of a natural image). As shown in Table 3, BIP-CENet achieved the best NIQE performance among all current methods on the LIME, NPE, and MEF datasets; on the VV and DICM datasets, its results consistently ranked second, only slightly inferior to the corresponding best methods. To facilitate observation of the experimental results, this study visualized the NIQE scores. Figure 9 Furthermore, considering the average NIQE across the five datasets, BIP-CENet remains the best overall, indicating that even under unreferenced, cross-scene, and unpaired testing conditions, this method can still generate augmented results that more closely approximate the statistical distribution of natural images and have a more natural subjective appearance, demonstrating good generalization ability and robustness. Due to space limitations, in... Figure 10 The paper presents typical visual comparison examples on the LIME, NPE, and MEF datasets. Taking the MEF dataset as an example, BIP-CENet not only brightens the scene and lampshade highlights but also better preserves details of the desktop and book while suppressing noise, resulting in a more natural and comfortable overall visual effect. The above subjective observations further validate the visual quality and generalization ability of BIP-CENet in unpaired scenes.

[0072] Table 3. NIQE evaluation results on the five datasets: LIME, VV, DICM, NPE, and MEF. In this embodiment of the invention, an ablation experiment example of the above-mentioned BIP-CENet is provided.

[0073] This embodiment evaluates the contribution of different components in the proposed BIP-CENet model to low-light image enhancement performance through a series of ablation experiments. The ablation experiments are divided into two parts: first, the impact of LSP and CSP is evaluated, and then the effects of each key component (such as MCPF, GB-CCA, and GSE-IRB) are evaluated. All experiments are conducted under the same experimental settings as the main experiment and standard evaluation metrics such as PSNR, SSIM, and LPIPS are used on the LOLV1 dataset.

[0074] Ablation of LSP and CSP priors: Table 4 presents the ablation results for the combined on / off operation of the luminance statistical prior LSP and chrominance structure prior CSP on the LOLv1 dataset. In implementation, disabling a prior does not change the network backbone and branch structure. Instead, it replaces the prior features originally calculated adaptively by Y or Cr / Cb with a learnable feature map independent of the input. This ensures that the corresponding prior path only provides a global bias and no longer carries sample-level luminance or chrominance statistics, thus guaranteeing that performance differences between different settings primarily reflect whether data-driven LSP / CSP priors are used, rather than the offset caused by changes in model capacity.

[0075] When both types of priors are disabled simultaneously (w / o LSP & CSP), the model can only rely on the HVI backbone and weakened prior branches for enhancement. PSNR and SSIM are the lowest in this group, while LPIPS is the highest. This indicates that relying solely on I=max(R,G,B) and its geometrically reparameterized H and V, plus content-independent channel scaling, is insufficient to achieve stable brightening and reliable noise reduction in complex low-light scenes. In the setting that only retains the luminance statistics prior (w / o CSP), PSNR and SSIM are significantly improved compared to w / o LSP & CSP, and LPIPS also decreases. This shows that the multi-view luminance statistics stack constructed by the Y channel does provide more reasonable overall exposure and local contrast constraints for the intensity branch, which can alleviate structural blur caused by underexposure. However, since the chromaticity structure is still controlled by the content-independent bias, color noise in dark areas and slight color cast still exist. The setting that retains only the chromaticity structure prior (w / o LSP) outperforms w / o CSP in both objective metrics and perceived quality: while PSNR and SSIM are further improved, LPIPS decreases more significantly. Subjectively, this manifests as a significant reduction in chromatic noise in dark areas, smoother color transitions, and more continuous chromatic boundaries at object edges. This is consistent with the design of CSP, which explicitly models chromaticity amplitude and gradient structure on the Cr / Cb plane and suppresses chromaticity noise in dark areas under luminance-sensitive gating, indicating that structured chromaticity priors are particularly crucial for improving visual quality in low-light scenes. When both LSP and CSP (BIP-CENet) are enabled, all three metrics reach their best levels in this group. Compared to the baseline of w / o LSP & CSP, PSNR and SSIM are significantly improved, and LPIPS decreases significantly. This indicates that luminance statistical priors and chromaticity structure priors have complementary and synergistic effects in intensity recovery and color noise suppression: the former stabilizes intensity estimation from the perspective of global luminance distribution and local contrast, while the latter provides reliable chromaticity amplitude and edge structure information for the HV branch under luminance gating. Overall, the data-driven LSP / CSP significantly enhances the adaptability of the prior branch to the current image content, which is one of the important factors for BIP-CENet to obtain high-quality low-light enhancement results.

[0076] Table 4 Ablation experimental results of LSP and CSP priors on the LOLv1 dataset Ablation of BIP-CENet components: Table 5 presents the ablation results of each component of BIP-CENet on the LOLv1 dataset. Under different settings, the core functions of MCPF, GB-CCA, and GSE-IRB were disabled, while the remaining network structures and training strategies remained unchanged. Units with weaker expressive power but structural compatibility were used as substitutes to ensure that the network could still feed forward normally and the channel size remained basically the same, so that the performance difference mainly reflected the gain brought by the component itself. Specifically, when removing MCPF, the original prior fusion unit with spatial gating and channel recalibration is replaced with a linear fusion form that concatenates the channels of prior features and backbone features and then performs 1×1 convolution for dimensionality reduction. When removing GB-CCA, the original luminance-chrominance bidirectional cross-attention is replaced with independent channel enhancement or local feedforward modules within each branch, and the interaction between the I branch and the HV branch is no longer explicitly modeled. When removing GSE-IRB, the original lightweight prior enhancement block containing GLU gating and SE channel recalibration is replaced with ordinary depthwise separable convolution with ReLU activation, so that the prior only undergoes a coarse convolution transformation before participating in subsequent collaboration.

[0077] Table 5. Ablation experiments on the BIP-CENet component on the LOLv1 dataset. The ablation results show that the performance of all three pruned configurations is significantly lower than that of the complete BIP-CENet, but the degree of degradation varies. The removal of GB-CCA has the most significant impact: in the w / o GB-CCA configuration, PSNR decreases significantly while LPIPS increases significantly. This indicates that once the prior-guided bidirectional luminance-chrominance cross-attention is lost, the I-branch and HV-branch can only be enhanced independently. Luminance restoration and color reconstruction lack consistent constraints, easily leading to phenomena such as normal brightness but distorted colors, or clear edges but unstable overall brightness relationships. Both objective metrics and perceptual quality degrade considerably. In the w / o MCPF configuration, the model still retains GB-CCA and GSE-IRB, but the shallow layers no longer use prior-guided QKV fusion; instead, they only perform simple concatenation and 1×1 convolutional linear mixing. At this point, both PSNR and SSIM are lower than the full model, and LPIPS also shows significant degradation, indicating that the shallow prior injection of "fusion before encoding" is crucial for overall enhancement: without spatial alignment and channel recalibration of LSP and CSP, noise and bias in the prior are more likely to be injected into the backbone, leading to fluctuations in brightness in dark areas and instability in local colors, weakening the positive guiding role of the dual-domain prior in the encoding of backbone features. In contrast, the degradation of removing GSE-IRB is relatively small: with the w / o GSE-IRB configuration, the prior branch still exists, but it is no longer gated and recalibrated. It is only fed into subsequent modules after ordinary depthwise separable convolution followed by ReLU activation. This results in increased residual graininess in dark areas, slightly coarser local texture depiction, and a decrease in perceptual quality indicators, but the overall structure recovery can still be basically maintained.

[0078] In addition, such as Figure 11 As shown, the visual comparison results of the ablation of each component of BIP-CENet are presented. The figure shows magnified results in the local colored pom-pom and color swatch areas, with pie charts of the RGB components of the corresponding areas plotted alongside. It can be seen that the color ratios of the three ablation configurations deviate from the GT to varying degrees in local areas, while the RGB ratio of the complete BIP-CENet is closest to that of the GT, indicating that the proposed module helps to maintain color balance and tonal consistency while restoring brightness.

[0079] In summary, MCPF, GB-CCA, and GSE-IRB in BIP-CENet respectively play complementary roles in shallow prior injection, luma-chroma cross-branch collaborative modeling, and lightweight prior enhancement and noise suppression. Weakening any one of these components, even with functionally equivalent but simpler alternatives, would lead to performance degradation directly corresponding to its role. However, in the complete configuration, their collaborative work enables the model to achieve optimal performance on PSNR, SSIM, and LPIPS simultaneously, further validating the effectiveness and necessity of the proposed component design and its collaborative mechanism.

[0080] Specific embodiments have been used to illustrate the principles and implementation methods of this invention. The descriptions of the embodiments above are only for the purpose of helping to understand the method and core ideas of this invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this invention. Therefore, the content of this specification should not be construed as a limitation of this invention.

[0081] Those skilled in the art will recognize that the embodiments described herein are intended to help the reader understand the principles of the invention, and should be understood that the scope of protection of the invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific modifications and combinations based on the technical teachings disclosed in this invention without departing from the spirit of the invention, and these modifications and combinations are still within the scope of protection of this invention.

Claims

1. A low-light image enhancement method based on dual-domain prior collaboration, characterized in that, Includes the following steps: For low-light images, the RGB color space is mapped to the HVI color space, and two types of prior information are constructed in the YCrCb domain: the luminance statistics prior LSP and the chromaticity structure prior CSP. Construct and train an end-to-end low-light enhancement network BIP-CENet; the end-to-end low-light enhancement network BIP-CENet is an HV-I dual-branch encoding and decoding structure, including a priori gated fusion module MCPF, a bidirectional channel cross attention module GB-CCA, and a guided enhancement module GSE-IRB; In the end-to-end low-light enhancement network BIP-CENet: Based on the conditional injection method of fusion before encoding, at the shallow scale, the luminance statistical prior LSP and chrominance structure prior CSP are injected into the luminance branch I and chrominance branch HV respectively through the prior gated fusion module MCPF. During the multi-scale encoding process, the bidirectional cross-branch collaborative interaction between the luma branch I and the chroma branch HV is realized through the bidirectional channel cross-attention module GB-CCA to stabilize chroma and enhance texture edges. At the same time, the guided enhancement module GSE-IRB is introduced into the LSP / CSP prior branch to perform lightweight enhancement and noise reduction on the prior features, generating robust guided features for use by the bidirectional channel cross-attention module GB-CCA. In the decoding stage, the luminance branch I and chrominance branch HV are reconstructed, multi-scale upsampling is performed, and the enhanced HVI is fused with the initial HVI through skip-connection fusion. The enhanced low-light image is then approximately reversibly mapped back to RGB via PHVIT.

2. The low-light image enhancement method based on dual-domain prior collaboration according to claim 1, characterized in that, The luminance statistical prior (LSP) is a multi-view luminance statistical stack constructed from the Y channel of a low-light image in the YCrCb domain, including normalized linear luminance Y. lin Logarithmic compression Y log Weber comparison, single-scale Retinex and HVI-gated isomorphism of the luminance sensitivity term Y sin_k ; The chromaticity structure prior (CSP) is constructed from the Cr / Cb channels in the YCrCb domain for low-light images, including: The Cr and Cb channels of low-light images were residualized with a neutral value of 0.5 and then standardized and normalized according to the sample standard deviation. The chromaticity magnitude and its gradient energy obtained by the Sobel operator are calculated on the normalized chromaticity plane. The gradient energies corresponding to the Cr and Cb channels are weighted and fused into the chromaticity structure energy, which is used as the chromaticity structure prior structure term. By adaptively gating and coupling the luminance sensitivity term with the chromaticity structure energy, the chromaticity radius is automatically reduced and structural noise is suppressed in the dark area, while color and edges are preserved in the medium and high luminance areas, thus constructing the chromaticity structure prior CSP.

3. The low-light image enhancement method based on dual-domain prior collaboration according to claim 1, characterized in that, The luminance statistical prior LSP is expressed as follows: The chromaticity structure prior CSP is represented as follows: In the formula, Indicates the brightness channel. Represents the local mean. This represents the low-pass approximation, where α>0, ε>0. , They represent C respectively b C r The chromaticity residuals after being residualized with a neutral value of 0.5 and normalized to the sample standard deviation are... Let k represent the Sobel operator. hvi An exponential parameter that is identical to the HVI brightness gating.

4. The low-light image enhancement method based on dual-domain prior collaboration according to claim 1, characterized in that, The prior gating fusion module MCPF injects the luminance statistical prior (LSP) and chrominance structure prior (CSP) into the luminance branch I and chrominance branch HV of the low-light image, respectively, including: The luminance statistical prior (LSP) and chrominance structure prior (CSP) are encoded and aligned to obtain the corresponding luminance guiding features and chrominance guiding features. The luminance features of the luminance branch I of the low-light image are used as the main path features, and the corresponding luminance guiding features are used as the auxiliary path features to form a luminance feature pair; the chrominance features of the chrominance branch HV of the low-light image are used as the main path features, and the corresponding chrominance guiding features are used as the auxiliary path features to form a chrominance feature pair. The luminance feature pair and the chrominance feature pair are respectively input into two prior gated fusion modules MCPF with identical structures for prior injection; wherein, one prior gated fusion module MCPF is used to inject the luminance guiding feature corresponding to the luminance statistical prior LSP into the luminance branch I, and the other prior gated fusion module MCPF is used to inject the chrominance guiding feature corresponding to the chrominance structure prior CSP into the chrominance branch HV. In the prior gating fusion module MCPF, the processing of luminance / chrominance feature pairs includes: The main road features and auxiliary road features are processed by 1×1 convolution for channel projection and 3×3 depth convolution respectively to obtain the local response features of the main road and the local response features of the auxiliary road. Extract value features from the auxiliary path features through a branch containing a 1×1 convolution; The local response features of the main road and the local response features of the auxiliary road are added element by element in spatial location, and then activated by Sigmoid to form a spatial gate; By using spatial gate and value features to perform element-wise multiplication to achieve position-wise modulation, intermediate gating results are obtained. The intermediate gating results are fused through a linear bottleneck consisting of two 1×1 convolutions, and ReLU activation is introduced between the two 1×1 convolutions before being fed back to the original number of channels to obtain the output features after spatial gating and bottleneck fusion. The channel gate is extracted from the main path features using an SE-style channel gate extraction branch. The SE-style channel gate extraction branch includes: performing global average pooling on the main path features, then sequentially passing them through two 1×1 convolution layers with ReLU activation introduced in the middle, and finally obtaining the channel gate through a Sigmoid function. The output features are channel scaled using a channel gate and added to the main features to form a residual output, thus obtaining the updated main features, thereby completing the injection of the luminance statistical prior (LSP) / chrominance structure prior (CSP).

5. The low-light image enhancement method based on dual-domain prior collaboration according to claim 1, characterized in that, Through cross-branch collaborative attention via the bidirectional channel cross-attention module GB-CCA, bidirectional collaborative interaction between the luma branch I and the chroma branch HV is achieved at multiple scales to stabilize chroma and enhance texture edges, including: The input luminance feature and its corresponding first guide map are set, as well as the chrominance feature and its corresponding second guide map; wherein, the first guide map is used to guide the collaborative interaction direction from "chrominance to luminance", and the second guide map is used to guide the collaborative interaction direction from "luminance to chrominance". Channel-independent weight coefficient pairs and spatial guidance maps are generated based on the first guidance map and the second guidance map, respectively; wherein, the channel-independent weight coefficient pairs are generated by global average pooling of the first guidance map and the second guidance map and then obtained by convolutional mapping, and then split; the spatial guidance maps are generated by convolutional mapping of the first guidance map and the second guidance map, respectively. Residual element-wise gating is performed on luminance features and chrominance features using the weighted coefficient pairs and spatial guided mapping respectively; wherein, in the "chrominance to luminance" direction, the weighted coefficient pairs generated by the first guided map and spatial guided mapping are used, and in the "luminance to chrominance" direction, the weighted coefficient pairs generated by the second guided map and spatial guided mapping are used. After gating, bidirectional channel cross-attention is performed: using the luminance branch as the query and the chrominance branch as the key and value, the channel cross-attention from chrominance to luminance is calculated to obtain the channel attention luminance feature; using the chrominance branch as the query and the luminance branch as the key and value, the channel cross-attention from luminance to chrominance is calculated to obtain the channel attention chrominance feature; wherein, the channel cross-attention includes: normalizing the gating features; generating queries, keys, and values ​​through convolutional projection; rearranging the queries, keys, and values ​​in a multi-head manner based on the generated queries, keys, and values, normalizing them in the spatial dimension, calculating channel correlation, introducing a learnable scaling factor to stabilize channel energy, performing weight normalization in the channel dimension to aggregate value features; and setting residual connections to stabilize feature responses. Lightweight refinement and enhancement are performed on the channel attention luminance features and channel attention chrominance features respectively, including: normalizing and convolutional transforming the respective channel attention luminance / chrominance features and splitting them into two sub-features along the channel dimension; performing lightweight convolution and nonlinear refinement on the two sub-features respectively to obtain two refined sub-features; and then multiplying the two refined sub-features element by element to obtain the fusion enhancement result. The fusion enhancement results of the luminance branch and the chrominance branch are respectively subjected to channel back projection through 1×1 convolution, and fused with the corresponding channel attention features to obtain the refined output features of the luminance branch I and the chrominance branch HV, so as to stabilize the chrominance and enhance the texture edge.

6. The low-light image enhancement method based on dual-domain prior collaboration according to claim 5, characterized in that, The expressions for generating channel-independent weight vectors and spatially guided mappings are as follows: In the formula, and Indicates The two scalar gates cut off This represents the function of splitting in half. This represents the Sigmoid activation function. Represents the ReLU activation function. express convolution, express convolution, Indicates global average pooling. This represents a guide mask with fine spatial granularity. The superscript X indicates the guide map type, X=A indicates a luminance guide map, and X=B indicates a chrominance guide map.

7. The low-light image enhancement method based on dual-domain prior collaboration according to claim 5, characterized in that, Perform channel attention on luminance / chrominance features, including: The gated luminance / chrominance features are pre-normalized using LayerNorm in the channel dimension; local projection is performed on the luminance branch I / chrominance branch HV, and then passed through 1×1 convolution and 3×3 depth convolution to obtain the query Q; local projection is performed on the chrominance branch HV / luminance branch I, and then passed through 1×1 convolution and 3×3 depth convolution to obtain the features, which are then split into two in the channel dimension to obtain the key K and value V; Rearrange Q, K, and V into a bullish pattern, and perform L2 normalization on Q and K in the last dimension N; The channel-dimensional attention weights are calculated based on the normalized Q and K, and V is weighted and aggregated. The heads are then merged and projected by 1×1 to obtain the channel attention output. The channel attention output is added to the corresponding gated luminance / chrominance feature by taking the residual, and the channel attention luminance / chrominance feature is obtained.

8. The low-light image enhancement method based on dual-domain prior collaboration according to claim 1, characterized in that, The guiding features of the prior branch LSP / CSP are enhanced with steady-state details and contrast through the guiding enhancement module GSE-IRB, including: The input features are first expanded by 1×1 point convolution to obtain intermediate features, and then 3×3 depth convolution and element-wise SiLU activation are performed to obtain local nonlinear representations. The local nonlinear representation is halved along the channel dimension into two sub-features; the second sub-feature is gated by Sigmoid to obtain a gating mask, and then multiplied element-wise with the first sub-feature to form a position-dependent channel-gated feature. The channel-gated features are projected onto the target number of channels through a 1×1 convolution, and recalibrated features are obtained by performing channel recalibration through an SE structure. The recalibrated features and input features are scaled and residually fused to obtain the output features; wherein, when the number of input / output channels is inconsistent, the residual branches are aligned by 1×1 convolution, and the scaling factor is a preset constant.

9. The low-light image enhancement method based on dual-domain prior collaboration according to claim 8, characterized in that, Channel recalibration is performed via the SE structure, as follows: In the formula, Indicates recalibration features, This indicates the characteristics after being fed back to the target number of channels; ⊙ indicates element-wise multiplication. Introducing parameter integration with coefficients to perform scaled residual fusion of recalibrated features and input features, expressed as: In the formula, Indicates the output of the residual branch. Let χ represent the input features of the i-th scale and branch. Indicates the number of input channels. Indicates the number of output channels. This represents a 1×1 convolution used for channel alignment. This represents the residual scaling factor.

10. The low-light image enhancement method based on dual-domain prior collaboration according to claim 1, characterized in that, The training loss function of the end-to-end low-light enhancement network BIP-CENet is: In the formula, ℓ(·,·) represents the composite reconstruction loss calculated in the image domain, and λ is the weight hyperparameter. This represents the enhanced HVI feature map output by the end-to-end low-light enhancement network BIP-CENet. This represents the HVI truth mapping constructed under the guidance of the YCrCb–HVI dual-domain prior. express The reconstructed image obtained by inverse transformation Φ_RGB mapping back to the sRGB domain This represents a reference, normally exposed sRGB image.

Citation Information

Patent Citations

  • Low-illumination image enhancement method combining multi-channel parallel attention and cross fusion

    CN121235928A

  • Low-light image enhancement method based on multi-branch feature fusion

    CN121235967A

  • System and method for fusion of images of different spectrum

    IN201941012915A