Low-illumination image dynamic enhancement method and device, medium and program product

By using an adaptive intensity parameter prediction module and a low-light image enhancement network to dynamically predict the η value and combining it with multi-scale feature fusion, the problem of poor adaptability of existing low-light image enhancement methods in diverse scenarios is solved, and adaptive and robust image enhancement effects are achieved.

CN121073804AActive Publication Date: 2025-12-05WEIFANG UNIVERSITY +3

Patent Information

Application Number
CN202511620494.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-07
Publication Date
2025-12-05
Estimated Expiration
2045-11-07

AI Technical Summary

Technical Problem

Existing low-light image enhancement methods rely on fixed parameters or manual annotation, which makes it difficult to adapt to diverse scenarios and limits automation and usability.

Method used

An adaptive enhancement intensity parameter prediction module and a low-light image enhancement network are adopted. The illumination features of the brightness channel are learned through residual blocks and channel attention mechanism to dynamically predict the enhancement intensity η value. Image enhancement is achieved by combining multi-scale feature fusion.

Benefits of technology

It improves the intelligence and robustness of image brightness restoration, adapts to different lighting conditions, avoids over-enhancement and color distortion, and achieves adaptive image enhancement effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121073804A_ABST
    Figure CN121073804A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of image data processing, and particularly discloses a low-illumination image dynamic enhancement method and device, a medium and a program product, and the method comprises the steps: obtaining a to-be-processed low-illumination image, converting the to-be-processed low-illumination image into a YCbCr color space, and extracting a brightness channel of the to-be-processed low-illumination image; the brightness channel is input into an adaptive enhancement intensity parameter prediction module, the adaptive enhancement intensity parameter prediction module learns illumination features of the brightness channel through a residual block and a channel attention mechanism, and an adaptive enhancement intensity eta value is output; and inputting the original RGB image, the brightness channel and the eta value into the low-illumination image enhancement network, realizing refined enhancement through multi-scale feature fusion, and outputting an enhanced image. According to the method, the self-adaptive enhancement intensity parameter prediction module and the low-illumination image enhancement network are fused, a new low-illumination image enhancement network is formed, automatic regression estimation and de-interactive image enhancement of the self-adaptive enhancement intensity eta value are realized, and the intelligence and robustness of image brightness recovery are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image data processing, and particularly relates to a low-illumination image dynamic enhancement method and device, a medium and a program product. BACKGROUND

[0002] In the image acquisition process, insufficient light caused by factors such as time and weather often causes problems such as insufficient brightness, low contrast and loss of details. These problems not only affect human visual perception, but also reduce the accuracy of subsequent computer vision tasks, so low-illumination images are difficult to use as input for visual tasks such as object detection and recognition.

[0003] To meet the requirements of computer vision applications, people have conducted in-depth research on low-illumination image enhancement methods. Existing low-illumination image enhancement methods mainly include traditional enhancement algorithms and low-illumination image algorithms based on deep learning. Traditional low-illumination image enhancement methods mainly rely on histogram equalization, adaptive filtering, Retinex theory, dark channel prior, etc. These methods usually improve the visual effect by adjusting the brightness and contrast of the image, but most of them rely on manually designed features and parameters, lack of adaptive ability, and are prone to problems such as over-enhancement and color distortion. With the rapid development of deep learning technology, researchers have gradually applied it to low-illumination image enhancement tasks. Deep learning models, especially convolutional neural networks (CNN), generative adversarial networks (GAN) and self-attention mechanism architectures, can automatically learn multi-level feature representations of images through large-scale data training, and have shown significant advantages in low-illumination image processing. Since the first low-illumination image enhancement method based on deep learning, LLNet, was proposed, methods such as RetinexNet and KinD have emerged and been widely used. However, most of these methods rely on fixed parameters or manually annotated enhancement targets, making it difficult to adapt to diverse scenarios, and interactive methods such as IceNet rely on user-drawn graffiti as guidance, limiting automation and usability. SUMMARY

[0004] The present application aims to provide a low-illumination image dynamic enhancement method to solve the problem of relying on fixed parameters or manual annotation in the prior art, which makes it difficult to adapt to diverse scenarios and limits automation and usability.

[0005] To achieve the above object, the application provides a low-illumination image dynamic enhancement method, which comprises the following steps: S100, acquiring a low-illumination image to be processed, converting the acquired RGB image into a YCbCr color space, and extracting a luminance channel; S200, inputting the extracted luminance channel into an adaptive enhancement intensity parameter prediction module, learning the illumination characteristics of the luminance channel by the residual block and the channel attention mechanism of the adaptive enhancement intensity parameter prediction module, and outputting an adaptive enhancement intensity η value; S300, inputting the original RGB image, the luminance channel and the η value into a low-illumination image enhancement network, realizing fine enhancement by multi-scale feature fusion of the low-illumination image enhancement network, and outputting an enhanced image. The low-illumination image dynamic enhancement method based on adaptive parameter estimation proposed in the application combines the adaptive enhancement intensity parameter prediction module and the low-illumination image enhancement network to form a new low-illumination image enhancement network, realizes automatic regression estimation of the adaptive enhancement intensity η value and de-interactive image enhancement, and effectively improves the intelligence and robustness of image brightness recovery.

[0006] The input of the adaptive enhancement intensity parameter prediction module in the step S200 is a single-channel luminance image, which is obtained by weighted fusion of the red, green and blue channel components of the RGB image in the step S100, wherein the green component is assigned the maximum weight and the blue component is assigned the minimum weight. Based on the physiological characteristics that the human visual system is more sensitive to the brightness than to the color, the use of the single-channel luminance image as the input not only effectively simplifies the input dimension and improves the robustness of the enhancement estimation, but also further matches the spectral response characteristics that the human eye is most sensitive to green light and least sensitive to blue light by the design of assigning the maximum weight to the green component and the minimum weight to the blue component in the weighted fusion strategy, so that the visual saliency area is more accurately focused on in the feature extraction stage, the color information interference is avoided, the learning efficiency of the model on the illumination distribution and the contrast feature is enhanced, and the visual naturalness and detail preservation ability of the low-illumination image enhancement are improved.

[0007] The adaptive enhancement intensity parameter prediction module comprises a feature extraction module, a multi-scale feature fusion module and an η value prediction module; the feature extraction module extracts hierarchical features from the luminance image based on the residual block and the channel attention mechanism; the multi-scale feature fusion module captures multi-granularity feature information by different convolution kernel sizes and fuses features of different scales; and the η value prediction module maps the fused features into the η value, which is used to guide the enhancement intensity of the subsequent low-illumination image enhancement network. The application only needs a single-channel luminance image as the input, reduces the input dimension and the computational complexity, improves the feature expression ability by fusing the residual connection and the channel attention mechanism, and enhances the adaptability of the model to different illumination conditions by using the multi-scale feature fusion strategy.

[0008] The feature extraction module comprises a residual block feature transformation unit, a channel attention unit, and a residual block and channel attention fusion unit. The residual block feature transformation unit performs at least two feature transformation operations on the input feature map, extracts high-level features through nonlinear activation, and obtains a transformed feature map. The channel attention unit learns the channel weight of the transformed feature map, and realizes adaptive filtering of the feature channel. The residual block and channel attention fusion unit fuses the input feature and the feature processed by the channel attention mechanism through residual connection, and outputs the fused feature through an activation function, thereby strengthening the expression of the key feature channel while retaining the original information of the input feature. The fusion of the residual block and the channel attention realizes double gain. The residual connection reduces the optimization difficulty of the deep network through "input-output residual learning", so that the model can maintain gradient stability when stacking multiple residual blocks, and suppresses the noise-dominant channel to reduce noise amplification in the enhancement process. The synergistic effect of the two makes the feature retain useful information and strengthen key details in the propagation process, providing a more robust feature basis for subsequent η value prediction.

[0009] The residual block feature transformation unit realizes feature transformation through the following steps: including two convolution operations, wherein the first convolution maps the input channel number from Cin to Cout, followed by a batch normalization layer and a ReLU activation function; the second convolution keeps the channel number Cout unchanged, followed by a batch normalization layer. The first convolution with the ReLU activation function completes the channel dimension mapping and nonlinear transformation, and improves the feature expression ability; the second convolution further extracts abstract features under the same dimension, and combines the batch normalization operation to stabilize the feature distribution and alleviate the internal covariate shift problem. The channel attention unit realizes channel weight learning through the following steps: a global feature compression step, which aggregates the spatial information of each channel into a channel descriptor through adaptive average pooling; a channel weight learning step, which takes the channel descriptor as input, constructs a channel weight learning network through at least two fully connected layers and nonlinear activation, realizes nonlinear modeling of channel importance, and outputs channel weight; and an attention weighting step, which applies the learned channel weight to the feature map through broadcast operation to enhance key channel features and suppress redundant channels. The channel weight learning step comprises: constructing a weight learning network through two fully connected layers, wherein the first fully connected layer reduces the dimension of the channel descriptor from Cout to Cout / r, and r is a dimension reduction coefficient; the second fully connected layer restores the dimension to Cout; a ReLU activation function is used between the two fully connected layers to introduce nonlinear transformation; and finally, the learned channel weight is constrained in the range of [0, 1] through a Sigmoid function.

[0010] The multi-scale feature fusion module extracts features of different scales through multiple parallel convolution branches, fuses features of different scales, and forms a feature map containing multi-scale information.

[0011] The multi-scale feature fusion module extracts features of different scales through two parallel convolution branches of local detail feature branch and global illumination feature branch; the local detail feature branch uses a 3x3 convolution kernel to extract fine-grained features, focusing on capturing local brightness gradient and detail texture; the global illumination feature branch uses a 5x5 convolution kernel to extract medium-grained features, focusing on global illumination distribution and regional brightness statistics; the output feature maps of the two branches are spliced in the channel dimension to form a fusion feature map containing multi-scale information.

[0012] The η value prediction module predicts the η value based on the fusion feature map, specifically including: performing global average pooling processing on the fusion feature map; converting the pooled features into η value prediction results through a fully connected layer; and using a Sigmoid function to constrain the η value prediction results in the range of [0, 1].

[0013] The low-illumination image enhancement network realizes image enhancement through the following steps: step S310, based on the original RGB image, its brightness channel information and the adaptive enhancement intensity parameter η value, extracting the brightness feature and the color feature of the image; step S320, using the η value to adaptively regulate the enhancement intensity of the brightness feature and the color feature; step S330, fusing the modulated brightness feature and the color feature, introducing a residual connection to combine the fused feature with the original RGB feature to output the enhanced image feature; step S340, mapping the enhanced image feature to a standard pixel value range to obtain the final enhanced image. Under the dynamic regulation of the η value, the low-illumination image enhancement network balances the brightness improvement, contrast optimization and detail preservation, and can meet the enhancement goal of natural brightness, clear details and coordinated colors.

[0014] The step S320 includes: a feature scaling modulation operation, which adjusts the activation intensity of the brightness feature through the η value, and enhances the expression of the illumination feature when the η value increases; and a weight modulation operation, which adjusts the weight parameter of the RGB feature extraction process through the η value, and reduces the extraction intensity of the color feature when the η value decreases. Through the feature scaling modulation operation, the overall brightness of the image can be accurately controlled, avoiding over-enhancement leading to overexposure; through the weight modulation operation, color noise amplification can be suppressed under low-illumination conditions, keeping the color natural; the two operations realize balanced regulation and adaptive scene adaptation of brightness and color.

[0015] The feature scaling modulation operation is implemented by the formula , where γ is a scaling coefficient controlling the influence amplitude of η on the feature.

[0016] The weight modulation operation is implemented by the formula , where W0 is the initial weight of the convolution kernel and β is a reference coefficient.

[0017] The step S330 comprises: fusing the brightness feature modulated by the η value and the RGB feature through channel splicing to integrate illumination and color information; combining the fused enhanced feature and the original image feature through a residual connection mechanism to retain the original image information.

[0018] The residual connection is realized by: reducing the fused feature to 3 channels through 3x3 convolution, and projecting the original RGB image feature to 3 channels through 1x1 convolution; element-wise adding the reduced fused feature and the projected original image feature to generate the final enhanced feature.

[0019] The step S340 comprises: normalizing the enhanced feature to the standard pixel value range of [0, 1] through a Sigmoid activation function, and outputting an enhanced image meeting the display requirements.

[0020] The steps S200 and S300 are realized based on a neural network model optimized by a multi-objective joint loss function, which comprises: an η value prediction loss function for optimizing η value accuracy; an enhanced image loss function for optimizing image enhancement effect; and a collaborative constraint loss function for ensuring the consistency of η value prediction and enhancement effect. Through multi-objective optimization, the model can simultaneously learn how to predict a reasonable η value and how to generate a high-quality enhanced image based on the η value, realizing the closed-loop collaboration of low-illumination image "perception-enhancement".

[0021] The η value prediction loss function adopts a mean square error loss function to optimize the prediction accuracy of the adaptive enhancement intensity parameter prediction module by calculating the mean square error between the predicted η value and the true η value. MSE is selected because it has more significant punishment for large errors, so that the model reduces extreme errors in η value prediction.

[0022] The enhanced image loss function comprises: a pixel-level L1 loss function that measures pixel-level differences by calculating the average absolute error of pixel values between the enhanced image and the reference image; a structural similarity loss function that measures the similarity of brightness, contrast and structural features between the enhanced image and the reference image through a structural similarity index; and a color consistency loss function that constrains the hue and saturation components of the enhanced image and the original image in the HSV color space to prevent color drift during the enhancement process. By designing a multi-dimensional loss function, pixel-level differences, structural similarity, perceptual quality and color consistency are comprehensively optimized.

[0023] The synergistic constraint loss function ensures the consistency of the prediction of the value of the enhancement intensity parameter and the image enhancement effect by comparing the difference between the predicted value of the enhancement intensity parameter and the quantitative indicator of the actual enhancement effect; and the specific calculation manner of the synergistic constraint loss function is as follows: the mean value of the absolute error between the predicted value of the enhancement intensity parameter of each sample in a batch and the enhancement intensity quantitative indicator is calculated; wherein the enhancement intensity quantitative indicator is defined as the average brightness difference value of the enhanced image and the original image in the brightness channel, and the mean value of the brightness difference of the corresponding pixel points is obtained by converting the enhanced image and the original image to the brightness space.

[0024] The total loss function of the multi-target joint loss function is the weighted sum of the loss functions.

[0025] The application further provides a computer device comprising a memory, a processor and a computer program stored in the memory, wherein the processor executes the computer program to realize the steps of the above method.

[0026] The application further provides a computer readable storage medium, wherein the computer program is executed by the processor to realize the steps of the above method.

[0027] The application further provides a computer program product, wherein the computer program is executed by the processor to realize the steps of the above method.

[0028] To sum up, the application has the beneficial effects that: the application proposes a low-illumination image dynamic enhancement method based on adaptive parameter estimation to solve the problem that the traditional image enhancement method excessively relies on artificial parameters and has poor scene adaptability. The method dynamically predicts the adaptive enhancement intensity value through the adaptive enhancement intensity parameter prediction module, and realizes fine image enhancement in combination with the low-illumination image enhancement network, and excellent performance is shown in low-light, complex light and other scenes. Specifically, the method proposes an adaptive value prediction mechanism based on the brightness channel, only uses the image brightness channel, learns the mapping relationship between the illumination feature and the enhancement intensity through the synergistic effect of the residual block and the channel attention, realizes the dynamic prediction of the value, and gets rid of the dependence on artificial parameter adjustment; the channel attention mechanism is embedded in the residual block, so that the model can automatically enhance the weight of the key feature channel and suppress the noise channel, and combines multi-scale feature fusion to consider local details and global light distribution, and realizes uniform enhancement of the image. BRIEF DESCRIPTION OF DRAWINGS

[0029] Figure 1 is the E-IceNet network model structure diagram in the embodiment of the application; Figure 2 is the residual block and channel attention fusion process diagram of the feature extraction module in the embodiment of the application; Figure 3This is a flowchart of the multi-scale feature fusion module and η value prediction in an embodiment of the present invention; Figure 4 This is a diagram of the improved IceNet network structure in an embodiment of the present invention; Figure 5 This is a schematic diagram of some of the results obtained by running the algorithm of this invention on dataset L; Figure 6 This is a comparison chart of the augmentation results of different algorithms on the LIME dataset; Figure 7 This is a comparison chart of the augmentation results of different algorithms on the Fusion dataset; Figure 8 This is a comparison chart of the augmentation results of different algorithms on the LOWL dataset; Figure 9 This is a radar chart comparing experimental data under four evaluation criteria: SSIM, StdDev, PIQE, and NIQE. Figure 10 This is a flowchart illustrating the low-light image dynamic enhancement method provided by the present invention. Detailed Implementation

[0030] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0031] In the process of implementing the embodiments of this application, the inventors of this application discovered that in the prior art, the low-light image dynamic enhancement method based on deep learning model relies on fixed parameters or manually labeled enhancement targets, which is difficult to adapt to diverse scenarios.

[0032] To address the aforementioned issues, this application provides a dynamic low-light image enhancement method based on adaptive parameter estimation, referred to as the low-light image dynamic enhancement method. Figure 10 The flowchart illustrates the low-light image dynamic enhancement method provided in the embodiments of this application, as shown below. Figure 10As shown, the method comprises: step S100, acquiring a low-illumination image to be processed, converting the acquired RGB image to a YCbCr color space, and extracting a luminance channel thereof; step S200, inputting the extracted luminance channel into an adaptive enhancement intensity parameter prediction module, the adaptive enhancement intensity parameter prediction module learning illumination features of the luminance channel through a residual block and a channel attention mechanism, and outputting an adaptive enhancement intensity η value; and step S300, inputting the original RGB image, the luminance channel, and the η value into a low-illumination image enhancement network, the low-illumination image enhancement network realizing fine enhancement through multi-scale feature fusion, and outputting an enhanced image.

[0033] The input of the adaptive enhancement intensity parameter prediction module in step S200 is a single-channel luminance image, which is obtained by weighted fusion of red, green, and blue channel components of the RGB image in step S100, wherein the green component is assigned the maximum weight and the blue component is assigned the minimum weight.

[0034] The adaptive enhancement intensity parameter prediction module comprises a feature extraction module, a multi-scale feature fusion module, and an η value prediction module; the feature extraction module extracts hierarchical features from the luminance image based on a residual block and a channel attention mechanism; the multi-scale feature fusion module captures multi-granularity feature information through different convolution kernel sizes and fuses features of different scales; and the η value prediction module maps the fused features to an η value, which is used to guide the enhancement intensity of the subsequent low-illumination image enhancement network.

[0035] The feature extraction module comprises a residual block feature transformation unit, a channel attention unit, and a residual block and channel attention fusion unit; the residual block feature transformation unit performs at least two feature transformation operations on the input feature map, extracts high-level features through a nonlinear activation, and obtains a transformed feature map; the channel attention unit learns channel weights of the transformed feature map, realizing adaptive selection of feature channels; and the residual block and channel attention fusion unit fuses the input feature and the feature processed through the channel attention mechanism by introducing a residual connection, outputs the fused feature through an activation function, and preserves the original information of the input feature while strengthening the expression of key feature channels.

[0036] The residual block feature transformation unit realizes feature transformation by the following steps: including two convolution operations, wherein the first convolution maps the input channel number from Cin to Cout, followed by a batch normalization layer and a ReLU activation function; the second convolution keeps the channel number Cout unchanged, followed by a batch normalization layer. The channel attention unit realizes channel weight learning by the following steps: a global feature compression step, which aggregates the spatial information of each channel into a channel descriptor through adaptive average pooling; a channel weight learning step, which takes the channel descriptor as input, constructs a channel weight learning network through at least two fully connected layers and a nonlinear activation, realizes nonlinear modeling of channel importance, and outputs channel weights; an attention weighting step, which applies the learned channel weights to the feature map through a broadcast operation, enhances key channel features, and suppresses redundant channels. The channel weight learning step includes: constructing a weight learning network through two fully connected layers, wherein the first fully connected layer reduces the dimension of the channel descriptor from Cout to Cout / r, r being a dimension reduction coefficient; the second fully connected layer restores the dimension to Cout; a ReLU activation function is used between the two fully connected layers to introduce a nonlinear transformation; finally, the learned channel weights are constrained in the range of [0, 1] through a Sigmoid function.

[0037] The multi-scale feature fusion module extracts features of different scales through multiple parallel convolution branches, fuses features of different scales, and forms a feature map containing multi-scale information.

[0038] The multi-scale feature fusion module extracts features of different scales through two parallel convolution branches of local detail feature branch and global illumination feature branch; the local detail feature branch adopts a 3x3 convolution kernel to extract fine-grained features, focusing on capturing local brightness gradients and detailed textures; the global illumination feature branch adopts a 5x5 convolution kernel to extract medium-grained features, focusing on global illumination distribution and regional brightness statistics; the output feature maps of the two branches are spliced in the channel dimension to form a fusion feature map containing multi-scale information.

[0039] The η value prediction module predicts the η value based on the fusion feature map, specifically including: performing global average pooling processing on the fusion feature map; converting the pooled features into η value prediction results through a fully connected layer; and using a Sigmoid function to constrain the η value prediction results in the range of [0, 1].

[0040] The low-illumination image enhancement network realizes image enhancement through the following steps: step S310, based on the original RGB image, the luminance channel information thereof, and the adaptive enhancement intensity parameter η value, extracting the luminance feature and the color feature of the image; step S320, using the η value to adaptively regulate the enhancement intensity of the luminance feature and the color feature; step S330, fusing the modulated luminance feature and the color feature, introducing a residual connection to combine the fused feature with the original RGB feature to output the enhanced image feature; and step S340, mapping the enhanced image feature to a standard pixel value range to obtain a final enhanced image.

[0041] Step S320 includes: a feature scaling modulation operation, which adjusts the activation intensity of the luminance feature through the η value, and enhances the expression of the illumination feature when the η value increases; and a weight modulation operation, which adjusts the weight parameter of the RGB feature extraction process through the η value, and reduces the extraction intensity of the color feature when the η value decreases. The feature scaling modulation operation is implemented by the formula wherein γ is a scaling coefficient, controlling the influence amplitude of η on the feature. The weight modulation operation is implemented by the formula wherein W0 is the initial weight of the convolution kernel, and β is a reference coefficient.

[0042] Step S330 includes: fusing the luminance feature modulated by the η value and the RGB feature through channel splicing to integrate the illumination and color information; and combining the fused enhanced feature with the original image feature through a residual connection mechanism to retain the original image information. The residual connection is realized by the following ways: reducing the fused feature to 3 channels through 3x3 convolution, and projecting the original RGB image feature to 3 channels through 1x1 convolution; and element-wise adding the reduced fused feature and the projected original image feature to generate the final enhanced feature.

[0043] Step S340 includes: normalizing the enhanced feature to the standard pixel value range of [0, 1] through a Sigmoid activation function, and outputting an enhanced image meeting the display requirements.

[0044] The steps S200 and S300 are implemented based on a neural network model optimized by a multi-objective joint loss function, and the multi-objective joint loss function comprises: an η value prediction loss function for optimizing η value accuracy; an enhanced image loss function for optimizing image enhancement effect; and a collaborative constraint loss function for ensuring consistency of η value prediction and enhancement effect. The η value prediction loss function adopts a mean square error loss function, and the prediction accuracy of the adaptive enhancement intensity parameter prediction module is optimized by calculating the square error mean value between the predicted η value and the real η value. The enhanced image loss function comprises: a pixel-level L1 loss function, which measures the pixel-level difference between the enhanced image and the reference image by calculating the average absolute error of the pixel values; a structural similarity loss function, which measures the similarity between the enhanced image and the reference image in terms of brightness, contrast and structural features by a structural similarity index; and a color consistency loss function, which performs consistency constraint on the hue and saturation components of the enhanced image and the original image in the HSV color space, so as to prevent color deviation in the enhancement process. The collaborative constraint loss function ensures the consistency of η value prediction and image enhancement effect by comparing the difference between the predicted enhancement intensity parameter η value and the quantitative index of the actual enhancement effect; and the specific calculation method of the collaborative constraint loss function is: calculating the absolute error mean value between the predicted η value of each sample in a batch and the enhancement intensity quantitative index; wherein the enhancement intensity quantitative index is defined as the average brightness difference value between the enhanced image and the original image in the brightness channel, and the average brightness difference value of the corresponding pixel points is obtained by converting the enhanced image and the original image into the brightness space.

[0045] The total loss function of the multi-objective joint loss function is the weighted sum of each loss function.

[0046] The following is a specific embodiment of the above low-illumination image dynamic enhancement method: 1. Overall framework The low-illumination image dynamic enhancement method based on adaptive parameter estimation proposed in the application fuses an adaptive enhancement intensity parameter prediction module (also known as an automatic enhancement factor estimator, English abbreviation EtaEstimator) and a low-illumination image enhancement network (in this embodiment, an improved non-interactive image enhancement network, English abbreviation IceNet), forms a new low-illumination image enhancement network (English abbreviation E-IceNet), realizes automatic regression estimation of adaptive enhancement intensity η value (abbreviation enhancement factor), and non-interactive image enhancement, and effectively improves the intelligence and robustness of image brightness recovery. The E-IceNet network model diagram is shown in Figure 1 .

[0047] Referring to Figure 1, the E-IceNet network model for low-light image enhancement processing is divided into three stages. In the first stage, the low-light image is input first, the input image is preprocessed, the RGB image is converted to YCbCr color space, and the luminance channel (also known as Y channel) is extracted as the input of EtaEstimator. The Y channel component is highly matched with the sensitivity of the human visual system to brightness, and based on the physiological characteristics that the human visual system is more sensitive to brightness than color, the input dimension can be effectively simplified and the robustness of the enhancement estimation can be improved, and the formula is: (1).

[0048] The second stage is an adaptive enhancement intensity parameter prediction module specially designed for the image enhancement task, which in this embodiment is a lightweight estimation model: EtaEstimator. The design integrates residual connection and channel attention mechanism, learns the illumination features of the Y channel through residual blocks and channel attention mechanism, and outputs the η value to guide the subsequent image enhancement process.

[0049] The third stage is image enhancement. The low-light image enhancement network is based on an improved network model framework: IceNet, which takes the original RGB image, the luminance channel (Y channel) and the η value as input, realizes fine enhancement through multi-scale feature fusion, and outputs the enhanced image.

[0050] 2. EtaEstimator: EtaEstimator is an adaptive parameter model specially designed for image enhancement tasks. By analyzing the luminance channel (Y channel) information of the input image, it learns the key features of the image such as illumination characteristics and contrast, and automatically predicts the η value. This value can be used to adjust the enhancement intensity, making the enhancement result more consistent with the characteristics of the image itself, effectively solving the problem of relying on manual setting of enhancement parameters in traditional image enhancement methods and difficulty in adapting to diversified scenes.

[0051] Compared with existing methods, the EtaEstimator module has the following characteristics: 1) Only single-channel luminance image (Y channel) is required as input, which greatly reduces the input dimension and computational complexity; 2) Fusion of residual connection and channel attention mechanism effectively improves the feature expression ability; 3) Multi-scale feature fusion strategy is adopted to enhance the adaptability of the model to different illumination conditions; 4) The output η value can be directly used to guide the enhancement intensity input of the subsequent low-light image enhancement network IceNet.

[0052] EtaEstimator is mainly composed of three modules: feature extraction module, multi-scale feature fusion module and η value prediction module. The feature extraction module extracts hierarchical features from the brightness map based on the residual block and channel attention mechanism; the multi-scale feature fusion module captures multi-granularity feature information through different convolution kernel sizes; the η value prediction module maps the fused features to the enhancement intensity parameter η in the range of [0, 1], and the η value is used to guide the enhancement intensity of the subsequent low-light image enhancement network IceNet.

[0053] 2.1 Feature extraction module

[0054] To improve the feature expression ability and gradient propagation stability, the channel attention mechanism is embedded in the residual block in the embodiment of the application. This design not only alleviates the gradient vanishing problem of deep network through residual connection, but also dynamically enhances the contribution of key feature channels and suppresses noise or redundant information. The process is as shown in Figure 2 Figure 2 The residual block and channel attention fusion process diagram of the feature extraction module.

[0055] (1) Residual block feature transformation unit: composed of two 3x3 convolutions and batch normalization, nonlinear transformation is introduced to extract high-level features. For input feature Figure X ∈R B×Cin×H×W (B is the batch size, C in is the input channel number, H, W is the spatial size), the feature preliminary transformation process is: (2) (3) Among them, Conv1 is a 3x3 convolution, which maps the input channel number from C in to C out ; ReLU is an activation function, which introduces nonlinear expression ability and suppresses the interference of negative feature values on subsequent transformation; Conv2 is the second 3x3 convolution, which keeps the channel number C out unchanged and further extracts abstract features.

[0056] BN1, BN2 are batch normalization operations, which accelerate the training convergence by standardizing the feature distribution, and the formula is: (4) Among them, uz, is the mean and variance of the feature Z, γ, β are learnable parameters, is a small value to prevent the denominator from being 0.

[0057] (2) Channel attention unit: X2∈R B×Cout×H×W ​The input channel attention module realizes feature screening by learning channel weights. The calculation of channel attention includes three steps: global feature compression, channel weight learning, and attention weighting.

[0058] The first step is global feature compression. The spatial information of each channel is aggregated into a scalar through adaptive average pooling to capture the global response of the channel. The formula is: (5) where Zc is the global average response of the cth channel of the bth sample, X 2,b,c,i,j is the feature value of X2 at position (b, c, i, j), and the output channel descriptor Z = [Z1, Z2,..., Z Cout ] ∈ R B×Cout .

[0059] The second step is channel weight learning. A two-layer fully connected layer (FC) and a nonlinear activation are used to build a channel weight learning subnetwork to realize nonlinear modeling of channel importance. The formula is: (6) where FC1(R Cout → R Cout / r , r is the dimension reduction coefficient) reduces the parameter quantity and enhances the generalization by dimension reduction, FC2(R Cout / r → R Cout ) restores the dimension to the input channel number, and σ is the Sigmoid activation function, which ensures that the weight s = [s1, s2,..., s Cout ] ∈ [0, 1] B×Cout , i.e., the weight of each channel is within the range of 0~1.

[0060] The third step is attention weighting. The learned channel weight s is applied to the feature Figure X 2 through broadcast operation to enhance key channel features and suppress redundant channels. The formula is: (7) where ⊙ is element-wise multiplication, is broadcast operation (extending the weight B × C out to the feature map size B × C out × H × W), and 1 H×W is an H × W all-1 matrix.

[0061] (3) Residual block and channel attention fusion unit: To ensure the dimension matching of the input and output features of the residual block, a residual connection adapter is introduced. Assuming that the input of the residual block is X and the output is H(X), the formula is: (8) The feature H(X) fused with attention and residual connection is output through a ReLU activation function, which not only strengthens the expression of key feature channels, but also retains the original information of the input features through the residual term, avoiding feature degradation in the deep network.

[0062] The fusion of the residual block and the channel attention realizes double gain. The residual connection reduces the optimization difficulty of the deep network through "input-output residual learning", so that the model can still maintain gradient stability when stacked with 3 residual blocks, while suppressing noise-dominant channels and reducing noise amplification in the enhancement process. The synergistic effect of the two makes the features retain useful information and strengthen key details during propagation, providing a more robust feature basis for subsequent η value prediction.

[0063] 2.2 Multi-scale feature fusion module

[0064] In the image enhancement task, a single scale of features cannot meet the needs of local details and global light distribution. Therefore, the present application designs a multi-scale feature fusion module to extract multi-granularity features through different size convolution kernels and splice them to provide more comprehensive feature support for η value prediction. The multi-scale feature fusion module and η value prediction process are as shown in Figure 3 . Figure 3 The multi-scale feature fusion module and η value prediction process diagram.

[0065] The multi-scale feature fusion module takes the output (F ∈ R B×64×H′×W′ ) of the feature extraction module as input, extracts features of different scales through two parallel convolution branches, which are the fine-grained feature branch and the medium-grained feature branch, respectively, and uses 3 × 3 convolution kernel and 5 × 5 convolution kernel respectively.

[0066] The fine-grained feature branch focuses on capturing local brightness gradient and detail texture, and its formula is: (9) Where Conv 3×3 is a 3 × 3 convolution layer that ensures the output feature map size is consistent with the input (H' × W'), F1 ∈ R B ×32×H′×W′ .

[0067] The medium-grained feature branch focuses on global light distribution and regional brightness statistics, and its formula is: (10) Where Conv 5×5 is a 5 × 5 convolution layer that also maintains the same feature map size, and F2 ∈ R B×32×H′×W′ .

[0068] The fine-grained feature F1 and the medium-grained feature F2 are fused by channel dimension splicing to form a feature map containing multi-scale information, and its formula is: (11) where concat(·) is a channel dimension concatenation operation, and F fused ∈R B×32×H′×W′,保 retains the local detail features of F1 and integrates the global illumination features of F2. F1 can capture the subtle textures in dark areas, ensuring that the details are fully enhanced, and F2 can capture the global illumination distribution with a larger receptive field, avoiding overexposure of highlights. Through the combination of the two, the final output η value considers both local and global.

[0069] 2.3 η value prediction module The η value prediction module takes the fused features F fused as input and outputs the η value prediction result. The η value prediction result is constrained in the range [0, 1] by using the Sigmoid function, providing adaptive enhancement guidance for IceNet.

[0070] 3. IceNet The low-light image enhancement network IceNet is the core of image enhancement, guided by the η value predicted by EtaEstimator, combining the brightness features and RGB information of the input image to achieve fine and adaptive image enhancement without human intervention. The design goal is to balance brightness enhancement, contrast optimization, and detail preservation under the dynamic regulation of η value, avoiding noise amplification or color distortion caused by over-enhancement. The improved IceNet network structure is shown in Figure 4 .

[0071] The input information of IceNet is the brightness channel (Y channel), the original RGB image I rgb , and the enhancement parameter η value, and the output is the enhanced RGB image I enhanced , which meets the enhancement goal of natural brightness, clear details, and coordinated colors. The four-stage structure of feature collaborative extraction-η value dynamic modulation-residual enhancement-output mapping is adopted, and the key is to convert η value into quantitative regulation of enhancement intensity, realized through two ways of "feature scaling" and "weight modulation", ensuring the linear correlation between enhancement strategy and η value.

[0072] Feature scaling modulation, η value directly acts on the brightness feature F Y , and the enhancement amplitude of the illumination feature is controlled by scaling, and its formula is: (12) where γ is the scaling coefficient, controlling the influence amplitude of η on the feature.

[0073] In the convolutional layer of RGB feature extraction, the value of η affects the extraction intensity of color features by adjusting the weight, and its formula is: (13) Where W0 is the initial weight of the convolution kernel, and β is the reference coefficient. When η is large, W η is closer to W0, enhancing the extraction of color details; when η is small, W η is reduced, reducing the over-activation of color features.

[0074] The brightness feature modulated by η is fused with the RGB feature F RGB through channel splicing, integrating illumination and color information, and its formula is: (14) The fused feature contains both η-regulated illumination information and original color features, laying the foundation for enhanced brightness and color consistency.

[0075] To avoid the loss of original image information during the enhancement process, a residual connection is introduced to combine the fused feature with the original RGB feature, and its formula is: (15) Where Conv fusion is a 3x3 convolution (reducing 128 channels to 3 channels), and Conv rgb_proj is a 1x1 convolution (projecting the original RGB feature to 3 channels) to ensure the dimensionality of the residual connection.

[0076] Finally, the enhanced feature is mapped to the range [0, 1] through the Sigmoid activation function to obtain the enhanced image, and its formula is: (16) Where σ is the Sigmoid function, ensuring that the output pixel value is in the range [0, 1].

[0077] 4. Loss function 4.1 EtaEstimator η value prediction loss To ensure the accuracy of the η value predicted by EtaEstimator and the quality of the IceNet enhancement effect, the invention designs a multi-objective joint loss function, including η value prediction loss, enhanced image loss, and collaborative constraint loss. Through multi-objective optimization, the model can simultaneously learn how to predict reasonable η values and how to generate high-quality enhanced images based on η values, achieving a closed-loop collaboration of low-light image "perception-enhancement".

[0078] The core objective of EtaEstimator is to learn the mapping relationship between brightness features and the optimal η value. Therefore, it uses mean squared error (MSE) loss to measure the difference between the predicted and true values. MSE is chosen because it penalizes larger errors more significantly, reducing extreme errors in the model's prediction of the η value. Let η... pred ∈R B×1 The value of η predicted by the model, η gt ∈R B×1 Let η be the optimal value for manual annotation, where B is the batch size. Its loss function formula is: (17).

[0079] 4.2 IceNet Enhanced Image Loss IceNet aims to enhance low-light images, generating enhanced images that are pixel-accurate, structurally complete, and have natural colors. Therefore, it designs a multi-dimensional loss function to comprehensively optimize pixel-level differences, structural similarity, perceptual quality, and color consistency.

[0080] To reduce the difference between the enhanced image and the reference image (I gt To measure the pixel value deviation, L1 loss is used to avoid overfitting. The formula is as follows: (18) Among them, I enhanced ∈R B×3×H×W The image is the enhanced image output by IceNet, where C=3, H, and W are the image dimensions. L1 loss ensures that the enhanced image closely approximates the reference image at the pixel level, avoiding significant deviations in overall brightness or contrast.

[0081] Pixel-level loss cannot fully reflect the structural integrity of an image; therefore, structural similarity (SSIM) loss is introduced to penalize the structural deviation between the enhanced image and the reference image. The formula is as follows: (19) Where SSIM(·) is the structural similarity index, with a value range of [0, 1]. SSIM By optimizing the consistency of brightness, contrast, and structure, we ensure that the structural features of the enhanced image, such as edges and textures, are not lost.

[0082] To prevent color shift during enhancement, loss is designed in the HSV color space, constraining the enhanced image to match the original image. rgb The formula for color consistency is: (20) where HSV(·) is the RGB to HSV space conversion, H is hue, S is saturation, and V is value. This loss ensures that the enhancement optimizes only the value (V channel) while hue H and saturation S remain consistent with the original image, avoiding the problem of "enhancement is color cast".

[0083] 4.3 Synergistic constraint loss To ensure the synergy of EtaEstimator and IceNet, i.e., the consistency of η value prediction and enhancement effect, an enhancement sensitivity loss is designed to penalize the mismatch between η value and enhancement intensity, whose formula is: (21) where EnhanceScore(·) is the enhancement intensity quantification indicator, defined as the average brightness difference between the enhanced image and the original image, whose formula is: (22) Y enhanced and Y rgb are the brightness channel values of the enhanced image and the original image, respectively, normalized to [0, 1]. L sync This forces the model to learn the logic that the higher the η value, the greater the enhancement intensity, avoiding the disconnection between η value and actual enhancement effect.

[0084] 4.4 Total loss function The total loss of the model is the weighted sum of the above losses, and the importance of each objective is balanced through hyperparameters, whose formula is: (23) In the experiment, the optimal weights are determined through grid search, λ1=1.0, η value prediction is the core premise; λ2=0.8, λ3=0.2, keep the balance between pixel-level and structure loss; λ4=0.1, auxiliary to improve subjective quality; λ5=0.1, weakly constrain color, allow necessary saturation adjustment; λ6=0.5, ensure the consistency of η value and enhancement intensity.

[0085] The application also provides a computer device, comprising a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to realize the steps of the above method.

[0086] The application also provides a computer readable storage medium, wherein the computer program is executed by the processor to realize the steps of the above method.

[0087] The application also provides a computer program product, wherein the computer program is executed by the processor to realize the steps of the above method.

[0088] Those skilled in the art can clearly understand from the description of the above embodiments that the embodiments can be implemented by means of software and necessary universal hardware platforms, and of course can also be implemented by hardware. Based on such understanding, the above technical solutions can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the method of each embodiment or some parts of the embodiment.

[0089] In order to verify the performance of the method provided by the application, a series of experiments are designed, and detailed analysis is made according to the experimental results.

[0090] 5. Experiment and analysis of the low-illumination image dynamic enhancement method provided by the application 5.1 Experimental environment All algorithms used in the experiment of the application are completed on a Windows 11 operating system, equipped with a 12th Gen Intel(R) Core(TM) i7-12700F 2.10GHz processor, with a running memory of 32.0GB. The experimental environment runs a 64-bit operating system based on an x64 processor. The software platform is PyCharm 2023.2, and the Python version is Python 3.11.4.

[0091] 5.2 Subjective evaluation In order to verify the effectiveness of the method provided by the application, a plurality of advanced image enhancement algorithms are run on low-illumination images in a plurality of data sets. The data set LOWL used by the application is collected from four public low-illumination image data sets, LIME, DICM, LOL, and FUSION, and contains 1000 low-illumination images. The EnlightenGAN, RetinexNet, RRDNet, Zero-DCE algorithms and the method proposed by the application are run on the self-built data set LOWL and the public data sets LIME and FUSION for experimental comparison.

[0092] First, the algorithm results are analyzed subjectively from the aspect of visual perception. The following are some results obtained by running the algorithm of the application on the data set L, as shown in Figure 5 . Figure 5 There are 4 rows and 6 columns, a total of 24 images, the first and second rows are a group of 6 input images and their detail images, and the third and fourth rows are enhanced images and their detail images after the input images of the group are processed by the algorithm of the application.

[0093] The subjective evaluation is mainly analyzed by human eyes, and through intuitive analysis of the experimental results, it is found that the image brightness is obviously improved after the low-illumination image is enhanced by the method, and the brightness distribution is uniform; and the details and colors of the image are better preserved.

[0094] In order to verify the effectiveness of the present application, the present application is compared with a plurality of advanced low-illumination image enhancement algorithms on the LIME data set, the Fusion data set and the self-built LOWL data set. Some experimental results are shown in Figure 6 , Figure 7 , Figure 8 . Figure 6 is a comparison chart of enhancement results of different algorithms on the LIME data set, Figure 6 , which has 6 rows and 6 columns, a total of 36 images; the images are divided into three groups according to every two rows as a group, and the image in the lower row in each group is the detail image of the image in the upper row; according to the column order, the first column is the input image, and the second to sixth columns are the enhancement results of the EnlightenGAN algorithm, the RRDNet algorithm, the RetinexNet algorithm, the Zero-DCE algorithm and the present application algorithm (E-IceNet) respectively. Figure 7 is a comparison chart of enhancement results of different algorithms on the Fusion data set, Figure 7 , which has 6 rows and 6 columns, a total of 36 images; the images are divided into three groups according to every two rows as a group, and the image in the lower row in each group is the detail image of the image in the upper row; according to the column order, the first column is the input image, and the second to sixth columns are the enhancement results of the EnlightenGAN algorithm, the RRDNet algorithm, the RetinexNet algorithm, the Zero-DCE algorithm and the present application algorithm (E-IceNet) respectively. Figure 8 is a comparison chart of enhancement results of different algorithms on the LOWL data set, Figure 8 , which has 4 rows and 6 columns, a total of 24 images; the images are divided into two groups according to every two rows as a group, and the image in the lower row in each group is the detail image of the image in the upper row; according to the column order, the first column is the input image, and the second to sixth columns are the enhancement results of the EnlightenGAN algorithm, the RRDNet algorithm, the RetinexNet algorithm, the Zero-DCE algorithm and the present application algorithm (E-IceNet) respectively.

[0095] 5.3 Objective evaluation The data test of four evaluation standards of SSIM, StdDev, PIQE and NIQE of five methods of EnlightenGAN algorithm, RRDNet algorithm, RetinexNet algorithm, Zero-DCE algorithm and the algorithm (E-IceNet) of the application on LIME data set, Fusion data set and self-built data set LOWL is shown in Tables 1, 2 and 3 respectively, which is the intuitive analysis of the enhancement effect of the algorithm. The radar chart of the evaluation result data is shown in Figure 9 . Figure 9 is the experimental data comparison radar chart under four evaluation standards of SSIM, StdDev, PIQE and NIQE. Each small graph in the figure corresponds to one standard, and the four small graphs correspond to four standards.

[0096] Table 1: Enhancement results of different algorithms on LIME data set

[0097] Table 2: Enhancement results of different algorithms on Fusion data set

[0098] Table 3: Enhancement results of different algorithms on LOWL data set

[0099] The entropy value analysis shows that the SSIM of the algorithm in the Fusion data set is 0.64, only less than 0.69 of the RRDNet algorithm; the SSIM in the LIME data set and the self-built data set LOWL is 0.66 and 0.63 respectively, higher than other algorithms, which shows that the enhanced image is closer to the reference image, and the enhancement effect is good. The standard deviation analysis shows that the standard deviation value of the algorithm in the LIME data set is 46.71, only less than 51.46 of the EnlightenGAN algorithm; the standard deviation in the Fusion data set is 81.15, only less than 81.55 of the RRDNet algorithm; the standard deviation value in the LOWL data set is the largest, which is 49.46, higher than other algorithms, which shows that the contrast of the enhanced image is strong. The PIQE analysis result shows that the PIQE value of the algorithm in the LIME data set and the Fusion data set is lower than that of other algorithms, which is 20.93 and 29.31 respectively; in the self-built LOWL data set, it is 35.98, only less than the RRDNet algorithm, which shows that the enhanced image is better in contrast, brightness, sharpness, color saturation and other aspects, and is more in line with human subjective perception. The NIQE analysis further confirms that the NIQE value of the algorithm in the LIME data set and the Fusion data set is lower than that of other algorithms, which is 2.45 and 3.31 respectively; in the self-built LOWL data set, it is 2.62, the same as the RRDNet algorithm and lower than other algorithms, which shows that the enhanced image is closer to the natural image, and the quality is excellent. Through the above index analysis, the algorithm has obvious advantages in low-light image enhancement, and has excellent performance in noise suppression, information preservation, definition and contrast improvement. The excellent performance of multiple indicators in different data sets fully proves the universality and robustness of the algorithm in multiple scenarios.

[0100] 5.4 Ablation experiment In order to verify the effectiveness of the channel attention, residual connection, multi-scale feature fusion and adaptive η value prediction proposed in the application, the ablation experiment is designed on the self-built LOWL data set. The experiment constructs model variants by removing core components one by one, compares the differences in objective indicators and η value prediction accuracy, and determines the role of each component.

[0101] The ablation experiment is divided into four steps, which are removing the channel attention module in the residual block, only keeping the residual connection and the basic convolution; removing the residual connection, changing the ResidualBlock to the series structure of "conv1→bn1→ReLU→conv2→bn2"; removing the multi-scale feature fusion module, only keeping the fine-grained features extracted by 3x3 convolution; without using EtaEstimator to predict η value, directly fixing the value η=0.5 to input the enhancement network.

[0102] To ensure the fairness of the experiment, all variant models use the same training data, optimizer, learning rate, training rounds and the same low-light pictures as the baseline model. The pictures are selected from the self-built LOWL dataset for the experiment, and the evaluation indicators use entropy, StdDev, PIQE and NIQE. The experimental results are shown in Table 4.

[0103] Table 4 Ablation experiment results of the self-built LOWL dataset

[0104] Through the ablation experiment data analysis, by using the EtaEstimator to predict the η value, the enhanced image quality is the best. Removing the channel attention module in the residual block, the entropy decreases by 0.1, the StdDev decreases by 3.1, the PIQE decreases by 1.89, and the NIQE decreases by 0.003. The reason is that the lack of channel attention causes the model to be unable to distinguish between “dark texture channels” and “noise channels”, and the noise is amplified simultaneously during enhancement, proving that the channel attention module can effectively enhance the discrimination of the key channel. Removing the residual connection, the performance decreases significantly. Through “input-output residual learning”, the perception of the slight changes in brightness is strengthened, and the feature extraction capability is greatly weakened after the removal. Removing the multi-scale feature fusion module will lose the global illumination information, lack of global illumination guidance, resulting in the overall brightness being too low after enhancement. The fixed value η = 0.5 has the worst low-light image enhancement effect, thereby verifying the core role of adaptive η value prediction.

[0105] In summary, the present application proposes a low-light image dynamic enhancement method based on adaptive parameter estimation to solve the problem of excessive dependence on artificial parameters and poor scene adaptability of traditional image enhancement methods. The model dynamically predicts the adaptive enhancement intensity η value through the EtaEstimator, and realizes fine image enhancement in combination with the IceNet, and performs excellently in low-light, complex lighting and other scenes. The present application proposes an adaptive η value prediction mechanism based on the brightness channel, which only takes the image brightness channel (Y channel) as the input, learns the mapping relationship between the illumination feature and the enhancement intensity through the cooperation of the residual block and the channel attention, and realizes the dynamic prediction of the η value, which breaks the dependence on artificial parameter adjustment. The channel attention mechanism is embedded in the residual block, so that the model can automatically enhance the weight of the key feature channel and suppress the noise channel at the same time. The multi-scale feature fusion takes into account the local details and the global illumination, and realizes uniform image enhancement. Through running on the self-built dataset LOWL and the public dataset LIME and FUSION and comparing with multiple mainstream image enhancement algorithms, it is found that the experimental data of the present application is better than other algorithms, verifying the superiority and practical value of the algorithm of the present application. The present application provides a new paradigm for the “adaptive parameter regulation” of image enhancement, and can be popularized to the fields of video enhancement, remote sensing image optimization and the like.

[0106] To ensure the objectivity and verifiability of the present application, the contrast pictures, effect pictures and data charts shown in the embodiment part of the present application are selected from the public test database widely used in the industry (including LIME data set and Fusion data set). Such database is usually published by research institutions or industry alliances, aiming to promote technology exchange and fair comparison, and explicitly allows the relevant content to be quoted in academic papers, technical reports and patent documents. The LOWL data set collected by our company does not involve any copyright dispute.

[0107] Finally, it should be pointed out that: the above embodiments are only used to illustrate the technical solutions of the present application, but not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present application.

Claims

1. A low-light image dynamic enhancement method, characterized in that, The method comprises the following steps: Step S100, acquiring a low-illumination image to be processed, converting the acquired RGB image into a YCbCr color space, and extracting a luminance channel thereof; Step S200, inputting the extracted luminance channel into an adaptive enhancement intensity parameter prediction module, learning the illumination characteristics of the luminance channel by the residual block and the channel attention mechanism of the adaptive enhancement intensity parameter prediction module, and outputting an adaptive enhancement intensity η value; Step S300, inputting the original RGB image, the luminance channel and the η value into a low-illumination image enhancement network, realizing fine enhancement by multi-scale feature fusion of the low-illumination image enhancement network, and outputting an enhanced image.

2. The low-light image dynamic enhancement method of claim 1, wherein, The input of the adaptive enhancement intensity parameter prediction module in the step S200 is a single-channel luminance image, which is obtained by weighted fusion of red, green and blue channel components of the RGB image in the step S100, wherein the green component is assigned the maximum weight and the blue component is assigned the minimum weight.

3. The low-light image dynamic enhancement method of claim 1, wherein, The adaptive enhancement intensity parameter prediction module comprises a feature extraction module, a multi-scale feature fusion module and an η value prediction module; The feature extraction module extracts hierarchical features from the luminance image based on the residual block and the channel attention mechanism; The multi-scale feature fusion module captures multi-granularity feature information by different convolution kernel sizes and fuses features of different scales; The η value prediction module maps the fused features into the η value, which is used to guide the enhancement intensity of the subsequent low-illumination image enhancement network.

4. The low-light image dynamic enhancement method of claim 3, wherein, The feature extraction module comprises a residual block feature transformation unit, a channel attention unit and a residual block and channel attention fusion unit. The residual block feature transformation unit performs at least two feature transformation operations on the input feature map, extracts high-level features through nonlinear activation, and obtains a transformed feature map; The channel attention unit learns the channel weight of the transformed feature map, and realizes adaptive selection of the feature channel; The residual block and channel attention fusion unit fuses the input feature and the feature processed by the channel attention mechanism through the introduction of residual connection, and outputs the fused feature through the activation function, which strengthens the expression of the key feature channel while retaining the original information of the input feature.

5. The low-light image dynamic enhancement method of claim 4, wherein, The residual block feature transformation unit realizes feature transformation by the following steps: including two convolution operations, wherein the first convolution maps the input channel number from Cin to Cout, followed by a batch normalization layer and a ReLU activation function; the second convolution keeps the channel number Cout unchanged, followed by a batch normalization layer.

6. The low-light image dynamic enhancement method of claim 4, wherein, The channel attention unit realizes channel weight learning by the following steps: a global feature compression step, which aggregates the spatial information of each channel into a channel descriptor through adaptive average pooling; a channel weight learning step, which constructs a channel weight learning network by at least two fully connected layers and nonlinear activation with the channel descriptor as input, realizes nonlinear modeling of the importance of the channel, and outputs the channel weight; An attention weighting step, which applies the learned channel weight to the feature map through a broadcast operation, enhances the key channel feature, and suppresses the redundant channel.

7. The low-light image dynamic enhancement method of claim 6, wherein, The channel weight learning step comprises: The weight learning network is constructed by two full connection layers, wherein the first full connection layer reduces the dimension of the channel descriptor from Cout to Cout / r, r being a dimension reduction coefficient; the second full connection layer restores the dimension to Cout; a ReLU activation function is used between the two full connection layers to introduce a nonlinear transformation; and finally, the learned channel weight is constrained in the range of [0, 1] through a Sigmoid function.

8. The low-light image dynamic enhancement method of claim 3, wherein, The multi-scale feature fusion module extracts features of different scales through multiple parallel convolution branches, fuses the features of different scales, and forms a feature map containing multi-scale information.

9. The low-light image dynamic enhancement method of claim 8, wherein, The multi-scale feature fusion module extracts features of different scales through two parallel convolution branches of local detail feature branch and global illumination feature branch; the local detail feature branch adopts a 3*3 convolution kernel to extract fine-grained features and focuses on capturing local brightness gradient and detail texture; the global illumination feature branch adopts a 5*5 convolution kernel to extract medium-grained features and focuses on global illumination distribution and regional brightness statistics; the output feature maps of the two branches are spliced in the channel dimension to form a fusion feature map containing multi-scale information.

10. The low-light image dynamic enhancement method of claim 3, wherein, The η value prediction module predicts the η value based on the fusion feature map output by the multi-scale feature fusion module, specifically including: performing global average pooling processing on the fusion feature map; converting the pooled features into an η value prediction result through a full connection layer; constraining the η value prediction result in the range of [0, 1] using a Sigmoid function.

11. The low-light image dynamic enhancement method of claim 1, wherein, The low-illumination image enhancement network realizes image enhancement through the following steps: Step S310, based on the original RGB image, its luminance channel information and the adaptive enhancement intensity parameter η value, extracting the luminance features and color features of the image; Step S320, using the η value to adaptively regulate the enhancement intensity of the luminance features and color features; Step S330, fusing the modulated luminance features and color features, introducing a residual connection to combine the fused features with the original RGB features to output enhanced image features; Step S340, mapping the enhanced image features to a standard pixel value range to obtain the final enhanced image.

12. The low-light image dynamic enhancement method of claim 11, wherein, The step S320 includes: a feature scaling modulation operation, which adjusts the activation intensity of the luminance features through the η value, and enhances the expression of the illumination features when the η value increases; a weight modulation operation, which adjusts the weight parameter of the RGB feature extraction process through the η value, and reduces the extraction intensity of the color features when the η value decreases.

13. The low-light image dynamic enhancement method of claim 12, wherein, The feature scaling modulation operation employs the formula is implemented, where γ is a scaling coefficient that controls the magnitude of the effect of η on the features.

14. The low-light image dynamic enhancement method of claim 12, wherein, The weight modulation operation employs the formula is implemented, where W0is the initial weight of the convolution kernel, and β is a reference coefficient.

15. The low-light image dynamic enhancement method of claim 11, wherein, The step S330 includes: fusing the luminance features modulated by the η value and the RGB features through channel splicing to integrate illumination and color information; combining the fused enhanced features with the original image features through a residual connection mechanism to retain the original image information.

16. The low-light image dynamic enhancement method of claim 15, wherein, The residual connection is realized by the following way: reducing the fused features to 3 channels through a 3*3 convolution, and projecting the original RGB image features to 3 channels through a 1*1 convolution; element-wise adding the reduced fused features and the projected original image features to generate the final enhanced features.

17. The low-light image dynamic enhancement method of claim 11, wherein, The step S340 includes: The enhanced features are normalized to the standard pixel value range of [0, 1] by a Sigmoid activation function, and an enhanced image meeting the display requirements is output.

18. The low-light image dynamic enhancement method of claim 1, wherein, The steps S200 and S300 are implemented based on a neural network model optimized by a multi-objective joint loss function, and the multi-objective joint loss function includes: an η value prediction loss function for optimizing η value accuracy; an enhanced image loss function for optimizing image enhancement effect; a collaborative constraint loss function for ensuring consistency between η value prediction and enhancement effect.

19. The low-light image dynamic enhancement method of claim 18, wherein, The η value prediction loss function adopts a mean square error loss function to optimize the prediction accuracy of the adaptive enhancement intensity parameter prediction module by calculating the mean square error between the predicted η value and the true η value.

20. The low-light image dynamic enhancement method of claim 18, wherein, The enhanced image loss function includes: a pixel-level L1 loss function for measuring pixel-level differences by calculating the mean absolute error of pixel values between the enhanced image and the reference image; a structural similarity loss function for measuring the similarity between the enhanced image and the reference image in terms of brightness, contrast and structural features through a structural similarity index; a color consistency loss function for consistency constraint on the hue and saturation components of the enhanced image and the original image in the HSV color space to prevent color deviation in the enhancement process.

21. The low-light image dynamic enhancement method of claim 18, wherein, The collaborative constraint loss function ensures the consistency between η value prediction and image enhancement effect by comparing the difference between the predicted enhancement intensity parameter η value and the quantitative indicator of the actual enhancement effect. The specific calculation method of the collaborative constraint loss function is to calculate the mean absolute error between the predicted η value of each sample in a batch and the enhancement intensity quantitative indicator. The enhancement intensity quantitative indicator is defined as the average brightness difference between the enhanced image and the original image in the brightness channel, which is obtained by converting the enhanced image and the original image to the brightness space and then calculating the mean brightness difference of the corresponding pixel points.

22. The low-light image dynamic enhancement method of claim 18, wherein, The total loss function of the multi-objective joint loss function is the weighted sum of each loss function.

23. A computer apparatus comprising a memory, a processor, and a computer program stored on the memory, wherein the computer program, when executed by the processor, causes the processor to perform the method of any one of claims 1-22. The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 22.

24. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method according to any one of claims 1 to 22.

25. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method according to any one of claims 1 to 22.

Citation Information

Patent Citations

  • Enhancement method and system for low-illumination image

    CN112927162A

  • Low-illumination image enhancement method based on multi-scale stacked attention network

    CN114972107A

  • Mine image enhancement method

    CN118314056A

  • Low-light image decoupling enhancement method based on Retinex model

    CN119090788A

  • Low-illumination image enhancement method based on adaptive brightness enhancement and high-fidelity color correction

    CN119477733A

Cited By

  • Low-illumination image dynamic reconstruction method and device

    CN121329842A

  • Low-illumination image dynamic reconstruction method and device

    CN121329842B