A dark light image enhancement model and its training method
Through a dark-light image enhancement model combining local fine-grained and global coarse-grained features, using Retinex decomposition theory and multi-level feature aggregation technology, the problem of difficulty in coordinating brightness, color and exposure balance in the prior art is solved, and a natural and high-quality image enhancement effect is achieved.
Patent Information
- Application Number
- CN202510330503.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-03-20
AI Technical Summary
Existing dark-light image enhancement technologies are difficult to maintain balance of brightness, color reproduction and exposure levels, resulting in unnatural image enhancement effects and limited generalization capabilities.
A dark light image enhancement model is proposed. Through the feature extraction module, local fine-grained and global coarse-grained features are integrated, and the refinement module uses Retinex decomposition theory to adjust brightness and contrast, and performs multi-level feature aggregation and residual mapping through the brightness adjustment module to generate natural and high-quality enhanced images.
It achieves the improvement of the overall quality of the image while carefully protecting and enhancing the image details information, ensuring excellent performance in the enhanced image color fidelity, detail clarity and global exposure balance, and enhancing the model's adaptability to different lighting conditions.
Smart Images

Figure CN119831917B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image enhancement, and in particular to a dark-light image enhancement model and a training method thereof. Background Art
[0002] As a key preprocessing step, dark light image enhancement is of great significance for improving visual experience and facilitating image analysis tasks. It not only improves the user's visual experience, but also provides support for hardware and software design in systems such as autonomous driving and safety monitoring. Although early methods such as histogram equalization, Retinex theory-based and defogging model-based technologies have improved image brightness and contrast to a certain extent, they often ignore the preservation of detail information and global exposure balance, resulting in color distortion or over-enhancement. These problems limit the effectiveness of these methods in practical applications.
[0003] In recent years, data-driven models, especially methods based on convolutional neural networks (CNNs), have made significant progress in the field of low-light image enhancement. Such methods improve image quality by learning prior knowledge on large-scale datasets. However, most existing techniques tend to extract deep features from the entire image, while ignoring local details and changes in brightness range, which affects the effective use of features. Although some methods attempt to improve the results through illumination balance, they fail to fully consider the specific transformation of image details and the accuracy of brightness adjustment, making the final image enhancement effect appear unnatural, and these methods have limited generalization capabilities and are difficult to adapt to the needs of low-light image enhancement under different illumination conditions.
[0004] In summary, although current technologies have improved the quality of low-light images in certain aspects, they are still insufficient in coordinating brightness, color reproduction, and exposure levels. Existing methods lack an effective synergistic mechanism and it is difficult to maintain a balance between these factors during low-light enhancement. Summary of the invention
[0005] In order to more comprehensively address the challenges in low-light image enhancement and ensure that the image details are carefully protected and enhanced while improving the overall image quality, so as to achieve a more natural and realistic enhancement effect, the present invention proposes a low-light image enhancement model, including:
[0006] A feature extraction module is used to integrate the local fine-grained features and the global coarse-grained features of the input original dark-light image to generate a feature map, reallocate weights to the channels in the feature map to obtain a shallow feature map, extract deep features in the original dark-light image based on the feature map and the shallow feature map, restore the deep features to the resolution of the original image to obtain target deep features, and generate a supervised feature map based on the target deep features and the feature map;
[0007] A joint refinement module is used to estimate the brightness information of each color channel in the supervised feature map, and adjust the brightness and contrast of each color channel using the Retinex decomposition theory to obtain a corrected supervised feature map, i.e., an intermediate feature map. Each feature in the intermediate feature map is scaled and shifted using a modulation parameter to obtain a space-variant feature map. The intermediate feature map and the space-variant feature map are integrated to obtain a joint feature.
[0008] The brightness adjustment module is used to aggregate the joint features and the original dark light image to obtain an aggregated image, predict a normal light image from the aggregated image, and predict a low-light image through the normal light image, obtain the residual difference between the low-light image and the original dark light image, and obtain a first residual map; based on the first residual map, estimate the residual difference between the predicted normal light image and the original normal light image corresponding to the original dark light image to obtain a second residual map, and fuse the second residual map with the joint features to obtain an enhanced image.
[0009] Furthermore, the feature extraction module specifically includes:
[0010] The residual unit is used to integrate the local fine-grained features and the global coarse-grained features of the input original dark-light image to generate a feature map;
[0011] Residual channel attention unit, used to redistribute weights of channels in feature maps to obtain shallow feature maps;
[0012] An encoder, used to extract deep features in the original dark light image through the feature map output by the residual unit and the shallow feature map output by the residual channel attention unit;
[0013] A decoder is used to restore the deep-level features to the resolution of the original image to obtain the target deep-level features;
[0014] The self-supervisory block is used to generate a supervised feature map through the target deep-level features and the feature map output by the residual unit.
[0015] Furthermore, the joint refinement module specifically includes:
[0016] A feature correction unit is used to continuously perform multiple convolution and ReLU activation operations on each color channel in the supervised feature map to estimate the brightness information of each color channel, and adjust the brightness and contrast of each color channel based on the brightness information using the Retinex decomposition theory, merge the adjusted color channels to obtain the target channel, and reconstruct the intermediate feature map through the target channel;
[0017] The spatial feature transformation unit is used to scale and shift each feature in the intermediate feature map using the modulation parameter to achieve spatial feature transformation, thereby obtaining a space-varying feature map; the formula for obtaining the space-varying feature map is: ;
[0018] In the formula, represents the modulation parameters, where represents the scale parameter obtained through prior learning, represents the shift parameter obtained by prior learning, represents the intermediate feature map, represents point-wise multiplication, represents the feature conversion function;
[0019] The joint refinement unit is used to integrate the intermediate feature map and the space-variant feature map to obtain the joint feature; the integration formula is: ;
[0020] In the formula, represents the supervised feature map, represents the corrected supervised feature map, represents the spatial feature transformation, represents color normalization of the spatial-variant feature map, Represents a joint feature.
[0021] Furthermore, the brightness adjustment module specifically includes:
[0022] Aggregation block, used to aggregate the joint features and the original dark light image to obtain an aggregated image;
[0023] A first light block is used to predict a normal light image from the aggregated image;
[0024] A dark block is used to predict a low-light image through the predicted normal-light image;
[0025] A first residual block is used to obtain a residual difference between the predicted low-light image and the original low-light image to obtain a first residual map;
[0026] The second light block;
[0027] A second residual block, configured to estimate a residual difference between a predicted normal light image and an original normal light image corresponding to the original dark light image by using the first residual map through the second lit block, to obtain a second residual map;
[0028] The fusion block is used to fuse the second residual map with the joint feature to obtain an enhanced image.
[0029] Furthermore, the estimation formula of the second residual mapping is:
[0030] ;
[0031] In the formula, ∈R, is the preset weight, R represents the real number set, represents an aggregate image, Indicates the first light block, Indicates dark blocks, Indicates the second light block; : represents the intermediate residual mapping; represents the second residual mapping; Represents the preset weight coefficient, represents the first residual map.
[0032] Further, the lighting block includes:
[0033] Coded convolution is used to reduce the number of feature channels of the image to be processed and obtain the key feature image;
[0034] Offset convolution is used to learn the difference between the original dark light image and the original normal light image, obtain the brightness offset of each pixel between the two, remove the negative brightness offset, and increase the brightness offset corresponding to each pixel in the key feature image to adjust the brightness of the pixel, thereby obtaining the adjusted key feature image;
[0035] Decoding convolution is used to increase the number of feature channels of the adjusted key feature image to obtain the target feature; the light-up block is the first light-up block or the second light-up block;
[0036] When the light-on block is the first light-on block, the image to be processed is an aggregated image, and the target feature is a predicted normal light image;
[0037] When the lit block is a second lit block, the image to be processed is an intermediate residual map, and the target feature is a second residual map.
[0038] The present invention also proposes a training method for a dark light image enhancement model, comprising the steps of:
[0039] S1: Acquire a data set; the data set includes a plurality of training samples; each of the training samples includes an original dark light image and its corresponding original normal light image;
[0040] S2: Set a loss function, and train the dark light image enhancement model as described above using the dataset and the set loss function.
[0041] Furthermore, the loss function includes:
[0042] A first loss function is used to evaluate the difference between the predicted low-light image and the original low-light image, and between the predicted normal-light image and the original normal-light image;
[0043] The second loss function is used to enhance the model’s ability to capture edge details of each image;
[0044] The total loss function is constructed by combining the first loss function and the second loss function.
[0045] Furthermore, the formula expression of the first loss function is:
[0046] ;
[0047] The formula expression of the second loss function is:
[0048] ;
[0049] In the formula, represents the Charbonnier loss, represents the enhanced image output by the dark light image enhancement model, represents the original normal light image in the training sample, is a constant; represents the edge loss, represents the Laplace operator.
[0050] Furthermore, the total loss function is expressed as:
[0051] ;
[0052] In the formula, For control and The weight parameters, Represents the total loss value.
[0053] The beneficial effects of the embodiments of the present invention include:
[0054] (1) The present invention integrates the local fine-grained features and the global coarse-grained features of the input image and redistributes the weights of the channels in the feature map, thereby ensuring that the model can effectively capture and maintain the key details in the image. This design not only helps to improve the overall quality of the image, but also ensures that the detail information is not lost during the enhancement process. In addition, the process of generating a supervised feature map based on the target deep-level features and the original feature map further ensures the accurate transmission of the detail information. At the same time, this also provides guidance for the subsequent processing stage and helps maintain the global exposure balance.
[0055] (2) In the present invention, the encoder is responsible for capturing deep information, while the decoder restores this information to the original resolution, preserving a large number of local details in the process. The existence of jump connections ensures that low-level features can be directly transmitted to deep layers, enhancing the expressiveness of local details. At the same time, the residual unit extracts features at different levels of abstraction, with special emphasis on the importance of local details. The residual channel attention unit allows the model to reallocate weights according to the importance of each channel, thereby more effectively processing changes in local details and brightness ranges, ultimately improving the overall quality and visual effects of the image.
[0056] (3) In order to more accurately consider the specific transformation of image details and brightness adjustment, the joint refinement module plays a key role in the present invention. This module uses the Retinex decomposition theory to estimate the brightness information of each color channel, and adjusts the brightness and contrast accordingly to obtain the corrected supervised feature map. This method pays special attention to the specific transformation of image details and ensures the accuracy of brightness adjustment. The spatial feature transformation unit realizes spatial feature transformation by modulating parameters to scale and shift each feature in the intermediate feature map. This step further optimizes the brightness adjustment process, making the enhanced image more natural and in line with actual scene requirements.
[0057] (4) The brightness adjustment module of the present invention generates an aggregated image by aggregating the joint features and the original dark light image, and predicts the normal light image therefrom, and then further predicts the low light image based on the normal light image, and calculates the residual difference between the low light image and the original dark light image to obtain a first residual map. Then, the residual difference between the predicted normal light image and the original normal light image is estimated based on the first residual map to obtain a second residual map, and finally the second residual map and the joint features are fused to generate an enhanced image. The beneficial effect of this process is that it can not only effectively improve the overall brightness and contrast of the image, but also carefully correct the local details and brightness range, ensuring that the enhanced image has excellent performance in color fidelity, detail clarity and global exposure balance. Through this multi-level feature aggregation and residual mapping mechanism, the model can significantly reduce the problems of color distortion and over-enhancement while maintaining natural visual effects, thereby providing more realistic and high-quality image output. In addition, the model's adaptability to different lighting conditions is enhanced, and its generalization performance is improved.
[0058] The details of one or more embodiments of the invention are set forth in the following drawings and description so that other features, objects, and advantages of the invention are more readily apparent. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art descriptions. Obviously, the drawings described below are only some embodiments of the present invention, and for ordinary technicians in this field, other embodiments can be obtained based on these drawings without creative work.
[0060] Figure 1 is a structural diagram of a dark light image enhancement model according to an embodiment of the present invention;
[0061] Figure 2 is a structural diagram of a feature extraction module according to an embodiment of the present invention;
[0062] Figure 3 is a structural diagram of a joint refinement module according to an embodiment of the present invention;
[0063] Figure 4 is a structural diagram of a brightness adjustment module according to an embodiment of the present invention;
[0064] Figure 5 It is a flow chart of a training method for a dark light image enhancement model according to an embodiment of the present invention. DETAILED DESCRIPTION
[0065] Embodiments of the present embodiment will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present embodiment are shown in the accompanying drawings, it should be understood that the present embodiment can be implemented in various forms and should not be construed as being limited to the embodiments set forth herein, which are instead provided for a more thorough and complete understanding of the present embodiment. It should be understood that the drawings and embodiments of the present embodiment are only for exemplary purposes and are not intended to limit the scope of protection of the present embodiment.
[0066] In order to more comprehensively address the challenges of low-light image enhancement, we need to ensure that while improving the overall image quality, we can also carefully protect and enhance the image details to achieve a more natural and realistic enhancement effect. Figure 1 As shown, the embodiment of the present invention proposes a dark light image enhancement model, including:
[0067] A feature extraction module is used to integrate the local fine-grained features and the global coarse-grained features of the input original dark-light image to generate a feature map, reallocate weights to the channels in the feature map to obtain a shallow feature map, extract deep features in the original dark-light image based on the feature map and the shallow feature map, restore the deep features to the resolution of the original image to obtain target deep features, and generate a supervised feature map based on the target deep features and the feature map;
[0068] like Figure 2 As shown, the feature extraction module specifically includes:
[0069] The residual unit is used to integrate the local fine-grained features and the global coarse-grained features of the input original dark-light image to generate a feature map;
[0070] In this embodiment, the residual unit includes two residual blocks (residual block 1 and residual block 2). The integration formula of the local fine-grained features and the global coarse-grained features of the original dark light image is:
[0071] ;
[0072] In the formula, represents the original dark-light image, represents residual block 1, represents residual block 2, Represents the feature map; both residual block 1 and residual block 2 include 3×3 convolution and PReLU activation function.
[0073] Residual channel attention unit, used to redistribute weights of channels in feature maps to obtain shallow feature maps;
[0074] After being processed by the residual unit, the receptive field of the neuron is significantly expanded, so that more global information can be captured. This design not only enhances the model's ability to understand the overall structure of the image, but also enables the effective combination of local fine-grained features with global coarse-grained features, thereby improving the enhancement effect of low-light images. Specifically, local fine-grained features help to retain the detail information in the image, while global coarse-grained features provide a wider range of contextual information. The combination of the two ensures that the enhanced image has both rich details and good overall consistency. In addition, the residual channel attention unit further optimizes the information extraction efficiency of each channel by redistributing weights to the channels in the feature map, ensuring that the importance of different feature channels is reasonably evaluated and utilized, and ultimately improving the overall expressiveness and robustness of the model. In this way, the model can not only better restore the brightness and contrast of images under low-light conditions, but also effectively avoid color distortion and over-enhancement.
[0075] In this embodiment, the process of reallocating weights of channels in the feature map includes three parts: squeezing (Fsq), excitation (Fex) and attention (Fscale). The corresponding formulas are as follows:
[0076] ;
[0077] ;
[0078] ;
[0079] In the formula, Represents the output after the squeeze operation, that is, the global average pooling result of each channel; , represents the feature map; represents the squeezing operation function, which is used to compress the information of the spatial dimension into a single value; Representation feature map Middle The location of the channel The eigenvalue at ; and Represent the height and width of the feature map respectively;
[0080] Represents the first The weight of each channel; represents the excitation operation function, which is used to generate the importance score of each channel; Represents a set of weight parameters, which includes and ,in:
[0081] , ; r represents the dimensionality reduction ratio, which is used to reduce the computational complexity; Indicates the total number of channels; represents the set of real numbers; Represents the Sigmoid function, which is used to map the output value to between (0,1) as the weighting coefficient of the channel; ReLU represents the ReLU activation function, which is used to introduce nonlinear elements;
[0082] represents the shallow feature map obtained after the final scaling operation, and each channel is reweighted according to its importance; Represents the scaling function used to adjust each channel of the feature map.
[0083] An encoder, used to extract deep features in the original dark light image through the feature map output by the residual unit and the shallow feature map output by the residual channel attention unit;
[0084] A decoder is used to restore the deep-level features to the resolution of the original image to obtain the target deep-level features;
[0085] The self-supervisory block is used to generate a supervised feature map through the target deep-level features and the feature map output by the residual unit. The main function of the self-supervisory block is to extract information useful for the task in a supervised manner and then pass it to the joint refinement module.
[0086] In the present invention, the encoder is responsible for capturing deep information, while the decoder restores this information to the original resolution, and a large number of local details are retained in the process. The existence of jump connections (the output of the residual unit is transmitted to the encoder via jump connections) ensures that low-level features can be directly transmitted to the deep layer, enhancing the expressiveness of local details. At the same time, the design of multiple residual blocks in the residual unit realizes the extraction of features at different levels of abstraction, with special emphasis on the importance of local details. The residual channel attention unit allows the model to reallocate weights according to the importance of each channel, thereby more effectively processing changes in local details and brightness ranges, ultimately improving the overall quality and visual effects of the image.
[0087] In addition, the present invention ensures that the model can effectively capture and maintain key details in the image by integrating the local fine-grained features and the global coarse-grained features of the input image and reallocating the weights of the channels in the feature map. This design not only helps to improve the overall quality of the image, but also ensures that detail information is not lost during the enhancement process. In addition, the process of generating a supervised feature map based on the target deep-level features and the original feature map further ensures the accurate transmission of detail information. At the same time, this also provides guidance for subsequent processing stages and helps maintain global exposure balance.
[0088] The joint refinement module is used to estimate the brightness information of each color channel (R, G, B) in the supervised feature map, and adjust the brightness and contrast of each color channel using the Retinex decomposition theory to obtain the corrected supervised feature map, i.e., the intermediate feature map. The modulation parameter is used to scale and shift each feature in the intermediate feature map to obtain the space-variant feature map, and the intermediate feature map and the space-variant feature map are integrated to obtain the joint feature.
[0089] like Figure 3 As shown, the joint refinement module specifically includes:
[0090] A feature correction unit is used to continuously perform multiple convolution and ReLU activation operations on each color channel in the supervised feature map to estimate the brightness information of each color channel, and adjust the brightness and contrast of each color channel based on the brightness information using the Retinex decomposition theory, merge the adjusted color channels to obtain the target channel, and reconstruct the intermediate feature map through the target channel;
[0091] Based on the Retinex theory, the feature correction unit assumes that the illumination of each color channel (R, G, B) is estimated independently to ensure that the reflectance map can preserve the original hue of the source image. To this end, the feature correction unit first performs a series of convolution operations on each color channel in the supervised feature map and applies the ReLU activation function after each convolution to accurately estimate the brightness information of each color channel. This multi-layer processing method effectively captures and adjusts the brightness and contrast within each channel.
[0092] Specifically, the feature correction unit consists of multiple sequentially connected convolutional layers, which include 3×3 convolution kernels, normalization functions, and ReLU activation functions, which are used to extract local features of the supervised feature map. The operation of each layer enhances the nonlinear expression ability of the model and accelerates the training process. On this basis, in order to further improve the expressiveness of the model, two sequentially connected residual blocks are designed after the last convolutional layer. Each residual block also uses a 3×3 convolution kernel, which helps to capture more complex image features and improve image quality. The design of the residual block enables the network to build a deeper architecture without increasing the difficulty of training, effectively alleviating the gradient vanishing problem in deep networks.
[0093] In this process, the intermediate layer image features refer to the image feature maps obtained after a series of convolutions, normalization functions, activation functions, and residual blocks. These intermediate layer image features are set to have 128 feature channels to ensure sufficient feature representation. These feature representations contain key information extracted from the supervised feature maps and are regarded as a high-dimensional abstract expression of the image, capturing detail information such as edges and textures. This process not only restores the details lost due to uneven lighting, but also maintains the tonal consistency of the original dark-light image, significantly improving the overall quality and visual effect of the image.
[0094] The brightness information is estimated through the intermediate layer image features; finally, based on the estimated brightness information, the Retinex decomposition theory is used to adjust the brightness and contrast of each color channel, and then these adjusted color channels are merged to form the target channel (RGB format). Through this target channel, the intermediate feature map is reconstructed, providing a high-quality foundation for further image enhancement.
[0095] The spatial feature transformation unit is used to scale and shift each feature in the intermediate feature map using the modulation parameter to achieve spatial feature transformation, thereby obtaining a space-varying feature map; the formula for obtaining the space-varying feature map is:
[0096] ;
[0097] In the formula, represents the modulation parameters, where represents the scale parameter obtained through prior learning, represents the shift parameter obtained by prior learning, represents the intermediate feature map, represents point-wise multiplication, represents the feature conversion function; , have the same dimensions.
[0098] This embodiment uses the modulation parameters and To scale and shift the values in each intermediate feature map, we perform spatial feature transformation to precisely adjust the response values in the intermediate feature maps. Specifically, As a scale parameter used to scale each element in the intermediate feature map, As shift parameters are used to adjust the positions of these elements. Through this mechanism, the model is able to highlight and sharpen subtle image features, significantly improving the quality of the output image. In addition, this approach helps to reduce the difference between the predicted image and the real image, ensuring that the enhanced image is closer to the original normal light image. The spatial feature transformation introduces learnable modulation parameters, allowing the model to adaptively learn how to apply these transformations during training, so as to flexibly respond to different types of input images. This flexibility not only enhances the model's ability to adapt to various lighting conditions, but also enables it to effectively capture and enhance key details in the image, ultimately achieving a more natural and higher quality image enhancement effect. In this way, the model is able to better handle complex image structures and generate images with excellent visual effects under low light conditions.
[0099] The joint refinement unit is used to integrate the intermediate feature map and the space-variant feature map to obtain the joint feature; the integration formula is: ;
[0100] In the formula, represents the supervised feature map, represents the corrected supervised feature map, represents the spatial feature transformation, represents color normalization of the spatial-variant feature map, Represents a joint feature.
[0101] It should be noted that after the feature extraction stage, the model can learn the main information features of the original dark light image, including local fine-grained and global coarse-grained features. However, in some images, the brightness distribution is uneven and there are extremely dark areas, which leads to the failure to fully restore the details during the enhancement process, and ultimately causes the enhanced image to have blurred details and color distortion. In order to solve this challenge, this embodiment introduces a joint refinement module. The module first estimates the brightness information of each color channel and adjusts the brightness and contrast of each channel by applying the Retinex decomposition theory, thereby correcting the brightness distribution in the supervised feature map. Then, each feature in the intermediate feature map is scaled and shifted using modulation parameters to achieve spatial feature transformation and further optimize the brightness adjustment process. Through this multi-level refinement processing, the joint refinement module can not only effectively improve the overall brightness and contrast of the image, but also carefully restore and enhance the detail information in the image, ensuring that the final output image has higher clarity and more realistic color performance. In this way, the model can significantly improve the quality of images under low light conditions while maintaining natural visual effects.
[0102] That is to say, in order to more accurately consider the specific transformation and brightness adjustment of image details, the joint refinement module plays a key role in the present invention. This module uses the Retinex decomposition theory to estimate the brightness information of each color channel, and adjusts the brightness and contrast accordingly to obtain the corrected supervised feature map. This method pays special attention to the specific transformation of image details and ensures the accuracy of brightness adjustment. The spatial feature transformation unit realizes spatial feature transformation by modulating parameters to scale and shift each feature in the intermediate feature map. This step further optimizes the brightness adjustment process, making the enhanced image more natural and in line with actual scene requirements.
[0103] Ideally, if the brightening and dimming operations are perfect, the predicted image should be exactly the same as the real image. However, in practical applications, there are differences in the conversion process from normal light images to low light images or vice versa, and these differences reflect the errors in the brightening or dimming operations. In order to compensate for these errors, and based on the inspiration of the back projection theory: "Using the dimming operation, the normal light image can be turned into a low light image. Conversely, the purpose of low light image enhancement is to perform appropriate brightening operations to predict the normal light image", this embodiment sets a brightness adjustment module.
[0104] The brightness adjustment module is used to aggregate the joint features and the original dark light image to obtain an aggregated image, predict a normal light image from the aggregated image, and predict a low-light image through the normal light image, obtain the residual difference between the low-light image and the original dark light image, and obtain a first residual map; based on the first residual map, estimate the residual difference between the predicted normal light image and the original normal light image corresponding to the original dark light image to obtain a second residual map, and fuse the second residual map with the joint features to obtain an enhanced image.
[0105] like Figure 4 As shown, the brightness adjustment module specifically includes:
[0106] An aggregation block is used to aggregate (including connecting, calibrating and integrating the joint features with the features in the original dark light image) the joint features and the original dark light image to obtain an aggregated image;
[0107] A first light block is used to predict a normal light image from the aggregated image;
[0108] A dark block is used to predict a low-light image through the predicted normal-light image;
[0109] A first residual block is used to obtain a residual difference between the predicted low-light image and the original low-light image to obtain a first residual map;
[0110] The second light block;
[0111] A second residual block, configured to estimate a residual difference between a predicted normal light image and an original normal light image corresponding to the original dark light image by using the first residual map through the second lit block, to obtain a second residual map;
[0112] The estimation formula of the second residual mapping is:
[0113] ;
[0114] In the formula, , is the preset weight, R represents the real number set, represents an aggregate image, Indicates the first light block, Indicates dark blocks, Indicates the second light block; : represents the intermediate residual mapping; represents the second residual mapping; Represents the preset weight coefficient, represents the first residual map.
[0115] The fusion block is used to fuse the second residual map with the joint feature to obtain an enhanced image.
[0116] Light up blocks include:
[0117] The encoded convolution (which includes a convolution layer and a parameterized ReLU layer) is used to reduce the number of feature channels of the image to be processed (reducing the number of feature channels from 64 to 32 to extract more representative features) to obtain the key feature image;
[0118] Offset convolution is used to learn the difference between the original dark light image and the original normal light image, obtain the brightness offset of each pixel between the two, remove the negative brightness offset, and increase the brightness offset corresponding to each pixel in the key feature image to adjust the brightness of the pixel, thereby obtaining the adjusted key feature image;
[0119] Decoding convolution is used to increase the number of feature channels of the adjusted key feature image to obtain the target feature (increasing the number of feature channels from 32 to 64);
[0120] The lighting block is a first lighting block or a second lighting block;
[0121] When the light-on block is the first light-on block, the image to be processed is an aggregated image, and the target feature is a predicted normal light image;
[0122] When the lit block is a second lit block, the image to be processed is an intermediate residual map, and the target feature is a second residual map.
[0123] The brightness adjustment module of the present invention generates an aggregated image by aggregating the joint features and the original dark light image, and predicts the normal light image therefrom, and then further predicts the low light image based on the normal light image, and calculates the residual difference between the low light image and the original dark light image to obtain the first residual map. Then, the residual difference between the predicted normal light image and the original normal light image is estimated based on the first residual map to obtain the second residual map, and finally the second residual map and the joint features are fused to generate the enhanced image. The beneficial effect of this process is that it can not only effectively improve the overall brightness and contrast of the image, but also carefully correct the local details and brightness range, ensuring that the enhanced image is excellent in color fidelity, detail clarity and global exposure balance. Through this multi-level feature aggregation and residual mapping mechanism, the model can significantly reduce the problems of color distortion and over-enhancement while maintaining natural visual effects, thereby providing a more realistic and high-quality image output. In addition, the model's adaptability to different lighting conditions is enhanced, and its generalization performance is improved.
[0124] like Figure 5 As shown, the embodiment of the present invention also proposes a training method for a dark light image enhancement model, comprising the steps of:
[0125] S1: Acquire a data set; the data set includes a plurality of training samples; each of the training samples includes an original dark light image and its corresponding original normal light image;
[0126] S2: Set a loss function, and train the dark light image enhancement model as described above using the dataset and the set loss function.
[0127] The loss function includes:
[0128] A first loss function is used to evaluate the difference between the predicted low-light image and the original low-light image, and between the predicted normal-light image and the original normal-light image;
[0129] The second loss function is used to enhance the model’s ability to capture edge details of each image;
[0130] The formula expression of the first loss function is:
[0131] ;
[0132] The formula expression of the second loss function is:
[0133] ;
[0134] In the formula, represents the Charbonnier loss, represents the enhanced image output by the dark light image enhancement model, represents the original normal light image in the training sample, is a constant; represents the edge loss, represents the Laplace operator.
[0135] The total loss function is constructed by combining the first loss function and the second loss function.
[0136] The formula expression of the total loss function is:
[0137] ;
[0138] In the formula, For control and The weight parameters, Represents the total loss value.
[0139] The training method of the dark light image enhancement model proposed in the present invention uses a data set containing the original dark light image and its corresponding original normal light image, and sets a total loss function including Charbonnier loss and edge loss, thereby effectively improving the performance of the model in the dark light image enhancement task. This method first uses a feature extraction module to capture and retain the key detail information of the input image to ensure that the details are not lost during the enhancement process. Then, the joint refinement module adjusts the brightness and contrast based on the Retinex decomposition theory, optimizes the detail performance in each color channel, and further improves the image quality. The brightness adjustment module ultimately generates a high-quality enhanced image through a residual mapping mechanism. This training method improves the model's ability to handle image brightness, contrast, and structural similarity.
[0140] In order to demonstrate the beneficial effects of the present invention, the following experiments were performed in this embodiment:
[0141] A variety of evaluation indicators are used to quantify and compare the performance of different methods:
[0142] For datasets such as LOL, COCO, and MIT, they provide low-light images and their corresponding high-quality ground truth pairs, which provide a benchmark for evaluating the performance of low-light image enhancement algorithms in the present invention. This embodiment uses a variety of evaluation indicators to quantify the performance of different methods, including reference evaluation indicators (such as PSNR, SSIM, and LPIPS) and no-reference evaluation indicators (such as BRISQUE and NIQE) to comprehensively evaluate the performance of various algorithms in terms of enhancement effect and naturalness of output images.
[0143] Due to the large amount of data, in order to clearly show the comparison results, the following three tables (Table 1, Table 2 and Table 3) are used to show the comparison results:
[0144] Table 1:
[0145]
[0146] Table 2:
[0147]
[0148] Table 3:
[0149]
[0150] In Tables 1 to 3, the upward arrows indicate that the larger the value, the better, and the downward arrows indicate that the smaller the value, the better. In the table, OursPT and OursMS represent the models of the present invention developed in different software frameworks (PT represents the PyTorch framework, and MS represents the multi-scale architecture).
[0151] The English explanations of the rows and columns in Tables 1 to 3 are as follows: PSNR means peak signal-to-noise ratio; SSIM means structural similarity; BRISQUE means reference-free spatial domain image quality evaluation; NIQE means natural image quality; LPIPS means image similarity; HE means histogram equalization; CLAHE means limited contrast adaptive histogram equalization; LDR means low dynamic range; LRS means local region segmentation; SCI means self-calibrated illumination framework; RUAS means low-light image enhancement using collaborative prior architecture search; Zero-DCE means zero-reference deep learning curve estimation; KinD means network to light the darkness; MIRNet means multi-scale residual network; EnGAN means efficient unsupervised generative adversarial network; DLN means deep lightning network; SDD means semi-decoupled decomposition; JED means joint enhancement and denoising method through sequential decomposition; Uretinex means deep unfolding network based on Retinex; LLFlow means low-light image enhancement pipeline; IAT means illumination adaptive transform; SNR stands for: Signal-to-Noise Ratio Aware Low-Light Images; CSDNet stands for: Learning Deep Context-Dependent Decomposition Networks; KinD++ stands for: Beyond Brightening Low-Light Images; FLW stands for: Fast and Lightweight Networks; DDNet stands for: Dual Domain Guided Networks.
[0152] Specifically, the above tables (Table 1, Table 2, Table 3) show the quantitative comparison results between the dark light image enhancement model proposed by the present invention and other state-of-the-art methods. It can be seen from the data in each table that the images enhanced by the method of the present invention have excellent performance in terms of brightness, contrast and structure, and have also achieved very competitive results in terms of perceptual quality and naturalness (measured by the three indicators of BRISQUE, NIQE and LPIPS). In particular, on the COCO dataset, the present invention not only surpasses most existing methods based on deep learning in traditional evaluation criteria such as peak signal-to-noise ratio (PSNR), image brightness, contrast and structural similarity (reflected by the SSIM indicator), but also on the MIT dataset, the present invention is also superior to most other methods in the measurement indicator of structural similarity index (SSIM). This means that compared with the prior art, the model of the present invention can not only effectively improve the overall visual quality of the image, but also better maintain the original structural information of the image, providing a more realistic and natural enhancement effect. These results further verify the superiority and applicability of the present invention in the field of dark light image enhancement.
[0153] It should be noted that the term "including" and its variations used in the embodiments of the present invention are open inclusions, that is, "including but not limited to". The term "based on" means "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one other embodiment"; the term "some embodiments" means "at least some embodiments". The modifications of "one" and "multiple" mentioned in the embodiments of the present invention are illustrative rather than restrictive. Those skilled in the art should understand that unless the context clearly indicates otherwise, it should be understood as "one or more".
[0154] The term "embodiment" in this specification refers to specific features, structures or characteristics described in conjunction with the embodiment that can be included in at least one embodiment of the present invention. The appearance of this phrase in various places in the specification does not necessarily mean the same embodiment, nor does it mean that it is mutually exclusive with other embodiments and is independent or optional. The various embodiments in this specification are described in a related manner, and the same or similar parts between the various embodiments refer to each other. In particular, for the device, equipment, and system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts refer to the partial description of the method embodiment.
[0155] The above-mentioned embodiments only express several implementation methods of the present invention, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the scope of patent protection. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present invention, and these all belong to the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the attached claims.
Claims
1. A dark light image enhancement model, characterized in that: include: A feature extraction module is used to integrate the local fine-grained features and the global coarse-grained features of the input original dark-light image to generate a feature map, reallocate weights to the channels in the feature map to obtain a shallow feature map, extract deep features in the original dark-light image based on the feature map and the shallow feature map, restore the deep features to the resolution of the original image to obtain target deep features, and generate a supervised feature map based on the target deep features and the feature map; A joint refinement module is used to estimate the brightness information of each color channel in the supervised feature map, and adjust the brightness and contrast of each color channel using the Retinex decomposition theory to obtain a corrected supervised feature map, i.e., an intermediate feature map. Each feature in the intermediate feature map is scaled and shifted using a modulation parameter to obtain a space-variant feature map. The intermediate feature map and the space-variant feature map are integrated to obtain a joint feature. a brightness adjustment module, configured to aggregate the joint features and the original dark-light image to obtain an aggregated image, predict a normal-light image from the aggregated image, and predict a low-light image through the normal-light image, obtain a residual difference between the low-light image and the original dark-light image, and obtain a first residual mapping; A residual difference between the predicted normal light image and the original normal light image corresponding to the original dark light image is estimated based on the first residual map to obtain a second residual map, and the second residual map is fused with the joint feature to obtain an enhanced image.
2. A dark light image enhancement model according to claim 1, characterized in that: The feature extraction module specifically includes: The residual unit is used to integrate the local fine-grained features and the global coarse-grained features of the input original dark-light image to generate a feature map; Residual channel attention unit, used to redistribute weights of channels in feature maps to obtain shallow feature maps; An encoder, used to extract deep features in the original dark light image through the feature map output by the residual unit and the shallow feature map output by the residual channel attention unit; A decoder is used to restore the deep-level features to the resolution of the original image to obtain the target deep-level features; The self-supervisory block is used to generate a supervised feature map through the target deep-level features and the feature map output by the residual unit.
3. A dark light image enhancement model according to claim 2, characterized in that: The joint refinement module specifically includes: A feature correction unit is used to continuously perform multiple convolution and ReLU activation operations on each color channel in the supervised feature map to estimate the brightness information of each color channel, and adjust the brightness and contrast of each color channel based on the brightness information using the Retinex decomposition theory, merge the adjusted color channels to obtain the target channel, and reconstruct the intermediate feature map through the target channel; The spatial feature transformation unit is used to scale and shift each feature in the intermediate feature map using the modulation parameter to achieve spatial feature transformation, thereby obtaining a space-varying feature map; the formula for obtaining the space-varying feature map is: ; In the formula, represents the modulation parameters, where represents the scale parameter obtained through prior learning, represents the shift parameter obtained by prior learning, represents the intermediate feature map, represents point-wise multiplication, represents the feature conversion function; The joint refinement unit is used to integrate the intermediate feature map and the space-variant feature map to obtain the joint feature; the integration formula is: ; In the formula, represents the supervised feature map, represents the corrected supervised feature map, represents the spatial feature transformation, represents color normalization of the spatial-variant feature map, Represents a joint feature.
4. A dark light image enhancement model according to claim 3, characterized in that: The brightness adjustment module specifically includes: Aggregation block, used to aggregate the joint features and the original dark light image to obtain an aggregated image; A first light block is used to predict a normal light image from the aggregated image; A dark block is used to predict a low-light image through the predicted normal-light image; A first residual block is used to obtain a residual difference between the predicted low-light image and the original low-light image to obtain a first residual map; The second light block; A second residual block, configured to estimate a residual difference between a predicted normal light image and an original normal light image corresponding to the original dark light image by using the first residual map through the second lit block, to obtain a second residual map; The fusion block is used to fuse the second residual map with the joint feature to obtain an enhanced image.
5. A dark light image enhancement model according to claim 4, characterized in that: The estimation formula of the second residual mapping is: ; In the formula, ∈R, is the preset weight, R represents the real number set, represents an aggregate image, Indicates the first light block, Indicates dark blocks, Indicates the second light block; : represents the intermediate residual mapping; represents the second residual mapping; Represents the preset weight coefficient, represents the first residual map.
6. A dark light image enhancement model according to claim 5, characterized in that: Light up blocks include: Coded convolution is used to reduce the number of feature channels of the image to be processed and obtain the key feature image; Offset convolution is used to learn the difference between the original dark light image and the original normal light image, obtain the brightness offset of each pixel between the two, remove the negative brightness offset, and increase the brightness offset corresponding to each pixel in the key feature image to adjust the brightness of the pixel, thereby obtaining the adjusted key feature image; Decoding convolution is used to increase the number of feature channels of the adjusted key feature image to obtain the target feature; the light-up block is the first light-up block or the second light-up block; When the light-on block is the first light-on block, the image to be processed is an aggregated image, and the target feature is a predicted normal light image; When the lit block is a second lit block, the image to be processed is an intermediate residual map, and the target feature is a second residual map.
7. A training method for a dark light image enhancement model, characterized in that: Includes steps: S1: Acquire a data set; the data set includes a plurality of training samples; each of the training samples includes an original dark light image and its corresponding original normal light image; S2: Setting a loss function, and training the dark-light image enhancement model as described in any one of claims 1 to 6 using the data set and the set loss function.
8. The method for training a dark light image enhancement model according to claim 7, characterized in that: The loss function includes: A first loss function is used to evaluate the difference between the predicted low-light image and the original low-light image, and between the predicted normal-light image and the original normal-light image; The second loss function is used to enhance the model’s ability to capture edge details of each image; The total loss function is constructed by combining the first loss function and the second loss function.
9. The method for training a dark light image enhancement model according to claim 8, characterized in that: The formula expression of the first loss function is: ; The formula expression of the second loss function is: ; In the formula, represents the Charbonnier loss, represents the enhanced image output by the dark light image enhancement model, represents the original normal light image in the training sample, is a constant; represents the edge loss, represents the Laplace operator.
10. The method for training a dark light image enhancement model according to claim 9, characterized in that: The formula expression of the total loss function is: ; In the formula, For control and The weight parameters, Represents the total loss value.
Citation Information
Patent Citations
Weak light image enhancement method based on tone mapping and regularization model
CN107316279A
Image processing method, electronic device and non-transitory computer-readable recording medium
US20190378247A1