A Low-Light Image Enhancement Method Using Long Exposure Compensation

By introducing brightness and color information of blurred long-exposure images, the problem of low-light image enhancement is solved by introducing constraints on the problem of low-light image enhancement, which solves the problem of difficulty in dealing with local overexposure and signal noise in traditional methods, and achieves more efficient low-light image enhancement performance.

CN115240022BActive Publication Date: 2025-06-27PEKING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210651629.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-09
Publication Date
2025-06-27
Estimated Expiration
2042-06-09

AI Technical Summary

Technical Problem

Traditional low-light image enhancement methods are difficult to solve local overexposure and signal noise problems, cannot meet general automation requirements, and are difficult to generalize to images that process various lighting conditions.

Method used

By introducing easily obtained blurred long-exposure images, using their brightness, color and other information to add constraints on low-light enhancement problems, reducing uncertainty and improving low-light enhancement performance. Specific methods include collecting low-light training data sets, generating synthetic data sets, training low-light enhancement models, and using feature alignment and brightening modules for image enhancement.

Benefits of technology

It significantly improves the performance of low-light picture enhancement, and can increase the Peak Signal to Noise Ratio from 14.93 to 25.15 on the LEC-LOL-Real low-light enhancement reference dataset.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115240022B_ABST
    Figure CN115240022B_ABST
Patent Text Reader

Abstract

The present invention discloses a low-light image enhancement method using long-exposure compensation. The method is as follows: 1) Collect a low-light training data set, wherein each training sample in the low-light training data set includes a low-light image and a normal-light image of the same scene; generate a corresponding set of short-exposure images, long-exposure images, and true-light images according to each training sample to obtain a synthetic data set S; 2) Use the synthetic data set S to train a low-light enhancement model, and the low-light enhancement model includes M-1 feature alignment modules and M-1 brightening modules; 3) Input the short-exposure image to be brightened and the corresponding blurred long-exposure image into the trained low-light enhancement model to obtain the corresponding low-light enhanced image. The present invention can significantly improve the low-light image enhancement performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of low-light enhancement of digital images, and relates to a low-light image enhancement method using long-exposure compensation. Background Art

[0002] Low light is a common image degradation, and insufficient light is usually caused by low-light shooting environments, camera malfunctions, incorrect parameter settings, etc. The enhancement of low-light images has always attracted the attention of the industrial and academic communities.

[0003] Traditional low-light image enhancement methods can be divided into three categories. Methods based on uniformly adjusting the image brightness brighten low-light images by uniformly adjusting the global brightness of the entire image. Methods based on the retino-cortical theory decompose the image into a reflectance layer and an illumination layer, and use prior knowledge to manually set constraints for adjustment to achieve the purpose of enhancing low-light images. Deep learning-based methods design data-driven convolutional models and perform end-to-end training on large datasets, and only require one parameter forward pass of the low-light image during inference.

[0004] However, the enhancement of low-light images is an ill-posed problem. A low-light image can correspond to multiple ideal normal-light images, and this uncertainty of the optimization target poses challenges for accurate and flexible low-light image enhancement. Methods based on uniformly adjusting the image brightness cannot solve the problems of local overexposure and signal noise. Methods based on the retino-cortical theory cannot meet the requirements of generalization and automation. Deep learning-based methods are difficult to generalize to images with various illumination conditions. Therefore, traditional low-light image enhancement methods are all difficult to generalize to low-light pictures under various illumination conditions and cannot meet the needs of practical applications. Summary of the Invention

[0005] Aiming at the above problems, the purpose of the present invention is to provide a low-light image enhancement method using long-exposure compensation. By introducing easily obtained blurred long-exposure images, using the brightness, color and other information of the blurred long-exposure images to add constraints to the low-light enhancement problem, reducing the uncertainty of the low-light image enhancement problem, making the optimization target more clear, and comprehensively improving the low-light enhancement performance. The overall framework is as shown in the appendix Figure 1 as follows.

[0006] The technical solution adopted by the present invention is as follows:

[0007] A low-light image enhancement method using long-exposure compensation, the steps of which include:

[0008] 1) Collect a low-light training dataset, where each training sample in the low-light training dataset includes a low-light image and a normal-light image of the same scene; generate a corresponding set of short-exposure images, long-exposure images, and true-light images for each training sample to obtain a synthetic dataset S;

[0009] 2) Use the synthetic dataset S to train a low-light enhancement model, where the low-light enhancement model includes M - 1 feature alignment modules and M - 1 brightening modules; among them, for the long-exposure image I long and the short-exposure image I short in the same set of images in the synthetic dataset S, the low-light enhancement model maps the long-exposure image I long and the short-exposure image I short to the feature space respectively to obtain the corresponding short-exposure feature and long-exposure feature and input them into the first feature alignment module;

[0010] 3) The i-th feature alignment module aligns the input i-th scale long-exposure feature with the i-th scale short-exposure feature ; among them, the i-th feature alignment module performs convolution processing on the i-th scale short-exposure feature to obtain an attention map A i , and then uses the attention map A i to perform a soft-threshold filtering operation on the i-th scale long-exposure feature to obtain where "⊙" represents element-wise multiplication; then and are jointly downsampled and passed into a convolutional layer to predict and output the (i + 1)-th scale long-exposure feature and is downsampled alone and passed into a convolutional layer to predict and output the (i + 1)-th scale short-exposure feature The (M - 1)-th scale long-exposure feature predicted and output by the (M - 1)-th feature alignment module and the (M - 1)-th scale short-exposure feature are concatenated as the (M + 1)-th scale long-exposure feature and the (M + 1)-th scale short-exposure feature where, i = 1 ~ M - 1,

[0011] 4) The i-th brightening module concatenates the (M + i)-th scale long-exposure feature and the (M + i)-th scale short-exposure feature , performs upsampling on the concatenated feature, and then combines the upsampled feature with the (M - i)-th scale short-exposure feature After being connected, the short-exposure features at the (M + i + 1)-th scale are obtained through a convolutional layer For The features obtained by upsampling are connected to the long-exposure features at the (M - i)-th scale After being connected, the long-exposure features at the (M + i + 1)-th scale are obtained through a convolutional layer Using the short-exposure features at the 2M-th scale output by the (M - 1)-th brightening module As optimization target I normal And the long-exposure features at the 2M-th scale As an auxiliary, optimize the low-light enhancement model; where the total loss function for training and optimizing the low-light enhancement model is L = L rec + λ SSIM L SSIM + λ LPIPS L LPIPS + λ a L a ; λ SSIM 、λ LPIPS And λ a Are weight terms, L rec Is the mean absolute error loss function between optimization target I normal And the true value I under normal light GT ; L SSIM Is the structural similarity loss function between optimization target I normal And the true value I under normal light GT ; L LPIPS Is the perceptual image patch similarity learning loss function; L a Is the mean absolute error loss function between the auxiliary output I assist And the true value I under normal light GT ;

[0012] 5) Input the short-exposure image to be brightened and the corresponding blurred long-exposure image into the trained low-light enhancement model to obtain the corresponding low-light enhanced image

[0013] Furthermore, the building of the low-light enhancement model further includes a detail removal module for eliminating the detail features of the long-exposure image I long And then mapping it to the feature space to obtain the long-exposure features

[0014] Furthermore, obtain a real-shot dataset R, where each group of images in the real-shot dataset R includes three images taken of the same scene, namely a short-exposure image, a long-exposure image, and a real-light image; use the real-shot dataset R to evaluate the trained low-light enhancement model

[0015] Further, the long-exposure images in the synthetic dataset S are synthesized from the normal illumination images in the training samples; the short-exposure images in the synthetic dataset S are the low-illumination images in the training samples, and the real illumination images in the synthetic dataset S are the normal illumination images in the training samples.

[0016] Further, the long-exposure images in the synthetic dataset S are obtained by processing the normal illumination images in the training samples using a blur kernel space model.

[0017] Further, L rec =‖I normal -I GT ‖; L a =‖I assist -I GT ‖.

[0018] Further, L SSIM =1 - SSIM(I normal , I GT ); where SSIM(x, y) represents the structural similarity between two images x and y; μ x is the average value of x, μ y is the average value of y, is the variance of x, is the variance of y, σ xy is the covariance of x and y, and c1, c2 are constants used to maintain stability.

[0019] Further, where φ l (I normal ) represents the feature of the l-th layer of the image I normal extracted by the VGG network, and H l and W l respectively represent the width and height of φ l (I normal ).

[0020] A server, comprising a memory and a processor, wherein the memory stores a computer program configured to be executed by the processor, and the computer program includes instructions for performing the steps in the above method.

[0021] A computer-readable storage medium, on which a computer program is stored, characterized in that the steps of the above method are implemented when the computer program is executed by a processor.

[0022] Compared with the prior art, the positive effects of the present invention are:

[0023] The present invention significantly improves the performance of low-light image enhancement. On the LEC-LOL-Real low-light enhancement benchmark dataset, it can increase the Peak Signal to Noise Ratio of the general low-light enhancement model AGLLNet from 14.93 to 25.15. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 It is a training framework diagram of a low-light image enhancement network using long exposure compensation.

[0025] Figure 2 It is a framework diagram of the feature alignment sub-module.

[0026] Figure 3 It is a framework diagram of the brightening sub-module. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0027] To make the above features and advantages of the present invention more obvious and understandable, specific embodiments are given below and described in detail in conjunction with the accompanying drawings. It should be noted that the specific number of layers, number of modules, number of functions, and settings of certain layers given in the following embodiments are only a preferred implementation manner and are not used for limitation. Those skilled in the art can select the number and set certain layers according to actual needs, which should be understandable.

[0028] This embodiment discloses a low-light enhancement method using long exposure compensation, which is specifically described as follows:

[0029] Step 1: Collect a low-light training dataset. For each low-light / normal-light image data pair in it, use the blur kernel space model to synthesize multiple different long-exposure blurred images, and build a long-exposure compensation low-light enhancement synthesis dataset S composed of normal-light / short-exposure / synthesized long-exposure image pairs for network training and testing; among them, the short-exposure / real-light image is directly obtained from the collected low-light / normal-light image pair, that is, the collected low-light image is the short-exposure image in set S, and the collected normal-light image is the real-light image in set S. For the same scene different from the training dataset S, take three short-exposure / long-exposure / normal-light images to form a data pair, and build a long-exposure compensation low-light enhancement actual shooting dataset R for network testing.

[0030] Step 2: Build a low-light enhancement training framework.

[0031] The structure of the network is as Figure 1 shown, including a feature alignment module S2L, a brightening module L2S, and a detail removal module DRP.

[0032] Next, take a pair of long-exposure images I long and short-exposure images Ishort Take it as an example to introduce the network framework. The long-exposure image I long and the short-exposure image I short First, each passes through a convolutional layer, followed by a normalization layer and a rectified linear unit (ReLU) to map the image to the feature space and obtain the initial short-exposure features and long-exposure features Specifically, since the long-exposure image provides information about brightness and illumination, it is not desirable for the detailed information of the long-exposure to interfere with the model. Before the long-exposure input enters the long-exposure feature decoding module, a 16-fold downsampling and 16-fold upsampling module DRP is added to eliminate the detailed features.

[0033] To align the brightness features of the long-exposure image in each layer with the detailed features of the short-exposure image and achieve effective feature interaction, a feature alignment module S2L is added to the model. The structure of the feature alignment module is as Figure 2 shown. The short-exposure features first pass through a convolutional layer to obtain an attention map Then, use this attention map A i to perform a soft-threshold filtering operation on the long-exposure features to obtain where "⊙" represents element-wise multiplication. After that, and are jointly downsampled and passed into a convolutional layer to predict the next long-exposure feature is separately downsampled and passed into a convolutional layer to predict the next short-exposure feature The next feature is half the size of the previous feature in the spatial dimension but twice the size of the previous feature in the channel dimension.

[0034] After passing through (M - 1) feature alignment modules, we obtain multi-scale short-exposure features and long-exposure features Next, the model decodes them from the feature space to the image space through the feature decoding stage. In the feature decoding stage, the guidance is in the opposite direction. The long-exposure features will guide the decoding of the short-exposure features because we need the brightness features to guide the enhancement of the short-exposure image, and M is an integer greater than 2.

[0035] Similar to the feature alignment module S2L, a brightening module L2S is also added to the model, and its structure is as Figure 3 shown. The brightening module takes as input the long-exposure features of the previous size and the short-exposure features as well as the input features connected by skip connections Long and short exposure features in the size - equivalent encoding stage More specifically, the long - exposure feature is first connected to the short - exposure feature Then, the feature obtained by upsampling is connected to the feature of the skip connection through the upsampling module and passed through a convolutional layer to obtain the feature of the next scale Upsampling is performed separately, and the feature obtained by upsampling is connected to the feature of the skip connection and passed through a convolutional layer to obtain the feature of the next scale Compared with the feature of the previous scale, the feature of the next scale is twice as large in the spatial dimension but half as large in the channel dimension

[0036] Step 3: Train the model using the constructed dataset S, with the output of the model in the short - exposure decoding module as the optimization target I normal , and the output of the long - exposure decoding module as the auxiliary output I assist to train the model. The total loss function term of the low - light image enhancement model using long - exposure compensation is:

[0037] L = L rec + λ SSIM L SSIM + λ LPIPS L LPIPS + λ a L a

[0038] where λ SSIM , λ LPIPS and λ a are weight terms. Usually, λ SSIM is set to 0.4, λ LPIPS is set to 1, and λ a is set to 1. The training batch size of the model is 16. The Adam optimizer is used, and the initial learning rate is 1×10 -4 , and the optimizer hyperparameters are set to β1 = 0.9, β2 = 0.999, and the weight decay parameter is 1×10 -4 . In addition, to avoid gradient explosion, the gradient value in the gradient backpropagation will be truncated in the interval [- 0.1, 0.1]. During the training process, randomly crop 256×256 - pixel blocks and use a two - stage training strategy. In the first stage, train for 1.5×10 5 iterations without adding the attention mechanism of the feature alignment module S2L, and then, after adding the attention mechanism, train for 3×10 -5 iterations with the initial learning rate of 1×10 4 .

[0039] 1)L recFor the optimization objective I normal and the true value I under normal illumination GT the mean absolute error loss function between them is:

[0040] L rec =‖I normal - I GT ‖,

[0041] 2) L SSIM For the optimization objective I normal and the true value I under normal illumination GT the structural similarity loss function between them is:

[0042] L SSIM = 1 - SSIM(I normal , I GT ),

[0043] where SSIM(x, y) represents the structural similarity between two images x and y, which can be obtained in the following way:

[0044]

[0045] where μ x is the average value of x, μ y is the average value of y, is the variance of x, is the variance of y, σ xy is the covariance of x and y. c1 = (k1L) 2 , c2 = (k2L) 2 are constants used to maintain stability. L is the dynamic range of pixel values. k1 = 0.01, k2 = 0.03.

[0046] 3) L LPIPS is the perceptual image patch similarity learning loss function:

[0047]

[0048] Here, the image features used to calculate L LPIPS are extracted using the pre-trained VGG network on ImageNet. By inputting the output result I normal of the network and the image I GT under normal illumination in the image pair into the pre-trained VGG network model respectively to obtain image features. Among them, φ l (I) represents the l-th layer feature of the image I extracted by the VGG network, and H l and W l represent the width and height of φ l (I) respectively.

[0049] 4) L aFor the auxiliary output I assist and the true value I under normal illumination GT The mean absolute error loss function between:

[0050] L a =‖I assist -I GT ‖.

[0051] Step 4: Use the real-shot dataset R to evaluate the trained low-light enhancement model.

[0052] Step 5: In the inference stage, input the short-exposure image to be brightened and the corresponding blurred long-exposure image, and finally output the desired low-light enhancement result.

[0053] The above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit them. Those of ordinary skill in the art can modify or equivalently replace the technical solutions of the present invention without departing from the spirit and scope of the present invention. The protection scope of the present invention shall be subject to what is described in the claims.

Claims

1. A low-light image enhancement method using long-exposure compensation, the steps of which include: 1) Collect a low-light training data set, where each training sample in the low-light training data set includes a low-light image and a normal-light image of the same scene; Generate a corresponding set of short-exposure images, long-exposure images, and true-light images according to each training sample to obtain a synthetic data set S; 2) Train a low-light enhancement model using the synthetic dataset S, where the low-light enhancement model includes M - 1 feature alignment modules and M - 1 brightening modules; among them, for the long-exposure image I and the short-exposure image I in the same group of images in the synthetic dataset S long and the short-exposure image I short , the low-light enhancement model maps the long-exposure image I long and the short-exposure image I short to the feature space respectively to obtain the corresponding short-exposure feature and long-exposure feature and input them into the first feature alignment module; 3) The i-th feature alignment module aligns the input i-th scale long-exposure feature with the i-th scale short-exposure feature ; specifically, the i-th feature alignment module performs convolution processing on the i-th scale short-exposure feature to obtain an attention map A i , and then uses the attention map A i to perform a soft-threshold filtering operation on the i-th scale long-exposure feature to obtain where "⊙" represents element-wise multiplication; then and are jointly downsampled and fed into a convolutional layer to predict and output the (i + 1)-th scale long-exposure feature and is separately downsampled and fed into a convolutional layer to predict and output the (i + 1)-th scale short-exposure feature The (M)-th scale long-exposure feature predicted and output by the (M - 1)-th feature alignment module and the (M)-th scale short-exposure feature are concatenated as the (M + 1)-th scale long-exposure feature and the (M + 1)-th scale short-exposure feature where i = 1 to M - 1, 4) The i-th brightening module splices the long-exposure feature of the M+i-th scale and the short-exposure feature of the M+i-th scale , upsamples the spliced feature, and then connects the upsampled feature with the short-exposure feature of the M-i-th scale and passes it through a convolutional layer to obtain the short-exposure feature of the M+i+1-th scale Connect the feature obtained by upsampling with the long-exposure feature of the M-i-th scale and pass it through a convolutional layer to obtain the long-exposure feature of the M+i+1-th scale Use the short-exposure feature of the 2M-th scale output by the M-1-th brightening module as the optimization target I normal and the long-exposure feature of the 2M-th scale as an auxiliary to optimize the low-light enhancement model; where the total loss function for training and optimizing the low-light enhancement model is L = L rec + λ SSIM L SSIM + λ LPIPS L LPIPS + λ a L a ; λ SSIM , λ LPIPS and λ a are weight terms, L rec is the mean absolute error loss function between the optimization target I normal and the true value I under normal light GT ; L SSIM is the structural similarity loss function between the optimization target I normal and the true value I under normal light GT ; L LPIPS is the perceptual image patch similarity learning loss function; L a is the mean absolute error loss function between the auxiliary output I assist and the true value I under normal light GT ; 5) Input the short-exposure image to be brightened and the corresponding blurred long-exposure image into the trained low-light enhancement model to obtain the corresponding low-light enhanced image.

2. The method according to claim 1, characterized in that, The construction of the low-light enhancement model further includes a detail removal module for eliminating the detail features of the long-exposure image I long and mapping them to the feature space to obtain long-exposure features 3. The method according to claim 1, characterized in that, Obtain a real-shot data set R, where each set of images in the real-shot data set R includes three images taken of the same scene, namely a short-exposure image, a long-exposure image, and a true-light image; use the real-shot data set R to evaluate the trained low-light enhancement model.

4. The method according to claim 1 or 2 or 3, characterized in that, The long-exposure image in the synthetic data set S is synthesized from the normal-light image in the training sample; the short-exposure image in the synthetic data set S is the low-light image in the training sample, and the true-light image in the synthetic data set S is the normal-light image in the training sample.

5. The method according to claim 1 or 2 or 3, characterized in that, Process the normal-light image in the training sample using a blur kernel space model to obtain the long-exposure image in the synthetic data set S.

6. The method according to claim 1 or 2 or 3, characterized in that, L rec = ‖I normal - I GT ‖; L a = ‖I assist - I GT ‖。 7. The method according to claim 1, characterized in that L SSIM = 1 - SSIM(I normal , I GT ); where SSIM(x, y) represents the structural similarity between two images x and y; μ x is the average value of x, μ y is the average value of y, is the variance of x, is the variance of y, σ xy is the covariance of x and y, and c1, c2 are constants used to maintain stability.

8. The method according to claim 1, characterized in that Among them, φ l (I normal ) represents the l-th layer feature of the image I normal extracted by the VGG network, and H l and W l respectively represent the width and height of φ l (I normal ).

9. A server, characterized in that, It includes a memory and a processor, the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program includes instructions for executing the steps in any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • No-reference low-illumination image enhancement method and system based on generative adversarial network

    CN111798400A

  • Implicit edge prior-based scale progressive image completion method

    CN113298733A