Low-light Image Enhancement and Denoising Method Combining NSST Domain, GAN and Scale Correlation Coefficient
Through the method of combining the NSST domain with GAN and scale correlation coefficient, the existing low-illumination image enhancement method has been solved and the generalization ability is poor at single scale, and the effect of efficiently enhancing low-illumination images at multiple scales is achieved, which enhances noise resistance and edge enhancement capabilities.
Patent Information
- Application Number
- CN202211168684.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-24
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2042-09-24
AI Technical Summary
The existing low-illumination image enhancement methods are carried out at a single scale, with poor generalization capabilities and requires precise pairing of training sets, resulting in poor results in real low-illumination image processing.
The method of combining the NSST domain with GAN and scale correlation coefficient is adopted to obtain low-frequency and high-frequency subband images through NSST multi-scale decomposition, and a low-frequency subband image enhancement model LF-EnlightenGAN based on GAN is constructed, and the noise and enhance edges are removed by combining the scale correlation coefficient.
It improves the overall brightness, clarity and information entropy of low-illumination images, retains more texture details, enhances noise resistance and edge enhancement capabilities, and is suitable for subsequent tasks such as image recognition, image classification, and object detection.
Smart Images

Figure CN115908155B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly to a low-light image enhancement and denoising method combining NSST domain, GAN and scale correlation coefficient. Background Art
[0002] The lighting condition in a scene is one of the important factors affecting image quality. During the process of image acquisition, due to insufficient ambient light, the captured low-light images have characteristics such as low recognition, low brightness, low contrast, low resolution, and low signal-to-noise ratio, resulting in poor usability of low-light images and posing more severe challenges to subsequent image analysis and processing. Image enhancement is an important technology in image processing, which can improve the visual effect of images and lay a foundation for further tasks such as image recognition, image classification, and target detection.
[0003] Currently, common low-light image enhancement methods at home and abroad are mainly divided into four types: The first is the histogram equalization enhancement method. This algorithm performs contrast-limited enhancement on the grid areas in the image and performs interpolation processing on the original image, thus significantly improving the contrast of the image. The second is the Retinex enhancement method, such as the LIME enhancement method. This method searches for the maximum value in the RGB channels of the image to estimate the illumination of each pixel, and then uses the structural prior to reconstruct the illumination map. However, the generalization ability of the above two methods is poor, and for real low-light images, noise is often generated. The third is the pseudo-haze image enhancement method. This method enhances the inverted image of the low-light image using a dehazing algorithm. However, when dealing with complex scenes, this method is prone to noise and block effects. The fourth is the neural network-based enhancement method. For example, a method that combines the learned global features and local features and transforms them into a bilateral network, and adds local affine to guide the bilateral network to perform interpolation in terms of space and color depth. However, this network has poor effects in terms of colorization, dehazing, etc. because this method is based on learning paired supervision, and in real life, there are few accurately paired training sets. With the proposal of the generative adversarial network, image enhancement technology has developed by leaps and bounds. EnlightenGAN uses a dual discriminator to balance global and local low-light enhancement, eliminates the dependence on paired training data, and proposes a self-feature retention loss method to constrain the feature distance between the low-light input image and the enhanced image. It uses the illumination information of the low-light input as the self-regular attention map at each depth feature level to regularize unsupervised learning, and establishes an unpaired mapping between the low-light and normal-light image spaces without relying on accurately paired images.
[0004] Existing low-light image enhancement methods mainly rely on paired supervision based on learning. However, in real life, there are few accurately paired training sets, and common low-light image enhancement methods mainly operate at a single scale. Nevertheless, due to the defects of low resolution, low contrast, and low signal-to-noise ratio in low-light images, the enhancement accuracy at a single scale is not high. Summary of the Invention
[0005] In view of this, the purpose of the present invention is to provide a low-light image enhancement and denoising method that combines NSST domain, GAN, and scale correlation coefficient, which can lay a foundation for subsequent tasks such as image recognition, image classification, and target detection, and has a significant improvement both in terms of visual effect and objective evaluation indicators of image quality.
[0006] To achieve the above purpose, the present invention adopts the following technical solutions: A low-light image enhancement and denoising method that combines NSST domain, GAN, and scale correlation coefficient, including the following steps:
[0007] Step1: Collect a dataset of low-light images and normal-light images, convert the images from the RGB space to the HSV space, keep the H and S component values unchanged, perform NSST transformation on the luminance V component, and obtain 1 low-frequency subband image and k scale high-frequency subbands respectively. Each scale high-frequency subband is further decomposed into l directional subbands, and use the obtained low-frequency subband image to construct a training set.
[0008] Step2: Construct a low-frequency subband image enhancement model LF-EnlightenGAN based on GAN, and use the constructed training set of low-frequency subband images to train the LF-EnlightenGAN model to generate an enhanced model for low-frequency subband images.
[0009] Step3: Convert the low-light image to be processed from the RGB space to the HSV space, keep the H and S component values unchanged, perform NSST transformation on the luminance V component, and obtain 1 low-frequency subband image and k scale high-frequency subbands respectively. Each scale high-frequency subband is further decomposed into l directional subbands, and use the LF-EnlightenGAN enhancement model to enhance the low-frequency subband image, while retaining more texture details while improving the overall brightness, clarity, and information entropy.
[0010] Step4: Calculate the noise coefficient threshold and the scale correlation coefficient for each high-frequency directional subband coefficient, and remove the noise coefficient and enhance the edge coefficient.
[0011] Step5: Perform NSST reconstruction on the enhanced low-frequency subband and high-frequency subband images to obtain the enhanced V component, use it to replace the original V component, and finally restore the image from the HSV space to the RGB space to obtain the final enhanced and denoised image.
[0012] In a preferred embodiment: the image is subjected to k-level non-subsampled pyramid NSP multiscale decomposition to obtain 1 low-frequency image and k high-frequency images with different scales. The high-frequency images are then subjected to l-level multi-directional decomposition to obtain 2l + 2 directional sub-band images. The low-frequency image is denoised to retain the contour information and most of the energy information of the image, and the high-frequency sub-band images contain the edge, texture features, gradient information, and noise coefficient of the image.
[0013] In a preferred embodiment: after the image is decomposed by NSST, the low-frequency sub-band image contains the contour information and energy information of the image; collect the low-light image and normal-light image datasets for NSST multiscale decomposition, and construct a training set from the obtained low-light low-frequency sub-band images and normal-light low-frequency sub-band images, and construct a low-frequency sub-band image enhancement model LF-EnlightenGAN based on GAN. The LF-EnlightenGAN model includes the following modules:
[0014] (1) Self-regularized guided U-Net network
[0015] The LF-EnlightenGAN model uses a self-regularized guided U-Net network as the generator, which consists of 8 convolutional blocks in total. The U-Net network is used as the backbone of the generator, and a self-regularized attention map is added for regularization; for regularization, the input luminance image I is taken, normalized, and then 1 - I is used as the self-regularized attention map. Finally, the size of the attention map is adjusted to multiply with all the feature maps and the output image of the upsampling part of the U-Net.
[0016] (2) Global-local discriminator
[0017] The LF-EnlightenGAN model adopts a global-local discriminator structure; both the global and local discriminators use PatchGAN for true-false discrimination. Among them, the global discriminator uses a relativistic discriminator structure to estimate the probability that the real data is more real than the fake data, and guides the generator to synthesize a fake image that is more real than the real image. The LSGAN loss is used to replace the sigmoid function. Assume that C is the discriminator network, x r and x f represent the distributions of real data and fake data respectively, and σ represents the sigmoid activation function; D Ra (x r , x f ) and D Ra (x f , x r ) are the standard functions of the relativistic discriminator. Then, for the global discriminator, the loss function of the generator G is:
[0018]
[0019]
[0020]
[0021] The local discriminator randomly crops 5 local patches from the output image and the real image each time to learn to distinguish whether they are real or fake. Using the original LSGAN as the adversarial loss, for the local discriminator, the loss function of the generator G is defined as:
[0022]
[0023] (3) Self-feature preservation loss
[0024] The LF-EnlightenGAN model adopts self-feature preservation loss, uses the pre-trained VGG to model the feature space distance between images, and restricts the VGG feature distance between the input low-light image and its enhanced normal-light output image; assuming I L represents the input low-light image, G(I L ) represents the enhanced output of the generator, φ i,j represents the feature map extracted from the pre-trained VGG-16 model on ImageNet, i represents the i-th max pooling, j represents the j-th convolutional layer after the i-th max pooling, W i,j and H i,j are the dimensions of the extracted feature map, taking i = 5, j = 1; then the self-feature preservation loss L SFP is defined as:
[0025]
[0026] For the local discriminator, the local patches cropped from the input and output images are also regularized by the defined self-feature preservation loss ; Therefore, the overall loss function of this model is:
[0027]
[0028] In a preferred embodiment: assuming is the coefficient of the sub-band at (m, n), is the mean of the sub-band coefficients, L is the total number of sub-bands in the k-th scale direction, is the sub-band coefficient energy in the l-th direction of the k-th scale, and the noise threshold in the l-th direction of the k-th scale is defined as:
[0029]
[0030] Assuming is the product of coefficients at the (m,n) position for different scales, is for the coefficient energy of the sub-band in the l-th direction at the k-th scale, is the normalization process to facilitate subsequent coefficient comparison. Define the scale-related coefficient of (m,n) on the sub-band in the l-th direction at the k-th scale as:
[0031]
[0032] Adjust the edge coefficients greater than according to the enhancement function, where a is the control intensity, taken as 20 here, and b is the enhancement range, between [0,1]. Assume is the maximum coefficient of this sub-band. Define the enhancement function as:
[0033]
[0034]
[0035] Directly remove the noise coefficients less than and enhance the edge coefficients greater than When the coefficient is between combine the inter-scale correlation coefficients to enhance the weak edge coefficients and remove the noise coefficients, is the adjusted coefficient of the sub-band in the l-th direction at the k-th scale at the point (m,n), defined as:
[0036]
[0037] Compared with the prior art, the present invention has the following beneficial effects: The present invention proposes a low-light image enhancement and denoising method combining NSST domain, GAN and scale-related coefficients. Calculate the noise coefficient threshold and scale-related coefficient in the NSST domain to locate the noise and edge coefficients of the image, remove the noise while enhancing the edge coefficients; use the LF-EnlightenGAN enhancement model to enhance the low-frequency image, eliminating the limitation of requiring paired training data sets, improving the overall brightness, clarity and information entropy of the image while retaining more texture details, and there will be no overexposure phenomenon. Compared with several existing low-light image enhancement methods, the present invention has better anti-noise performance and edge enhancement ability, and has a greater improvement both in terms of visual effect and objective evaluation index of image quality, laying a foundation for subsequent tasks such as image recognition, image classification, and target detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 is a schematic diagram of the NSST multi-scale decomposition of the low-light image of the preferred embodiment of the present invention;
[0039] Figure 2 Schematic diagram of the LF-EnlightenGAN enhancement model structure for the preferred embodiment of the present invention;
[0040] Figure 3 Schematic diagram of the implementation process for low-light image enhancement and denoising in the preferred embodiment of the present invention;
[0041] Figure 4 Schematic diagram of the subjective visual comparison of different algorithms on the synthetic low-light image test set in the preferred embodiment of the present invention;
[0042] Figure 5 Schematic diagram of the comparison of denoising effects and edge detection effects of different algorithms on the synthetic low-light image test set in the preferred embodiment of the present invention;
[0043] Figure 6 Schematic diagram of the subjective visual comparison of different algorithms on real low-light images in the preferred embodiment of the present invention. Detailed implementation manners
[0044] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0045] It should be noted that the following detailed description is illustrative and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.
[0046] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary embodiments according to the present application; as used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0047] A low-light image enhancement and denoising method combining NSST domain, GAN and scale correlation coefficient. First, collect the low-light image and normal-light image datasets, convert the images from the RGB space to the HSV space, keep the values of the H and S components unchanged, perform NSST multi-scale decomposition on the luminance V component, and construct a training set using the low-pass subband images obtained from the decomposition. Secondly, construct a low-frequency subband image enhancement model LF-EnlightenGAN based on GAN, and train the model using the low-frequency subband image training set. Then, perform NSST decomposition on the low-light image to be processed, use the trained LF-EnlightenGAN model to enhance the low-frequency subband image, remove the noise from each high-frequency direction subband using the scale correlation coefficient first, and then enhance the edge coefficients through a non-linear gain function. Finally, perform NSST reconstruction on the processed high- and low-frequency subband images, restore the images to the RGB space, and obtain the enhanced and denoised images. The present invention has better anti-noise performance and edge enhancement ability, and has a great improvement both in terms of visual effect and objective evaluation index of image quality, laying a foundation for subsequent tasks such as image recognition, image classification, and object detection.
[0048] The detailed technical solution is as follows:
[0049] NSST multi-scale decomposition of low-light images
[0050] The HSV color space can well separate the chromaticity (H), saturation (S) and luminance (V) of the image, bringing great convenience to the enhancement of color images. Therefore, the input low-light images are converted from the RGB space to the HSV space for processing. Since the human visual system is more sensitive to luminance changes than to hue and saturation changes, the V component is extracted for NSST multi-scale decomposition, and the values of the H and S components are kept unchanged.
[0051] NSST decomposition includes two parts: multi-scale decomposition and multi-direction decomposition. As Figure 1 shown, perform k-level non-subsampled pyramid (NSP) multi-scale decomposition on the image to obtain 1 low-frequency image and k high-frequency images with different scales. The high-frequency images are then subjected to l-level multi-direction decomposition to obtain 2l + 2 directional subband images. The low-frequency image removes noise and retains the contour information and most of the energy information of the image. The high-frequency subband images contain the edge, texture features, gradient information and noise coefficients of the image.
[0052] Construct the LF-EnlightenGAN enhancement model for low-frequency subband images
[0053] After the image is decomposed by NSST, the low-frequency sub-band image mainly contains the contour information and most of the energy information of the image. The overall contrast and clarity of the low-illumination image can be improved by enhancing the contour details and brightness of the low-frequency sub-band image. The present invention collects weak-light image and normal-light image data sets for NSST multi-scale decomposition, constructs a training set from the obtained weak-light low-frequency sub-band images and normal-light low-frequency sub-band images, and constructs a low-frequency sub-band image enhancement model LF-EnlightenGAN based on GAN. The overall architecture of LF-EnlightenGAN is as shown in Figure 2 and mainly includes the following three modules:
[0054] (1) Self-regularized guided U-Net network
[0055] To make the enhancement of the dark area greater than that of the bright area, and the output image is neither overexposed nor underexposed, the model uses a self-regularized guided U-Net network as the generator, which consists of 8 convolutional blocks in total. The U-Net network is used as the backbone of the generator, and a self-regularized attention map is added for regularization. The regularization is achieved by taking the input luminance image I, normalizing it, then using 1 - I as the self-regularized attention map, and finally adjusting the size of the attention map to multiply with all the feature maps and the output image in the upsampling part of the U-Net.
[0056] (2) Global-local discriminator
[0057] To enhance the global illumination while adaptively enhancing the local area, the model adopts a global-local discriminator structure, which ensures that all local areas of the enhanced image look like real natural light, effectively avoiding local overexposure or underexposure.
[0058] Both the global and local discriminators use PatchGAN for real / fake discrimination. Among them, the global discriminator uses the relativistic discriminator structure to estimate the probability that real data is more real than fake data, and guides the generator to synthesize a pseudo-image that is more real than the real image. The LSGAN loss is used instead of the sigmoid function. Assuming C is the discriminator network, x r and x f represent the distributions of real data and fake data respectively, and σ represents the sigmoid activation function. D Ra (x r ,x f ) and D Ra (x f ,x r ) are the standard functions of the relativistic discriminator. Then, for the global discriminator, the loss function of the generator G is:
[0059]
[0060]
[0061]
[0062] The local discriminator randomly crops 5 local patches from the output image and the real image each time to learn to distinguish whether they are real or fake. Using the original LSGAN as the adversarial loss, for the local discriminator, the loss function of the generator G is defined as:
[0063]
[0064] (3) Self-feature preservation loss
[0065] To keep the content features of the image unchanged before and after enhancement, the model adopts self-feature preservation loss, uses the pre-trained VGG to model the feature space distance between images, and restricts the VGG feature distance between the input low-light image and its enhanced normal-light output image. Assume I L represents the input low-light image, G(I L ) represents the enhanced output of the generator, φ i,j represents the feature map extracted from the pre-trained VGG-16 model on ImageNet, i represents the i-th max pooling, j represents the j-th convolutional layer after the i-th max pooling, W i,j and H i,j are the dimensions of the extracted feature map, taking i = 5, j = 1. Then the self-feature preservation loss L SFP is defined as:
[0066]
[0067] For the local discriminator, the local patches cropped from the input and output images are also regularized by the defined self-feature preservation loss . Therefore, the overall loss function of the model is:
[0068]
[0069] High-frequency subband denoising and edge enhancement
[0070] To avoid the influence of noise on subsequent processing, noise must be removed before edge enhancement, combining the threshold and scale correlation coefficient calculated from the energy characteristics in the high-frequency domain. Assume is the coefficient of the subband at (m, n), is the mean of the subband coefficients, L is the total number of subbands in the k-th scale direction, is the energy of the subband coefficient in the l-th direction of the k-th scale, and the noise threshold in the l-th direction of the k-th scale is defined as:
[0071]
[0072] If the coefficients less than the threshold in the image are directly removed, it is easy to cause some weak edge coefficients to be eliminated as noise. If the coefficients greater than the threshold are directly enhanced, it is easy to cause a part of the noise to be enhanced as weak edge coefficients. Since after the image is decomposed by NSST, as the decomposition scale becomes finer and finer, it shows the characteristics of strong correlation of edge coefficients and weak correlation of noise coefficients. According to this characteristic, the noise coefficients with weak correlation can be further removed, and the edge coefficients with strong correlation can be enhanced. Assume is the product of the coefficients at the position (m,n) of different scales, is the coefficient energy of the sub-band in the l direction at the kth scale, is the normalization process to facilitate subsequent coefficient comparison. Define the scale correlation coefficient of (m,n) on the sub-band in the l direction at the kth scale as:
[0073]
[0074] Adjust the edge coefficients greater than according to the enhancement function, where a is the control intensity, here take 20, and b is the enhancement range, between [0,1]. Assume is the maximum coefficient of this sub-band. Define the enhancement function as:
[0075]
[0076]
[0077] Directly remove the noise coefficients less than , enhance the edge coefficients greater than . When the coefficient is between , combine the inter-scale correlation coefficient to enhance the weak edge coefficients and remove the noise coefficients. is the coefficient after adjustment at the point (m,n) of the sub-band in the l direction at the kth scale, defined as:
[0078]
[0079] Specific implementation process and steps
[0080] In summary, the implementation process of the low-light image enhancement and denoising method combining GAN and scale correlation coefficient in the NSST domain of the present invention is as Figure 3 shown, and the specific implementation steps are as follows:
[0081] Step1: Collect the low-light image and normal-light image datasets. Convert the images from the RGB space to the HSV space, keep the values of the H and S components unchanged, perform the NSST transform on the luminance V component, and obtain 1 low-frequency subband image and k scale high-frequency subbands respectively. Each scale high-frequency subband is further decomposed into l directional subbands. Use the obtained low-frequency subband image to construct a training set.
[0082] Step2: Construct a low-frequency subband image enhancement model LF-EnlightenGAN based on GAN. Use the constructed low-frequency subband image training set to train the LF-EnlightenGAN model to generate an enhanced model for low-frequency subband images.
[0083] Step3: Convert the low-illumination image to be processed from the RGB space to the HSV space, keep the values of the H and S components unchanged, perform the NSST transform on the luminance V component, and obtain 1 low-frequency subband image and k scale high-frequency subbands respectively. Each scale high-frequency subband is further decomposed into l directional subbands. Use the LF-EnlightenGAN enhancement model to enhance the low-frequency subband image, while retaining more texture details while improving the overall brightness, clarity, and information entropy.
[0084] Step4: For each high-frequency directional subband coefficient, calculate the noise coefficient threshold and the scale correlation coefficient Then combine Equation (9) and Equation (10) to remove the noise coefficient and enhance the edge coefficient.
[0085] Step5: Perform NSST reconstruction on the enhanced low-frequency subband and high-frequency subband images to obtain the enhanced V component. Use it to replace the original V component. Finally, restore the image from the HSV space to the RGB space to obtain the final enhanced and denoised image.
[0086] Specific embodiments and descriptions
[0087] To evaluate the enhancement effect of the method of the present invention on low-illumination images, compare the enhancement results of the present invention with those of common low-illumination image enhancement methods, including MSRCR, LIME, MSRNet, RetinexNet, DUAL, and EnlightenGAN. Conduct comparative experimental analyses using the synthetic low-illumination image test set and the real low-illumination image test set respectively to verify their performance.
[0088] 1. Comparative experiment on enhancing synthetic low-illumination images
[0089] Select an underwater image, a normal-light image, and a night image respectively as the synthetic low-illumination image test set, and enhance them using the method of the present invention and common low-illumination image enhancement methods. The enhancement results are as Figure 4As shown, structural similarity (SSIM) and mean squared error (MSE) are used as performance indicators to measure the test results of synthesized low-light images. The statistical results of SSIM and MSE for various methods are shown in Table 1: Although MSRCR, MSRNet, and RetinexNet have improved the illumination problem, the colors of the enhanced images are severely distorted, and noise and blurring effects appear; the brightness of the images enhanced by EnlightenGAN has been significantly improved, but the enhancement effect on underwater images is poor, and there are some artifacts in the enhanced images; although LIME performs well in terms of contrast and has a good enhancement effect on underwater images, there is regional blurring in the enhanced images. The enhanced image of the present invention is closest to the real image in terms of visual effect, and the objective evaluation indicators are the best among other methods, with a wide range of applications. Among them, for the enhancement of underwater images, the SSIM of the present invention has been increased by an average of 0.27, and the MSE has been reduced by an average of 2.74%; for the enhancement of normal light images, the SSIM of the present invention has been increased by an average of 0.17, and the MSE has been reduced by an average of 3.00%; for the enhancement of night images, the SSIM of the present invention has been increased by an average of 0.21, and the MSE has been reduced by an average of 4.11%; for the overall enhancement of the synthesized low-light image test set, the SSIM of the present invention has been increased by an average of 0.22, and the MSE has been reduced by an average of 3.29%.
[0090] To further objectively verify the anti-noise performance and edge enhancement effect of the present invention, Gaussian white noise with a mean of 0 and different variances was added to the synthesized low-light images for enhancement experiments. The canny operator was used to detect the edges of the enhanced low-light images, the PSNR was used to evaluate the noise reduction performance, and the continuous edge pixel ratio P was used to measure the edge enhancement effect. P is defined as:
[0091] P = γ / η (12)
[0092] Where γ is the total number of continuous edge pixels in the edge image, and η is the total number of pixels in the edge image. The larger P is, the better the continuity of the detected edges and the better the edge enhancement effect.
[0093] The enhancement results of the present invention were compared and analyzed with the results of common low-light image enhancement methods. The enhancement results and edge detection effects of each method are as Figure 5As shown, where the first row is the enhancement effect with a noise variance of 10%, the second row is the enhancement effect with a noise variance of 30%, and the third row is the edge detection effect of the image after enhancement with a noise variance of 10%. The statistical results of PSNR and P are shown in Table 2. Due to the influence of noise, the edges detected in the original low-light image are discontinuous and there are a large number of noise points. When the noise variance is 10%: the P value of the noisy image is 84.57%. The PSNR values of MSRCR, MSRNet, and RetinexNet are relatively low, and the noise reduction effect is poor. There are still a large number of speckles on the image. Although the edges detected after algorithm enhancement are relatively complete, many edge detail information is filtered out; although LIME and DUAL have better noise reduction performance than MSRCR, MSRNet, and RetinexNet, their P values are relatively low, the detected edges are incomplete, and there are a small number of noise points; EnlightenGAN has higher PSNR and P values than the above five algorithms, but there are artifacts in the enhanced image and a large number of noise points near the edges; the image enhanced by the present invention obtains the best PSNR value, has better noise reduction ability, the detected edges are relatively clear and complete, and there are fewer noise points, and has the best P value. When the noise variance is 30%: the P value of the noisy image is 68.89%. The enhancement performance of the other six algorithms has decreased significantly, but the PSNR value of the algorithm of the present invention remains at 20.9697, and the P value remains at 87.02%, having better anti-noise and edge enhancement abilities.
[0094] Table 1 Comparison of objective evaluation indexes of different algorithms on the synthetic low-light image test set
[0095]
[0096] Table 2 Denoising effect and edge detection effect of different algorithms on the synthetic low-light image test set
[0097]
[0098] 2. Real low-light image enhancement comparison experiment
[0099] To verify the enhancement effect of the present invention on real low-light images, 100 images were selected from the common low-light image databases SICE and DICM and the collected real underwater images to form a real low-light image test set, and were enhanced by the present invention and common low-light image enhancement methods. The enhancement results are as Figure 6As shown, the entropy, blind referenceless image spatial quality evaluator (BRISQUE), entropy-based no-reference image quality assessment (ENIQA), and hypernetwork-based no-reference image quality assessment (HyperIQA) are used to evaluate the quality of the test results of real low-light images. The statistical results of the objective evaluation indicators of various methods are shown in Table 3: The results of the MSRCR method are relatively smooth, but there are a large number of blocking effects and noises, and the visual and objective index results are poor; the results of the LIME method are more colorful, but the enhancement effect on dark areas is not good and the local color enhancement is excessive, and there is local overexposure in the enhancement results of images with uneven illumination; although the results of the MSRNet method improve the brightness, the color restoration is poor, and the enhancement effect on images with uneven illumination is not good; the results of the RetinexNet method have phenomena such as noise, blurring effect, and color deviation; the results of the DUAL method have insufficient brightness enhancement, especially in the backlight area with uneven illumination; the results of the EnlightenGAN method perform well in terms of brightness and color restoration and can well handle the enhancement of images with uneven illumination, but the effect in detail processing is not good; the results of the present invention are the best in all objective evaluation indicators except that HyperIQA is inferior to DUAL. From the visual effect, the present invention can effectively improve the brightness and contrast, has a good color enhancement effect, is superior to other algorithms in detail enhancement, can well handle the enhancement of underwater images and images with uneven illumination, and has a wide range of applications.
[0100] Table 3 Comparison of objective evaluation indicators of different algorithms on real low-light images
[0101]
[0102]
Claims
1. A low - illumination image enhancement and denoising method combining NSST domain, GAN and scale - related coefficients, characterized in that, It includes the following steps: Step1: Collect the low-light image and normal-light image datasets. Convert the images from the RGB space to the HSV space, keep the values of the H and S components unchanged, perform the NSST transform on the luminance V component, and obtain 1 low-frequency subband image and k scale high-frequency subbands respectively. Each scale high-frequency subband is further decomposed into l directional subbands, and use the obtained low-frequency subband image to construct the training set; Step2: Construct the low-frequency subband image enhancement model LF-EnlightenGAN based on GAN, and use the constructed low-frequency subband image training set to train the LF-EnlightenGAN model to generate the enhancement model of the low-frequency subband image; Step3: Convert the low-illumination image to be processed from the RGB space to the HSV space, keep the values of the H and S components unchanged, perform the NSST transform on the luminance V component, and obtain 1 low-frequency subband image and k scale high-frequency subbands respectively. Each scale high-frequency subband is further decomposed into l directional subbands, and use the LF-EnlightenGAN enhancement model to enhance the low-frequency subband image, while improving the overall brightness, clarity and information entropy and retaining more texture details; Step4: Calculate the noise coefficient threshold for each high-frequency direction subband coefficient and the scale correlation coefficient Remove the noise coefficient and enhance the edge coefficient; Step5: Perform NSST reconstruction on the enhanced low-frequency subband and high-frequency subbands to obtain the enhanced V component, use it to replace the original V component, and finally restore the image from the HSV space to the RGB space to obtain the final enhanced and denoised image.
2. The low - illumination image enhancement and denoising method combining NSST domain, GAN and scale - related coefficients according to claim 1, characterized in that: Perform k-level non-subsampled pyramid NSP multiscale decomposition on the image to obtain 1 low-frequency image and k high-frequency images with different scales. The high-frequency images are further decomposed in l-level multi-directions to obtain 2l + 2 directional subband images. The low-frequency image removes noise and retains the contour information and most of the energy information of the image. The high-frequency subbands contain the edge, texture features, gradient information and noise coefficients of the image.
3. The low - illumination image enhancement and denoising method combining NSST domain, GAN and scale - related coefficients according to claim 1, characterized in that: After the image is decomposed by NSST, the low-frequency subband image contains the contour information and energy information of the image; Collect the low-light image and normal-light image datasets for NSST multiscale decomposition, and construct the training set with the obtained low-light low-frequency subband image and normal-light low-frequency subband image. Construct the low-frequency subband image enhancement model LF-EnlightenGAN based on GAN. The LF-EnlightenGAN model includes the following modules: (1) Self-regularized guided U-Net network The LF-EnlightenGAN model uses the self-regularized guided U-Net network as the generator, which consists of 8 convolutional blocks in total. Take the U-Net network as the backbone of the generator and add the self-regularized attention map for regularization; for regularization, take the input luminance image I, normalize it, then use 1 - I as the self-regularized attention map, and finally adjust the size of the attention map to multiply with all the feature maps and the output image in the upsampling part of the U-Net; (2) Global-local discriminator The LF-EnlightenGAN model adopts a global-local discriminator structure; both the global and local discriminators use PatchGAN to distinguish between real and fake. Among them, the global discriminator uses the relativistic discriminator structure to estimate the probability that real data is more real than fake data, and guides the generator to synthesize pseudo-images that are more real than real images. The LSGAN loss is used to replace the sigmoid function. Assume that C is the discriminator network, x r and x f represent the distributions of real data and fake data respectively, and σ represents the sigmoid activation function; D Ra (x r ,x f ) and D Ra (x f ,x r ) are the standard functions of the relativistic discriminator. Then, for the global discriminator, the loss function of the generator G is: The local discriminator randomly crops 5 local patches from the output image and the real image each time to learn to distinguish whether they are real or fake. Use the original LSGAN as the adversarial loss. Then, for the local discriminator, the loss function of the generator G is defined as: (3) Self-Feature Preservation Loss The LF-EnlightenGAN model adopts a self-feature retention loss, uses a pre-trained VGG to model the feature space distance between images, and restricts the VGG feature distance between the input low-light image and its enhanced normal-light output image; assuming I L represents the input low-light image, G(I L ) represents the enhanced output of the generator, φ i,j represents the feature map extracted from the pre-trained VGG-16 model on ImageNet, i represents the i-th max pooling, j represents the j-th convolutional layer after the i-th max pooling, W i,j and H i,j are the dimensions of the extracted feature map, taking i = 5 and j = 1; then the self-feature retention loss L SFP is defined as: For the local discriminator, the local patches cropped from the input and output images are also regularized by the defined self-feature retention loss ; thus, the overall loss function of the model is:
4. The low - illumination image enhancement and denoising method combining NSST domain, GAN and scale - related coefficients according to claim 1, characterized in that: Assume is the coefficient of the sub - band at (m,n), is the average value of the sub - band coefficients, L is the total number of sub - bands in the k - th scale direction, is the energy of the sub - band coefficients in the l - th direction of the k - th scale. Define the noise threshold as follows: Hypothesis is the product of coefficients at the (m, n) position for different scales, is the coefficient energy of the sub-band in the l-th direction at the k-th scale, is the normalization process to facilitate subsequent coefficient comparison. Define the scale-dependent coefficient at (m, n) on the sub-band in the l-th direction at the k-th scale as: For edge coefficients greater than , they are adjusted according to the enhancement function, where a is the control intensity, taken as 20 here, and b is the enhancement range, between [0, 1]. Assuming is the maximum coefficient of this sub-band, the enhancement function is defined as: Directly remove noise coefficients less than , enhance edge coefficients greater than . When the coefficient is between , combine the inter-scale correlation coefficient to enhance weak edge coefficients and remove noise coefficients. is the adjusted coefficient of the sub-band in the l-th direction at the k-th scale at the point (m, n), defined as:
Citation Information
Patent Citations
NSST domain flotation froth image enhancement and denoising method based on quantum harmony search fuzzy set
CN110246106A
Foam infrared image segmentation method based on NSST saliency detection and image segmentation
CN110648342A