Low-light image enhancement method based on adaptive frequency decomposition and related equipment
By adopting an adaptive frequency decomposition network in the low-illumination image enhancement model, using the Laplace pyramid layer and the adaptive frequency decomposition layer, the problems of large training volume and complex frequency extraction of existing models are solved, and a more efficient low-illumination image enhancement effect is achieved.
Patent Information
- Application Number
- CN202210763940.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-29
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2042-06-29
AI Technical Summary
The existing frequency decomposition enhancement model has problems such as large training volume and complex frequency band extraction process in low-illumination image enhancement scenarios.
Adaptive frequency decomposition network is adopted to build a generative adversarial network through the Laplace pyramid layer and the adaptive frequency decomposition layer to perform low illumination enhancement of images.
The complexity of model training and the number of parameter adjustments is reduced, the effect of low illumination enhancement of image is improved, and the strategy of adaptive adjustment of receptive fields is realized.
Smart Images

Figure CN115063318B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision image technology, and in particular relates to a low-light image enhancement method with adaptive frequency decomposition and related equipment. Background Art
[0002] Due to unavoidable environmental or technical limitations, many images are often taken under undesirable lighting conditions. Such images often have problems such as overall darkness, high noise, and poor contrast. On the one hand, such images affect the visual effect, and on the other hand, they bring difficulties to the advanced visual processing of the computer in the later stage. An efficient low-light image algorithm can make up for the shortcomings of the equipment. By improving the imaging quality of the image through the algorithm, the viewing experience can be improved, and it can also provide preprocessing for subsequent advanced visual tasks, such as target recognition and target tracking. Therefore, studying low-light image enhancement algorithms is a task with practical needs and wide applications.
[0003] General low-light image enhancement methods have predictable quality problems. For example, while brightening the overall brightness and contrast of the image, the noise in the dark area of the image will be amplified, causing the dark details to be lost. The research on related algorithms initially had two directions: one algorithm is based on physical models, such as histogram equalization (HE), which mainly improves the contrast of the image by expanding the dynamic range of the entire image; the other algorithm is based on Retinex theory, which mainly filters out low-frequency information through single-scale SSR and leaves high-frequency information, thereby enhancing the edge information of the image. On this basis, multi-scale Retinex (MSR) and multi-scale Retinex with color restoration (MSRCR) methods have emerged. However, the above methods are limited to the way the image is output, and often some areas in the image are over-enhanced, causing the image to look unnatural. With the development of deep learning technology, some low-light image enhancement algorithms based on deep learning have also been proposed. The Low-Light Network (LLNet) proposed by Lore et al. built a deep network to enhance and denoise low-light images. However, the data set used by the network is a synthetic data set, which does not produce a good effect on images in real scenes; Shen et al. designed the traditional multi-scale Retinex (MSR) as a feedforward neural network with multiple Gaussian convolution feedforwards, and proposed MSR-Net following the MSR process to achieve end-to-end image enhancement. The above methods are relatively early and are all supervised methods, which makes the training process more complicated. Some researchers have adopted the unsupervised method of generative adversarial networks (GANs) to construct a network EnlightenGAN for low-light image enhancement. It does not require paired datasets, but only needs to provide unpaired low-light datasets and normal-light datasets to enable the network to learn the nonlinear mapping between low-light images and normal-light images, which can achieve good results both subjectively and objectively. Li et al. proposed a reference-free low-light image method Zero-DCE, which learns the mapping relationship between low-light images and curve parameters through a set of reference-free losses, and enhances image brightness and contrast in an iterative manner. Subsequently, Zero-DCE++ was proposed based on deep separable convolution, but the disadvantage is that the image output by its model is still not enough to achieve a high-contrast effect, and there are some noise points.
[0004] In a more cutting-edge study, the paper "Learning to Restore Low-Light Images via Decomposition and Enhancement" designed a frequency-based decomposition enhancement model for enhancing low-light images. In the first stage, it extracts low-frequency information for noise suppression and low-frequency layer information enhancement, and in the second stage, it extracts high-frequency information for detail enhancement. The problem is that this model requires a lot of experiments to determine the optimal parameters to control how much high and low frequency information of the receptive field is extracted, so it is not possible to achieve a good strategy for adaptively adjusting the receptive field, and the extraction of frequency band information in stages greatly increases the difficulty of model training. Summary of the invention
[0005] The embodiments of the present invention provide a low-light image enhancement method and related equipment based on adaptive frequency decomposition, aiming to solve the problems of large training amount and complex frequency band extraction process in the existing frequency decomposition enhancement model in the low-light image enhancement scenario.
[0006] In a first aspect, an embodiment of the present invention provides a low-light image enhancement method based on adaptive frequency decomposition, the method comprising the following steps:
[0007] S101, obtaining a LOL data set including a plurality of images of different brightness, and preprocessing the LOL data set to obtain a training data set and a test data set;
[0008] S102, constructing an adaptive frequency decomposition network including a Laplace pyramid layer, a feature extraction layer, and an adaptive frequency decomposition layer, wherein the feature extraction layer includes an encoding branch and a decoding branch; the adaptive frequency decomposition network is used as a generation network of a generative adversarial network structure, and a discriminator network corresponding to the adaptive frequency decomposition network is constructed, wherein the discriminator network includes a global discriminator and a local discriminator;
[0009] S103, introducing a generator loss function and a discriminator loss function, and using the training data set as the input of the adaptive frequency decomposition network and the discriminator network as a whole for training, until the training is completed and a low-light enhancement model is output, and then using the test data set as the input of the low-light enhancement model to perform low-light enhancement on the image, and calculate quantitative indicators.
[0010] Furthermore, the method for preprocessing the LOL dataset in step S101 includes at least one of normalization, random cropping and random horizontal flipping.
[0011] Furthermore, in the adaptive frequency decomposition network, the input image is processed by the Laplacian pyramid layer to obtain a Laplacian residual map, and the Laplacian residual map has shallow features and deep features, and the shallow features and the deep features satisfy the following expressions (1) and (2) respectively:
[0012] I k+1 =f↓(I k ) (1)
[0013] L k =I k -f↑(L k+1 ) (2)
[0014] Where k∈{1,2,3}, f↓() represents the downsampling of the bilinear difference, and f↑() represents the upsampling of the bilinear difference.
[0015] Furthermore, the adaptive frequency decomposition layer includes a low-frequency feature branch and a high-frequency feature branch. The encoding branch extracts features from the Laplace residual graph to obtain encoding features. The encoding features are defined as x en The low-frequency feature branch and the high-frequency feature branch respectively extract perceptual features from the coding features to obtain two sets of features with different receptive fields, and further combine the features of different receptive fields to obtain two sets of perceptual feature maps C a , the perceptual feature map C a The following relationship (3) is satisfied:
[0016]
[0017] Among them, i takes the value of 1 or 2, and uses f d1 () and f d2 () calculates two sets of features with different receptive fields respectively. When i takes the value of 1, f 1 d1 and f 1 d2 Both represent convolution operations with a kernel size of 3×3 and dilation rates of 1 and 6. When i takes the value of 2, f 2 d1 and f 2 d2 Both represent convolution operations with a kernel size of 3×3 and dilation rates of 1 and 12, and σ represents the linear activation function Leakyrelu;
[0018] The different perceptual feature maps are concatenated with the coding features in the channel dimension to obtain high-frequency features and low-frequency features, and the high-frequency features and the low-frequency features satisfy the following equations (4) and (5) respectively:
[0019]
[0020]
[0021] Furthermore, after the adaptive frequency decomposition layer obtains the high-frequency features and the low-frequency features, the high-frequency features and the low-frequency features are input into a SE attention mechanism to obtain a global vector.
[0022] Furthermore, the generator loss function is defined as L total , and the generator loss function satisfies the following expression (6):
[0023] L total =L content +L quelity +5×L mc +L tv (6)
[0024] Among them, L content is the content loss, which is composed of the reconstruction loss L rec and the perceptual loss L vgg Composition, L mc is the mutual consistency loss, L quelity is the perceptual quality index, L tv is the total variational loss, the perceptual quality index L quelity Satisfies the following expression (7):
[0025] L quelity =L Gg +L Gl
[0026]
[0027]
[0028] In expression (7), L Gg and L Gl Denote the global adversarial loss and local adversarial loss of the generative adversarial network, respectively. g represents the global discriminator, D l represents the local discriminator, E() is the mean calculation, x r and x f is a preset data sample;
[0029] The discriminator loss satisfies the following expression (8):
[0030] L D =L Dg +L Dl
[0031]
[0032]
[0033] Furthermore, when the adaptive frequency decomposition network and the discriminator network are trained as a whole, Adam is used as the optimizer, and the number of training rounds is 200 rounds, wherein the training learning rate is set to le-4 in the first 100 rounds, and the training learning rate is linearly decayed to 0 in the next 100 rounds.
[0034] In a second aspect, an embodiment of the present invention further provides a low-light image enhancement system with adaptive frequency decomposition, comprising:
[0035] A data acquisition module is used to acquire a LOL data set containing multiple images of different brightness, and preprocess the LOL data set to obtain a training data set and a test data set;
[0036] A network construction module is used to construct an adaptive frequency decomposition network including a Laplace pyramid layer, a feature extraction layer, and an adaptive frequency decomposition layer, wherein the feature extraction layer includes an encoding branch and a decoding branch; the adaptive frequency decomposition network is used as a generation network of a generative adversarial network structure, and a discriminator network corresponding to the adaptive frequency decomposition network is constructed, wherein the discriminator network includes a global discriminator and a local discriminator;
[0037] The network training module is used to introduce a generator loss function and a discriminator loss function, and use the training data set as the input of the adaptive frequency decomposition network and the discriminator network as a whole for training until the training is completed and a low-light enhancement model is output. After that, the test data set is used as the input of the low-light enhancement model to perform low-light enhancement on the image and calculate quantitative indicators.
[0038] In a third aspect, an embodiment of the present invention further provides a computer device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps in the low-light image enhancement method with adaptive frequency decomposition as described in any one of the above embodiments are implemented.
[0039] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the low-light image enhancement method with adaptive frequency decomposition as described in any one of the above embodiments are implemented.
[0040] The beneficial effect achieved by the present invention is that, due to the use of Laplace pyramid branches and adaptive frequency decomposition modules in the low-light enhancement network, the potential information of the image can be mined to the greatest extent, and at the same time, multiple experiments are not required to determine the model parameters, the amount of training is reduced, and the low-light enhancement effect of the image is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 is a flowchart of the steps of the low-light image enhancement method using adaptive frequency decomposition provided by an embodiment of the present invention;
[0042] Figure 2 is a schematic diagram of a framework of an adaptive frequency decomposition network provided by an embodiment of the present invention;
[0043] Figure 3 is a schematic diagram of the structure of an adaptive frequency decomposition layer provided by an embodiment of the present invention;
[0044] Figure 4 It is a schematic diagram of the training data flow of the adaptive frequency decomposition network provided by an embodiment of the present invention;
[0045] Figure 5 is a structural schematic diagram of a low-light image enhancement system 200 with adaptive frequency decomposition provided by an embodiment of the present invention;
[0046] Figure 6 It is a schematic diagram of the structure of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0047] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0048] Please refer to Figure 1 , Figure 1 : is a flowchart of the steps of a method for low-light image enhancement by adaptive frequency decomposition provided by an embodiment of the present invention, the method comprising the following steps:
[0049] S101, obtaining a LOL dataset including multiple images of different brightness, and preprocessing the LOL dataset to obtain a training dataset and a test dataset.
[0050] Specifically, the method for preprocessing the LOL dataset in step S101 includes at least one of normalization, random cropping and random horizontal flipping. The LOL (Low-Light Enhancement) dataset is an open source dataset, which includes 500 low-brightness and high-brightness paired data. The size of each image is 400×600, and the image format is PNG. Exemplarily, in an embodiment of the present invention, after preprocessing the LOLO dataset, 485 paired low-illumination paired images and normal-illumination paired images are obtained, and the two images are equally divided into the training dataset and the test dataset.
[0051] S102, constructing an adaptive frequency decomposition network including a Laplace pyramid layer, a feature extraction layer, and an adaptive frequency decomposition layer, wherein the feature extraction layer includes an encoding branch and a decoding branch; using the adaptive frequency decomposition network as a generation network of a generative adversarial network structure, and constructing a discriminator network corresponding to the adaptive frequency decomposition network, wherein the discriminator network includes a global discriminator and a local discriminator.
[0052] For details, please refer to Figure 2 , Figure 2 : is a schematic diagram of the framework of the adaptive frequency decomposition network provided in an embodiment of the present invention. The adaptive frequency decomposition network used in the embodiment of the present invention is based on U-Net. U-Net is a semantic segmentation deep model. On the basis of the existing U-Net model, the embodiment of the present invention uses it as a feature extraction layer, and additionally adds the Laplacian pyramid layer between its input layer and the encoding branch, so that the image obtains more obvious hierarchical features.
[0053] Furthermore, in the adaptive frequency decomposition network, the input image is processed by the Laplacian pyramid layer to obtain a Laplacian residual map, and the Laplacian residual map has shallow features and deep features, and the shallow features and the deep features satisfy the following expressions (1) and (2) respectively:
[0054] I k+1 =f↓(I k ) (1)
[0055] L k =I k -f↑(L k+1 ) (2)
[0056] Wherein, k∈{1,2,3}, f↓() represents the downsampling of the bilinear difference, and f↑() represents the upsampling of the bilinear difference. In the embodiment of the present invention, L4 is equal to I4, which makes the original image upsampled or downsampled to output an image of 16 times the size.
[0057] The convolution kernel size used in the feature extraction layer used in the embodiment of the present invention is 3×3, and the encoding branch uses the convolution kernel to extract features from the Laplacian residual map. Figure 3 , Figure 3 is a schematic diagram of the structure of an adaptive frequency decomposition layer provided in an embodiment of the present invention, wherein the adaptive frequency decomposition layer comprises a low-frequency feature branch and a high-frequency feature branch, wherein the encoding branch extracts features from the Laplace residual graph to obtain encoding features, and the encoding features are defined as x en The low-frequency feature branch and the high-frequency feature branch respectively extract perceptual features from the coding features to obtain two sets of features with different receptive fields, and further combine the features of different receptive fields to obtain two sets of perceptual feature maps C a , the perceptual feature map C a The following relationship (3) is satisfied:
[0058]
[0059] Among them, i takes the value of 1 or 2, and uses f d1 () and f d2 () calculates two sets of features with different receptive fields respectively. When i takes the value of 1, f 1 d1 and f 1 d2 Both represent convolution operations with a kernel size of 3×3 and dilation rates of 1 and 6. When i takes the value of 2, f 2 d1 and f 2 d2 Both represent convolution operations with a kernel size of 3×3 and dilation rates of 1 and 12, and σ represents the linear activation function Leakyrelu;
[0060] The different perceptual feature maps are concatenated with the coding features in the channel dimension to obtain high-frequency features and low-frequency features, and the high-frequency features and the low-frequency features satisfy the following equations (4) and (5) respectively:
[0061]
[0062]
[0063] Specifically, the perceptual feature map C a It is the pixel contrast information. The difference between the high-frequency and low-frequency perceptual feature maps is their different contrasts. a With the encoding feature x en High-frequency information can be extracted using 1-C aThe low-frequency information is also extracted by extracting frequency-based perceptual features of different scales in a self-driven manner. When splicing in the channel dimension, (1-C a ) extracts low-frequency information and uses C a To extract high-frequency information, so that when it is finally applied to low-light image enhancement, the low-frequency content of the image can be enhanced and noise suppressed at a low scale, and the high-frequency content of the image can be restored in detail at a high scale.
[0064] Furthermore, after the adaptive frequency decomposition layer obtains the high-frequency features and the low-frequency features, the high-frequency features and the low-frequency features are input into an SE attention mechanism to obtain a global vector, and the global vector is finally weighted multiplied with the image originally input into the adaptive frequency decomposition network to reflect the importance of different channels in the image. In an embodiment of the present invention, by controlling the weights during weighting, the performance of the receptive field during image enhancement can be adjusted. Exemplarily, an embodiment of the present invention sets a weight template for the weights of different branches to achieve a network training strategy for adaptively selecting the image receptive field.
[0065] Combination Figure 2 and Figure 3 The adaptive frequency decomposition layer provided in the embodiment of the present invention extracts the perceptual feature map C by dilated convolutions of different branches and different convolution rates. a After subtraction, the features are multiplied with the input features to obtain frequency-based features. The features of the two branches are then concatenated with the upsampled features of the decoding branch by channel, and then a global vector is obtained through an SE module to adaptively weight all channels. Finally, the decoding branch outputs the upsampled and restored residual image. In an embodiment of the present invention, the residual image output by the feature extraction layer is multiplied by the original input and a learnable parameter α to obtain the final enhanced result of the image. Exemplarily, the learnable parameter α used in the embodiment of the present invention is initialized to 1, and its requires_grad attribute is set to True, and its parameter value is saved in the final result of the network training.
[0066] Please refer to Figure 4 , Figure 4It is a schematic diagram of the training data flow of the adaptive frequency decomposition network provided in an embodiment of the present invention. The adaptive frequency decomposition network in the embodiment of the present invention includes a structure of a generative adversarial network, which is used to improve the final image visual effect. Exemplarily, the global discriminator used in the embodiment of the present invention is a full convolutional network composed of 7 convolutional layers, and the local discriminator is a full convolutional network composed of 6 convolutional layers. As a discriminator, the output channel of its discrimination result is 1, which is used to discriminate whether the image generated by the global or local generator is an image of normal brightness or an image enhanced by low illumination.
[0067] S103, introducing a generator loss function and a discriminator loss function, and using the training data set as the input of the adaptive frequency decomposition network and the discriminator network as a whole for training, until the training is completed and a low-light enhancement model is output, and then using the test data set as the input of the low-light enhancement model to perform low-light enhancement on the image, and calculate quantitative indicators.
[0068] Furthermore, based on the data set used in the embodiment of the present invention and the structure of the generative adversarial network, the generator loss function is defined as L total , and the generator loss function satisfies the following expression (6):
[0069] L total =L content +L quelity +5×L mc +L tv (6)
[0070] Among them, L content is the content loss, which is composed of the reconstruction loss L rec and the perceptual loss L vgg Composition, perceptual loss L vgg It is used to calculate the VGG feature distance between the enhanced image and the reference image to encourage the enhanced image feature to be as close to the reference image as possible. In order to restore the details of the local area of the image, the embodiment of the present invention randomly extracts five local areas of size 3232 in the image to calculate the perceptual loss, thereby constraining the network to learn local information. The content loss L content Satisfies the following expression:
[0071] L content =L rec +L vgg
[0072]
[0073]
[0074] Among them, I lowis the image after low illumination enhancement, I normal is the reference image, is the local area of the image after low illumination enhancement, is the local region of the reference image, It is a feature map with depth i and width j extracted by the VGG-16 model pre-trained on ImageNet.
[0075] L quelity is a perceptual quality indicator, wherein the perceptual quality indicator L quelity Satisfies the following expression (7):
[0076] L quelity =L Gg +L Gl
[0077]
[0078]
[0079] In expression (7), L Gg and L Gl Denote the global adversarial loss and local adversarial loss of the generative adversarial network, respectively. g represents the global discriminator, D l represents the local discriminator, E() is the mean calculation, x r and x f is a preset data sample;
[0080] L mc is the mutual consistency loss, the mutual consistency loss L mc satisfy:
[0081] L mc =||M*exp(-c*M)||1
[0082]
[0083] Among them, c is the penalty factor, which is a parameter used to control the shape of the function. The smaller the penalty factor c is, the more significant the proportional relationship between M and L is. The larger c is, the stronger the nonlinearity is.
[0084] L tv is the total variational loss, the total variational loss L tv satisfy:
[0085]
[0086] in, is the gradient of the image on the x-axis after low illumination enhancement, is the gradient of the image on the y-axis after low-light enhancement, and N is the batch size;
[0087] The discriminator loss satisfies the following expression (8):
[0088] L D =L Dg +L Dl
[0089]
[0090]
[0091] Furthermore, when the adaptive frequency decomposition network and the discriminator network are trained as a whole, Adam is used as the optimizer, and the number of training rounds is 200 rounds, wherein the training learning rate is set to le-4 in the first 100 rounds, and the training learning rate is linearly decayed to 0 in the next 100 rounds.
[0092] Exemplarily, the embodiment of the present invention calculates quantitative indicators for the low-light enhancement model obtained after training, and compares it with various existing low-light enhancement neural network models, specifically including: LIME, MBLLEN, Retinex-Net, Zero-DCE, EnlightenGAN, Kind and Kind++. The quantitative indicators calculated by the embodiment of the present invention include: MAE (mean average error), MSE (root mean square error), PSNR (peak signal-to-noise ratio), SSIM (structural similarity), AB (brightness mean), LPIPS (learning perceptual image block similarity), NIQE (natural image quality). In order to obtain image enhancement comparison results from different images, the embodiment of the present invention performs image enhancement comparison on the following five public natural low-light image data sets: DICM, Fusion, LIME, low, MEF, NPE. Specifically, the indicator results of the low-light enhancement model provided by the embodiment of the present invention and the existing models in the above environment are shown in Table 1 below.
[0093] Table 1 Index results of low-light enhancement model and existing models in the above environments
[0094]
[0095] Since the above data set used in the comparison of the embodiment of the present invention has no paired reference image, a reference-free evaluation index NIQE is used when comparing the indicators with the low-light enhancement model of the embodiment of the present invention. The smaller the NIQE, the more natural the image is and the closer it is to the real light image distribution. The index comparison results of the low-light enhancement model provided by the embodiment of the present invention and the existing model in the above environment are shown in Table 2 below. It can be seen that the indicators of our method on the data set are better than other methods, which proves the effectiveness of the method proposed in the present invention. The index comparison results of the low-light enhancement model provided by the embodiment of the present invention and the existing model in the above environment are shown in Table 2 below.
[0096] Table 2 Comparison results of the low-light enhancement model and the existing model in the above environment
[0097] NIQE(↓) DCIM Fusion LIME low MEF MBLLEN 3.6940 4.7166 4.6265 3.9725 4.5147 Retinex-Net 4.4972 4.3378 4.8011 4.0007 5.6886 Kind 3.8612 4.1223 4.3540 3.6267 4.6410 Kind++ 3.1143 3.7137 5.0014 3.3849 4.1043 Ours 3.0008 3.6647 4.1820 3.3429 3.3321
[0098] Based on the above data, it can be seen that the low-light enhancement model provided by the embodiment of the present invention is superior to other neural network models in indicators on the comparison data set.
[0099] It should be noted that the low-light enhancement model provided by the embodiment of the present invention uses U-Net as the underlying network for feature extraction when it is constructed, but the structure of the underlying network itself does not limit the use of the Laplacian pyramid layer and the adaptive frequency decomposition layer added in the embodiment of the present invention. Exemplarily, the structure of the Laplacian pyramid layer and the adaptive frequency decomposition layer provided by the embodiment of the present invention can also be applied to network structures for feature extraction such as ResNet, DenseNet, MobileNets, etc. At the same time, it can also be applied to neural network models for image restoration and image segmentation, and similar technical effects can be obtained.
[0100] The beneficial effect achieved by the present invention is that, due to the use of Laplace pyramid branches and adaptive frequency decomposition modules in the low-light enhancement network, the potential information of the image can be mined to the greatest extent, and at the same time, multiple experiments are not required to determine the model parameters, the amount of training is reduced, and the low-light enhancement effect of the image is improved.
[0101] The embodiment of the present invention also provides a low-light image enhancement system with adaptive frequency decomposition, please refer to Figure 5 , Figure 5 is a structural schematic diagram of a low-light image enhancement system 200 of adaptive frequency decomposition provided by an embodiment of the present invention. The low-light image enhancement system 200 of adaptive frequency decomposition includes:
[0102] The data acquisition module 201 is used to acquire a LOL data set including a plurality of images of different brightness, and pre-process the LOL data set to obtain a training data set and a test data set;
[0103] A network construction module 202 is used to construct an adaptive frequency decomposition network including a Laplace pyramid layer, a feature extraction layer, and an adaptive frequency decomposition layer, wherein the feature extraction layer includes an encoding branch and a decoding branch; use the adaptive frequency decomposition network as a generation network of a generative adversarial network structure, and construct a discriminator network corresponding to the adaptive frequency decomposition network, wherein the discriminator network includes a global discriminator and a local discriminator;
[0104] The network training module 203 is used to introduce the generator loss function and the discriminator loss function, and use the training data set as the input of the adaptive frequency decomposition network and the discriminator network as a whole for training until the training is completed and the low-light enhancement model is output. After that, the test data set is used as the input of the low-light enhancement model to perform low-light enhancement on the image and calculate quantitative indicators.
[0105] The adaptive frequency decomposition low-light image enhancement system 200 can implement the steps in the adaptive frequency decomposition low-light image enhancement method in the above embodiment, and can achieve the same technical effects. Please refer to the description in the above embodiment and will not be repeated here.
[0106] The embodiment of the present invention also provides a computer device, please refer to Figure 6 , Figure 6 300 is a schematic diagram of the structure of a computer device provided in an embodiment of the present invention. The computer device 300 includes: a memory 302, a processor 301, and a computer program stored in the memory 302 and executable on the processor 301.
[0107] The processor 301 calls the computer program stored in the memory 302 to execute the steps of the low-light image enhancement method of adaptive frequency decomposition provided by the embodiment of the present invention. Figure 1 , specifically including:
[0108] S101, obtaining a LOL dataset including multiple images of different brightness, and preprocessing the LOL dataset to obtain a training dataset and a test dataset.
[0109] Furthermore, the method for preprocessing the LOL dataset in step S101 includes at least one of normalization, random cropping and random horizontal flipping.
[0110] S102, constructing an adaptive frequency decomposition network including a Laplace pyramid layer, a feature extraction layer, and an adaptive frequency decomposition layer, wherein the feature extraction layer includes an encoding branch and a decoding branch; using the adaptive frequency decomposition network as a generation network of a generative adversarial network structure, and constructing a discriminator network corresponding to the adaptive frequency decomposition network, wherein the discriminator network includes a global discriminator and a local discriminator.
[0111] Furthermore, in the adaptive frequency decomposition network, the input image is processed by the Laplacian pyramid layer to obtain a Laplacian residual map, and the Laplacian residual map has shallow features and deep features, and the shallow features and the deep features satisfy the following expressions (1) and (2) respectively:
[0112] I k+1 =f↓(I k ) (1)
[0113] L k =I k -f↑(L k+1 ) (2)
[0114] Where k∈{1,2,3}, f↓() represents the downsampling of the bilinear difference, and f↑() represents the upsampling of the bilinear difference.
[0115] Furthermore, the adaptive frequency decomposition layer includes a low-frequency feature branch and a high-frequency feature branch. The encoding branch extracts features from the Laplace residual graph to obtain encoding features. The encoding features are defined as x en The low-frequency feature branch and the high-frequency feature branch respectively extract perceptual features from the coding features to obtain two sets of features with different receptive fields, and further combine the features of different receptive fields to obtain two sets of perceptual feature maps C a , the perceptual feature map C a The following relationship (3) is satisfied:
[0116]
[0117] Among them, i takes the value of 1 or 2, and uses f d1 () and f d2 () calculates two sets of features with different receptive fields respectively. When i takes the value of 1, f 1 d1 and f 1 d2 Both represent convolution operations with a kernel size of 3×3 and dilation rates of 1 and 6. When i takes the value of 2, f 2 d1 and f 2 d2Both represent convolution operations with a kernel size of 3×3 and dilation rates of 1 and 12, and σ represents the linear activation function Leakyrelu;
[0118] The different perceptual feature maps are concatenated with the coding features in the channel dimension to obtain high-frequency features and low-frequency features, and the high-frequency features and the low-frequency features satisfy the following equations (4) and (5) respectively:
[0119]
[0120]
[0121] Furthermore, after the adaptive frequency decomposition layer obtains the high-frequency features and the low-frequency features, the high-frequency features and the low-frequency features are input into a SE attention mechanism to obtain a global vector.
[0122] S103, introducing a generator loss function and a discriminator loss function, and using the training data set as the input of the adaptive frequency decomposition network and the discriminator network as a whole for training, until the training is completed and a low-light enhancement model is output, and then using the test data set as the input of the low-light enhancement model to perform low-light enhancement on the image, and calculate quantitative indicators.
[0123] Furthermore, the generator loss function is defined as L total , and the generator loss function satisfies the following expression (6):
[0124] L total =L content +L quelity +5×L mc +L tv (6)
[0125] Among them, L content is the content loss, which is composed of the reconstruction loss L rec and the perceptual loss L vgg Composition, L mc is the mutual consistency loss, L quelity is the perceptual quality index, L tv is the total variational loss, the perceptual quality index L quelity Satisfies the following expression (7):
[0126] L quelity =L Gg +L Gl
[0127]
[0128]
[0129] In expression (7), L Gg and L Gl Denote the global adversarial loss and local adversarial loss of the generative adversarial network, respectively. g Denotes the global discriminator, D l represents the local discriminator, E() is the mean calculation, x r and x f is a preset data sample;
[0130] The discriminator loss satisfies the following expression (8):
[0131] L D =L Dg +L Dl
[0132]
[0133]
[0134] Furthermore, when the adaptive frequency decomposition network and the discriminator network are trained as a whole, Adam is used as the optimizer, and the number of training rounds is 200 rounds, wherein the training learning rate is set to le-4 in the first 100 rounds, and the training learning rate is linearly decayed to 0 in the next 100 rounds.
[0135] The computer device 300 provided in the embodiment of the present invention can implement the steps in the low-light image enhancement method of adaptive frequency decomposition in the above embodiment, and can achieve the same technical effect. Please refer to the description in the above embodiment and will not be repeated here.
[0136] An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the various processes and steps in the low-light image enhancement method with adaptive frequency decomposition provided in the embodiment of the present invention are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0137] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes of the embodiments of the above-mentioned methods. The storage medium can be a disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM).
[0138] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.
[0139] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, a magnetic disk, or an optical disk), and includes a number of instructions for enabling a terminal (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in each embodiment of the present invention.
[0140] The embodiments of the present invention are described above in conjunction with the accompanying drawings. What is disclosed is only the preferred embodiment of the present invention. However, the present invention is not limited to the above-mentioned specific implementation manner. The above-mentioned specific implementation manner is only illustrative rather than restrictive. Under the enlightenment of the present invention, ordinary technicians in this field can also make many forms and equivalent changes without departing from the scope of protection of the purpose of the present invention and the claims, all of which are within the protection of the present invention.
Claims
1. A low-light image enhancement method based on adaptive frequency decomposition, characterized in that: The method comprises the following steps: S101, obtaining a LOL data set including a plurality of images of different brightness, and preprocessing the LOL data set to obtain a training data set and a test data set; S102, constructing an adaptive frequency decomposition network including a Laplace pyramid layer, a feature extraction layer, and an adaptive frequency decomposition layer, wherein the feature extraction layer includes an encoding branch and a decoding branch; using the adaptive frequency decomposition network as a generation network of a generative adversarial network structure, and constructing a discriminator network corresponding to the adaptive frequency decomposition network, wherein the discriminator network includes a global discriminator and a local discriminator; In the adaptive frequency decomposition network, the input image is processed by the Laplacian pyramid layer to obtain a Laplacian residual map, and the Laplacian residual map has shallow features. I k and deep features L k , the shallow features and the deep features satisfy the following expressions (1) and (2) respectively: (1) (2) in, k ϵ{1,2,3}, f↓() represents the downsampling of the bilinear interpolation, f↑() represents the upsampling of bilinear interpolation; The adaptive frequency decomposition layer includes a low-frequency feature branch and a high-frequency feature branch. The encoding branch extracts features from the Laplace residual graph to obtain encoding features. The encoding features are defined as x en The low-frequency feature branch and the high-frequency feature branch respectively extract perceptual features from the coding features to obtain two sets of features with different receptive fields, and further combine the features of different receptive fields to obtain two sets of perceptual feature maps C a , the perceptual feature map C a The following relationship (3) is satisfied: (3) in ,i Take the value 1 or 2 and use f d1 () and f d2 () Calculate two sets of features with different receptive fields respectively. i When the value is 1, and Both represent convolution operations with a kernel size of 3×3 and dilation rates of 1 and 6. i When the value is 2, and Both represent convolution operations with a kernel size of 3×3 and dilation rates of 1 and 12. σ Represents the linear activation function Leakyrelu; The different perceptual feature maps are concatenated with the coding features in the channel dimension to obtain high-frequency features and low-frequency features, which respectively satisfy the following equations (4) and (5): (4) (5); After the adaptive frequency decomposition layer obtains the high-frequency features and the low-frequency features, the high-frequency features and the low-frequency features are input into an SE attention mechanism to obtain a global vector; S103, introducing a generator loss function and a discriminator loss function, and using the training data set as the input of the adaptive frequency decomposition network and the discriminator network as a whole for training, until the training is completed and a low-light enhancement model is output, and then using the test data set as the input of the low-light enhancement model to perform low-light enhancement on the image, and calculate quantitative indicators.
2. The method for low-light image enhancement based on adaptive frequency decomposition as claimed in claim 1, characterized in that: The method for preprocessing the LOL dataset in step S101 includes at least one of normalization, random cropping and random horizontal flipping.
3. The method for low-light image enhancement based on adaptive frequency decomposition as claimed in claim 1, characterized in that: Define the generator loss function as L total , and the generator loss function satisfies the following expression (6): L total = L content + L quelity + 5×L mc + L tv (6) in, L content is the content loss, which is composed of the reconstruction loss L rec and perceived loss L vgg composition, L mc is the mutual consistency loss, L quelity is the perceived quality indicator, L tv is the total variational loss, the perceptual quality index L quelity Satisfies the following expression (7): = + = = (7) In expression (7), L Gg and L Gl Represent the global adversarial loss and local adversarial loss of the generative adversarial network, D g represents the global discriminator, D l represents the local discriminator, E() is the mean calculation, x r and x f is the preset data sample, is the image after low-light enhancement. I normal is the reference image, It is the local area of the image after low illumination enhancement; The discriminator loss satisfies the following expression (8): L D = L Dg + L Dl = (8); In expression (8), is a local region of the reference image.
4. The method for low-light image enhancement based on adaptive frequency decomposition as claimed in claim 1, characterized in that: When the adaptive frequency decomposition network and the discriminator network are trained as a whole, Adam is used as the optimizer, and the number of training rounds is 200 rounds, wherein the training learning rate is set to le-4 in the first 100 rounds, and the training learning rate is linearly decayed to 0 in the next 100 rounds.
5. An adaptive frequency decomposition low-light image enhancement system, characterized in that: include: A data acquisition module is used to acquire a LOL data set containing multiple images of different brightness, and preprocess the LOL data set to obtain a training data set and a test data set; A network construction module, used to construct an adaptive frequency decomposition network including a Laplace pyramid layer, a feature extraction layer, and an adaptive frequency decomposition layer, wherein the feature extraction layer includes an encoding branch and a decoding branch; Using the adaptive frequency decomposition network as a generative network of a generative adversarial network structure, and constructing a discriminator network corresponding to the adaptive frequency decomposition network, wherein the discriminator network includes a global discriminator and a local discriminator; In the adaptive frequency decomposition network, the input image is processed by the Laplacian pyramid layer to obtain a Laplacian residual map, and the Laplacian residual map has shallow features. I k and deep features L k , the shallow features and the deep features satisfy the following expressions (1) and (2) respectively: (1) (2) in, k ϵ{1,2,3}, f↓() represents the downsampling of the bilinear interpolation, f↑() represents the upsampling of bilinear interpolation; The adaptive frequency decomposition layer includes a low-frequency feature branch and a high-frequency feature branch. The encoding branch extracts features from the Laplace residual graph to obtain encoding features. The encoding features are defined as x en The low-frequency feature branch and the high-frequency feature branch respectively extract perceptual features from the coding features to obtain two sets of features with different receptive fields, and further combine the features of different receptive fields to obtain two sets of perceptual feature maps C a , the perceptual feature map C a The following relationship (3) is satisfied: (3) in ,i Take the value 1 or 2 and use f d1 () and f d2 () Calculate two sets of features with different receptive fields respectively. i When the value is 1, and Both represent convolution operations with a kernel size of 3×3 and dilation rates of 1 and 6. i When the value is 2, and Both represent convolution operations with a kernel size of 3×3 and dilation rates of 1 and 12. σ Represents the linear activation function Leakyrelu; The different perceptual feature maps are concatenated with the coding features in the channel dimension to obtain high-frequency features and low-frequency features, which respectively satisfy the following equations (4) and (5): (4) (5); After the adaptive frequency decomposition layer obtains the high-frequency features and the low-frequency features, the high-frequency features and the low-frequency features are input into an SE attention mechanism to obtain a global vector; The network training module is used to introduce a generator loss function and a discriminator loss function, and use the training data set as the input of the adaptive frequency decomposition network and the discriminator network as a whole for training until the training is completed and a low-light enhancement model is output. After that, the test data set is used as the input of the low-light enhancement model to perform low-light enhancement on the image and calculate quantitative indicators.
6. A computer device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps in the low-light image enhancement method with adaptive frequency decomposition as described in any one of claims 1 to 4 are implemented.
7. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the low-light image enhancement method using adaptive frequency decomposition as claimed in any one of claims 1 to 4 are implemented.
Citation Information
Patent Citations
No-reference low-illumination image enhancement method and system based on generative adversarial network
CN111798400A
Photographing method based on low-illumination image enhancement algorithm of brightness attention mechanism
CN111915526A