Image Enhancement Method Based on Improved Multi-Scale Fusion Generative Adversarial Network

Through the improved multi-scale fusion generation adversarial network image enhancement method, the problem of single light source and poor effect of data sets in low-illumination image enhancement is solved, image brightness improvement and detail retention are achieved, and image quality is improved.

CN115223004BActive Publication Date: 2025-06-10CHANGAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210692241.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-17
Publication Date
2025-06-10
Estimated Expiration
2042-06-17

AI Technical Summary

Technical Problem

In the enhancement of low-illumination images taken at night, the data set light source is single, and it is impossible to simulate real night scenes. The traditional methods cannot effectively improve image quality, and there are problems such as local area mutations and noise amplification.

Method used

Using an improved multi-scale fusion generation adversarial network image enhancement method, by establishing image data sets of different brightness, gamma correction, camera response model and Photoshop manual adjustment, combined with generator and discriminator, the channel attention mechanism and residual dense blocks are used to optimize the loss function to improve the image enhancement effect.

Benefits of technology

It effectively improves the brightness and contrast of low-illumination images, maintains image details and color nature, solves local mutations and noise problems, and improves the effect and flexibility of image enhancement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115223004B_ABST
    Figure CN115223004B_ABST
Patent Text Reader

Abstract

The present invention discloses an image enhancement method based on an improved multi-scale fusion generative adversarial network: Step 1: Establish an image data set with different brightness levels; Step 2: Establish an improved generative adversarial network image enhancement model, and substitute the training set into the model for training to obtain a trained model; Step 3: Input the real low-light image to be processed into the trained model to obtain the enhanced image. The present invention introduces a channel attention mechanism and a residual dense block into the generator, improves the local mutation situation of the enhanced picture, enables the network to pay more attention to the information of interest, and enhances the flexibility of the network. Utilize a multi-layer network to extract picture features in multiple dimensions, allowing each layer of the network to transmit the information that needs to be retained to the subsequent network, fuse the shallow features with the deep features, and more detailed information can be extracted in image enhancement, solving the problem of local distortion after enhancement.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image enhancement, and particularly relates to an image enhancement method based on an improved multi-scale fusion generative adversarial network. Background Art

[0002] With the wide application of image processing technology in machine vision, high-quality images have become particularly important and at the same time higher requirements are put forward for image preprocessing operations. In real life, the images captured are often plagued by problems such as blurred details, low contrast, and color distortion due to insufficient exposure or poor lighting conditions. Moreover, the real outdoor scenes are complex and changeable, which not only causes the quality of the pictures to decline, but also seriously affects the subsequent use of the images. Therefore, it is of great theoretical research significance to enhance the images with complex light sources and low visibility taken at night. At present, there are problems that the light sources in the existing night enhancement datasets are single and cannot simulate real night scenes, and traditional image enhancement techniques cannot achieve ideal effects.

[0003] Non-model image enhancement mainly includes histogram equalization algorithm and spatial domain filtering. Among them, the histogram equalization algorithm compares the clear image with the images taken by imaging devices in complex weather conditions, where the overall brightness and contrast are low and the dynamic range is small. The spatial domain is used to expand the gray levels of the concentrated-range images, and the gray range of the images is expanded to the entire gray range. So that the pixels originally concentrated in a certain high-brightness area or low-brightness area are distributed in the entire gray range. The ordinary histogram equalization algorithm calculates the mapping function based on the entire image. It only redistributes the noise and does not achieve the purpose of noise reduction. It may also cause over-enhancement of such night environment images, amplifying the noise in the dark areas; Spatial domain filtering is to slide-process each pixel in the image using a filter, using mathematical statistical operations including inverse transformation, logarithmic transformation, power transformation, etc. The filter is moved to perform the same operation on each pixel neighborhood until each neighborhood is processed. However, it is difficult to select the size of the sliding window during the operation. Because when the number of pixels in the filtering window is much smaller than the number of noises in the image, the filtering will fail; Model-based enhancement methods include the Retinex image enhancement method and the dark channel prior algorithm. Retinex first estimates the light source strength of the low-resolution image, removes the light intensity component, reduces the influence on the image caused by the lighting factor, and then separates and retains the reflection part of the object in the image. Although a certain balance is achieved in color naturalness preservation and detail enhancement, Retinex does not consider that the actual situation often cannot meet the premise assumption of the smooth change of the illumination image, resulting in local area mutation effects; The dark channel prior algorithm obtains the atmospheric transmittance and ambient light intensity based on the observed image, and then calculates the real image. The problems of the dark channel prior algorithm are that the restored image has a dark color and fails for the sky, etc. Summary of the Invention

[0004] The object of the present invention is to provide an image enhancement method based on an improved multi-scale fusion generative adversarial network to solve the problems in the prior art.

[0005] To achieve the above task, the present invention adopts the following technical solutions:

[0006] An image enhancement method based on an improved multi-scale fusion generative adversarial network includes the following steps:

[0007] Step 1: Establish an image data set with different brightness levels;

[0008] Step 2: Establish an improved generative adversarial network image enhancement model, substitute the training set obtained in Step 1 into the improved generative adversarial network image enhancement model for training to obtain a trained model; and perform testing through a test set; wherein, the improved generative adversarial network image enhancement model includes a generator and a discriminator;

[0009] The processing process of the generator for each image in the low-light image data set input therein includes the following sub-steps:

[0010] Step 21: Use a 7×7 convolutional layer to extract the shallow feature information of the low-light images in the training set;

[0011] Step 22: The shallow feature information sequentially passes through two downsampling layers to obtain image information of different scales from the original image;

[0012] Step 23: Send the image information obtained in Step 22 into a deep feature extraction module, which includes two parallel branches. These two parallel branches obtain deep information with different receptive fields through convolutional kernels of different sizes; the two parallel branches have the same composition, including a convolutional layer, an activation function, an alternating module layer, and a Concat layer connected in sequence, and the output of the activation function and the output of the alternating module are jointly connected to the Concat layer; wherein, the alternating module includes three residual dense blocks and three channel attention modules connected alternately;

[0013] Step 24: Visually fuse the deep information with different receptive fields obtained by the two branches to obtain a fused feature map;

[0014] Step 25: Perform one convolution, one activation, two upsamplings, and another convolution operation on the fused feature map in sequence to obtain an output feature map; wherein, the first convolution is used to further extract feature information from the fused features, the two upsamplings are used to restore the image size, and the another convolution is used for image restoration;

[0015] Step 26: Concatenate the output feature map obtained in Step 25 with its corresponding original image in the training set through a Concat layer to obtain the final output result;

[0016] Step 3: Input the real low - illumination image to be processed into the trained model obtained in Step 3 to obtain the enhanced image.

[0017] Furthermore, Step 1 includes the following sub - steps:

[0018] Step 11: Process the original image I using gamma correction to obtain the first low - illumination image I 1 , and the calculation formula is as follows:

[0019]

[0020] In the formula, I represents the original image (ground truth), I 1 represents the first low - illumination image, and γ 1 represents the gamma value, taking 0.2;

[0021] Step 12: Process the original image I using the camera response function to reduce the image brightness to obtain the second low - illumination image I 2 , and the calculation formula is as follows:

[0022]

[0023] In the formula, I 2 represents the second low - illumination image, k represents the virtual exposure rate, k = - 5.33, and a and b are two parameters of the camera response function;

[0024] Step 13: Manually adjust the image brightness of the original image I using image - processing software to obtain the third low - illumination image I 3 .

[0025] Step 14: Divide the data set into a test set and a training set;

[0026] Furthermore, in Step 23, the residual dense block includes a dense connection block and a residual connection block. Among them, the dense connection block is used to transmit the information extracted by each convolutional layer to each subsequent convolutional layer respectively. After passing through the dense connection blocks corresponding to the three convolutional layers respectively, the features of the shallow layer and the deep layer are fused through a connection function, and then the number of channels of the feature map is restored using a convolutional layer with a convolution kernel of 1; the residual connection block adopts a skip - connection method to add the features input to the residual dense block and the output after convolution and activation of the dense connection block pixel by pixel.

[0027] Furthermore, in Step 23, the convolution kernel sizes of the residual dense blocks of the two parallel branches are 3*3 and 5*5 respectively;

[0028] Further, in step 23, the processing of the feature map obtained by the residual dense block by the channel attention module includes the following sub-steps:

[0029] Step 231: The feature maps output by the residual dense block located before the channel attention module enter the two branches of the channel attention module respectively;

[0030] Step 232: The first branch uses global average pooling to convert the feature map information obtained by the residual dense block into a channel descriptor, converting the feature image of C×H×W into a feature map of C×1×1;

[0031] Step 233: The second branch uses global max pooling to take the maximum value of all features in each channel of the feature map information obtained by the residual dense block as the representative feature of the channel, obtaining a feature map;

[0032] Step 234: The feature maps obtained in step 232 and step 233 are respectively passed through two convolutional layers to amplify and reduce the number of channels by the same multiple, obtaining the feature weights corresponding to different channels. Among them, an activation function layer is connected after the first convolutional layer to prevent the divergence of the extracted feature data during transmission. Then, these two different-sized channel features are fused to obtain channel feature information, and then activated through the Sigmoid activation function;

[0033] Step 235: The Element-wise product module is used to multiply the weights of each channel in the feature map obtained by the previous residual dense block with the feature map obtained in step 234 to obtain a weighted feature map.

[0034] Further, in step 2, the discriminator includes 6 convolutional layers. An instance normalization layer and a relu activation function are provided after each of the first 5 convolutional layers to prevent gradient disappearance; a Sigmoid activation function is provided after the last convolutional layer.

[0035] Further, in step 2, the loss function of the discriminator is:

[0036]

[0037] where, refers to the Wasserstein distance between the distribution of samples generated by the generator and the distribution of real images, denotes that the picture is taken from the set of generated pictures output by the generator network, denotes that the picture is taken from the set of normal illumination images in the training set, It indicates that the picture is taken from the area between the generated sample and the real sample. G(x) represents the image generated by the generator according to the original input image x, D(x) represents the discriminative evaluation of the discriminator on the original input image x, A represents the expected value expression, and λ is the constant coefficient of the gradient penalty term. is the discriminator gradient, is the WGAN-GP gradient penalty term, whose purpose is to make the discriminator gradient not exceed 1, to solve the problem of unstable gradient, and at the same time accelerate convergence.

[0038] Furthermore, in step 2, the overall loss of the generator:

[0039] L G = L condition + λ * L content

[0040] where λ is the correction coefficient, taking 100; L condition is the conditional loss, and L content is the content loss;

[0041] The loss function of the conditional loss is:

[0042]

[0043] where B represents the input low-illumination picture, represents the conditional probability that the output generated image belongs to the original input low-illumination image, represents the average value of whether the picture generated by the generator is a real picture;

[0044] The content loss adopts perceptual loss.

[0045] Compared with the prior art, the present invention has the following technical features:

[0046] (1) For the dataset of the present invention, the existing low-illumination images are introduced from the HDR dataset of a fixed scene, which has the problems of small quantity, difficulty in covering different scenes and single light source. By using different brightness adjustment functions and parameters to synthesize low-illumination images, mainly three methods of gamma correction, camera response model and manual adjustment in Photoshop are used. By covering more and more complex brightness transformation curves, the night scene is simulated.

[0047] (2) In the generator of the improved generative adversarial network model of the present invention, a channel attention mechanism is introduced to improve the local mutation situation of the enhanced picture. The channel attention mechanism shows that different channel features have completely different weighted information according to the brightness and the information contained in each area of the same picture. By treating different channel features unequally, the network pays more attention to the information of interest and enhances the flexibility of the network.

[0048] (3) In the generator of the improved generative adversarial network model of the present invention, a residual dense block is introduced to extract image features in multiple dimensions using a multi-layer network, allowing each layer of the network to transmit the information that needs to be retained to the subsequent network, fusing shallow features with deep features, and enabling more detailed information to be extracted in image enhancement, thereby solving the problem of local distortion after enhancement.

[0049] (4) In the improved generative adversarial network model of the present invention, the discriminator adopts a Markov discriminator network including 6 convolutional layers. After each of the first 5 convolutional layers, there is an instance normalization layer (IN) and a relu activation function to prevent gradient disappearance. After the last 1 convolutional layer, there is a Sigmoid activation function to map the range of the output image pixel values to between (0, 1), which is beneficial for the discriminative network to distinguish the authenticity of the generated image and the target image in a certain area, and solves the problem of local area denoising failure by scoring n small areas of the image.

[0050] (5) The improved generative adversarial network model of the present invention optimizes the loss function. A gradient penalty term WGAN-GP is added to the discriminator loss function to prevent GAN training from collapsing. The perceptual loss is used as the content loss in the generator's loss function to solve the problems of L1 loss and L2 loss. When optimizing with the average value in the pixel space as the only target, the generated image will be blurred. Description of the Drawings

[0051] Figure 1 It is the structure diagram of the generative adversarial network in the present invention;

[0052] Figure 2 It is the structure diagram of the channel attention module in the present invention;

[0053] Figure 3 It is the structure diagram of the residual dense connection block in the present invention;

[0054] Figure 4 It is the structure diagram of the generator network in the present invention;

[0055] Figure 5 It is the structure diagram of the discriminator network in the present invention;

[0056] Figure 6 It is the effect diagram of the gamma function in the embodiment of the present invention;

[0057] Figure 7 It is the camera response model diagram in the embodiment of the present invention;

[0058] Figure 8 It is the Photoshop adjustment diagram in the embodiment of the present invention;

[0059] Figure 9It is the contrast experiment 1 of the low-light image enhancement structure in the embodiment of the present invention, where:

[0060] (a) is the low-light image;

[0061] (b) is the enhanced effect diagram of the SRIE algorithm;

[0062] (c) is the enhanced effect diagram of the DC-GAN algorithm;

[0063] (d) is the enhanced effect diagram of the Cycle-GAN algorithm;

[0064] (e) is the enhanced effect diagram of the Lime algorithm;

[0065] (f) is the enhanced effect diagram of the method of the present invention;

[0066] Figure 10 It is the contrast experiment 2 of the low-light image enhancement structure in the embodiment of the present invention, where:

[0067] (a) is the low-light image;

[0068] (b) is the enhanced effect diagram of the SRIE algorithm;

[0069] (c) is the enhanced effect diagram of the DC-GAN algorithm;

[0070] (d) is the enhanced effect diagram of the Cycle-GAN algorithm;

[0071] (e) is the enhanced effect diagram of the Lime algorithm;

[0072] (f) is the enhanced effect diagram of the method of the present invention;

[0073] Figure 11 It is the example of the enhanced results of the low-light images with different levels of illumination in the example of the present invention, where:

[0074] (a) is the corresponding enhanced effect diagram of the low-light image generated by the gamma function method;

[0075] (b) is the corresponding enhanced effect diagram of the low-light image generated by the camera response function;

[0076] (c) is the corresponding enhanced effect diagram of the low-light image generated by manually adjusting through Photoshop;

[0077] Figure 12 It is the enhanced results of the low-light images with different levels of illumination in the example of the present invention:

[0078] (a) is the corresponding enhanced effect diagram of the low-light image generated by the gamma function method;

[0079] (b) is the enhanced effect diagram corresponding to the low-light image generated by the camera response function;

[0080] (c) is the enhanced effect diagram corresponding to the low-light image generated by manual adjustment in Photoshop;

[0081] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Specific Embodiments

[0082] First, the technical terms appearing in the present invention are explained:

[0083] Generative adversarial network: It is an excellent deep learning model proposed by Lan Goodfellow et al. in 2014, which consists of two parts: a generative network and a discriminative network, as Figure 1 shown. The generative adversarial network uses the idea of zero-sum game, and through the adversarial learning and mutual game between the generator network and the discriminator network, it reaches the Nash equilibrium point, at which time the network reaches the optimal effect.

[0084] Gamma correction: Also known as gamma non-linearity or gamma coding. It is a non-linear operation or inverse operation for the luminance of light or the tristimulus values in a film or imaging system. The purpose of performing gamma correction on an image is to compensate for the characteristics of human vision. If gamma correction is not performed on the image, the utilization of data bits or bandwidth will be unevenly distributed, resulting in abnormal visual perception of the image. Gamma correction aims to exclude the influence and correct the image by pre-increasing the RGB values.

[0085] Camera response model: Also known as the camera response function. When a camera captures an image, the luminance L of the objects in the scene is irradiance E on the surface of the image sensor after passing through the camera lens, and the relationship between the scene luminance L and E is:

[0086]

[0087] where h is the focal length of the camera lens, is the angle between the incident light and the perpendicular plane of the image sensor, and d is the size of the lens aperture. Since multiple frames of images are required to be captured with the camera stationary, during the capture process, h, If d are all invariant quantities, then the process of mapping L to E is a linear mapping process. After pressing the shutter, the total exposure amount reaching the image sensor within the exposure time Δt is converted into an analog signal after passing through the light spot conversion of the sensor, and then undergoes steps such as analog-to-digital conversion and quantization and rounding to obtain the pixel value Z. The process of mapping the irradiance E to the pixel value Z is a non-linear process, and this non-linear mapping relationship is called the camera response function.

[0088] The image enhancement method based on the improved multi-scale fusion generative adversarial network given in this embodiment includes the following steps:

[0089] Step 1: Establish a low-light image dataset with different brightness levels.

[0090] In this step, aiming at the problems of small data volume and single brightness in the existing dataset, three different methods are used to artificially generate data to construct a low-light image dataset, covering more and more complex brightness transformation curves. Specifically, it includes the following sub-steps:

[0091] Step 11: Process the original image I using gamma correction to obtain the first low-light image I 1 , and the calculation formula is as follows:

[0092]

[0093] In the formula, I represents the original image (ground truth), I 1 represents the first low-light image, and γ 1 represents the gamma value, taking 0.2.

[0094] Step 12: Process the original image I using the camera response function to reduce the image brightness to obtain the second low-light image I 2 , and the calculation formula is as follows:

[0095]

[0096] In the formula, I 2 represents the second low-light image, k represents the virtual exposure rate, k = -5.33, and a and b are two parameters of the camera response function, which can be calculated through different exposure natural image pairs. After experiments, a = -0.32 and b = 1.12 are adopted.

[0097] In this step, the camera response function uses an exposure rate different from that of the real image to simulate the generation of images under low exposure levels. The camera response function is more complex than the gamma function and can cover different categories of low-light images through different parameters, thus increasing the data coverage of the dataset.

[0098] Step 13: Manually adjust the brightness of the original image I using image processing software (preferably Photoshop) to obtain the third low-illumination image I 3 . Using the gamma correction function and the camera response model to change the image brightness, the obtained result may still be different from the real image. To increase the authenticity and reliability of the dataset, this step performs manual adjustment for each original image, and 100 low-illumination images I are obtained 3 .

[0099] Specifically, it establishes three different brightness low-illumination image datasets as Figure 6 、 7 、shown in Figure 8. In this embodiment, 100 original images are used, and 100 low-illumination images corresponding to three different brightness levels are generated respectively, totaling 300 pairs of images

[0100] Step 14: Divide the dataset into a test set and a training set in a ratio of 2:8

[0101] Step 2: Establish an improved generative adversarial network image enhancement model, substitute the training set obtained in Step 1 into the improved generative adversarial network image enhancement model for training to obtain a trained model; test it through the test set. Among them, the improved generative adversarial network image enhancement model includes a generator and a discriminator (see Figure 1 ), where the generator is a multi-scale residual dense connection network based on the attention mechanism, the discriminator is a discriminator based on PatchGan, and the loss function is an optimized loss function

[0102] Specifically, as Figure 4 shown, the processing process of the generator for each image in the low-illumination image dataset input into it includes the following sub-steps

[0103] Step 21: Use a 7*7 convolutional layer to extract the shallow feature information of the low-illumination images in the training set

[0104] Step 22: The shallow feature information passes through two downsampling layers in sequence to obtain image information with different scales from the original image

[0105] Step 23: Send the image information obtained in Step 22 into the deep feature extraction module, which includes two parallel branches, and these two parallel branches obtain deep information with different receptive fields through convolutional kernels of different sizes

[0106] The two parallel branches are identical, i.e., each includes a convolutional layer, an activation function, an alternating module layer, and a Concat layer connected in sequence. The output of the activation function and the output of the alternating module are jointly connected to the Concat layer. Among them, the alternating module includes three residual dense blocks and three channel attention modules connected alternately. In the two parallel branches, the convolutional kernel sizes of the residual dense blocks are 3*3 and 5*5 respectively; the residual dense blocks and the channel attention modules are used to learn complex detailed features; the Concat layer jointly connected by the output of the activation function and the output of the alternating module is used to copy shallow information to the deep layer to prevent the loss of low-dimensional information during the convolution process.

[0107] Specifically, as Figure 3 shown, the residual dense block includes a dense connection block and a residual connection block (the upper connection line in the figure represents the dense connection, and the lower connection line represents the residual connection). Among them, the dense connection block is used to transmit the information extracted by each convolutional layer to each subsequent convolutional layer. After passing through the dense connection blocks corresponding to the three convolutional layers respectively, the features of the shallow layer and the deep layer are fused through a connection function, and then the number of channels of the feature map is restored using a convolutional layer with a convolutional kernel of 1; the residual connection block adopts a skip connection method to add the features input to the residual dense block pixel by pixel to the output after convolution and activation by the dense connection block, which can reduce the training complexity while improving the accuracy by increasing the network depth. The feature map output by the residual dense block (with a size of C*H*W, where C is the number of channels, and H and W are the width and height of the feature map respectively)

[0108] Specifically, as Figure 2 shown, the processing of the feature map obtained by the channel attention module for the residual dense block includes the following sub-steps:

[0109] Step 231: The feature maps output by the residual dense block located before the channel attention module enter the two branches of the channel attention module respectively;

[0110] Step 232: The first branch uses global average pooling to convert the feature map information obtained by the residual dense block into a channel descriptor (representing the background features of the image), and converts the feature image of C×H×W into a feature map of C×1×1;

[0111] Step 233: The second branch uses global max pooling to take the maximum value of all features in each channel of the feature map information obtained by the residual dense block as the representative feature of the channel (representing the texture detail information of the image), and obtains a feature map;

[0112] Step 234: The feature maps obtained in Step 232 and Step 233 are respectively passed through two convolutional layers to amplify and reduce the number of channels by the same multiple, obtaining the feature weights corresponding to different channels. Among them, an activation function layer is connected after the first convolutional layer to prevent the extracted feature data from diverging during transmission. Then, these two channel features of different sizes are fused (i.e., added pixel by pixel) to obtain channel feature information, obtaining a new feature map weight. After that, it is activated through the Sigmoid activation function to prevent the feature data from diverging during transmission.

[0113] The purpose of this step is to achieve the characteristics of complete retention of enhanced image details and overall harmony and consistency.

[0114] Step 235: The Element-wise product module is adopted to multiply the feature map obtained from the previous residual dense block by the weights of each channel in the feature map obtained in Step 234, obtaining a weighted feature map.

[0115] This step can enhance the feature learning ability of the network.

[0116] Step 24: The deep information of different receptive fields obtained from the two branches is visually fused (i.e., added pixel by pixel) to obtain a fused feature map;

[0117] Step 25: The fused feature map is successively subjected to one convolution, one activation, two upsamplings, and another convolution operation to obtain an output feature map; among them, the first convolution is used to further extract feature information from the fused features, the two upsamplings are used to restore the image size, and the another convolution is used for image restoration;

[0118] Step 26: The output feature map obtained in Step 25 and the corresponding original image in the training set are jointly connected to the Concat layer to obtain the final output result.

[0119] Specifically, as Figure 5 shown, in Step 2, the discriminator optimization result adopts a PatchGan discriminator (Markov discriminator) network, including 6 convolutional layers. An instance normalization layer and a relu activation function are provided after each of the first 5 convolutional layers to prevent gradient disappearance; a Sigmoid activation function is provided after the last convolutional layer to map the range of the output image pixel values to between (0, 1), which is beneficial for the discriminator network to distinguish the authenticity of the generated image and the target image in a certain area, and solves the problem of local area denoising failure by scoring n small areas of the image.

[0120] Specifically, in Step 2, the loss function of the discriminator is:

[0121]

[0122] Among them, refers to the Wasserstein distance between the distribution of samples generated by the generator and the distribution of real images, indicates that the picture is taken from the set of generated pictures output by the generator network, indicates that the picture is taken from the set of normal illumination images in the training set, indicates that the picture is taken from the area between the generated sample and the real sample. G(x) represents the image generated by the generator according to the original input image x, D(x) represents the discriminative evaluation of the discriminator on the original input image x, A represents the expected value expression, and λ is the constant coefficient of the gradient penalty term. is the discriminator gradient, is the WGAN-GP gradient penalty term, whose purpose is to make the discriminator gradient not exceed 1, to solve the problem of unstable gradient, and at the same time accelerate convergence.

[0123] The overall loss of the generator is the sum of the conditional loss and the optimized content loss, and its expression is;

[0124] L G = L condition + λ * L content

[0125] λ is the correction coefficient, which is fixed at 100 in this embodiment. Among them, L condition is the conditional loss, which focuses on maintaining the dependence of the generator output on the low-illumination image of the input, and L content is the content loss, which focuses on ensuring the authenticity of the generated picture.

[0126] The goal of the conditional loss is to maximize the probability that the discriminator judges the generated image as a real image, and its loss function is:

[0127]

[0128] B represents the input low-illumination picture, represents the conditional probability that the output generated image belongs to the original input low-illumination image, represents the average value of whether the picture generated by the generator is judged as a real picture;

[0129] Since both the L1 loss and the L2 loss of the classical content loss need to take the average value in the pixel space, it will cause the generated image to be blurred when it is optimized as the only goal. In the present invention, the perceptual loss is used as the content loss. The perceptual loss is to calculate the difference between the feature maps after activation of the conv3-3 layer of the VGG-19 network when the real clear image and the generated image obtained by the generator are input, so as to ensure that the edge information of the enhanced image is sharper. The calculation formula of the perceptual loss is as follows:

[0130]

[0131] denotes the feature map obtained after the i-th convolutional layer and before the j-th max pooling layer of the VGG-19 network; W i,j and H i,j respectively represent the width and height of the feature map, S is the clear image, and G(B) is the image generated by the generator.

[0132] Step 3: Input the real low-light image to be processed into the trained model obtained in Step 3 to obtain the enhanced image.

[0133] To verify the feasibility and effectiveness of the method of the present invention, the process and results of testing using the test set in Step 3 are given:

[0134] Use the same training set in the dataset to train on two deep learning algorithm models, DC-Gan and Cycle-Gan, respectively, to obtain the trained network models. Then, use the same test set in the dataset to conduct comparative experiments on the SRIE, Lime algorithm, trained DC-Gan network, and Cycle-Gan network, proving that the method of the present invention has a significant impact on improving the image enhancement effect.

[0135] Specifically, the Adam optimizer is used for training. After every 3 gradient descents on the discriminator, the generator is updated once, and the total number of training batches is 300 times. The initial learning rates of both the generator and the discriminator are set to 1×10 -4 . The effect diagrams of the comparative experiments are as Figure 9 、 Figure 10 shown. From left to right are the results of the low-light image, SRIE, DC-Gan network, Cycle-Gan network, Lime algorithm, and finally the result of the method of the present invention.

[0136] Enhanced images of different levels of illuminance are as Figure 11 、 Figure 12 shown; from left to right are the enhancement results of three different low-light images of the gamma function, camera response model, and Photoshop.

[0137] From Figure 9 、 10 it can be seen that the method of the present invention can well maintain the original color of the image compared with other classical algorithms after enhancement, restore the detailed information of the image to the greatest extent, improve the image brightness and contrast while ensuring soft colors, proving that the model disclosed in the present invention effectively improves the visual effect of night enhanced images; Figure 11 、 12The low-light image enhancement effects reflecting three different illuminations are shown to simulate the different brightness phenomena of the images after imaging caused by complex light sources at night. It can be seen from the figure that the low-light image enhancement effects of different brightness are different. The low-light image with slightly stronger brightness on the far right has a better effect, but most details can also be restored after the enhancement of the darker image on the far left, proving that the model of the present invention can adapt to images of different brightness and has compatibility.

[0138] In addition, to test the superiority of the method of the present invention over general methods, the same dataset is used to test on SRIE, Lime algorithm, trained DC-Gan network and Cycle-Gan network, and the model of the present invention respectively. The average performance indicators of various algorithms are shown in the following table.

[0139] Table 1 Experimental results of image enhancement by different algorithms

[0140]

[0141] The PSNR of the improved multi-scale fusion generative adversarial network image enhancement algorithm of the present invention is as high as 23.43, and the SSIM is as high as 0.82. Compared with traditional image enhancement and image enhancement of ordinary Gan network, the PSNR of the improved multi-scale fusion generative adversarial network image enhancement algorithm is increased by 17.12%, and the SSIM is increased by 17.56%, laying a good foundation for the subsequent use of the pictures.

Claims

1. An image enhancement method based on an improved multi-scale fusion generative adversarial network, characterized in that, it includes the following steps: Step 1: Establish an image data set with different brightness levels; Step 2: Establish an improved generative adversarial network image enhancement model, substitute the training set obtained in Step 1 into the improved generative adversarial network image enhancement model for training to obtain a trained model; and test it through a test set; wherein, the improved generative adversarial network image enhancement model includes a generator and a discriminator; The processing process of the generator for each image in the low-light image data set input into it includes the following sub-steps: Step 21: Use a 7*7 convolutional layer to extract the shallow feature information of the low-light images in the training set; Step 22: The shallow feature information passes through two downsampling layers in sequence to obtain image information of different scales from the original image; Step 23: Send the image information obtained in Step 22 into a deep feature extraction module, which includes two parallel branches. These two parallel branches obtain deep information with different receptive fields through convolutional kernels of different sizes; the two parallel branches are composed of the same structure, including a convolutional layer, an activation function, an alternating module layer, and a Concat layer connected in sequence, and the output of the activation function and the output of the alternating module are jointly connected to the Concat layer; wherein, the alternating module includes three residual dense blocks and three channel attention modules connected alternately; Step 24: Visually fuse the deep information with different receptive fields obtained by the two branches to obtain a fused feature map; Step 25: Perform one convolution, one activation, two upsamplings, and another convolution operation on the fused feature map in sequence to obtain an output feature map; wherein, the first convolution is used to further extract feature information from the fused features, the two upsamplings are used to restore the image size, and the another convolution is used for image restoration; Step 26, jointly connect the output feature map obtained in Step 25 with its corresponding original image in the training set to the Concat layer to obtain the final output result; Step 3: Input the real low-light image to be processed into the trained model obtained in Step 2 to obtain an enhanced image.

2. The image enhancement method based on an improved multi-scale fusion generative adversarial network according to claim 1, characterized in that, Step 1 includes the following sub-steps: Step 11: Process the original image I using gamma correction to obtain the first low-illumination image I 1 , and the calculation formula is as follows: where I represents the original image (ground truth), I 1 represents the first low-light image, γ 1 represents the gamma value, taking 0.2; Step 12: Process the original image I using the camera response function to reduce the image brightness and obtain the second low-illumination image I 2 , and the calculation formula is as follows: where I 2 represents the second low-illumination image, k represents the virtual exposure rate, k = -5.33, and a and b are two parameters of the camera response function; Step 13: Manually adjust the brightness of the original image I using image processing software to obtain the third low-illumination image I 3 ; Step 14, divide the data set into a test set and a training set.

3. The image enhancement method based on an improved multi-scale fusion generative adversarial network according to claim 1, characterized in that, In Step 23, the residual dense block includes a dense connection block and a residual connection block. Among them, the dense connection block is used to transmit the information extracted by each convolutional layer to each subsequent convolutional layer respectively. After passing through the dense connection blocks corresponding to the three convolutional layers, the features of the shallow layer and the deep layer are fused through a connection function, and then the number of channels of the feature map is restored by a convolutional layer with a convolution kernel of 1; the residual connection block adopts a skip connection method to add the features input into the residual dense block pixel by pixel to the output after convolution and activation by the dense connection block.

4. The image enhancement method based on an improved multi-scale fusion generative adversarial network according to claim 1, It is characterized in that In step 23, the convolution kernel sizes of the residual dense blocks of the two parallel branches are 3*3 and 5*5 respectively.

5. The method for enhancing an image based on an improved multi-scale fusion generative adversarial network according to claim 1, It is characterized in that In step 23, the processing of the feature map obtained by the residual dense block by the channel attention module includes the following sub-steps: Step 231: The feature maps output by the residual dense blocks located before the channel attention module respectively enter the two branches of the channel attention module; Step 232: The first branch uses global average pooling to convert the feature map information obtained by the residual dense block into a channel descriptor, converting the feature image of C×H×W into a feature map of C×1×1; Step 233: The second branch uses global max pooling to use the maximum value of all features in each channel of the feature map information obtained by the residual dense block as the representative feature of the channel, obtaining a feature map; Step 234: The feature maps obtained in step 232 and step 233 are respectively passed through two convolutional layers to amplify and reduce the number of channels by the same multiple, obtaining the feature weights corresponding to different channels. Among them, an activation function layer is connected after the first convolutional layer to prevent the divergence of the extracted feature data during transmission. Then, these two channel features of different sizes are fused to obtain channel feature information, and then activated through a Sigmoid activation function; Step 235: An Element-wise product module is adopted to multiply the weights of each channel of the feature map obtained by the previous residual dense block and the feature map obtained in step 234 to obtain a weighted feature map.

6. The method for enhancing an image based on an improved multi-scale fusion generative adversarial network according to claim 1, It is characterized in that In step 2, the discriminator includes 6 convolutional layers. After the first 5 convolutional layers, an instance normalization layer and a relu activation function are provided respectively to prevent gradient disappearance; a Sigmoid activation function is provided after the last convolutional layer.

7. The method for enhancing an image based on an improved multi-scale fusion generative adversarial network according to claim 1, It is characterized in that In step 2, the loss function of the discriminator is: Among them, refers to the Wasserstein distance between the distribution of samples generated by the generator and the distribution of real images, represents that the images are taken from the generated image set output by the generator network, represents that the images are taken from the normal illumination image set in the training set, represents that the images are taken from the region between the generated samples and the real samples. G(x) represents the image generated by the generator according to the original input image x, D(x) represents the discriminant evaluation of the discriminator on the original input image x, A represents the expected value expression, and λ is the constant coefficient of the gradient penalty term. is the discriminator gradient, is the WGAN-GP gradient penalty term, whose purpose is to make the discriminator gradient not exceed 1, to solve the problem of unstable gradients, and to accelerate convergence at the same time.

8. The method for enhancing an image based on an improved multi-scale fusion generative adversarial network according to claim 1, It is characterized in that In step 2, the overall loss of the generator: L G = L condition + λ * L content where λ is a correction coefficient, taking 100; L condition is the conditional loss, and L content is the content loss; The loss function of the conditional loss is: Among them, B represents the input low-light image, indicating the conditional probability that the output generated image belongs to the original input low-light image, represents the average value of whether the image generated by the generator is a real image; the content loss adopts perceptual loss.