Method for enhancing low-light images using generative network, training method and device for generative network
By adopting the residual structure of the generation network and the channel attention mechanism in low-light image enhancement, combined with the training method of unsupervised generation adversarial networks, problems such as noise amplification meeting and feature details distortion in the prior art are solved, and high-quality low-light image enhancement effect is achieved.
Patent Information
- Application Number
- CN202310396091.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-13
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2043-04-13
AI Technical Summary
The existing low-light image enhancement methods have problems such as noise amplification, feature details distortion, color errors, and limited generalization capabilities.
Using a method based on the generation network, through the residual structure and channel attention mechanism, the main branch of the generation network generates the residual between the low-light image and the generated image. The shortcut branch transmits the low-light image to the output end of the main branch and adds the residuals to obtain the generated image. At the same time, the generative network is trained using an unsupervised generative adversarial network, and the parameters are adjusted by semantic consistency loss and adversarial loss to improve the quality of the generated images.
It effectively reduces noise, improves the authenticity of image details and color accuracy, enhances the focus of the generation network on the areas that need to be enhanced, and improves the quality of the generated images.
Smart Images

Figure CN116703792B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing, and more specifically, to a method for enhancing a low-light image using a generative network, and a training method and device for the generative network. Background Art
[0002] Low-light images refer to images acquired under low illumination conditions. They have low quality, poor recognition performance, contain a lot of noise, are difficult to discern details, and have low use value. Low-light images have low visual quality due to insufficient illumination, which affects the performance of many computer vision systems when applied or processed. Low-light enhancement refers to the process of processing low-light images, with the aim of improving the perceived quality of images captured in low-light environments so that they can basically achieve the visual perception effect of images with sufficient illumination.
[0003] Solutions for low-light image enhancement include: traditional image methods and methods based on deep learning, among which the traditional image methods include methods based on HE histogram equalization and methods based on Retinex theoretical models.
[0004] The theory of HE histogram equalization believes that the three-channel RGB pixel values of a normal light intensity image have equal probability of appearing in the entire dynamic range. Except for some prominent pixel values, the distribution of the entire three-channel RGB pixel values is close to uniform distribution. Such an image has a large color dynamic range and high contrast. Based on this, after obtaining the histogram information of the input image, a transformation function can be applied to the input image to make the three-channel RGB pixel distribution of the original input image closer to uniform distribution, thereby completing HE histogram equalization and achieving low-light image enhancement. However, some areas of the image enhanced by this method will be over-enhanced, and the generated image will have a lot of noise.
[0005] The Retinex theory is based on the color constancy of objects. It believes that the color of an object is determined by the object's ability to reflect light of different wavelengths, that is, the different colors presented by different wavelengths, rather than by the intensity of the reflected light, and has color constancy. The basic assumption of the Retinex theory is that the original image is the product of the image illumination component and the image reflectance component. The image enhancement method based on the Retinex theory is to estimate the illumination component from the original image component, thereby decomposing the image reflectance component. This method can usually perform a certain degree of adaptive enhancement on images that are not particularly dark. However, it is not very reasonable for the Retinex theory to a priori use the image reflectance component as the enhancement result, especially under different ambient light conditions. Such a priori false assumptions may lead to unrealistic enhancements, such as distortion of feature details and color errors. Moreover, the influence of noise is often ignored in the Retinex theoretical model, and a problem that follows is that these noises are amplified in the enhanced image.
[0006] In recent years, deep learning-based algorithms have been widely used in low-light image enhancement tasks. The solutions for low-light image enhancement based on deep learning mainly involve: supervised learning (SL) and unsupervised learning (UL). However, in actual scenarios, it is difficult to obtain paired images of low-light and full-light images in the same scene, which leads to certain limitations in the application of supervised learning in low-light image enhancement. For unsupervised learning, its generalization ability is limited and there are still problems such as network instability, complex network structure, low learning efficiency, and unrealistic local details of the enhanced image.
[0007] Based on this, the technology of low-light image enhancement needs to be further improved. Summary of the invention
[0008] In order to solve the problems existing in the above-mentioned low-light image enhancement method, the present invention provides a method for enhancing a low-light image using a generative network, comprising: inputting a low-light image into a generative network 1; the generative network 1 enhances the low-light image to obtain a generated image; the generative network 1 outputs the generated image; the generative network 1 adopts a residual structure, the main branch of the generative network 1 generates the residual between the low-light image and the generated image, and the shortcut branch of the generative network 1 transmits the low-light image to the output end of the main branch of the generative network 1 and adds the residual to obtain the generated image; wherein the main branch of the generative network 1 is at least two residual learning units 11 arranged in cascade, and each residual learning unit generates the residual step by step and progressively.
[0009] According to one embodiment of the present invention, the residual learning unit 11 adopts a residual structure, the main branch of the residual learning unit 11 is the generator 111, and the shortcut branch of the residual learning unit 11 connects the input and output ends of the generator 111.
[0010] According to one embodiment of the present invention, the generation network 1 also includes a channel attention branch, which is connected to the input end of the first-level generator and the output end of the last-level generator 111; the channel attention branch takes out the RGB three channels of the low-light image and inverts them, and then multiplies them with the output of the last-level generator 111 in channel order.
[0011] According to one embodiment of the present invention, the generator 111 adopts an encoder-residual block-decoder architecture, including a downsampling layer 1111, a residual block 1112, an upsampling layer 1113 and a dimensionality reduction layer 1114, and the encoder and the decoder are symmetrically arranged.
[0012] According to one embodiment of the present invention, the generator 111 includes three downsampling layers 1111, nine residual blocks 1112, two upsampling layers 1113 and one dimensionality reduction layer 1114; the three downsampling layers 1111 constitute an encoder, and the two upsampling layers 1113 and one dimensionality reduction layer 1114 constitute a decoder; the downsampling layer 1111 is a CBR structure; the upsampling layer 1113 is a TBR structure.
[0013] According to another aspect of the present invention, a method for training a generative network based on an unsupervised generative adversarial network is provided, comprising: obtaining a low-light image and a sufficient-light image as training samples, wherein the semantic contents of the low-light image and the sufficient-light image do not need to be strictly consistent; inputting the low-light image into the generative network to obtain a generated image; inputting the generated image and the sufficient-light image into the discriminant network to obtain a discrimination result; iteratively adjusting the parameters of the generative network or the discriminant network until the generative network and the discriminant network reach a Nash equilibrium; wherein the parameters of the generative network are adjusted based on the semantic consistency loss between the low-light image and the generated image and the adversarial loss between the generated image and the sufficient-light image; and the parameters of the discriminant network are adjusted based on the adversarial loss; wherein the generative network adopts a residual structure, and a shortcut branch of the generative network 1 connects the input and output of the generative network 1; the main branch of the generative network 1 is at least two residual learning units 11 arranged in cascade; the residual learning unit 11 adopts a residual structure, and the main branch of the residual learning unit 11 is a generator 111, and the shortcut branch of the residual learning unit 11 connects the input and output of the generator 111.
[0014] According to one embodiment of the present invention, the consistency of VGG19 features is used to maintain the semantic content without deviation, and the semantic content consistency loss Loss VGGThe definition is as follows:
[0015]
[0016] WH represents the width and length of the image; Wi represents the i-th pixel in the width direction; Hj represents the j-th pixel in the length direction; represents the VGG features of the input low-light image; v(G(X)) represents the VGG features of the enhanced generated image; ||||2 represents the L2 norm measurement.
[0017] According to one embodiment of the present invention, the adversarial loss is defined as follows:
[0018]
[0019]
[0020]
[0021] in, is the loss function of the discriminator at the global scale; is the loss function of the discriminator at the local scale; is the loss function of the generator at the global scale; is the loss function of the generator at the local scale; X r ~real means that the input image follows the distribution of the real image domain; X g ~generate means that the input image obeys the distribution of the generated image domain; D represents the discriminator network; ||||2 represents the L2 norm measurement; E represents expectation; the local discriminator uses the average discriminator results of n random sub-blocks, where n is an integer greater than 1.
[0022] According to one embodiment of the present invention, n=6.
[0023] For the generative network in the present invention, by setting a channel attention mechanism, the generative network is helped to automatically focus on the information of the area that needs to be enhanced from the input. By setting the local scale of the discriminant network, the area that needs to be enhanced in the input low-light image is adaptively enhanced when the generative network is trained. By setting both the architecture of the generative network and the training process of the generative network, the emphasis on the area that needs to be enhanced when the generative network performs image enhancement is increased, thereby improving the quality of the generated image.
[0024] In the present invention, for the generative network, by setting a residual structure, the task of generating residuals is decomposed, the tasks of each generator are reduced, and the training purpose is easier to achieve during training. By setting a cascade structure, each generator only bears part of the task, further reducing the difficulty of generation and learning. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 It is a structural schematic diagram of a low-light image enhancement model based on an unsupervised generative adversarial network of the present invention;
[0026] Figure 2 This is a schematic diagram of the structure of the generative network part of the low-light image enhancement model;
[0027] Figure 3 It is a schematic diagram of the structure of the generative network and residual learning unit;
[0028] Figure 4 is a schematic diagram of a generative network including a channel attention branch;
[0029] Figure 5 It is a schematic diagram of the structure of the generator;
[0030] Figure 6 is a schematic diagram of the global discriminator;
[0031] Figure 7 is a schematic diagram of a local discriminator;
[0032] Figure 8 It is a schematic diagram of the steps of the method for training a generative network based on an unsupervised generative adversarial network;
[0033] Fig. 9 It is a schematic diagram of the steps of a method for enhancing low-light images using a generative network;
[0034] Fig.10 is an exemplary structural block diagram of a device for training a model. DETAILED DESCRIPTION
[0035] The principles and spirit of the present invention will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided only to enable those skilled in the art to better understand and implement the present invention, and are not intended to limit the scope of the present invention in any way. On the contrary, these embodiments are provided to make the present disclosure more thorough and complete, and to fully convey the scope of the present invention to those skilled in the art.
[0036] Figure 1 A structural schematic diagram of a low-light image enhancement model based on an unsupervised generative adversarial network of the present invention is shown.
[0037] like Figure 1As shown in the figure, the generative adversarial network contains two sub-networks, the generator network 1 and the discriminator network 2. The task of the generator network 1 is to enhance the input low-light image and generate a generated image that looks natural and real and similar to the original data. The task of the discriminator network 2 is to judge whether it is true or false, that is, taking the generated image and the well-lit image as input, judging whether the generated image is close to the well-lit image, and outputting the discrimination result.
[0038] In the present invention, low-light images and sufficient-light images are a set of relative concepts, and the visual effect of sufficient-light images is better than that of low-light images. The present invention does not quantitatively limit the illumination intensity of low-light images and sufficient-light images. In the image enhancement process, any picture that needs to be enhanced can be used as a low-light image regardless of the absolute value of its illumination intensity.
[0039] The training process of the generative adversarial network includes first training the discriminant network 2: input the fully illuminated image with the true label and the generated image generated by the generative network 1 with the false label into the discriminant network 2, and train the discriminant network 2. When calculating the loss, the discriminant network 2 makes the judgment of the fully illuminated image close to true, and the judgment of the generated image generated by the generative network 1 close to false. In this process, only the parameters of the discriminant network 2 are updated, and the parameters of the generative network 1 are not updated. Then train the generative network 1: input the low-light image into the generative network 1, and then put the generated image generated by the generative network 1 into the discriminant network 2 with the true label. When calculating the loss, the discriminant network 2 makes the judgment of the generated image close to true. In this process, only the parameters of the generative network 1 are updated, and the parameters of the discriminant network 2 are not updated. Then, according to the above steps, the discriminant network 2 and the generative network 1 are updated cyclically, and after multiple iterations, the generative network 1 and the discriminant network 2 reach a Nash equilibrium.
[0040] Based on the unsupervised learning network model, low-light images and sufficient-light images can use image pairs whose semantic content does not need to be strictly consistent. That is, when inputting a low-light image, a sufficient-light image whose semantic content does not need to be strictly consistent is assigned as a reference image for the generated domain.
[0041] Figure 2 A schematic diagram of the structure of the generative network part of the low-light image enhancement model is shown.
[0042] like Figure 2As shown, in the low-light image enhancement model based on an unsupervised generative adversarial network, the generative network 1 of the generative adversarial network adopts a residual structure, and the main branch of the generative network 1 is at least two residual learning units 11 arranged in cascade; the shortcut branch of the generative network 1 connects the input and output of the generative network 1; the residual learning unit 11 adopts a residual structure, and the main branch of the residual learning unit 11 is a generator 111; the shortcut branch of the residual learning unit 11 connects the input and output of the generator 111.
[0043] When using a generative network to enhance a low-light image, the low-light image is input into a generative network 1; the generative network 1 enhances the low-light image to obtain a generated image; the generative network 1 outputs the generated image; the main branch of the generative network 1 generates a residual between the low-light image and the generated image, and the shortcut branch of the generative network 1 transmits the low-light image to the output end of the main branch of the generative network 1 and adds the residual to obtain a generated image.
[0044] The residual structure includes a main branch and a shortcut branch. The shortcut branch provides a shortcut connection to superimpose the input of the main branch on the output of the main branch. In the present invention, at least two residual learning units 11 are cascaded to form a main branch of the generative network 1. The input of this main branch is a low-light image, and the output is a residual image of the low-light image and the sufficient light image. The shortcut branch of the generative network 1 superimposes the features of the low-light image with the output of the main branch to obtain a generated image. Since the main branch part of the generative network 1 only generates residuals during the enhancement process and only learns residuals during the training process, the learning difficulty during the training process and the task intensity during the enhancement process are reduced.
[0045] The cascade structure of two or more residual learning units 11 enables each residual learning unit 11 to complete only part of the residual learning and generation tasks. For example, the input of the first residual learning unit is a low-light image, and its output is a first intermediate image. The input of the second residual learning unit is the first intermediate image, and its output is a second intermediate image. It gradually progresses to the last level of residual learning units and outputs a residual image. This structure reduces the difficulty of the learning and generation tasks of each residual learning unit 11, and can simplify the complexity of the generator 111 in the residual learning unit 11. Preferably, the generator 111 in each residual learning unit 11 adopts the same structure, so that the residual learning and generation proceed steadily.
[0046] Figure 3 A schematic diagram of the structure of a generative network including a residual learning unit is shown.
[0047] like Figure 3As shown, the generator 111 is a key component for feature extraction, conversion and reconstruction of low-light images. In the present invention, the residual learning unit 11 is constructed with the generator 111 as the main branch, and the shortcut branch connects the input and output ends of the generator 111, so that the learning and generation tasks of the generator 111 account for a small part of the learning and generation tasks of the corresponding residual learning unit. The residual structure of the residual learning unit 11 greatly reduces the complexity of the generator 111, making it easier to train the generator 111 and obtain ideal results.
[0048] In the present invention, the residual structure of the generator network 1 as a whole forms a nested residual structure with multiple residual learning units 11. Multiple residual learning units 11 are cascaded to form the main branch of the generator network 1, which learns the residuals of low-light images and sufficient-light images as a whole, further reducing the learning tasks of the generator 111 and simplifying the structure of the generator 111.
[0049] Figure 4 A schematic diagram of a generative network including a channel attention branch is shown.
[0050] like Figure 4 As shown, the generation network 1 also includes a channel attention branch, which connects the input end of the first-level generator and the output end of the last-level generator; the channel attention branch is used to take out the RGB three channels of the input low-light image and invert them, and then multiply them with the output of the last-level generator in channel order.
[0051] The channel attention branch extracts the features of the input low-light image according to the RGB channels, and after inverting them, it pays more attention to the darker areas in the low-light image, thereby strengthening the learning of the darker areas and making the output results more realistic.
[0052] Figure 5 A schematic diagram of the structure of the generator is shown.
[0053] like Figure 5 As shown, the generator 111 adopts an encoder-residual block-decoder architecture, including a downsampling layer 1111, a residual block 1112, an upsampling layer 1113 and a dimensionality reduction layer 1114, and the encoder and the decoder are symmetrically arranged.
[0054] Each generator 111 uses an encoder-decoder architecture with a residual block 1112, where the downsampling layer 1111 constitutes the encoder, and the upsampling layer 1113 and the dimension reduction layer 1114 constitute the decoder. The encoder extracts features from the input image to obtain a feature map, which is input into the residual block for conversion to obtain a new feature representation, and finally the decoder reconstructs these features into a new image as the output of the generator 111.
[0055] like Figure 5As shown, the generator 111 includes three downsampling layers 1111, nine residual blocks 1112, two upsampling layers 1113 and one dimensionality reduction layer 1114; the three downsampling layers 1111 constitute an encoder, and the two upsampling layers 1113 and one dimensionality reduction layer 1114 constitute a decoder; the downsampling layer 1111 is a CBR structure; the upsampling layer 1113 is a TBR structure, and the dimensionality reduction layer 1114 restores the output image to the same format as the input image.
[0056] The CBR structure consists of two-dimensional convolution, batch normalization and ReLu activation function. This part of the network structure learns to extract the features of the input image through downsampling in the form of convolution.
[0057] The addition of residual block 1112 alleviates the drawbacks of using only conventional convolutional networks, namely, many parameters and high network complexity, which makes training difficult, and the problem of gradient vanishing is prone to occur in the back propagation of gradient descent, making it difficult to continue training. The local jump line of residual block 1112 allows the gradient to be directly transmitted back to the previous structure without blocking.
[0058] The TBR structure consists of two-dimensional deconvolution, batch normalization and ReLu activation function. This part of the network structure maps the feature layer back to the three-channel image domain through upsampling in the form of deconvolution.
[0059] Figure 6 A schematic diagram of the global discriminator is shown.
[0060] Figure 7 A schematic diagram of a local discriminator is shown.
[0061] like Figure 6 and Figure 7 As shown, the discriminant network 2 of the generative adversarial network includes a global discriminator 21 and a local discriminator 22. The global discriminator 21 is used to discriminate the entire image; the local discriminator 22 is used to discriminate the local area of the image.
[0062] The discriminator consists of two scales, the global scale and the local small scale. The network is composed of a CBR structure, and the last convolutional layer is a single-channel output spectrum, which completes the 0 / 1 judgment of the true and false Boolean value of the image. Under the global scale discrimination, the input of the discriminator is the entire image. Under the local small scale discrimination mode, the input of the discriminator is multiple small image blocks at corresponding positions, which are taken from the entire image according to the corresponding positions.
[0063] In the present invention, the residual learning unit 11 of the cascade structure only generates residual images, and the complexity of the learning and generation tasks is greatly reduced, making it easier for the generation network to obtain ideal results during training. Using the generator 111 with the same structure, each learning and generation of residuals progresses steadily, enhancing the stability of the generation network. The discriminator adopts a multi-scale form, and the enhanced image is closer to the real image, and the local details of the image can also be more realistic. The channel attention mechanism helps the generation network 1 to automatically focus on the information of the area that needs to be enhanced from the input, and there are more learnable parameters when extracting features, so that the model can better complete the visual task.
[0064] Figure 8 A schematic diagram of the steps of a method for training a generative network based on an unsupervised generative adversarial network is shown.
[0065] like Figure 8 As shown, a method for training a generative network based on an unsupervised generative adversarial network includes: step S1, obtaining a low-light image and a sufficient-light image as training samples, wherein the semantic contents of the low-light image and the sufficient-light image do not need to be strictly consistent; step S2, inputting the low-light image into the generative network to obtain a generated image; step S3, inputting the generated image and the sufficient-light image into the discriminant network to obtain a discriminant result; step S4 iteratively adjusting the parameters of the generative network or the discriminant network until the generative network and the discriminant network reach a Nash equilibrium; wherein, based on the relationship between the low-light image and the generated image, the discriminant network is trained. Semantic consistency loss, adversarial loss between the generated image and the adequately illuminated image adjusts the parameters of the generation network; based on the adversarial loss, the parameters of the discriminative network are adjusted; wherein, the generation network adopts a residual structure, and the shortcut branch of the generation network 1 connects the input and output of the generation network 1; the main branch of the generation network 1 is at least two residual learning units 11 arranged in cascade; the residual learning unit 11 adopts a residual structure, and the main branch of the residual learning unit 11 is a generator 111, and the shortcut branch of the residual learning unit 11 connects the input and output of the generator 111.
[0066] In the present invention, an alternating training method is adopted for training the generator network and the discriminator network, that is, the parameters of the two are adjusted iteratively.
[0067] When training the generative network 1, its loss function includes two parts. One is the adversarial loss part, and the other is the semantic consistency loss part. According to the semantic consistency loss between the low-light image and the generated image, and the adversarial loss between the generated image and the sufficient light image, the parameters of the generative network 1 are adjusted. Since the present invention belongs to an unsupervised scheme, there is no strict pairing in image content between the input low-light image and the sufficient light image during the training process, which will result in the inability to limit the semantic consistency between the low-light image and the generated image during the discrimination process. Based on this, in the present invention, by limiting the consistency of the VGG19 features, the semantic content of the enhanced generated image and the input low-light image is limited to remain consistent.
[0068] VGG19 uses a pre-trained model on ImageNet. Since the VGG19 feature is insensitive to the pixel dynamic range of the input image, it can be used to limit the image content features of the low-light image and the enhanced generated image to be consistent, that is, to maintain the consistency of the semantic content of the two. Semantic content consistency loss Loss VGG The definition is as follows:
[0069]
[0070] Wherein, WH represents the width and length of the feature-limited loss image; Wi represents the i-th pixel in the width direction; Hj represents the j-th pixel in the length direction; Respectively represent the VGG features of the input low-light image and the enhanced generated image, and ||||2 represents the L2 norm measurement.
[0071] For the training of the discriminative network, training is performed only based on the adversarial loss.
[0072] In the present invention, the discriminant network adopts a global and local multi-scale discrimination scheme. Specifically, in addition to the discrimination of the size of the entire image, n random regions are additionally discriminated for local size. n is an integer greater than 1. For example, n=6, which can basically cover the main features when the number of random regions selected is small.
[0073] The adversarial loss is defined as follows:
[0074]
[0075]
[0076]
[0077] in, is the loss function of the discriminator at the global scale; is the loss function of the discriminator at the local scale; is the loss function of the generator at the global scale; is the loss function of the generator at the local scale; X r ~real means that the input image follows the distribution of the real image domain; X g ~generate means that the input image obeys the distribution of the generated image domain; D represents the discriminator network; ||||2 represents the L2 norm measurement; the local discriminator uses the average discriminator results of n random sub-blocks; E represents expectation.
[0078] N random regions are cropped from the generated image and the well-illuminated image according to corresponding positions to determine whether these small blocks belong to the real distribution or the distribution of the generated image.
[0079] The global and local multi-scale structure helps to ensure that the overall and local areas of the enhanced image are as close as possible to the real well-illuminated image. It can also adaptively enhance the areas that need to be enhanced in the input low-light image instead of simply enhancing the whole image.
[0080] For the generative network in the present invention, by setting a channel attention mechanism, the generative network is helped to automatically focus on the information of the area that needs to be enhanced from the input. By setting the local scale of the discriminant network, the area that needs to be enhanced in the input low-light image is adaptively enhanced when the generative network is trained. Through the combination of the above two solutions, the emphasis on the area that needs to be enhanced when the generative network performs image enhancement is increased, and the quality of the generated image is better improved.
[0081] In the present invention, for the generative network, by setting a nested residual structure, the learning and generation tasks of a single generator are reduced, the task difficulty is reduced, and it is easier to achieve the training purpose. By setting a cascade structure, the learning difficulty is further reduced during the training process, making it easier to achieve the training purpose.
[0082] In the present invention, when training the generative network, not only the adversarial loss is used for constraint, but also the semantic consistency loss is used to limit the consistency of the content of the front and back images, thereby enhancing the robustness of the model and overcoming the problem of poor generalization performance caused by the unsupervised training mode.
[0083] Fig. 9 A schematic diagram of the steps of a method for enhancing low-light images using a generative network is shown.
[0084] like Fig. 9As shown, it includes: step S10 inputting the low-light image into the generation network 1; step S20 the generation network 1 enhances the low-light image, the main branch of the generation network 1 generates the residual between the low-light image and the generated image, and the shortcut branch of the generation network 1 transmits the low-light image to the output end of the main branch of the generation network 1 and adds the residual to obtain the generated image. Step 30, the generation network 1 outputs the generated image.
[0085] Fig.10 The present invention is an exemplary structural block diagram of a device for training a generative network or a device for enhancing a low-light image using a generative network.
[0086] It will be appreciated that device 300 may be a single device (eg, a computing device) or a multi-function device including various peripheral devices.
[0087] like Fig.10As shown, the device 300 may include a central processing unit or central processing unit ("CPU") 311, which may be a general-purpose CPU, a dedicated CPU, or other information processing and program execution unit. Further, the device 300 may also include a large-capacity memory 312 and a read-only memory ("ROM") 313, wherein the large-capacity memory 312 may be configured to store various types of data, including various programs required for image enhancement, algorithm data, intermediate results, and operation of the device 300. ROM313 may be configured to store data and instructions required for power-on self-test of the device 300, initialization of each functional module in the system, basic input / output drivers of the system, and booting the operating system. Optionally, the device 300 may also include other hardware platforms or components, such as the tensor processing unit ("TPU") 314, the graphics processing unit ("GPU") 315, the field programmable gate array ("FPGA") 316, and the machine learning unit ("MLU") 317 shown. It is understood that although a variety of hardware platforms or components are shown in the device 300, this is only exemplary and not restrictive, and those skilled in the art can add or remove corresponding hardware according to actual needs. For example, the device 300 may only include a CPU, related storage devices, and interface devices to implement the method of using a generation network to enhance low-light images and the method of training a generation network of the present invention. In some embodiments, in order to facilitate the transmission and interaction of data with an external network, the device 300 also includes a communication interface 318, so that it can be connected to a local area network / wireless local area network ("LAN / WLAN") 305 through the communication interface 318, and then connected to a local server 306 or to the Internet ("Internet") 307 through the LAN / WLAN. Alternatively or additionally, the device 300 can also be directly connected to the Internet or a cellular network based on wireless communication technology through the communication interface 318, such as wireless communication technology based on the third generation ("3G"), the fourth generation ("4G"), or the fifth generation ("5G") generation. In some application scenarios, the device 300 may also access a server 308 and a database 309 on an external network as needed to obtain various known algorithms, data and modules, and may remotely store various data, such as various data or instructions for image enhancement.
[0088] The peripheral devices of the device 300 may include a display device 302, an input device 303, and a data transmission interface 304. In one embodiment, the display device 302 may include, for example, one or more speakers and / or one or more visual displays, which are configured to enhance low-light images using a generation network, and to perform voice prompts and / or image video display for image enhancement model training. The input device 303 may include other input buttons or controls such as a keyboard, a mouse, a microphone, a gesture capture camera, etc., which are configured to receive input of audio data and / or user instructions. The data transmission interface 304 may include, for example, a serial interface, a parallel interface or a universal serial bus interface ("USB"), a small computer system interface ("SCSI"), a serial ATA, a FireWire ("FireWire"), a PCI Express, and a high-definition multimedia interface ("HDMI"), etc., which are configured for data transmission and interaction with other devices or systems. According to the solution of the present invention, the data transmission interface 304 can receive low-light images and sufficient light images, and transmit various data or results to the device 300.
[0089] The CPU 311, mass storage 312, ROM 313, TPU 314, GPU 315, FPGA 316, MLU 317 and communication interface 318 of the device 300 can be interconnected through a bus 319, and data interaction with peripheral devices can be achieved through the bus. In one embodiment, through the bus 319, the CPU 311 can control other hardware components in the device 300 and its peripheral devices.
[0090] According to the above description in combination with the accompanying drawings, those skilled in the art can also understand that the embodiments of the present invention can also be implemented by a software program. Therefore, the present invention also provides a computer program product. The computer program product can be used to implement the present invention in combination with the accompanying drawings. Figure 8 The described method used to train the model.
[0091] It should be noted that although the operations of the method of the present invention are described in a particular order in the accompanying drawings, this does not require or imply that the operations must be performed in this particular order, or that all the operations shown must be performed to achieve the desired results. On the contrary, the steps depicted in the flow chart can be performed in a different order. Additionally or alternatively, some steps can be omitted, multiple steps can be combined into one step, and / or one step can be decomposed into multiple steps.
[0092] It should be understood that when the terms "first", "second", "third" and "fourth" are used in the claims, the specification and the drawings of the present invention, they are only used to distinguish different objects, rather than to describe a specific order. The terms "include" and "comprise" used in the specification and claims of the present invention indicate the presence of the described features, wholes, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or their collections.
[0093] It should also be understood that the terms used in this specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the claims, the singular forms of "a", "an" and "the" are intended to include the plural forms unless the context clearly indicates otherwise. It should also be further understood that the term "and / or" used in the specification of the present invention and the claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes these combinations.
[0094] Although the embodiments of the present invention are as described above, the contents are only examples used to facilitate understanding of the present invention, and are not intended to limit the scope and application scenarios of the present invention. Any technician in the technical field of the present invention can make any modifications and changes in the form and details of the implementation without departing from the spirit and scope disclosed by the present invention, but the scope of patent protection of the present invention shall still be based on the scope defined by the attached claims.
Claims
1. A method for enhancing low-light images using a generative network, characterized in that: include, Feed the low-light image into the generative network (1); The generation network (1) enhances the low-light image to obtain a generated image; The generation network (1) outputs the generated image; The generating network (1) adopts a residual structure, the main branch of the generating network (1) generates a residual between the low-light image and the generated image, and the shortcut branch of the generating network (1) transmits the low-light image to the output end of the main branch of the generating network (1) and adds the residual to obtain the generated image; The main branch of the generation network (1) is at least two residual learning units (11) arranged in cascade, and each residual learning unit generates the residual step by step; The residual learning unit (11) adopts a residual structure, the main branch of the residual learning unit (11) is a generator (111), the shortcut branch of the residual learning unit (11) connects the input end and the output end of the generator (111) of the residual learning unit (11), and the generators (111) in each residual learning unit (11) adopt the same structure; The generator (111) adopts an encoder-residual block-decoder architecture. It includes three downsampling layers (1111), nine residual blocks (1112), two upsampling layers (1113) and one dimension reduction layer (1114); The three downsampling layers (1111) constitute an encoder, the two upsampling layers (1113) and one dimension reduction layer (1114) constitute a decoder, and the encoder and the decoder are symmetrically arranged; The downsampling layer (1111) is a CBR structure; The upsampling layer (1113) is a TBR structure.
2. The method for enhancing low-light images using a generative network according to claim 1, characterized in that: The generation network (1) further includes a channel attention branch, wherein the channel attention branch is connected to an input end of the first-level generator and an output end of the last-level generator (111); The channel attention branch extracts the RGB three channels of the low-light image and inverts them, and then multiplies them with the output of the last-level generator (111) in sequence according to the channel order.
3. A method for training a generative network based on an unsupervised generative adversarial network, characterized in that: include, Obtain low-light images and sufficient-light images as training samples. The semantic contents of the low-light images and sufficient-light images do not need to be strictly consistent. Inputting the low-light image into a generation network to obtain a generated image; Inputting the generated image and the sufficient illumination image into a discrimination network to obtain a discrimination result; Iteratively adjust the parameters of the generating network or the discriminative network until the generating network and the discriminative network reach Nash equilibrium; Among them, based on the semantic consistency loss between the low-light image and the generated image, the adversarial loss between the generated image and the fully illuminated image adjusts the parameters of the generation network; Adjust the parameters of the discriminative network based on the adversarial loss; The generative network adopts a residual structure, and the shortcut branch of the generative network (1) connects the input end and the output end of the generative network (1); the main branch of the generative network (1) is at least two residual learning units (11) arranged in cascade; the residual learning unit (11) adopts a residual structure, the main branch of the residual learning unit (11) is a generator (111), and the shortcut branch of the residual learning unit (11) connects the input end and the output end of the generator (111) of the residual learning unit (11); The generator (111) adopts an encoder-residual block-decoder architecture. It includes three downsampling layers (1111), nine residual blocks (1112), two upsampling layers (1113) and one dimension reduction layer (1114); The three downsampling layers (1111) constitute an encoder, the two upsampling layers (1113) and one dimension reduction layer (1114) constitute a decoder, and the encoder and the decoder are symmetrically arranged; The downsampling layer (1111) is a CBR structure; The upsampling layer (1113) is a TBR structure.
4. The method for training a generative network based on an unsupervised generative adversarial network according to claim 3, characterized in that: The discriminant network (2) of the generative adversarial network includes a global discriminator (21) and a local discriminator (22). The global discriminator (21) is used to discriminate the entire image; The local discriminator (22) is used to discriminate the local area of the image.
5. The method for training a generative network based on an unsupervised generative adversarial network according to claim 4, characterized in that: The global discriminator (21) and the local discriminator (22) are CBR structures, and the last convolutional layer is a single-channel output spectrum.
6. The method for training a generative network based on an unsupervised generative adversarial network according to claim 3, characterized in that: The consistency of VGG19 features is used to maintain the semantic content without deviation, and the semantic content consistency loss Loss VGG The definition is as follows: WH represents the width and length of the image; Wi represents the i-th pixel in the width direction; Hj represents the jth pixel in the length direction; VGG features representing the input low-light image; Represents the enhanced VGG features of the generated image; || ||2 represents the L2 norm measure.
7. The method for training a generative network based on an unsupervised generative adversarial network according to claim 3, characterized in that: The adversarial loss is defined as follows: in, is the loss function of the discriminator at the global scale; is the loss function of the discriminator at the local scale; is the loss function of the generator at the global scale; is the loss function of the generator at the local scale; X r ~real means that the input image follows the distribution of the real image domain; X g ~generate means that the input image follows the distribution of the generated image domain; D represents the discriminator network; E stands for expectation; || ||2 represents L2 norm measurement; The local discriminator uses the average discriminator results of n random sub-blocks, where n is an integer greater than 1.
8. A device for enhancing low-light images using a generative network, characterized in that: include: processor; as well as A memory storing program instructions for a method for enhancing a low-light image using a generation network, wherein when the program instructions are executed by the processor, the device implements the method according to claim 1 or 2.
9. A computer-readable storage medium, characterized in that: Program instructions for a method for enhancing a low-light image using a generation network are stored thereon, and when the program instructions are executed by one or more processors, the method according to claim 1 or 2 is implemented.
10. A device for training a generative network based on an unsupervised generative adversarial network, characterized in that: include: processor; as well as A memory storing program instructions for training a generative network based on an unsupervised generative adversarial network, wherein when the program instructions are executed by the processor, the device implements the method according to any one of claims 3 to 7.
11. A computer-readable storage medium, characterized in that: Program instructions for training a generative network based on an unsupervised generative adversarial network are stored thereon, and when the program instructions are executed by one or more processors, the method according to any one of claims 3-7 is implemented.
Citation Information
Patent Citations
Unsupervised learning method and system for low-illumination image enhancement
CN113313657A