Low-illumination image enhancement method based on cyclic generative adversarial network

By adopting a cyclic generation adversarial network-based method in low-illumination image enhancement technology, combining cyclic consistency loss, MCFE, AESA and SELE modules, data shortage and noise problems are solved, and high-quality low-illumination image enhancement is achieved.

CN120219196AInactive Publication Date: 2025-06-27XIAN UNIV OF SCI & TECH
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
CN202510284540.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing low-illumination image enhancement technologies rely on paired data sets, resulting in insufficient data and insufficient generalization capabilities of the model, and the enhanced images are prone to noise and overexposure problems.

Method used

Using a circular generation adversarial network-based approach, the cyclic consistency loss and multi-scale convolution feature enhancement module (MCFE) are introduced to eliminate the dependence on paired training data, and dynamic effective self-attention aggregation module (AESA) and SE-Res2Block-based lighting enhancement module (SELE) are added to the generator to improve image feature extraction and detail recovery capabilities.

Benefits of technology

It effectively alleviates the problem of insufficient data sets, improves the quality of generated images and the robust performance of the model, reduces noise and overexposure, and enhances the contrast and detail recovery capabilities of the image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219196A_ABST
    Figure CN120219196A_ABST
Patent Text Reader

Abstract

The invention discloses a low-illumination image enhancement method based on a cyclic generative adversarial network, and the method comprises the following specific steps: 1, building a model, a low-illumination data set and a normal light data set; step 2, extracting a normal light image # imgabs0 # and a low illumination image # imgabs1 #; step 3, acquiring an image by using a generator # imgabs2 # and a generator # imgabs3 #; step 4, sending the generated image and the initial image into a discriminator, and calculating the adversarial loss # imgabs6 # and # imgabs7 # of the discriminator # imgabs4 # and # imgabs5 #; step 5, calculating generator adversarial loss functions # imgabs10 # and # imgabs11 # of the generator # imgabs8 # and # imgabs9 #, and calculating the generator adversarial loss functions # imgabs10 # and # imgabs11 # of the generator # imgabs9 #; the circulation consistency loss # imgabs 12 # and the circulation consistency loss # imgabs 13 # are calculated; 6, calculating the self-feature retention loss # imgabs14 # and the self-feature retention loss # imgabs15 #; step 7, calculating the total loss # imgabs18 # of the generator # imgabs16 # and the generator # imgabs17 #; and step 8, alternately training the discriminator and the generator, generating a real image, and completing enhancement. According to the low-illumination image enhancement method based on the cyclic generative adversarial network, the problems of dependence on a paired data set and noise after image enhancement in the prior art are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image enhancement, and in particular relates to a low-illumination image enhancement method based on a cyclic generative adversarial network. Background Art

[0002] Most of the current low-light deep learning algorithms are based on paired data. The collection of data sets requires images of the same scene during the day and at night, and requires that objects in the scene do not move or change over time. While ensuring quality, the large data volume requirements must also be met. The stringent standards for data sets result in a small amount of available data, making it difficult to cover all aspects of life scenes, limiting the generalization ability of neural networks, and resulting in poor low-light image enhancement effects. Faced with the shortage of data sets, conventional data enhancement methods have limited improvements and have failed to fundamentally solve the problem. In order to overcome this problem, existing technologies use high-dynamic image data sets to artificially distinguish between night images and daytime images; or through simulation experiments, nonlinearly reduce the brightness of normal images to simulate low-light scenes. Although the problem of insufficient data can be alleviated to a certain extent, the generated samples are different from those obtained in natural scenes, and it cannot be guaranteed that the trained model is suitable for real scenes.

[0003] In view of the above problems, inspired by unsupervised methods, we decided to introduce a generative adversarial network method for low-light image enhancement. This method establishes an unpaired mapping between low-light and normal-light image spaces without relying on precisely paired images, which can alleviate the problem of insufficient data sets, effectively improve the quality of generated images, and improve the robustness of the model. The generative adversarial network allows the generator and the discriminator to be trained in an adversarial manner. The generator strives to generate more realistic fake data, while the discriminator is responsible for distinguishing between real data and generated data. The two networks promote each other's training in adversarial ways. In terms of low-light image enhancement, the EnlightenGAN network structure adopted by the existing technology eliminates the dependence on paired training data for the first time, making the flexibility of generated images unprecedentedly improved and adaptable to a variety of scenarios. The proposal of RetinexGAN combines the generative adversarial network with the Retinex model, which includes a decomposition network and two discrimination networks. The decomposition network decomposes the Retinex model, and the discriminator discriminates between the illumination component and the reflection component. However, in order to adapt to images of different sizes from the dataset, the above methods need to preprocess the images before sending them to the network, including resizing and cropping, which can easily lead to distortion and information loss in the generated images, and are overly dependent on the paired dataset. Therefore, an enhancement method based on zero-reference depth curve estimation is adopted. Although it avoids the paired data required for training and brings convenience to image enhancement in complex environments, the enhanced images are overexposed and the noise problem after enhancement is not considered. Summary of the invention

[0004] The object of the present invention is to provide a low-light image enhancement method based on a cycle generative adversarial network, which solves the problems of dependence on paired data sets and noise in the image after enhancement in the prior art.

[0005] The technical solution adopted by the present invention is a low-light image enhancement method based on a cycle generative adversarial network, which is specifically implemented according to the following steps: Step 1, establish a low-light image enhancement model and a low-light and normal light data set; Step 2, extract a normal light image from the normal light data set , and extract a low-light image from the low-light data set ; Step 3, use generator and generator to obtain a generated image and a reconstructed image; Step 4, send the generated image and the initial image into the global-local discriminator and , calculate the adversarial loss of discriminator and the adversarial loss of discriminator , and update the discriminator parameters; Step 5, calculate the generator adversarial loss functions , of generator , ; and calculate the cycle consistency loss , between the initial image and the reconstructed image; Step 6, calculate the self-feature retention loss , between the initial image and the generated image; Step 7, calculate the total loss , of generator , and update the generator parameters; Step 8, alternately train the discriminator and the generator to generate a real image and complete the enhancement.

[0006] The feature of the technical solution of the present invention is further that, Step 3 is specifically as follows: Step 3.1, pass the normal light image through generator , extract the image features of the normal light image to obtain a generated low-light image ; pass the low-light image through generator , extract the low - illumination image of the image features to obtain the generated normal - light image , and the specific formula is as follows: (1) (2) In the formula, represents the low - illumination image generated by the generator , , represents the normal - light image generated by the generator ; ; Step 3.2, pass the generated low - illumination image through the generator , extract the image features of the generated low - illumination image to obtain the reconstructed normal - light image ; pass the generated normal - light image through the generator , extract the image features of this image to obtain the reconstructed low - illumination image , and the specific formula is as follows: (3) (4) In the formula, represents the reconstructed normal - light image obtained after passing the normal - light image through the generator to obtain the generated low - illumination image and then passing through the generator again; ; represents the reconstructed low - illumination image obtained after passing the low - illumination image through the generator to obtain the generated low - illumination image and then passing through the generator again .

[0007] The image features in Step 3 are all edges, textures, colors, corner points, and spatial hierarchies.

[0008] The generator and It combines the generator structure of the CycleGAN network with the U-Net network, MCFE multi-scale convolutional feature enhancement module, AESA dynamic effective self-attention aggregation module, and illumination enhancement module (SELE) based on SE-Res2Block. The MCFE multi-scale convolutional feature enhancement module is added during the encoding process of the generator to extract multi-scale image features; the AESA module is added during the decoding process to establish long-range dependencies and extract image feature information; the SELE module is also added to the generator to enhance the detail features of the image while adjusting the image brightness; finally, the self-feature retention loss is introduced to enhance the feature extraction ability and detail processing ability for the low-light image enhancement task.

[0009] The SELE module uses 9 SE-Res2Block modules, and the dilation rates of the 9 SE-Res2Block modules are 1, 2, 4, 8, 8, 4, 2, 1, 1 in sequence.

[0010] Step 4 is specifically as follows: The generated low-light image obtained and the low-light image L are sent into the global-local discriminator to discriminate the authenticity of the generated image , calculate the loss of the discriminator ; the generated normal-light image obtained and the normal-light image are sent into the global-local discriminator to discriminate the authenticity of the generated image, calculate the loss of the discriminator , update the discriminator parameters, and the adversarial loss of the discriminator and are specifically expressed as follows: and are specifically as follows: (5) (6) In the formula, is the prediction probability that the global-local discriminator judges the real normal-light image , and the prediction probability is 0 - 1, the closer to 1 is better; is the prediction probability that the discriminator judges the real low-light image , and the prediction probability is 0 - 1, the closer to 1 is better; is the prediction probability that the global-local discriminator judges the generated normal-light image , and the prediction probability is 0 - 1, the closer to 0 is better; is the discriminator Judge the predicted probability of generating a low - illumination image. The predicted probability ranges from 0 to 1, and the closer it is to 0, the better.

[0011] The generator adversarial loss function in step 5 and The specific formula is as follows: (7) (8) Among them, represents the discriminator 's predicted probability of the generated low - illumination image ( ). The goal of the generator is to maximize , that is, to make the discriminator judge the generated low - illumination image as a real image. The predicted probability ranges from 0 to 1, and the closer it is to 1, the better. represents the discriminator 's predicted probability of the generated normal - illumination image ( ). The goal of the generator is to maximize , that is, to make the discriminator judge the generated normal - light image as true. The predicted probability ranges from 0 to 1, and the closer it is to 1, the better.

[0012] The cycle - consistency loss between the initial image and the reconstructed image in step 5 is divided into: the cycle - consistency loss between the original normal - illumination image and the reconstructed normal - illumination image , the cycle - consistency loss between the original low - illumination image and the reconstructed low - illumination image , and Specifically, it is expressed as: (9) (10) Among them, is the reconstructed normal - light image obtained after the normal - light image passes through the generator to get the generated low - illumination image and then passes through the generator again; is the low - illumination image after passing through the generator , the resulting low-light image After that, pass through the generator again The reconstructed low-light image .

[0013] Step 6 is as follows: Calculate the normal image and generate low light images The self-feature preservation loss between , calculate low-light image and generate a normal light image The self-feature preservation loss between ; The specific formula is as follows: (11) (12) in, Representation Generator Generated low light image ; Representation Generator Generated normal light image ; Represents the feature map extracted from the pre-trained VGG16 model; Indicates Max pooling layer.

[0014] Step 7 is as follows: Generator , The total loss function combines the generator adversarial loss , , cycle consistency loss , and the self-feature preservation loss , , the total loss function of the generator The specific expressions are as follows: (13) Compared with the prior art, the present invention has the following beneficial effects: (1) The low-light image enhancement method based on the cyclic generative adversarial network provided by the present invention introduces a cycle consistency loss and eliminates the dependence on paired training data. In the practical application of low-light image enhancement, paired data sets are often difficult to obtain, and it is difficult to find one-to-one corresponding images. The present invention solves this problem by using cycle consistency loss to transform the image from the low-light domain to the Mapping to the normal ray domain , and then mapped back to the low-light domain , it should ultimately be as similar as possible to the original image. By introducing two generators and their corresponding discriminators, the reversibility of the conversion process from one domain to another is ensured, and even without paired data, the similarity and consistency between images can be maintained in this way.

[0015] (2) The low-light image enhancement method based on the cyclic generative adversarial network provided by the present invention utilizes the powerful feature extraction and aggregation ability of the illumination enhancement module based on SE-Res2Block in the generator to globally model the image feature information, improving the model's ability to capture global information and local details, accurately extracting features such as the color, texture, and shape of the image, thereby increasing the contrast of the image, reducing noise, and restoring details, and improving the drawback of the EnlightenGAN algorithm being limited by the monotonicity of the generative network, where the generated images are prone to obvious noise and overexposure.

[0016] (3) The low-light image enhancement method based on the cyclic generative adversarial network provided by the present invention incorporates MCFE and AESA attention mechanisms during the encoder-decoder process of the generator. The MCFE module enables the model to accurately capture image information at different scales by extracting multi-scale image features, improving the network's ability to enhance the illumination of objects of different sizes, avoiding the loss of illumination information, and combining attention mechanisms in the channel and spatial dimensions, enabling the model to adaptively adjust the attention to different regions and channels, enhancing the suppression of noise and interference information, making the edges of the generated images clearer and the detail restoration more realistic. The AESA module globally models the image feature information by effectively establishing long-range dependencies, improving the model's ability to extract image feature information, expanding the receptive field of the features after upsampling, and improving the drawback of the convolutional neural network having a limited receptive field and only being able to focus on local features. Additionally, using the kernel function optimization strategy avoids adding excessive computational burden, enabling the model to effectively utilize spatial-illumination information from different scales, thereby achieving more realistic low-light image enhancement. The MCFE and AESA modules solve the problems of blurry detail textures, easy generation of noise and artifacts in the generated normal-light images due to the easy loss of low-light image feature information during the encoder-decoder process.

[0017] (4) The low-light image enhancement method based on the cyclic generative adversarial network provided by the present invention improves the discriminator using a global-local network architecture, enhancing the illumination of the local area while improving the illumination of the global image, improving the global light and adaptively enhancing the local area, thereby enhancing the ability to perceive image details and improving the low-light enhancement performance of the network. It solves the problems of local over-enhancement and local under-enhancement in the images caused by the limited enhancement effect of the single discriminator in the cyclic generative adversarial network algorithm CycleGAN and the insufficiently comprehensive information fed back to the generator. Brief Description of the Drawings

[0018] Figure 1 This is the flow framework diagram of the low - illumination image enhancement method based on the cyclic generative adversarial network of the present invention; Figure 2 This is the structural schematic diagram of the generator of the present invention; Figure 3 This is the structural schematic diagram of the MCFE multi - scale convolutional feature enhancement module of the present invention; Figure 4 This is the structural schematic diagram of CBAM of the present invention; Figure 5 This is the structural schematic diagram of the AESA dynamic effective self - attention aggregation module of the present invention; Figure 6 This is the structural schematic diagram of SE - Res2Block of the present invention; Figure 7 This is the network structural schematic diagram of the global - local discriminator of the present invention; Figure 8 This is the forward propagation and reverse reconstruction flow chart of Embodiment 6 of the present invention; Figure 9 This is the input image schematic diagram of Embodiment 6 of the present invention; Figure 10 This is the image generated by the generator of Embodiment 6 of the present invention; Figure 11 This is the reconstructed image of the generator of Embodiment 6 of the present invention; Figure 12 This is the low - illumination image enhancement comparison schematic diagram of the present invention. Detailed Description of the Invention

[0019] The present invention will be described in detail below in conjunction with the drawings and specific embodiments.

[0020] Embodiment 1 The present invention provides a low - illumination image enhancement method based on a cyclic generative adversarial network, and the specific steps are as follows: Step 1, establish a low - illumination image enhancement model and a low - illumination and normal - light dataset containing 500 low - light / normal - light images.

[0021] Step 2, extract a normal - light image from the normal - light dataset , and extract a low - illumination image from the low - illumination dataset ; Step 3, use the generator and the generator to obtain the generated image and the reconstructed image; Step 4, send the generated image and the initial image into the global - local discriminator and Calculate the discriminator 's adversarial loss and the discriminator 's adversarial loss . Update the discriminator parameters; Step 5, calculate the generator , 's generator adversarial loss function , ; and calculate the cycle consistency loss between the initial image and the reconstructed image , ; Step 6, calculate the self-feature retention loss between the initial image and the generated image , ; Step 7, calculate the total loss , of the generator , and update the generator parameters; Step 8, alternately train the discriminator and the generator to generate real images and complete enhancement.

[0022] Example 2 Based on Example 1, Step 3 is specifically as follows: Step 3.1, pass the normal light image through the generator to extract the image features of the normal light image and obtain the generated low-light image ; pass the low-light image through the generator to extract the image features of the low-light image and obtain the generated normal light image , and the specific formula is as follows: (1) (2) In the formula, represents the low-light image generated by the generator , represents the normal light image generated by the generator ; Step 3.2, pass the generated low-light image through the generator to extract the image features of the generated low-light image and obtain the reconstructed normal light image ; pass the generated normal light image through the generator , extract features such as edges, textures, colors, corner points, and spatial hierarchies of the image to obtain the reconstructed low - illumination image , which is specifically expressed by the following formula: (3) (4) In the formula, represents the normal - light image after passing through the generator , the generated low - illumination image obtained , and then passing through the generator again to obtain the reconstructed normal - light image ; represents the low - illumination image after passing through the generator , the generated low - illumination image obtained , and then passing through the generator again to obtain the reconstructed low - illumination image ; The image features in step 3 are all edges, textures, colors, corner points, and spatial hierarchies.

[0023] Step 4 is specifically as follows: Send the obtained generated low - illumination image and the low - illumination image into the global - local discriminator to discriminate the authenticity of the generated image , and calculate the loss of the discriminator ; Send the obtained generated normal - light image and the normal - light image into the global - local discriminator to discriminate the authenticity of the generated image, and calculate the loss of the discriminator . Update the parameters of the discriminator. The adversarial losses and are specifically expressed as follows: (5) (6) In the formula, is the predicted probability that the global - local discriminator judges the real normal - light image . The predicted probability is between 0 and 1, and the closer it is to 1, the better; is the predicted probability that the discriminator judges the real low - illumination image . The predicted probability is between 0 and 1, and the closer it is to 1, the better; is the global - local discriminator Judge the predicted probability of generating a normal light image The predicted probability ranges from 0 to 1, and the closer it is to 0, the better; It is the discriminator Judge the predicted probability of generating a low-illumination image The predicted probability ranges from 0 to 1, and the closer it is to 0, the better; among them, and the larger the better.

[0024] Example 3 Based on the above example, in step 5, the generator adversarial loss function and The specific formula is as follows: (7) (8) Among them, represents the discriminator for the predicted probability of generating a low-illumination image ( ) The goal of the generator is to maximize , that is, to make the discriminator judge the generated low-illumination image as a real image. The predicted probability ranges from 0 to 1, and the closer it is to 1, the better; represents the discriminator for the predicted probability of generating an image with normal illumination ( ) The goal of the generator is to maximize , that is, to make the discriminator judge the generated normal light image as true. The predicted probability ranges from 0 to 1, and the closer it is to 1, the better, and the smaller the better.

[0025] In step 5, the cycle consistency loss between the initial image and the reconstructed image is divided into: the cycle consistency loss between the original normal illumination image and the reconstructed normal illumination image , the cycle consistency loss between the original low-illumination image and the reconstructed low-illumination image , and Specifically expressed as: (9) (10) Among them, is the normal light image After passing through the generator , the generated low-light image is obtained After that, it passes through the generator again The reconstructed normal light image obtained ; is the low-light image After passing through the generator , the generated low-light image is obtained After that, it passes through the generator again The reconstructed low-light image obtained . , The smaller the better.

[0026] Step 6 is specifically as follows: Calculate the self-feature retention loss between the normal image and the generated low-light image , and calculate the self-feature retention loss between the low-light image and the generated normal light image ; The specific formula is as follows: (11) (12) Among them, represents the low-light image generated by the generator ; ; represents the normal light image generated by the generator ; ; represents the feature map extracted from the pre-trained VGG16 model; represents the th max pooling layer.

[0027] Example 4 Based on the above example, step 7 is specifically as follows: The total loss function of the generators , combines the generator adversarial loss , , the cycle consistency loss , and the self-feature retention loss , , and the total loss function of the generator is specifically expressed as follows: (13) Step 8, continue the iteration and alternately train the discriminator and the generator. At each update, the discriminator is trained based on the images generated by the current generator, while the generator adjusts its generation strategy according to the feedback from the discriminator until the generated images are realistic enough.

[0028] Such as Figure 1As shown in the figure, this model is based on CycleGAN, and designs a cyclic image enhancement main network architecture, which is combined with a U-Net network, an MCFE multi-scale convolutional feature enhancement module, an AESA dynamic effective self-attention aggregation module, an illumination enhancement module (SELE) based on SE-Res2Block, and a global-local network architecture. More specifically, first, the MCFE and AESA modules are designed in the encoding-decoding process of the generator. The MCFE multi-scale convolutional feature enhancement module is added in the generator encoding process to extract multi-scale image features, enabling the model to accurately capture image information at different scales, improving the network's illumination enhancement ability for objects of different sizes, avoiding the loss of illumination information, and combining the attention mechanisms in the channel and spatial dimensions, enabling the model to adaptively adjust the attention to different regions and channels, enhancing the suppression of noise and interference information, making the edges of the generated image clearer and the details restored more realistically; and, by adding the AESA module in the decoding process, long-range dependencies are effectively established, enhancing the model's ability to extract image feature information, expanding the receptive field of the features after upsampling, improving the shortcoming of the limited receptive field of the convolutional neural network that can only focus on local features, and using the kernel function optimization strategy to avoid adding too much computational burden, enabling the model to effectively utilize the spatial-illumination information from different scales, thus achieving more realistic low-light image enhancement. The MCFE and AESA modules solve the problems such as the blurring of details and textures, the easy generation of noise and artifacts in the generated normal-light image due to the loss of low-light image feature information during the encoding-decoding process. Secondly, an illumination enhancement module based on SE-Res2Block is designed, which uses its powerful feature extraction and aggregation ability to globally model the image feature information, enhancing the model's ability to capture global information and local details, accurately extracting features such as the color, texture, and shape of the image, thereby improving the contrast of the image, reducing noise, and restoring details. In addition, the discriminator is improved using a global-local network architecture, while improving the illumination of the global image, ensuring that the illumination of the local area is effectively enhanced, improving the global light and adaptively enhancing the local area, thereby enhancing the ability to perceive image details and improving the low-light enhancement performance of the network. Finally, a self-feature retention loss is introduced to keep the image content features unchanged before and after enhancement, further enhancing the feature extraction ability and detail processing ability for the low-light image enhancement task. The present invention improves on the CycleGAN algorithm, making the details and textures of the enhanced image more real and clear, the illumination more uniform, the frequency of noise and artifacts lower, and the contrast higher.

[0029] Among them, due to the simple structure of the generator of the basic CycleGAN network, the effect of processing image details is not good, and for low-illumination images, there are generally disadvantages such as dark brightness, blurred detail features, a large amount of noise, and artifacts. Therefore, the generator of the CycleGAN network is improved to make the enhanced network more suitable for the task requirements of low-illumination image enhancement.

[0030] To reduce the problem of loss of image information features, as Figure 2 shown, an MCFE multi-scale convolutional feature enhancement module is designed to be added during the encoding process of the generator. This module enhances the feature representation ability of the convolutional neural network (CNN) when processing images, improves the expression of important features, and suppresses irrelevant information. As Figure 4 shown, by combining the Channel Attention and Spatial Attention mechanisms, that is, the CBAM structure, the module combines convolutional feature extraction at different scales with the attention mechanism, and can effectively capture important information in the image; during the decoding process, an AESA dual-channel attention module is designed to be added to improve the model's ability to extract image feature information, expand the receptive field of features, and reduce the computational burden, enabling the model to effectively utilize spatial-illumination information from different scales, thereby achieving more realistic low-illumination image enhancement. It solves the problems that due to the encoding-decoding process, it is easy to lose the feature information of low-illumination images, resulting in blurred detail textures, easy generation of noise, artifacts, etc. in the generated normal-light images; an illumination enhancement module based on SE-Res2Block is designed to be added to the generator, and by using its powerful feature extraction and aggregation ability, global modeling of image feature information is carried out, improving the model's ability to capture global information and local details, accurately extracting features such as the color, texture, and shape of the image, thereby improving the contrast of the image, reducing noise, and restoring details.

[0031] In the generator network, the downsampling part is mainly used for the extraction and compression of image features. And to better focus on the useful information in the image, the MCFE module is adopted. In the original CycleGAN, the downsampling module consists of two convolutional layers, but the original downsampling module has a simple structure and limited ability to extract image features. To better extract the depth features of the image, the MCFE module is introduced. First, it passes through the first convolutional layer with a kernel size of 7×7, which can effectively reduce the number of parameters while increasing the receptive field. Then it passes through two 3×3 convolutional layers and the MCFE module. The schematic diagram of the module structure is as Figure 3 shown.

[0032] The MCFE module extracts feature information at different scales through multi-scale convolutional kernels of 3x3, 5x5, and 7x7, enabling the module to adapt to image information at different scales, improving the network's enhancement ability for objects of different sizes, and avoiding the loss of detailed information. By combining the Channel Attention and Spatial Attention mechanisms, the module combines convolutional feature extraction at different scales with the attention mechanism, enabling the network to focus on important spatial regions and channel information. First, the input feature map is input into the channel attention mechanism, compressed in two dimensions for the feature map, obtaining two feature descriptions of different dimensions, which are transmitted through pooling to a multi-layer perceptron network, and weights are assigned to the feature vectors respectively to obtain weight coefficients , and finally, the channel attention is multiplied by the input to obtain the feature map F1 to achieve the transmission of feature information in the channel dimension. The calculation process is as follows.

[0033] (14) (15) Among them, is the sigmoid activation function, represents the multi-layer perceptron, is the element-wise multiplication. After passing through the channel attention mechanism, the spatial attention mechanism then processes the output feature map of the channel attention mechanism in the spatial domain, performs global max pooling and global average pooling on the feature map F1 again in the spatial dimension to obtain a two-dimensional feature map, and then performs a convolution operation to obtain the weight coefficient Ms to achieve feature enhancement in the spatial dimension. Finally, the spatial attention and are used to obtain the new feature map . The calculation process is as follows.

[0034] (16) (17) Among them, is the sigmoid activation function, represents the convolution operation, is the element-wise multiplication.

[0035] Since most of the current low-light image enhancement methods are based on CNN methods, due to some inherent excellent characteristics of CNN, they are naturally suitable for a variety of computer vision tasks. However, since the convolution operation in CNN can only capture local information and cannot establish long-range connections of the global image, this long-range global feature modeling ability is the key to reducing noise and artifacts in low-light images. To address the above problems, as Figure 6As shown in the figure, a lighting enhancement module based on SE-Res2Block (SELE) (Lighting enhancement module based on SE-Res2Block, hereinafter referred to as the SELE module) is proposed. A total of 9 SE-Res2Blocks are used in the SELE module. The SE-Res2Block module is as Figure 6 shown. The module extracts basic features such as edges and textures from the input image through convolutional operations, and applies dilated convolutions to expand the receptive field in the generator network, enabling the network to capture long-range dependencies across the entire image without increasing the number of parameters. The dilation rates of the 9 SE-Res2Blocks are 1, 2, 4, 8, 8, 4, 2, 1, 1, which helps to balance the extraction of local and global features and avoid the loss of details caused by the premature expansion of the network's receptive field. Finally, it enters the SE-block, which uses global information to selectively emphasize important features and suppress irrelevant information, enhancing the model's ability to focus on the most important feature enhancement. The entire module ensures that the network can retain the key information in the original low-light image through skip connections, preventing the network from overfitting during the enhancement process and losing necessary details.

[0036] During the generator decoding process, to address the problem of insufficient recovery of detailed features in the upsampling process of the original CycleGAN generator network, an AESA dynamic effective self-attention aggregation module is introduced. Its structure is as Figure 5 shown. The improved upsampling module consists of a transposed convolution block and an AESA module. The transposed convolution block consists of a 3×3 convolutional kernel and a non-linear activation function ReLU. After two upsamplings, in the last convolutional block, ReflectionPad is used to increase the image resolution, the convolutional kernel is 7×7 to restore the image size, the instance normalization layer is removed, and the activation function is replaced with tanh to maintain the non-linear monotonic increase and decrease relationship between the input and output. Introducing the dynamic effective self-attention aggregation module AESA can enhance the model's ability to extract image feature information, expand the receptive field of features, and reduce the computational burden, enabling the model to effectively utilize spatial-illumination information from different scales, thereby improving the ability to recover detailed features of the image during the upsampling process and achieving more realistic low-light image enhancement. The AESA dynamic effective self-attention aggregation module is as Figure 5 shown.

[0037] To reasonably enhance the local region brightness while increasing the global brightness of the image, a global-local discriminator is used to judge the "authenticity" of the overall image and the local region of the image respectively, realizing the enhancement of non-uniform low-light images. The network structure diagram is as follows Figure 7 shown.

[0038] In the network structure of the global-local discriminator, both the global discriminator and the local discriminator use PatchGAN, where: the global discriminator judges the style of the overall image; the local discriminator randomly crops 5 local region image patches from the overall image and judges their authenticity. Thus, while improving the illumination of the global image, it ensures that the illumination of the local region is effectively enhanced, avoiding shadow and overexposure phenomena.

[0039] Finally, the self-feature retention loss is introduced to keep the image content features unchanged before and after enhancement, further enhancing the feature extraction ability and detail processing ability for the low-light image enhancement task. This loss function mainly uses the VGG16 network to extract the features of the input image and the generated image, and the extracted features are used to calculate the loss between the real image and the generated image. The key idea of this method is to ensure that the generated image remains similar to the original image at the semantic level without losing important content information by restricting the distance between the real image and its corresponding generated image in the VGG feature space. Since the VGG16 network mainly focuses on the shape and structural features of the image and is less sensitive to brightness, calculating the difference between the low-light image and the enhanced image in the VGG feature space helps to retain the structural or semantic information of the original image that the enhanced image may lose.

[0040] Example 5 Taking the conversion of low-light to normal-light images in this example as an illustration, as Figure 8 shown, that is, converting from a low-light image to a normal-light image. During the forward propagation process, the input of the generator network is the low-light image , and this low-light image is passed through this network and converted into the corresponding generated normal-light image . First, in the encoding process of the generator, through three convolutional layers, where the MCFE multi-scale convolutional feature enhancement module is introduced to improve the network's enhancement ability for objects of different sizes and avoid the loss of detail information; then, through the illumination enhancement module based on SE-Res2Block, without increasing the number of parameters, it expands the receptive field in the generator network, enabling the network to capture the long-range dependencies of the entire image and enhancing the model's ability to focus on the most important feature enhancements; finally, in the decoding process, through three convolutional layers, where the AESA dynamic effective self-attention aggregation module is introduced to reduce the computational burden, improve the model's ability to extract image feature information, expand the receptive field of the features, and complete the generation of the fake normal-light image . And the original normal-light image and the generated fake normal-light image are input into the discriminator , by Identify its authenticity. After the forward propagation is completed, during the reverse reconstruction process of the network, the generated fake normal light image is used as the input. Through the generator with a similar network structure to that in the forward propagation process , a reconstructed low-light image similar to the original low-light image is obtained . In the network architecture, to control the network stability and regulate the performance of the discriminator and generator as well as the low-light enhancement performance, cycle consistency, adversarial loss, and self-feature retention are used as loss functions. As Figures 9 - 12 shown, the images under different light conditions are generated by using the low-light image enhancement method based on the cycle generative adversarial network of the present invention

[0041] Example 6 This example provides a low-light image enhancement method based on the cycle generative adversarial network. As Figures 9 - 12 shown, the specific steps are as follows Step 1: Establish a low-light image enhancement model based on the cycle generative adversarial network and prepare a low-light and normal light dataset

[0042] Step 2: Extract a pair of unmatched images. The first one is the normal light image in the normal light dataset , and the other one is the low-light image in the low-light dataset ; Step 3: Pass the normal light image through the generator to extract the edge, texture, color, corner, and spatial hierarchy features of this image, and obtain the generated low-light image ; Pass the low-light image through the generator to extract the edge, texture, color, corner, and spatial hierarchy features of this image, obtain the low-light image enhancement result, and generate the normal light image , and the formula is as follows (1) (2) In the formula, represents the low-light image generated by the generator , represents the normal light image generated by the generator ; Step 4: Pass the generated low-light image through the generator to extract the edge, texture, color, corner, and spatial features of this image, and obtain the reconstructed normal light image ; The normal light image will be generated Pass through the generator , extract features such as the edges, textures, colors, corner points, and spatial hierarchies of the image to obtain the reconstructed low-light image , the formula is as follows; (3) (4) In the formula, represents the normal light image Pass through the generator , and the generated low-light image obtained After that, pass through the generator again to obtain the reconstructed normal light image ; represents the low-light image Pass through the generator , and the generated low-light image obtained After that, pass through the generator again to obtain the reconstructed low-light image .

[0043] Step 5, send the obtained generated low-light image and the low-light image into the global-local discriminator , discriminate the authenticity of the generated image , and calculate the adversarial loss of the discriminator . Send the obtained generated normal light image and the normal light image into the global-local discriminator , discriminate the authenticity of the generated image, and calculate the adversarial loss of the discriminator . Update the parameters of the discriminator. The adversarial losses , are as follows: (5) (6) In the formula, is the prediction probability of the global-local discriminator for judging the real normal light image , and the prediction probability is from 0 to 1, and the closer it is to 1, the better; similarly, is the prediction probability of the discriminator for judging the real low-light image , and the prediction probability is from 0 to 1, and the closer it is to 1, the better. is the global-local discriminator Judge the generation of normal light images for the prediction probability, where the prediction probability ranges from 0 to 1, and the closer it is to 0, the better; similarly, is the discriminator judging the prediction probability of the generated low-light image . The prediction probability ranges from 0 to 1, and the closer it is to 0, the better. 、 The larger the better.

[0044] Step 6, calculate the generator adversarial loss function of the generator 、 . The goal of the generator is to generate fake images that can deceive the discriminator. For 、 , its goal is to generate a low-light image that the discriminator cannot recognize ; for , its goal is to generate a low-light image that the discriminator cannot recognize The formula is as follows: (7) (8) where represents the prediction probability of the discriminator for the generated low-light image ( ). The goal of the generator is to maximize , that is, to make the discriminator judge the generated low-light image as a real image. The prediction probability ranges from 0 to 1, and the closer it is to 1, the better; represents the prediction probability of the discriminator for the generated normal illumination image ( ). The goal of the generator is to maximize , that is, to make the discriminator judge the generated normal light image as true. The prediction probability ranges from 0 to 1, and the closer it is to 1, the better. 、 The smaller the better; Step 7, calculate the cycle consistency loss between the original normal light image and the reconstructed normal illumination image ; calculate the cycle consistency loss between the original low-light image and the reconstructed low-light image . The cycle consistency loss​ , is as follows: (9) (10) Among them, is the normal light image after passing through the generator , the generated low-light image and then passing through the generator again to obtain the reconstructed normal light image ; is the low-light image after passing through the generator , the generated low-light image and then passing through the generator again to obtain the reconstructed low-light image . , the smaller the better; Step 8, calculate the self-feature retention loss between the normal image and the generated low-light image , and calculate the self-feature retention loss between the low-light image and the generated normal light image ; The formula is as follows: (11) (12) Among them, represents the low-light image generated by the generator ; represents the normal light image generated by the generator ; represents the feature map extracted from the pre-trained VGG16 model; represents the th max pooling layer; represents the th convolutional layer after the th max pooling layer; and are the dimensions of the extracted feature map.

[0045] Step 9, calculate the total loss , of the generators , , and update the parameters of the generators , for optimization. The generator , The total loss function combines the generator adversarial loss , , the cycle consistency loss , and the self-feature retention loss , , the total loss function of the generator is as follows: (13) Step 10: Continue the iteration and alternately train the discriminator and the generator. Each time it is updated, the discriminator is trained according to the image generated by the current generator, while the generator adjusts its generation strategy according to the feedback of the discriminator until the generated image is realistic enough. Through the improvement of the present invention, the low-illumination image can restore a more realistic and clear image through the image enhancement algorithm, effectively solving the problems of color shift and color distortion, making the image color more vivid and realistic; and effectively eliminating artifacts and noise, making the enhanced image have higher clarity; in addition, by improving the detail extraction ability, the illumination of the enhanced image is more uniform, paying more attention to the restoration of local details of the image, and thus obtaining a more accurate image enhancement effect.

[0046] Note: represents the normal light image; represents the low-illumination image; represents generator A (i.e., the generator that generates a low-illumination image from a normal light image); represents generator B (i.e., the generator that generates a normal light image from a low-illumination image); represents the generated low-illumination image (i.e., the image generated by generator A); represents the generated normal light image (i.e., the image generated by generator B); represents discriminator A (i.e., the discriminator that judges the normal light real image); represents discriminator B (i.e., the discriminator that judges the low-illumination real image); represents the reconstructed normal light image; represents the reconstructed low-light image; represents the original normal light image and the cycle consistency loss of the reconstructed normal light image ; represents the original low-illumination image and the cycle consistency loss of the reconstructed low-illumination image ; represents the discriminator judging the predicted probability of the real low-illumination image ; Indicates the predicted probability that discriminator B judges the generated low-illumination image ; Indicates the discriminator 's predicted probability for the real normal-illumination image ; Indicates the discriminator 's predicted probability for the generated normal-light image ; Indicates the loss function of the discriminator ; Indicates the loss function of the discriminator ; Indicates the self-feature retention loss between the normal image and the generated low-illumination image ; Indicates the self-feature retention loss between the low-illumination image and the generated normal-light image ; Is the total loss function of the generator.

Claims

1. A low-light image enhancement method based on a cyclic generative adversarial network, characterized in that: Please follow the steps below to implement: Step 1: Establish a low-light image enhancement model, low-light and normal light data sets; Step 2: Extract a normal light map from the normal light dataset , extract a low-light image from the low-light dataset ; Step 3: Using the generator With the generator Obtaining generated images and reconstructed images; Step 4: Send the generated image and the initial image to the global-local discriminator and In the calculation of the discriminator The adversarial loss With the discriminator The adversarial loss , update the discriminator parameters; Step 5: Calculate the generator , The generator adversarial loss function , ; And calculate the initial image Cycle consistency loss with reconstructed image , ; Step 6: Calculate the self-feature preservation loss between the initial image and the generated image , ; Step 7: Calculate the generator , Total loss , update the generator parameters; Step 8: Alternately train the discriminator and generator to generate real images and complete the enhancement.

2. The low-light image enhancement method based on a cyclic generative adversarial network according to claim 1, characterized in that: The step 3 is specifically as follows: Step 3.1: Normal light image Through the generator , extract the normal light image The image features are used to generate low-light images ; Low light image Through the generator , extract low-light images The image features are used to generate the normal light image , the specific formula is as follows: (1) (2) In the formula, Representation Generator Generated low light image , Representation Generator Generated normal light image ; Step 3.2, a low-light image will be generated Through the generator , extract and generate low-light images The image features are used to reconstruct the normal light image. ; A normal ray image will be generated Through the generator , extract the image features of the image and obtain the reconstructed low-light image , the specific formula is as follows: (3) (4) In the formula, Represents a normal light image Through the generator , the resulting low-light image After that, pass through the generator again The reconstructed normal light image obtained ; Represents low-light images Through the generator , the resulting low-light image After that, pass through the generator again The reconstructed low-light image .

3. The low-light image enhancement method based on a cyclic generative adversarial network according to claim 2, characterized in that: The image features described in step 3 are edges, textures, colors, corners, and spatial levels.

4. The low-light image enhancement method based on a cyclic generative adversarial network according to claim 2, characterized in that: The generator and The generator structure of the CycleGAN network is adopted, and it is combined with the U-Net network, the MCFE multi-scale convolutional feature enhancement module, the AESA dynamic effective self-attention aggregation module, and the SE-Res2Block-based lighting enhancement module. The MCFE multi-scale convolutional feature enhancement module is added in the encoding process of the generator to extract multi-scale image features; the AESA module is added in the decoding process to establish long-distance dependencies and extract image feature information; the SELE module is also added to the generator to enhance the image detail features while adjusting the image brightness; finally, the self-feature preservation loss is introduced to enhance the feature extraction capability and detail processing capability for low-light image enhancement tasks.

5. The low-light image enhancement method based on a cyclic generative adversarial network according to claim 4, characterized in that: The SELE module adopts 9 SE-Res2Block modules, and the void rates of the 9 SE-Res2Block modules are 1, 2, 4, 8, 8, 4, 2, 1, 1 respectively.

6. The low-light image enhancement method based on a cyclic generative adversarial network according to claim 2, characterized in that: The step 4 is specifically as follows: The resulting low-light image With low light images Feed into the global-local discriminator , discriminate the generated image The authenticity of Loss ; Generate a normal light image Normal light image Feed into the global-local discriminator , determine the authenticity of the generated image and calculate the discriminator Loss ; Adversarial loss of the discriminator and The specific expressions are as follows: (5) (6) In the formula, is a global-local discriminator Determine the true normal light image The predicted probability is 0-1, the closer to 1, the better; It is a discriminator Determining true low-light images The predicted probability is 0-1, the closer to 1, the better; is a global-local discriminator Determine the generation of normal light images The predicted probability is 0-1, and the closer to 0, the better; It is a discriminator Determine and generate low-light images The predicted probability is 0-1, and the closer to 0, the better.

7. The low-light image enhancement method based on a cyclic generative adversarial network according to claim 6, characterized in that: The generator adversarial loss function described in step 5 and The specific formula is as follows: (7) (8) in, Representative Discriminator Generate low-light images ( ), the predicted probability of the generator The goal is to maximize , that is, let the discriminator A low light image will be generated The image is judged to be real, and the prediction probability is 0-1, the closer to 1, the better; Representative Discriminator Generate a normal illumination image ( ), the predicted probability of the generator The goal is to maximize , that is, let the discriminator A normal ray image will be generated The judgment is true, and the predicted probability is 0-1, the closer to 1 the better.

8. The low-light image enhancement method based on a cyclic generative adversarial network according to claim 7, characterized in that: Initial image as described in step 5 The cycle consistency loss with the reconstructed image is divided into: normal lighting original image and reconstruct the normal illumination image The cycle consistency loss With low light original image Reconstructing low-light images The cycle consistency loss , and Specifically expressed as: (9) (10) in, It is a normal light image Through the generator , the resulting low-light image After that, pass through the generator again The reconstructed normal light image obtained ; It is a low light image Through the generator , the resulting low-light image After that, pass through the generator again The reconstructed low-light image .

9. The low-light image enhancement method based on a cyclic generative adversarial network according to claim 8, characterized in that: The step 6 is specifically as follows: calculating the normal image and generate low light images The self-feature preservation loss between , calculate low-light image and generate a normal light image The self-feature preservation loss between ; The specific formula is as follows: (11) (12) in, Representation Generator Generated low light image ; Representation Generator Generated normal light image ; Represents the feature map extracted from the pre-trained VGG16 model; Indicates Max pooling layer.

10. The low-light image enhancement method based on a cyclic generative adversarial network according to claim 9, characterized in that: The step 7 is specifically as follows: , The total loss function combines the generator adversarial loss , , cycle consistency loss , and self-feature preservation loss , , the total loss function of the generator The specific expressions are as follows: (13)。

Citation Information

Patent Citations

  • Single low-light image enhancement method based on generative adversarial network

    CN111161178A

  • Unsupervised learning method and system for low-illumination image enhancement

    CN113313657A

  • Dark light image enhancement method based on attention mechanism

    CN114399431A

  • Robot out-of-order target pushing and grabbing method with domain self-adaption

    CN114918918A

  • Low-illumination image enhancement method based on multi-scale and context learning network

    CN114998145A