Dark light enhancement method based on neural network
By adopting a dual generator and a dual discriminator structure in dark light image enhancement, combining perceived loss and other loss functions, and using technical means such as hollow convolution and attention mechanism, the problems such as noise amplification meeting and detailed information loss in dark light image enhancement in the existing technology are solved, and high-quality dark light image enhancement effect is achieved.
Patent Information
- Application Number
- CN202510161430.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-05-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art has problems such as noise amplification, loss of detail information, color distortion and pattern crashes in dark light image enhancement, and the problem of gradient disappearance and difficulty in convergence of the model during training is prone to occur.
The dual generator and dual discriminator structure based on neural network are adopted, combining perceptual loss, edge preservation loss and noise suppression loss, unsupervised training is achieved through adversarial training of forward and reverse generators, and the structure, texture and semantic consistency of the image is enhanced through technical means such as hollow convolution, spatial attention mechanism, pyramid pooling and channel attention mechanism.
It effectively improves the structural, texture and semantic consistency of dark light enhances the image, retains edge details information, reduces noise and artifacts, improves the clarity and nature of the image, and conducts model training stably.
Smart Images

Figure CN120070285A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image enhancement, and more specifically, it relates to a low-light enhancement method based on a neural network. Background Art
[0002] Low-light enhancement refers to improving the brightness, contrast, and details of images taken under low-light conditions to make them close to the visual effects under normal light conditions.
[0003] Traditional image low-light enhancement mainly includes histogram equalization and the Retinex algorithm. Histogram equalization refers to enhancing the overall contrast of an image by adjusting the gray-level distribution of the image. However, this method will also amplify the noise in the image while increasing the brightness, and global equalization processing may lead to overexposure or underexposure of the image, resulting in the loss of its detail information; the Retinex algorithm assumes that an image can be decomposed into a reflection component (reflecting the characteristics of the object itself) and an illumination component (reflecting the influence of the light source), and enhances the brightness and contrast of the image by adjusting the illumination component. Common Retinex algorithms include single-scale Retinex (SSR), multi-scale Retinex (MSR), and multi-scale Retinex with color restoration (MSRCR). SSR is processed using a single-scale Gaussian filter, but it is prone to losing its detail information and blurring the edges. MSR can retain more image detail information while enhancing the contrast, but there is still a problem of color distortion. Although MSRCR has improved the color distortion problem, it is still difficult to fully restore natural colors under complex lighting conditions.
[0004] With the development of artificial intelligence and deep learning, currently, the mapping relationship from low-light images to enhanced images is learned through a convolutional neural network (CNN) to restore the brightness, contrast, and details of the images. However, the receptive field of the CNN is limited and it is difficult to handle global illumination changes, resulting in uneven enhancement effects. Currently, a generative adversarial network (GAN) is added on the basis of the CNN, and through the adversarial training of the generator and the discriminator, the limitation of the receptive field of the CNN is overcome, so that the generator generates enhanced images close to the real ones. However, during the training process of the GAN, mode collapse is likely to occur, that is, the generator can only generate a small number of images with insufficient diversity. In addition, the generator may introduce unnatural artifacts or noise when generating images, and there may also be a problem of vanishing gradients during the training process, resulting in difficulty for the model to converge. Summary of the Invention
[0005] The present invention provides a low-light enhancement method based on a neural network to solve the technical problems in the above background art.
[0006] The present invention provides a low-light enhancement method based on a neural network, including the following steps: Step S101, collect low-light images and normal-light images, and preprocess all images to construct a training dataset; Step S102, construct a forward generator, a forward discriminator, a reverse generator, and a reverse discriminator, and train them using the training dataset; The forward generator takes a low-light image as input and outputs an enhanced image that approximates normal light; The forward discriminator takes an enhanced image that approximates normal light and a normal-light image as input, and outputs the probability value that the enhanced image that approximates normal light belongs to a normal-light image; The reverse generator takes an enhanced image that approximates normal light as input and outputs a weakened image that approximates low light; The reverse discriminator takes a weakened image that approximates low light and a low-light image as input, and outputs the probability value that the weakened image that approximates low light belongs to a low-light image; The forward generator and the reverse generator have the same network structure, and the forward discriminator and the reverse discriminator have the same network structure; Step S103, retain the forward generator after training is completed, input the low-light image into the forward generator, and output an enhanced image that approximates normal light.
[0007] Further, the preprocessing includes: uniformly adjusting the sizes of the low-light images and the normal-light images to 512×512; performing data augmentation on the low-light images, and the data augmentation includes: random rotation, brightness adjustment, and contrast adjustment, the range of random rotation is between -15° and 15°, the range of brightness adjustment is between 0.6 and 1.3, and the range of contrast adjustment is between 0.8 and 1.5.
[0008] Further, the network structure of the forward generator includes: a normalization layer, a convolutional layer, an atrous convolutional layer, a spatial attention mechanism layer, a pyramid pooling layer, a channel attention mechanism layer, and a dimensionality reduction convolutional layer; The normalization layer is used to scale the pixel values of the low-light image to between -1 and 1, and the calculation formula of the normalization layer is as follows: (pixel / 127.5)-1, where pixel represents the pixel value of the low-light image; The convolutional layer consists of 64 3×3 convolutional kernels, and the stride of the convolutional kernels is 1, the edge padding is 1, the padding method is SAME, the convolutional layer inputs the normalized low-light image, and outputs a first feature map, the number of channels of the first feature map is 64, and the height and width of the first feature map are the same as the height and width of the normalized low-light image; The dilated convolutional layer consists of dilated convolutional blocks with three different dilation rates, a concatenation layer, and a dimensionality reduction convolutional layer. The three different dilation rates are 1, 2, and 4 respectively. The three dilated convolutional blocks have the same composition as the convolutional layer, and the size of the output feature map is the same as that of the first feature map. The concatenation layer is used to concatenate the feature maps output by the three dilated convolutional blocks, and then the dimensionality reduction convolutional layer reduces the number of channels from 3×64 = 192 to 64 using a 1×1 convolutional kernel as the output of the second feature map. The second feature map has the same size as the first feature map; The spatial attention mechanism layer takes the second feature map as input and outputs the third feature map; The pyramid pooling layer takes the third feature map as input and outputs the fourth feature map; The channel attention mechanism layer takes the fourth feature map as input and outputs the fifth feature map; The dimensionality reduction convolutional layer takes the fifth feature map as input and outputs an enhanced image approaching normal illumination; The third feature map, the fourth feature map, and the fifth feature map have the same size as the first feature map; The enhanced image approaching normal illumination has the same size as the normalized low-light image.
[0009] Furthermore, the spatial attention mechanism layer consists of a block embedding layer, N self-attention mechanism layers, and a block merging layer. The self-attention mechanism layer consists of a fixed window attention block and a sliding window attention block; The block embedding layer divides the second feature map into multiple non-overlapping regions through blocks of size P×P, unfolds each non-overlapping region into a vector representation, and transforms it into a vector with a dimensionality of d through a linear projection, outputting a token sequence of length (H / P)×(W / P), where H and W represent the height and width of the second feature map respectively, and both P and d are user-defined parameters; The fixed window size of the fixed window attention block and the sliding window size of the sliding window attention block are both M×M, and the stride of the sliding window is M / 2, where M is a user-defined parameter; The calculation formula for the number N of self-attention mechanism layers is as follows: ; where S represents the stride of the sliding window, and Round represents rounding up; Both the fixed window attention block and the sliding window attention block consist of 8 self-attention mechanism blocks and a fusion block. The length of the token sequence processed by each self-attention mechanism block is M×M, and the vector dimensionality corresponding to the sequence unit becomes d / 8; The calculation formula for the i-th self-attention mechanism block includes: ; ; ; ; wherein represents the feature map output by the i-th self-attention mechanism block, represents the token sequence input to the i-th self-attention mechanism block, represents the length of, , , , , , and respectively represent the first intermediate matrix, the second intermediate matrix, the third intermediate matrix, the first weight matrix, the second weight matrix, the third weight matrix and the bias matrix of the i-th self-attention mechanism block, T represents the transpose operation, and Softmax represents the Softmax activation function; The calculation formula of the fusion block includes: ; wherein represents the feature map output by the fusion block, respectively represent the feature maps output by the 1st to 8th self-attention mechanism blocks, LN represents the LayerNorm layer normalization operation, represents the fourth weight matrix, and GELU represents the GELU activation function; The block merging layer is used to merge the feature maps output after processing all non-overlapping regions through the N-th self-attention mechanism layer, and outputs the third feature map through a linear transformation.
[0010] Furthermore, the pyramid pooling layer is composed of pooling layers with four different pooling windows, an upsampling layer and a stacking layer; The sizes of the four different pooling windows are 64×64, 32×32, 16×16 and 8×8 respectively, the strides are 64, 32, 16 and 8 respectively, and the sizes of the output feature maps are 8×8×64, 16×16×64, 32×32×64 and 64×64×64 respectively; The upsampling layer is used to upsample the feature maps output by the four different pooling windows to the same size as the first feature map; The stacking layer is used to perform channel accumulation on the feature maps output by the upsampling layer and output the fourth feature map.
[0011] Furthermore, the calculation formula of the channel attention mechanism layer is as follows: ; wherein represents the fourth feature map input to the channel attention mechanism layer, Denote the fifth feature map output by the channel attention mechanism layer, and GAP denote global average pooling. 、 and denote the first, second, and third intermediate feature maps, with sizes of 1×1×64, 1×1×32, and 1×1×64 respectively. denote the first fully connected layer. denote the second fully connected layer, ReLU denote the ReLU activation function, and sigmoid denote the sigmoid activation function.
[0012] Furthermore, during the training process, in addition to the cycle consistency loss, it also includes perceptual loss, edge preservation loss, and noise suppression loss. Perceptual loss The calculation formula is as follows: ; where A denotes the low-light image input to the forward generator. denotes the enhanced image output by the forward generator approaching normal illumination, B denotes the normal illumination image. denotes extracting the feature map through the pre-trained network, Q denotes the number of feature maps extracted through the pre-trained network. 、 and respectively denote the number of channels, height, and width of the i-th feature map extracted through the pre-trained network. denotes the L2 norm, and the pre-trained network only includes the normalization layer and convolutional layer of the forward generator. Edge preservation loss The calculation formula is as follows: ; where denotes the Sobel edge detection operator. denotes the L1 norm. Noise suppression loss The calculation formula is as follows: ; where D denotes the wavelet transform denoising operator.
[0013] Furthermore, the network structures of the forward discriminator and the reverse discriminator are both PatchGAN discriminators.
[0014] The beneficial effects of the present invention are as follows: The present invention realizes unsupervised training through a dual generator and a dual discriminator, and at the same time introduces perceptual loss, edge-preserving loss, and noise suppression loss. While ensuring stable training, it improves the structural, textural, and semantic consistency of low-light enhanced images, and retains edge detail information. Moreover, the forward generator integrates dilated convolution, spatial attention mechanism, pyramid pooling, and channel attention mechanism. Dilated convolution expands the receptive field without increasing the number of parameters to capture global illumination information. The spatial attention mechanism and pyramid pooling can capture the relationships between pixels at long distances in the feature map, thus adapting to complex illumination scenarios. The channel attention mechanism can reduce the noise and artifacts generated during the image enhancement process, making the image clearer and more natural. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 is a flowchart of a low-light enhancement method based on a neural network according to the present invention; Figure 2 is a schematic diagram of the network structure of the forward generator according to the present invention; Figure 3 is a schematic diagram of the cyclic adversarial training according to the present invention; Figure 4 is a schematic diagram of a low-light image according to the present invention; Figure 5 is a schematic diagram of low-light enhancement through a convolutional neural network according to the present invention; Figure 6 is a schematic diagram of low-light enhancement through the forward generator according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0016] The subject matter described herein will now be discussed with reference to example embodiments. It should be understood that discussing these embodiments is only to enable those skilled in the art to better understand and thus implement the subject matter described herein. Without departing from the scope of protection of the content of this specification, changes can be made to the functions and arrangements of the elements discussed. Each example can omit, substitute, or add various processes or components as needed. Additionally, the features described relative to some examples can also be combined in other examples.
[0017] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in one or more embodiments of the present invention should have the ordinary meanings understood by those of ordinary skill in the field to which the present invention pertains. The "first", "second" and similar terms used in one or more embodiments of the present invention do not denote any order, quantity or importance, but are only used to distinguish different components. Words such as "including" or "comprising" mean that the elements or objects appearing before this word cover the elements or objects listed after this word and their equivalents, without excluding other elements or objects. Words such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Upper", "lower", "left", "right", etc. are only used to represent relative position relationships, and when the absolute position of the object being described changes, the relative position relationship may also change accordingly.
[0018] As Figures 1 to 6 shown, a low-light enhancement method based on a neural network includes the following steps: Step S101, collect low-light images and normal-light images, and preprocess all images to construct a training dataset; Step S102, construct a forward generator, a forward discriminator, a reverse generator, and a reverse discriminator, and train them using the training dataset; The forward generator inputs a low-light image and outputs an enhanced image approaching normal light; The forward discriminator inputs the enhanced image approaching normal light and the normal-light image, and outputs the probability value that the enhanced image approaching normal light belongs to the normal-light image; The reverse generator inputs the enhanced image approaching normal light and outputs a weakened image approaching low light; The reverse discriminator inputs the weakened image approaching low light and the low-light image, and outputs the probability value that the weakened image approaching low light belongs to the low-light image; The network structures of the forward generator and the reverse generator are the same, and the network structures of the forward discriminator and the reverse discriminator are the same; Step S103, retain the forward generator after training is completed, input the low-light image into the forward generator, and output an enhanced image approaching normal light.
[0019] It should be noted that traditional GANs require a large number of paired low-light images and normal-light images as training datasets, while CycleGAN can be trained on unpaired low-light image and normal-light image training datasets, that is, unsupervised learning can be achieved. Moreover, the CycleGAN cycle consistency loss can ensure that after the low-light image is converted to a normal-light image, it can be restored to the original image, thereby keeping the geometric structure and detailed information of the image unchanged. Through structural constraints, artifacts and halos that appear during the generation process can be effectively suppressed, and the naturalness of the generated image can be improved. In addition, the structural design of the dual generators and dual discriminators, as well as the introduction of the cycle consistency loss, make it more stable during the training process and reduce the phenomenon of mode collapse.
[0020] In one embodiment of the present invention, the preprocessing includes: uniformly adjusting the sizes of the low-light image and the normal-light image to 512×512; performing data augmentation on the low-light image, and the data augmentation includes: random rotation, brightness adjustment, and contrast adjustment. The range of random rotation is between -15° and 15°, the range of brightness adjustment is between 0.6 and 1.3, and the range of contrast adjustment is between 0.8 and 1.5.
[0021] In one embodiment of the present invention, the network structure of the forward generator includes: a normalization layer, a convolutional layer, an atrous convolutional layer, a spatial attention mechanism layer, a pyramid pooling layer, a channel attention mechanism layer, and a dimensionality reduction convolutional layer; The normalization layer is used to scale the pixel values of the low-light image to between -1 and 1, and the calculation formula of the normalization layer is as follows: (pixel / 127.5)-1, where pixel represents the pixel value of the low-light image; The convolutional layer consists of 64 3×3 convolutional kernels, and the stride of the convolutional kernel is 1, the edge padding is 1, and the padding method is SAME. The convolutional layer inputs the normalized low-light image and outputs a first feature map. The number of channels of the first feature map is 64, and the height and width of the first feature map are the same as those of the normalized low-light image; The atrous convolutional layer consists of atrous convolutional blocks with three different dilation rates, a splicing layer, and a dimensionality reduction convolutional layer. Among them, the three different dilation rates are 1, 2, and 4 respectively. The compositions of the three atrous convolutional blocks are the same as that of the convolutional layer, and the size of the output feature map is the same as that of the first feature map. The splicing layer is used to splice the feature maps output by the three atrous convolutional blocks, and then the dimensionality reduction convolutional layer reduces the number of channels from 3×64 = 192 to 64 through a 1×1 convolutional kernel and outputs it as a second feature map. The second feature map has the same size as the first feature map; The spatial attention mechanism layer inputs the second feature map and outputs a third feature map; The pyramid pooling layer inputs the third feature map and outputs a fourth feature map; The channel attention mechanism layer takes the fourth feature map as input and outputs the fifth feature map; The dimensionality reduction convolutional layer takes the fifth feature map as input and outputs an enhanced image approximating normal illumination; The third, fourth, and fifth feature maps have the same size as the first feature map; The enhanced image approximating normal illumination has the same size as the normalized low-light image.
[0022] It should be noted that the convolutional layer is used to extract the basic features of the low-light image, such as edge, texture, and color features; the dilated convolutional layer expands the receptive field without increasing the number of parameters, can capture the global illumination information of the feature map, thus improving the problem of uneven illumination distribution, and dilated convolutions with different dilation rates can extract multi-scale features from local details to global structures, which helps to enhance details and overall brightness simultaneously; the spatial attention mechanism layer can capture the relationships between distant pixels in the feature map and can adapt to complex illumination scenarios; the pyramid pooling layer extracts the global features of the feature map through pooling windows of different sizes, helps to understand the overall illumination distribution, avoids the problem of uneven local enhancement, and also helps to enhance details and overall brightness simultaneously; the channel attention mechanism layer is used to adaptively adjust the importance of different channel features, highlight key features, suppress useless information, reduce noise and artifacts generated during the enhancement process, and make the image clearer and more natural.
[0023] In an embodiment of the present invention, the spatial attention mechanism layer is composed of a block embedding layer, N self-attention mechanism layers, and a block merging layer, where the self-attention mechanism layer is composed of a fixed window attention block and a sliding window attention block; The block embedding layer divides the second feature map into multiple non-overlapping regions through blocks of size P×P, unfolds each non-overlapping region into a vector representation, and transforms it into a vector with a dimensionality of d through linear projection, outputting a token sequence of length (H / P)×(W / P), where H and W respectively represent the height and width of the second feature map, and both P and d are user-defined parameters. Preferably, P is set to 8 and d is set to 128; The fixed window size of the fixed window attention block and the sliding window size of the sliding window attention block are both M×M, and the stride of the sliding window is M / 2, where M is a user-defined parameter. Preferably, M is set to 8; The calculation formula for the number N of self-attention mechanism layers is as follows: ; where S represents the stride of the sliding window, and Round represents rounding up; Both the fixed-window attention block and the sliding-window attention block are composed of 8 self-attention mechanism blocks and a fusion block. The length of the token sequence processed by each self-attention mechanism block is M×M, and the number of vector dimensions corresponding to the sequence units becomes d / 8; The calculation formula of the i-th self-attention mechanism block includes: ; ; ; ; where represents the feature map output by the i-th self-attention mechanism block, represents the token sequence input to the i-th self-attention mechanism block, represents the length of, , , , , , and respectively represent the first intermediate matrix, the second intermediate matrix, the third intermediate matrix, the first weight matrix, the second weight matrix, the third weight matrix and the bias matrix of the i-th self-attention mechanism block. T represents the transpose operation, and Softmax represents the Softmax activation function; The calculation formula of the fusion block includes: ; where represents the feature map output by the fusion block, respectively represent the feature maps output by the 1st to 8th self-attention mechanism blocks. LN represents the LayerNorm layer normalization operation, represents the fourth weight matrix, and GELU represents the GELU activation function; The block merging layer is used to merge the feature maps output after processing all non-overlapping regions by the N-th self-attention mechanism layer and output the third feature map through a linear transformation.
[0024] It should be noted that P is set to 8, d is set to 128, both H and W are 512. The number of dimensions of the vector representation of each non-overlapping region expanded is equal to 8×8×64 (the number of channels of the second feature map) = 4096. Then the length of the token sequence is (512 / 8)×(512 / 8) = 4096, and the number of vector dimensions corresponding to each sequence unit of the token sequence is transformed to 128 through a linear projection.
[0025] It should be noted that the first weight matrix, the second weight matrix, the third weight matrix, the fourth weight matrix, and the bias matrix are all learnable hyperparameters. For example, if M is set to 8 and d is set to 128, then the size of the token sequence processed by each self-attention mechanism block is 64×16. The sizes of the first weight matrix, the second weight matrix, and the third weight matrix can all be designed as 16×64. Then the sizes of the first intermediate matrix, the second intermediate matrix, and the third intermediate matrix are 64×64. Then the size of the feature map output by each self-attention mechanism block is 64×64. The size of the feature map obtained by splicing the feature maps output by 8 self-attention mechanism blocks is 64×512. Then the size of the fourth weight matrix can be designed as 512×64. Then the size of the feature map output by the fusion block is 64×64.
[0026] In one embodiment of the present invention, the pyramid pooling layer is composed of pooling layers with four different pooling windows, an upsampling layer, and a stacking layer; The sizes of the four different pooling windows are 64×64, 32×32, 16×16, and 8×8 respectively, the strides are 64, 32, 16, and 8 respectively, and the sizes of the output feature maps are 8×8×64, 16×16×64, 32×32×64, and 64×64×64 respectively; The upsampling layer is used to upsample the feature maps output by the four different pooling windows to the same size as the first feature map; The stacking layer is used to perform channel accumulation on the feature maps output by the upsampling layer and output the fourth feature map.
[0027] In one embodiment of the present invention, the calculation formula of the channel attention mechanism layer is as follows: ; where represents the fourth feature map input to the channel attention mechanism layer, represents the fifth feature map output by the channel attention mechanism layer, GAP represents global average pooling, , and represent the first, second, and third intermediate feature maps, with sizes of 1×1×64, 1×1×32, and 1×1×64 respectively, represents the first fully connected layer, represents the second fully connected layer, ReLU represents the ReLU activation function, and sigmoid represents the sigmoid activation function.
[0028] In one embodiment of the present invention, during the training process, in addition to the cycle consistency loss, there are also perceptual loss, edge-preserving loss, and noise suppression loss; Perceptual loss The calculation formula of ; Where A represents the low-light image input to the forward generator, represents the enhanced image approximated to normal illumination output by the forward generator, B represents the normal illumination image, represents the extraction of the feature map through the pre-trained network, Q represents the number of feature maps extracted through the pre-trained network, , and respectively represent the number of channels, height, and width of the i-th feature map extracted through the pre-trained network, represents the L2 norm, and the pre-trained network only includes the normalization layer and convolutional layer of the forward generator; Edge-preserving loss The calculation formula is as follows: ; Where represents the Sobel edge detection operator, represents the L1 norm; Noise suppression loss The calculation formula is as follows: ; Where D represents the wavelet transform denoising operator.
[0029] It should be noted that the perceptual loss can improve the structural, texture, and semantic consistency of the image by comparing the differences between the generated image and the real image in the feature space rather than the pixel space, and pre-training can improve the model convergence speed; the edge-preserving loss can ensure the edge details and texture information of the image while maintaining the brightness and contrast of the enhanced image by imposing constraints on the edge features of the image; the noise suppression loss can ensure that the noise in the low-light image is not amplified during the image enhancement process by penalizing the noise components in the generated image.
[0030] Further, weights are assigned to the cycle consistency loss, perceptual loss, edge-preserving loss, and noise suppression loss respectively, so as to achieve multi-scale loss and ensure the effect of image enhancement.
[0031] In an embodiment of the present invention, the network structures of the forward discriminator and the reverse discriminator are both PatchGAN discriminators, which will not be elaborated here.
[0032] The above has described the embodiments of this embodiment, but this embodiment is not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative rather than restrictive. Under the inspiration of this embodiment, those of ordinary skill in the art can also make many forms, all of which fall within the protection scope of this embodiment.
Claims
1. A dark light enhancement method based on neural network, characterized in that: The following steps are involved: Step S101, collecting low-light images and normal-light images, and preprocessing all images to construct a training data set; Step S102, constructing a forward generator, a forward discriminator, a reverse generator and a reverse discriminator, and training them using a training data set; The forward generator inputs a low-light image and outputs an enhanced image close to normal light. The forward discriminator inputs the enhanced image close to normal illumination and the normal illumination image, and outputs the probability value of the enhanced image close to normal illumination belonging to the normal illumination image; The inverse generator inputs an enhanced image close to normal lighting and outputs a weakened image close to low lighting; The inverse discriminator inputs the weakened image approximating low light and the low light image, and outputs a probability value that the weakened image approximating low light belongs to the low light image; The network structure of the forward generator and the reverse generator is the same, and the network structure of the forward discriminator and the reverse discriminator is the same; Step S103, retaining the trained forward generator, inputting the low-light image into the forward generator, and outputting an enhanced image close to normal light.
2. The dark light enhancement method based on neural network according to claim 1, characterized in that: Preprocessing includes: Resize low-light images and normal-light images to a uniform size of 512×512; Data enhancement is performed on low-light images. The data enhancement includes random rotation, brightness adjustment, and contrast adjustment. The range of random rotation is between -15° and 15°, the range of brightness adjustment is between 0.6 and 1.3, and the range of contrast adjustment is between 0.8 and 1.
5.
3. The dark light enhancement method based on neural network according to claim 1, characterized in that: The network structure of the forward generator includes: normalization layer, convolution layer, hole convolution layer, spatial attention mechanism layer, pyramid pooling layer, channel attention mechanism layer and dimension reduction convolution layer; The normalization layer is used to reduce the pixel value of the low-light image to between -1 and 1. The calculation formula of the normalization layer is as follows: (pixel / 127.5)-1, where pixel represents the pixel value of the low-light image; The convolution layer consists of 64 3×3 convolution kernels, and the stride of the convolution kernel is 1, the edge padding is 1, and the padding mode is SAME. The convolution layer inputs the normalized low-light image and outputs the first feature map. The number of channels of the first feature map is 64, and the height and width of the first feature map are the same as the height and width of the normalized low-light image. The atrous convolution layer consists of three atrous convolution blocks with different expansion rates, a splicing layer, and a dimensionality reduction convolution layer. The three different expansion rates are 1, 2, and 4, respectively. The three atrous convolution blocks have the same composition as the convolution layer, and the size of the output feature map is the same as the size of the first feature map. The splicing layer is used to splice the feature maps output by the three atrous convolution blocks, and then the dimensionality reduction convolution layer uses a 1×1 convolution kernel to reduce the number of channels from 3×64=192 to 64 as the second feature map output. The size of the second feature map is the same as the first feature map. The spatial attention mechanism layer inputs the second feature map and outputs the third feature map; The pyramid pooling layer inputs the third feature map and outputs the fourth feature map; The channel attention mechanism layer inputs the fourth feature map and outputs the fifth feature map; The dimension reduction convolution layer inputs the fifth feature map and outputs an enhanced image close to normal illumination; The third characteristic map, the fourth characteristic map and the fifth characteristic map have the same size as the first characteristic map; The enhanced image close to normal lighting has the same size as the normalized low-light image.
4. The dark light enhancement method based on neural network according to claim 3, characterized in that: The spatial attention mechanism layer consists of a block embedding layer, N self-attention mechanism layers, and a block merging layer, where the self-attention mechanism layer consists of a fixed window attention block and a sliding window attention block; The block embedding layer divides the second feature map into multiple non-overlapping regions through P×P blocks, expands each non-overlapping region into a vector representation, and transforms it into a vector with a dimension of d through linear projection, and outputs a token sequence of length (H / P)×(W / P), where H and W represent the height and width of the second feature map, respectively, and P and d are custom parameters; The fixed window size of the fixed window attention block and the sliding window size of the sliding window attention block are both M×M, and the step size of the sliding window is M / 2, where M is a custom parameter; The number of self-attention mechanism layers N is calculated as follows: ; Where S represents the step size of the sliding window, and Round means rounding up; Both the fixed window attention block and the sliding window attention block are composed of 8 self-attention mechanism blocks and fusion blocks. The length of the token sequence processed by each self-attention mechanism block is M×M, and the number of vector dimensions corresponding to the sequence unit becomes d / 8; The calculation formula for the i-th self-attention mechanism block includes: ; ; ; ; in represents the feature map output by the i-th self-attention mechanism block, represents the token sequence input to the i-th self-attention mechanism block, express Length, , , , , , and Respectively represent the first intermediate matrix, the second intermediate matrix, the third intermediate matrix, the first weight matrix, the second weight matrix, the third weight matrix and the bias matrix of the i-th self-attention mechanism block, T represents the transpose operation, and Softmax represents the Softmax activation function; The calculation formula of the fusion block includes: ; in represents the feature map output by the fusion block, They represent the feature maps output by the 1st to 8th self-attention mechanism blocks respectively, and LN represents the LayerNorm layer normalization operation. represents the fourth weight matrix, GELU represents the GELU activation function; The block merging layer is used to merge the feature maps output by all non-overlapping areas after being processed by the Nth self-attention mechanism layer, and output the third feature map through linear transformation.
5. The dark light enhancement method based on neural network according to claim 3, characterized in that: The pyramid pooling layer consists of four pooling layers with different pooling windows, upsampling layers, and stacking layers; The sizes of the four different pooling windows are 64×64, 32×32, 16×16, and 8×8, with step sizes of 64, 32, 16, and 8, respectively. The output feature map sizes are 8×8×64, 16×16×64, 32×32×64, and 64×64×64, respectively. The upsampling layer is used to upsample the feature maps output by the four different pooling windows to the same size as the first feature map; The stacking layer is used to perform channel accumulation on the feature map output by the upsampling layer to output the fourth feature map.
6. The dark light enhancement method based on neural network according to claim 3, characterized in that: The calculation formula of the channel attention mechanism layer is as follows: ; in The fourth feature map representing the input of the channel attention mechanism layer, represents the fifth feature map output by the channel attention mechanism layer, GAP represents global average pooling, , and represents the first, second and third intermediate feature maps, with sizes of 1×1×64, 1×1×32 and 1×1×64 respectively. represents the first fully connected layer, represents the second fully connected layer, ReLU represents the ReLU activation function, and sigmoid represents the sigmoid activation function.
7. The dark light enhancement method based on neural network according to claim 3, characterized in that: In the training process, in addition to the cycle consistency loss, it also includes perceptual loss, edge preservation loss and noise suppression loss; Perceived loss The calculation formula is as follows: ; Where A represents the low-light image input to the forward generator, represents the enhanced image close to normal lighting output by the forward generator, B represents the normal lighting image, represents the feature map extracted by the pre-trained network, Q represents the number of feature maps extracted by the pre-trained network, , and They represent the number of channels, height, and width of the i-th feature map extracted by the pre-trained network, respectively. represents the L2 norm, and the pre-trained network only includes the normalization layer and convolution layer of the forward generator; Edge retention loss The calculation formula is as follows: ; in represents the Sobel edge detection operator, represents the L1 norm; Noise suppression loss The calculation formula is as follows: ; Where D represents the wavelet transform denoising operator.
8. The dark light enhancement method based on neural network according to claim 1, characterized in that: The network structures of the forward discriminator and the reverse discriminator are both PatchGAN discriminators.
Citation Information
Cited By
Forestry fire monitoring method based on super-resolution image processing
CN120997771A