Reference-Free Low-Light Image Enhancement Method Based on Local Scene Perception

By designing a low-illumination image enhancement network with local scene perception, the problem of insufficient attention to local scenes in the prior art is solved, dynamic lighting adjustment and high-quality enhancement of the image are achieved, and it is suitable for computer vision tasks.

CN115205160BActive Publication Date: 2025-07-11FUZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210960432.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-11
Publication Date
2025-07-11
Estimated Expiration
2042-08-11

AI Technical Summary

Technical Problem

The existing low-illumination image enhancement methods lack attention to local scenes, resulting in poor image enhancement effects and it is difficult to dynamically balance the lighting conditions in local areas.

Method used

A low-illumination image enhancement network based on local scene perception is designed, including local scene-aware branch network, enhanced branch network, attention module and iterative enhancement module. Through the training network without reference loss function, the local lighting characteristics of the image are dynamically adjusted to achieve high-quality image enhancement.

Benefits of technology

Through local scene perception and dynamic enhancement, the visibility and quality of low-illumination images can be effectively improved, and high-quality normal illumination images can be output, which is suitable for computer vision tasks such as autonomous driving and natural image analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115205160B_ABST
    Figure CN115205160B_ABST
Patent Text Reader

Abstract

The present invention relates to a no-reference low-light image enhancement method based on local scene perception, comprising the following steps: Step S1: Obtain low-light images and preprocess each image to obtain a training data set; Step S2: Construct a low-light image enhancement network based on local scene perception; Step S3: Design a no-reference loss function for training the network designed in Step S2; Step S4: Based on the no-reference loss function, use the training data set to train the low-light image enhancement network based on local scene perception; Step S5: Pass the image to be measured through the trained low-light image enhancement network based on local scene perception to obtain a normal illumination image. The present invention can effectively enhance low-light images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of image processing and computer vision, and particularly to a reference-free low-light image enhancement method based on local scene perception. Background Art

[0002] In recent years, the progress and development of science and technology have significantly promoted the improvement of society and people's lives. With the miniaturization of the volume of image acquisition devices and the high efficiency of acquisition capabilities, images and image systems are closely related to people's daily lives and production development. Image-based processing systems have a wide range of applications in life scenarios. Image systems can facilitate users to record and observe in real time intuitively, with high convenience. However, limited by the environment and usage status of the acquisition device, the obtained images lack ideal observability, which is reflected in phenomena such as motion blur and poor illumination of the images. Especially when the lighting conditions are poor, the situation of poor illumination is very common. At this time, the captured images often have multiple dark areas, which bring difficulties to human eye reading or machine vision processing. Therefore, designing an enhancement method for low-light images has important theoretical and application significance.

[0003] Low-light image enhancement aims to improve the visibility of the content of images with insufficient lighting conditions. In low-light images, the visibility and distinguishability of objects, scenes, and textures are poor, and it is difficult to be recognized by the human eye or used for advanced computer vision tasks. For enhancing low-light images, there are usually two typical scenarios: one is to enhance the entire image, which is usually seen when the overall lighting conditions are very poor or the exposure parameters (shutter time or white balance parameters) of the camera are incorrectly set; the other is to enhance some dark areas in the image, and this situation is more common in naturally captured images.

[0004] Low-light image enhancement has important practical significance. It can directly improve the visibility and perception of the human eye for images with poor lighting, and can also be reflected in specific computer vision tasks, such as for night scenes of autonomous driving, for the analysis of natural images or video content (which can include object detection, object tracking, human pose detection, etc.). Using the low-light image enhancement method to directly preprocess images with poor illumination enables the algorithms designed for specific tasks to directly process image data in low-light environments, which can greatly reduce the data annotation cost and model reconstruction cost of specific tasks, and thus save a large amount of human and material resources. Summary of the Invention

[0005] In view of this, the purpose of the present invention is to provide a reference-free low-light image enhancement method based on local scene perception, which can effectively enhance low-light images.

[0006] To achieve the above purpose, the present invention adopts the following technical solutions:

[0007] A no-reference low-light image enhancement method based on local scene perception, comprising the following steps:

[0008] Step S1: Obtain low-light images and preprocess each image to obtain a training dataset;

[0009] Step S2: Construct a low-light image enhancement network based on local scene perception;

[0010] Step S3: Design a no-reference loss function for training the network designed in Step S2;

[0011] Step S4: Based on the no-reference loss function, use the training dataset to train the low-light image enhancement network based on local scene perception;

[0012] Step S5: Pass the image to be tested through the trained low-light image enhancement network based on local scene perception to obtain a normal illumination image.

[0013] Further, the preprocessing is specifically: scale each image in the dataset to an image of the same size with dimensions H×W;

[0014] Normalize the image. Given an image I train , calculate the normalized image using the following formula:

[0015]

[0016] where I train is an image with 8-bit color depth and size H×W, and I bit_max is an image with size H×W and all pixel values being 255.

[0017] Further, the low-light image enhancement network based on local scene perception includes a local scene perception branch network, an enhancement branch network, an attention module, an iterative enhancement module, and a denoising module.

[0018] Further, the local scene perception branch network receives the normalized image X with dimensions H×W as input and performs a spatial 4-equal division cropping. The cropped image is denoted as The local scene perception branch network consists of 4 perception branch networks J1, J2, J3, J4 and feature transformation blocks T1, T2. The perception branch networks are used to extract the features of local images, and the feature transformation blocks are used to generate convolution parameters for image enhancement. The operation of obtaining the spatially 4-equal division cropped image is as follows:

[0019]

[0020] where (p, q) is the pixel position of the input normalized image X with size H×W.

[0021] Furthermore, the perception branch networks J1, J2, J3, and J4 all have the same network structure, and their trainable parameters are shared; the perception branch network consists of 4 convolutional blocks, and each convolutional block consists of a convolutional layer and an activation layer in sequence;

[0022] The feature transformation block T1 is composed of a convolutional layer with a kernel size of 1×1, a stride of 0, and a padding of 0, and receives the aggregated output features of the 3rd convolutional block of the perception branch network, and outputs the convolutional parameter k1 for image enhancement; the feature transformation block T2 is composed of a convolutional layer with a kernel size of 1×1, a stride of 0, and a padding of 0, and its number of output channels is 25% of the number of input channels, and receives the aggregated output features of the 4th convolutional block of the perception branch network, and outputs the convolutional parameter k2 for image enhancement, which is expressed by the formula as follows:

[0023]

[0024]

[0025] where Concat(·) represents concatenating features in the convolutional channel dimension, and T1 and T2 are feature transformation blocks.

[0026] Furthermore, the enhancement branch network includes 2 convolutional blocks, 2 variable parameter convolutional blocks D1, D2, an upsampling layer, and an upsampling layer; the convolutional block consists of a convolutional layer and an activation layer in sequence, the convolutional layer uses a convolution with a kernel size of 3×3, a stride of 1, and a padding of 0, and the activation layer uses the ReLU activation function;

[0027] The variable parameter convolutional blocks D1 and D2 consist of a variable parameter convolutional layer and an activation layer in sequence. The convolutional layer in the variable parameter convolutional block D1 uses a convolution kernel from the output k1 of the local scene perception branch network, a convolutional stride of 1, and a padding of 0, and the activation layer uses the ReLU activation function; the convolutional layer in the variable parameter convolutional block D2 uses a convolution kernel from the output k2 of the local scene perception branch network, a convolutional stride of 1, and a padding of 0, and the activation layer uses the ReLU activation function;

[0028] The input of the upsampling layer is the output features from the variable parameter convolutional block D2. The upsampling operation performs bilinear interpolation for each feature channel, where the output size of the upsampling is set to H×W, and the generated illumination feature for image enhancement is denoted as f.

[0029] Furthermore, the attention module includes an attention layer and a convolutional block;

[0030] The input of the attention layer is the image enhancement illumination feature f output by the enhancement branch network, and at the same time accepts the normalized image X of size H×W as input; the image X is converted into a grayscale image denoted as X gary , the grayscale image X gary passes through a Gaussian filter to obtain a smoothed grayscale image X smooth , where the Gaussian kernel size of the filter is set to 3, and the attention layer calculates the output in the following way:

[0031]

[0032] where is the element-wise multiplication of matrices with a broadcasting mechanism, and the broadcasting mechanism makes each channel dimension in f multiply element-wise with X smooth for element-wise matrix multiplication;

[0033] The convolution block consists of a convolution layer and an activation layer in sequence, and accepts f A′ as input, where the convolution layer is a convolution with a kernel size of 3×3, a stride of 1, and a padding of 0, and the activation layer uses the tanh activation function. The output f A of the convolution block is the illumination coefficient of the image enhancement output by the attention module.

[0034] Furthermore, the input of the enhancement module is the normalized image X of size H×W and the illumination coefficient f of the image enhancement generated by the attention module A , and the formula for generating the preliminarily enhanced image is as follows:

[0035]

[0036] The above formula is in an iterative form, n is the number of iterations, and E n represents the result obtained after iterating n times with f A as the illumination coefficient of image enhancement, and the default setting is n = 4. For convenience of representation, denote the preliminarily enhanced image E4 as E;

[0037] The input of the denoising module is the preliminarily enhanced image E and the normalized image X of size H×W, and the output of this module is the final image enhancement result

[0038] Furthermore, the optimization goal is to minimize the total loss function

[0039]

[0040] where represents the brightness control loss function, and λ bri represents the weight of the brightness control loss function; denotes the color enhancement loss function, λ col denotes the weight of the color enhancement loss function; denotes the global consistency constraint loss function, λ glo denotes the weight of the global consistency constraint loss function; denotes the illumination smoothing loss function, λ tv denotes the weight of the illumination smoothing loss function, denotes the denoising loss function, λ noi denotes the weight of the denoising loss function;

[0041] Brightness control loss function The calculation formula is as follows:

[0042]

[0043] Among them, the preliminarily enhanced image E is divided into N = 16×16, that is, 256 regions of the same size, and the average brightness value within the region is denoted as E b = 0.6 is a preset target brightness constant value, || ||1 is the absolute value operation;

[0044] Color enhancement loss function The calculation formula is as follows:

[0045]

[0046] In the above formula, c1 and c2 represent the color channels of the image, Ω ={(R, G), (R, B), (G, B)} represents different combinations of color channels, where R, G, and B respectively represent the red, green, and blue channels in the RGB color space; and represent the mean values of the c1 color channel and the c2 color channel in the preliminarily enhanced image E; similarly, represents the mean values of the c1 color channel and the c2 color channel in the input image X; || || 2 is the operation to obtain the Euclidean norm; the definition of G(·) is as follows:

[0047]

[0048] Among them represents the mean values of the c1 color channel and the c2 color channel in any input image I, || ||1 is the absolute value operation, Gamma γ (·) represents the gamma correction operation, and its calculation method is as follows:

[0049]

[0050] Among them represents the c1∈{R, G, B} color channel in any input image I, and γ is a preset parameter;

[0051] Global consistency constraint loss function The calculation formula is as follows:

[0052]

[0053] where, represents the global maximum consistency constraint loss, represents the global minimum consistency constraint loss;

[0054] represents the global maximum consistency constraint loss, and the calculation formula is as follows:

[0055]

[0056] where, for the input image X, it is divided into 16×16 regions, and the average pixel value is calculated for each region to obtain an image with a width and height of 16×16 composed of average pixel values. Then, max pooling (MaxPooling) is performed, and the size of the pooling operation is set to 8×8. The resulting image is of size 2×2, and each value in it is denoted as r1,r2)∈{(1,2),(1,3),(1,4),(2,3),(2,4),(3,4)} are different combinations of the extracted maximum values; || ||1 is the absolute value operation;

[0057] represents the global minimum consistency constraint loss, which makes the darkest regions in the 16×16 partition space of the image change as similarly as possible. The calculation formula is as follows:

[0058]

[0059] where, for the preliminarily enhanced image E, it is divided into 16×16 regions, and the average pixel value is calculated for each region to obtain an image with a width and height of 16×16 composed of average pixel values. Then, min pooling (MinPooling) is performed, and the size of the pooling operation is set to 8×8. The resulting image is of size 2×2, and each pixel value in it is denoted as Similarly, for the input image X, it is divided into 16×16 regions, and the average pixel value is calculated for each region to obtain an image with a width and height of 16×16 composed of average pixel values. Then, min pooling is performed, and the size of the pooling operation is set to 8×8. The resulting image is of size 2×2, and each value in it is denoted as (r1,r2)∈{(1,2),(1,3),(1,4),(2,3),(2,4),(3,4)} are different combinations of the extracted minimum values;

[0060] Lighting smoothing loss function The calculation formula is as follows:

[0061]

[0062] in, The illumination coefficient f is the image enhancement output by the attention module A The corresponding color channel c is taken from the RGB color space. || || 2 is the operation that finds the Euclidean norm. represents the operation of obtaining the first-order gradient, that is, and They are The first-order difference in the vertical and horizontal directions, || ||1 is the absolute value operation;

[0063] Denoising loss function The calculation formula is as follows:

[0064]

[0065] Here, Φ(·) represents the operation of extracting Conv4-1 layer features using the VGG-16 classification model pre-trained on ImageNet; therefore, Φ(E) represents the classification features extracted from the initially enhanced image E. Represents the final image enhancement result The extracted classification features; represents the operation of obtaining the first-order gradient, as well as in and They are The first-order differences in the vertical and horizontal directions, and are the first-order differences of E in the vertical and horizontal directions, respectively; where * represents the multiplication of the corresponding elements of the matrix, e is the base of the natural logarithm, and μ is a parameter that can control the perception of edge strength.

[0066] Furthermore, the step S4 is specifically as follows:

[0067] Step S41: Select a random training image X from the training data set;

[0068] Step S42: Input the image X, and obtain the preliminarily enhanced image E and the illumination coefficient f for image enhancement output by the attention module through the local scene perception branch network, enhancement branch network, attention module, and iterative enhancement module. A And the final image enhancement result Calculate the total loss function loss

[0069] Step S43: Use the backpropagation method to calculate the gradients of the parameters in the local scene perception branch network, enhancement branch network, attention module, and iterative enhancement module, and update the parameters using the Adam optimization method;

[0070] Step S44: The above steps are one iteration of the training process. The entire training process requires a preset number of iterations, and in each iteration, multiple image pairs are randomly sampled as a batch for training.

[0071] The present invention has the following beneficial effects compared with the prior art:

[0072] 1. Aiming at the problem of lack of attention to local scenes in the existing low-light image enhancement methods, the present invention aims to perceive local scenes and dynamically enhance images;

[0073] 2. The present invention designs a network to perceive the local information of the image scene and predict the parameters for enhancement, so that the network can dynamically balance the relationship between local scenes. The attention module in the network is used to balance the extreme brightness regions of the image, the iterative enhancement module outputs the preliminary enhancement result, and the denoising module further improves the image quality, and finally can output a high-quality normal illumination image. Description of the Drawings

[0074] Figure 1 It is a flowchart of the method of the embodiment of the present invention.

[0075] Figure 2 It is the low-light image enhancement process of local scene perception of the embodiment of the present invention.

[0076] Figure 3 It is the local scene perception branch network of the embodiment of the present invention.

[0077] Figure 4 It is the enhancement branch network of the embodiment of the present invention.

[0078] Figure 5 It is the attention module of the embodiment of the present invention. Detailed Embodiment

[0079] The following further describes the present invention with reference to the drawings and embodiments.

[0080] Please refer toFigure 1 , the present invention provides a no-reference low-light image enhancement method based on local scene perception, as Figure 1 shown, including the following steps:

[0081] Step S1: Obtain low-light images and preprocess each image to obtain a training dataset;

[0082] Step S2: Construct a low-light image enhancement network based on local scene perception;

[0083] Step S3: Design a no-reference loss function for training the network designed in Step S2;

[0084] Step S4: Based on the no-reference loss function, use the training dataset to train the low-light image enhancement network based on local scene perception;

[0085] Step S5: Pass the image to be tested through the trained low-light image enhancement network based on local scene perception to obtain a normal illumination image.

[0086] In this embodiment, the preprocessing in Step S1 is specifically: scale each image in the dataset to an image of the same size with dimensions H×W;

[0087] Normalize the image. Given an image I train , calculate the normalized image The formula is as follows:

[0088]

[0089] where, I train is an image with 8-bit color depth and size H×W, and I bit_max is an image with size H×W and all pixel values being 255.

[0090] In this embodiment, the low-light image enhancement network based on local scene perception includes a local scene perception branch network, an enhancement branch network, an attention module, an iterative enhancement module, and a denoising module.

[0091] Preferably, the local scene perception branch network receives the normalized image X with size H×W as input and performs a 4-equal division spatial crop. The cropped image is denoted as The local scene perception branch network consists of 4 perception branch networks J1, J2, J3, J4 and feature transformation blocks T1, T2. The perception branch networks are used to extract the features of local images, and the feature transformation blocks are used to generate convolution parameters for image enhancement. The operation of obtaining the 4-equal division spatial crop image is as follows:

[0092]

[0093] where (p, q) is the pixel position of the input normalized image X of size H×W.

[0094] The perception branch networks J1, J2, J3, J4, whose inputs are respectively The perception branch networks all have the same network structure, and their trainable parameters are shared. The structure is as follows: Each perception branch network consists of 4 convolutional blocks, and each convolutional block consists of a convolutional layer and an activation layer in sequence. The convolutional layer is a convolution with a kernel size of 3×3, a stride of 2, and a padding of 1, and the activation layer uses the ReLU activation function. The four equally divided cropped images are respectively input into the perception branch networks J1, J2, J3, J4. For each perception branch network, the features output by the 3rd convolutional block and the 4th convolutional block are respectively extracted. Among them, the output feature of the 3rd convolutional block is denoted as The output feature of the 4th convolutional block is denoted as They respectively correspond to the features of 4 local parts of the image.

[0095] The feature transformation block T1 is composed of a convolutional layer with a kernel size of 1×1, a stride of 0, and a padding of 0, and accepts the aggregated output feature of the 3rd convolutional block of the perception branch network, and outputs the convolutional parameter k1 for image enhancement. The feature transformation block T2 is composed of a convolutional layer with a kernel size of 1×1, a stride of 0, and a padding of 0, and its output channel number is 25% of the input channel number, and accepts the aggregated output feature of the 4th convolutional block of the perception branch network, and outputs the convolutional parameter k2 for image enhancement. Expressed by the formula as follows:

[0096]

[0097]

[0098] where Concat(·) represents concatenating the features in the convolutional channel dimension, and T1 and T2 are feature transformation blocks.

[0099] In this embodiment, as Figure 4 shown, the enhancement branch network includes 2 convolutional blocks, 2 variable parameter convolutional blocks D1, D2 and an upsampling layer, and consists of an upsampling layer; the convolutional block consists of a convolutional layer and an activation layer in sequence, the convolutional layer uses a convolution with a kernel size of 3×3, a stride of 1, and a padding of 0, and the activation layer uses the ReLU activation function;

[0100] The variable parameter convolution blocks D1 and D2 are composed of a variable parameter convolution layer and an activation layer in sequence. Among them, the convolution kernel used in the convolution layer of the variable parameter convolution block D1 is the output k1 from the local scene perception branch network, the convolution stride is 1, the padding is 0, and the activation layer uses the ReLU activation function; the convolution kernel used in the convolution layer of the variable parameter convolution block D2 is the output k2 from the local scene perception branch network, the convolution stride is 1, the padding is 0, and the activation layer uses the ReLU activation function;

[0101] The input of the upsampling layer is the output feature from the variable parameter convolution block D2. The upsampling operation performs bilinear interpolation on each feature channel. Among them, the output size of the upsampling is set to H×W, and the enhanced illumination feature of the generated image is denoted as f.

[0102] In this embodiment, as Figure 5 shown, the attention module receives the enhanced illumination feature f of the image output by the enhanced branch network, and at the same time receives the normalized image X of size H×W as input. Here, the image X is the same as the input of the enhanced branch network. The attention module consists of an attention layer and one convolution block.

[0103] The input of the attention layer is the enhanced illumination feature f of the image output by the enhanced branch network, and at the same time receives the normalized image X of size H×W as input; the image X is converted into a grayscale image denoted as X gary , the grayscale image X gary passes through a Gaussian filter to obtain a smoothed grayscale image X smooth , where the Gaussian kernel size of the filter is set to 3. The attention layer calculates the output in the following way:

[0104]

[0105] where is the element-wise multiplication of matrices with a broadcasting mechanism. The broadcasting mechanism performs element-wise multiplication of each channel dimension in f with X smooth ;

[0106] The convolution block is composed of a convolution layer and an activation layer in sequence, and receives f A′ as input. The convolution layer in it is a convolution with a convolution kernel size of 3×3, a stride of 1, and a padding of 0. The activation layer uses the tanh activation function. The output f A of the convolution block is the enhanced illumination coefficient of the image output by the attention module.

[0107] In this embodiment. The input of the enhancement module is the normalized image X of size H×W and the enhanced illumination coefficient f generated by the attention module A , and the formula for generating the preliminarily enhanced image is as follows:

[0108]

[0109] The above formula is in an iterative form, where n is the number of iterations, and E n represents the result obtained after iterating n times as the illumination coefficient for image enhancement, and the default setting is n = 4. For convenience of representation, the image E4 after preliminary enhancement is denoted as E; A As the result obtained by iterating n times as the illumination coefficient for image enhancement, the default setting is n = 4. For convenience of representation, the image E4 after preliminary enhancement is denoted as E;

[0110] The input of the denoising module is the preliminarily enhanced image E and the normalized image X of size H×W, and the output of this module is the final image enhancement result The denoising module consists of 3 convolutional blocks. Each convolutional block in the denoising module. The 3 convolutional blocks in the denoising module are each composed of a convolutional layer and an activation layer in sequence. The convolutional layer uses a convolution with a kernel size of 3×3, a stride of 1, and a padding of 1, and the activation layer uses the ReLU activation function. The final image enhancement result of size H×W is output

[0111] In this embodiment, the optimization objective is to minimize the total loss function

[0112]

[0113] where represents the brightness control loss function, and λ bri represents the weight of the brightness control loss function; represents the color enhancement loss function, and λ col represents the weight of the color enhancement loss function; represents the global consistency constraint loss function, and λ glo represents the weight of the global consistency constraint loss function; represents the illumination smoothness loss function, and λ tv represents the weight of the illumination smoothness loss function, represents the denoising loss function, and λ noi represents the weight of the denoising loss function;

[0114] Brightness control loss function The calculation formula is as follows:

[0115]

[0116] where the preliminarily enhanced image E is divided into N = 16×16, that is, 256 regions of the same size, and the average brightness value within the region is denoted as E b = 0.6 is the preset target brightness constant value, and || ||1 is the absolute value operation;

[0117] Color enhancement loss function The calculation formula is as follows:

[0118]

[0119] In the above formula, c1 and c2 represent the color channels of the image, Ω = {(R, G), (R, B), (G, B)} represents the combination of different color channels, where R, G, and B represent the red, green, and blue channels in the RGB color space respectively; and represent the mean values of the c1 color channel and the c2 color channel in the preliminarily enhanced image E; similarly, represent the mean values of the c1 color channel and the c2 color channel in the input image X; || || 2 is the operation to obtain the Euclidean norm; the definition of G(·) is as follows:

[0120]

[0121] where represent the mean values of the c1 color channel and the c2 color channel in any input image I, || ||1 is the absolute value operation, Gamma γ (·) represents the gamma correction operation, and its calculation method is as follows:

[0122]

[0123] where represent the c1 ∈ {R, G, B} color channel in any input image I, and the general empirical setting of the parameter is γ = 2.2.

[0124] Global consistency constraint loss function The calculation formula is as follows:

[0125]

[0126] where, represents the global maximum consistency constraint loss, represents the global minimum consistency constraint loss;

[0127] represents the global maximum consistency constraint loss, making the brightest regions in the 16×16 partition space of the image change as identically as possible

[0128]

[0129] Among them, the preliminarily enhanced image E is divided into 16×16 parts, and the average pixel value is calculated for each region to obtain an image with a width and height of 16×16 composed of the average pixel values. Then, max pooling (MaxPooling) is performed, and the size of the pooling operation is set to 8×8. The resulting image is of size 2×2, and each pixel value is respectively denoted as Similarly, the input image X is divided into 16×16 parts, and the average pixel value is calculated for each region to obtain an image with a width and height of 16×16 composed of the average pixel values. Then, max pooling (MaxPooling) is performed, and the size of the pooling operation is set to 8×8. The resulting image is of size 2×2, and each value therein is respectively denoted as (r1, r2) ∈ {(1, 2), (1, 3), (1, 4), (2, 3), (2, 4), (3, 4)} are different combinations of the extracted maximum values. || ||1 is the absolute value operation.

[0130] Denotes the global minimum consistency constraint loss, making the changes in the darkest regions in the 16×16 partition space of the image as similar as possible. The calculation formula is as follows:

[0131]

[0132] Among them, the preliminarily enhanced image E is divided into 16×16 parts, and the average pixel value is calculated for each region to obtain an image with a width and height of 16×16 composed of the average pixel values. Then, min pooling (MinPooling) is performed, and the size of the pooling operation is set to 8×8. The resulting image is of size 2×2, and each pixel value is respectively denoted as Similarly, the input image X is divided into 16×16 parts, and the average pixel value is calculated for each region to obtain an image with a width and height of 16×16 composed of the average pixel values. Then, min pooling (MinPooling) is performed, and the size of the pooling operation is set to 8×8. The resulting image is of size 2×2, and each value therein is respectively denoted as (r1, r2) ∈ {(1, 2), (1, 3), (1, 4), (2, 3), (2, 4), (3, 4)} are different combinations of the extracted minimum values. || ||1 is the absolute value operation.

[0133] Illumination smoothing loss function The calculation formula is as follows:

[0134]

[0135] Among them, The illumination coefficient f is the image enhancement output by the attention module A The corresponding color channel c is taken from the RGB color space. || || 2 is the operation that finds the Euclidean norm. represents the operation of obtaining the first-order gradient, that is, and They are The first-order difference in the vertical and horizontal directions, || ||1 is the absolute value operation;

[0136] Denoising loss function The calculation formula is as follows:

[0137]

[0138] Here, Φ(·) represents the operation of extracting Conv4-1 layer features using the VGG-16 classification model pre-trained on ImageNet; therefore, Φ(E) represents the classification features extracted from the initially enhanced image E. Represents the final image enhancement result The extracted classification features; represents the operation of obtaining the first-order gradient, as well as in and They are The first-order differences in the vertical and horizontal directions, and are the first-order differences of E in the vertical and horizontal directions, respectively; where * represents the multiplication of the corresponding elements of the matrix, e is the base of the natural logarithm, and μ is a parameter that can control the perception of edge strength.

[0139] In this embodiment, step S4 is specifically as follows:

[0140] Step S41: Select a random training image X from the training data set;

[0141] Step S42: Input image X, and obtain the initially enhanced image E after passing through the local scene perception branch network, the enhancement branch network, the attention module and the iterative enhancement module, and the illumination coefficient f of the image enhancement output by the attention module A And the final image enhancement result Calculate the total loss function loss

[0142] Step S43: using the back propagation method to calculate the gradients of the parameters in the local scene perception branch network, the enhancement branch network, the attention module and the iterative enhancement module, and using the Adam optimization method to update the parameters;

[0143] Step S44: The above steps are one iteration of the training process. The entire training process requires a preset number of iterations. And during each iteration, multiple image pairs are randomly sampled as a batch for training.

[0144] The above are only the preferred embodiments of the present invention. All equivalent changes and modifications made according to the scope of the patent application of the present invention shall fall within the scope covered by the present invention.

Claims

1. A no-reference low-light image enhancement method based on local scene perception, characterized in that, It includes the following steps: Step S1: Obtain low-light images and preprocess each image to obtain a training dataset; Step S2: Construct a low-light image enhancement network based on local scene perception; Step S3: Design a reference-free loss function for training the network designed in Step S2; Step S4: Based on the reference-free loss function, use the training dataset to train the low-light image enhancement network based on local scene perception; Step S5: Pass the image to be tested through the trained low-light image enhancement network based on local scene perception to obtain a normal illumination image; The low-light image enhancement network based on local scene perception includes a local scene perception branch network, an enhancement branch network, an attention module, an iterative enhancement module, and a denoising module; The local scene perception branch network receives the normalized image X of size H×W as input and performs a spatial four-equal-part cropping. The cropped image is represented as The local scene perception branch network consists of four perception branch networks J1, J2, J3, J4 and feature transformation blocks T1, T2. The perception branch networks are used to extract the features of local images, and the feature transformation blocks are used to generate convolution parameters for image enhancement. The operation of obtaining the spatially four-equal-part cropped image is as follows: where (p, q) is the pixel position of the input normalized image X with size H×W; The perception branch networks J1, J2, J3, J4 all have the same network structure, and their trainable parameters are shared; the perception branch network is composed of 4 convolutional blocks, and each convolutional block is composed of a convolutional layer and an activation layer in sequence; The feature transformation block T1 is composed of a convolutional layer with a kernel size of 1×1, a stride of 0, and a padding of 0, and accepts the aggregated output features of the 3rd convolutional block of the perception branch network, and outputs the convolutional parameter k1 for image enhancement; the feature transformation block T2 is composed of a convolutional layer with a kernel size of 1×1, a stride of 0, and a padding of 0, and its output channel number is 25% of the input channel number, and accepts the aggregated output features of the 4th convolutional block of the perception branch network, and outputs the convolutional parameter k2 for image enhancement, which is expressed by the following formula: where Concat(·) represents concatenating features in the convolutional channel dimension, and T1 and T2 are feature transformation blocks.

2. The method for enhancing low - illumination images without reference based on local scene perception according to claim 1, wherein, The preprocessing is specifically: Scale each image in the dataset to an image with the same size of H×W; Normalize the image. Given an image I train , calculate the normalized image using the following formula: where I train is an image of size H×W with an 8-bit color depth, and I bit_max is an image of size H×W with all pixel values being 255.

3. The method for enhancing a reference-free low-illumination image based on local scene perception according to claim 1, wherein The enhancement branch network includes 2 convolutional blocks, 2 variable parameter convolutional blocks D1, D2, and an upsampling layer, and consists of an upsampling layer; the convolutional block is composed of a convolutional layer and an activation layer in sequence, and the convolutional layer uses a convolutional kernel with a size of 3×3, a stride of 1, and a padding of 0, and the activation layer uses the ReLU activation function; The variable parameter convolutional blocks D1, D2 are composed of a variable parameter convolutional layer and an activation layer in sequence; among them, the convolutional layer in the variable parameter convolutional block D1 uses a convolutional kernel from the output k1 of the local scene perception branch network, a convolutional stride of 1, and a padding of 0, and the activation layer uses the ReLU activation function; the convolutional layer in the variable parameter convolutional block D2 uses a convolutional kernel from the output k2 of the local scene perception branch network, a convolutional stride of 1, and a padding of 0, and the activation layer uses the ReLU activation function; The input of the upsampling layer is the output features from the variable parameter convolutional block D2, and the upsampling operation is a bilinear interpolation operation for each feature channel, where the output size of the upsampling is set to H×W, and the generated illumination feature for image enhancement is denoted as f.

4. The method for enhancing a reference-free low-illumination image based on local scene perception according to claim 1, wherein The attention module includes an attention layer and a convolutional block; The input of the attention layer is the image enhancement illumination feature f output by the enhanced branch network, and at the same time receives the normalized image X of size H×W as input; the image X is converted into a grayscale image denoted as X gary , the grayscale image X gary passes through a Gaussian filter to obtain a smoothed grayscale image X smooth , where the Gaussian kernel size of the filter is set to 3, and the attention layer calculates the output in the following way: Among them is the element-wise multiplication of matrices with a broadcasting mechanism, and the broadcasting mechanism multiplies each channel dimension in f with X smooth for element-wise matrix multiplication; The convolutional block consists of a convolutional layer and an activation layer in sequence, and accepts f A' as input. The convolutional layer in it is a convolution with a kernel size of 3×3, a stride of 1, and a padding of 0. The activation layer uses the tanh activation function. The output f A of the convolutional block is the illumination coefficient for image enhancement output by the attention module.

5. The method for enhancing a reference-free low-illumination image based on local scene perception according to claim 1, wherein The input of the enhancement module is the normalized image X of size H×W, and the illumination coefficient f for image enhancement generated by the attention module A , and the formula for generating the preliminarily enhanced image is as follows: The above formula is in an iterative form, where n is the number of iterations and E n represents the result obtained by iterating n times as the illumination coefficient for image enhancement after passing through f A The default setting is n = 4; For convenience of representation, denote the preliminarily enhanced image E4 as E; The input of the denoising module is the preliminarily enhanced image E and the normalized image X with a size of H×W, and the output of this module is the final image enhancement result 6. The method for enhancing low - illumination images without reference based on local scene perception according to claim 1, characterized in that, The optimization objective is to minimize the total loss function Among them, represents the brightness control loss function, and λ bri represents the weight of the brightness control loss function; represents the color enhancement loss function, and λ col represents the weight of the color enhancement loss function; represents the global consistency constraint loss function, and λ glo represents the weight of the global consistency constraint loss function; represents the illumination smoothing loss function, and λ tv represents the weight of the illumination smoothing loss function, represents the denoising loss function, and λ noi represents the weight of the denoising loss function; Brightness control loss function The calculation formula is as follows: Among them, the preliminarily enhanced image E is divided into N = 16×16, that is, 256 regions of the same size, and the average luminance value within the region is denoted as E b = 0.6 is a preset target luminance constant value, and || ||1 is an absolute value operation; Color enhancement loss function The calculation formula is as follows: In the above formula, c1 and c2 represent the color channels of the image, and Ω = {(R, G), (R, B), (G, B)} represents the combinations of different color channels, where R, G, and B represent the red, green, and blue channels in the RGB color space respectively; and respectively represent the means of the c1 and c2 color channels in the preliminarily enhanced image E; similarly, respectively represent the means of the c1 and c2 color channels in the input image X; || || 2 is the operation to obtain the Euclidean norm; The definition of G(·) is as follows: where represents the mean values of the c1 and c2 color channels in any input image I, || ||1 is the absolute value operation, and Gamma γ (·) represents the gamma correction operation, and its calculation method is as follows: wherein represents the c1∈{R, G, B} color channel in any input image I, and γ is a preset parameter; Global consistency constraint loss function The calculation formula is as follows: Among them, represents the global maximum consistency constraint loss, represents the global minimum consistency constraint loss; Indicates the global maximum consistency constraint loss, and the calculation formula is as follows: Among them, the input image X is divided into 16×16 parts, and the average pixel value is calculated for each region to obtain an image with a width and height of 16×16 composed of the average pixel values. Then, max pooling is performed, and the size of the pooling operation is set to 8×8. The resulting image is of size 2×2, and each value in it is denoted as (r1, r2) ∈ {(1, 2), (1, 3), (1, 4), (2, 3), (2, 4), (3, 4)} are different combinations for extracting the maximum value; || ||1 is the absolute value operation; Indicates the global minimum consistency constraint loss, which makes the changes in the darkest regions in the 16×16 partition space of the image as identical as possible. The calculation formula is as follows: Among them, the preliminarily enhanced image E is divided into 16×16 regions, and the average pixel value of each region is calculated to obtain an image with a width and height of 16×16 composed of the average pixel values. Then, min pooling is performed, and the size of the pooling operation is set to 8×8. The resulting image is of size 2×2, and each pixel value is respectively denoted as Similarly, the input image X is divided into 16×16 regions, and the average pixel value of each region is calculated to obtain an image with a width and height of 16×16 composed of the average pixel values. Then, min pooling is performed, and the size of the pooling operation is set to 8×8. The resulting image is of size 2×2, and each value therein is respectively denoted as (r1, r2) ∈ {(1, 2), (1, 3), (1, 4), (2, 3), (2, 4), (3, 4)} are different combinations of the extracted minimum values; Illumination smoothing loss function The calculation formula is as follows: Among them, is the illumination coefficient f of the image enhancement output by the attention module A for the corresponding color channel c, where c is taken from the RGB color space; || || 2 is the operation of obtaining the Euclidean norm; represents the operation of obtaining the first-order gradient, that is and are respectively the first-order differences in the vertical and horizontal directions, and || ||1 is the absolute value operation; Denoising loss function The calculation formula is as follows: Among them, Φ(·) represents the operation of extracting the features of the Conv4-1 layer using the VGG-16 classification model pre-trained on ImageNet; thus, Φ(E) represents the classification features extracted from the preliminarily enhanced image E. represents the classification features extracted from the final image enhancement result represents the operation of obtaining the first-order gradient, and there is and where and are respectively the first-order differences in the vertical and horizontal directions, and are respectively the first-order differences of E in the vertical and horizontal directions; in the formula, * represents the element-wise multiplication of matrices, e is the base of the natural logarithm, and μ is a parameter that can control the edge intensity perception.

7. The method for enhancing a reference-free low-illumination image based on local scene perception according to claim 1, wherein The specific steps of step S4 are as follows: Step S41: Select a random training image X from the training dataset; Step S42: Input the image X, and obtain the preliminarily enhanced image E and the illumination coefficient f for image enhancement output by the attention module through the local scene perception branch network, the enhancement branch network, the attention module, and the iterative enhancement module A and the final image enhancement result Calculate the total loss function loss Step S43: Use the backpropagation method to calculate the gradients of the parameters in the local scene perception branch network, enhancement branch network, attention module, and iterative enhancement module, and update the parameters using the Adam optimization method; Step S44: The above steps are one iteration of the training process. The entire training process requires a preset number of iterations. Moreover, in each iteration process, multiple image pairs are randomly sampled as a batch for training.

Citation Information

Patent Citations

  • Image Semantic Segmentation Method Based on Deep Full Convolutional Network and Conditional Random Field

    AU2020103901A4

  • Space-based intelligent imaging system

    CN110207671A