Low-illumination image enhancement method based on multi-scale fusion U-Net network structure

Through a low-illumination image enhancement method based on multi-scale fusion U-Net network structure, combined with illumination coefficient estimation, multi-scale illumination extraction and attention mechanism, the problems of poor low-illumination image enhancement effect and difficulty in deploying edge devices in the existing technology are solved, and high-efficiency and low-power image quality improvement is achieved.

CN120013787AActive Publication Date: 2025-05-16NANJING UNIV OF POSTS & TELECOMM

Patent Information

Application Number
CN202510178406.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-05-16
Estimated Expiration
2045-02-18

AI Technical Summary

Technical Problem

The existing low-illumination image enhancement technology is difficult to achieve ideal visual effects under different lighting conditions, and it is difficult to deploy on edge devices and is inefficient.

Method used

The low-illumination image enhancement method based on multi-scale fusion U-Net network structure is adopted, including the lighting coefficient estimation module, the encoder network structure and the decoder network structure. The illumination of the image is improved through multi-scale illumination extraction blocks and local residual layers, and combined with the channel attention mechanism and spatial attention mechanism, the feature expression and image reconstruction quality are optimized.

Benefits of technology

It effectively improves the quality of low-illumination images, solves the problems of light unevenness and details loss, improves the structure and texture of the image, and is suitable for low-power and high-efficiency deployment of embedded devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120013787A_ABST
    Figure CN120013787A_ABST
Patent Text Reader

Abstract

The invention discloses a low-illumination image enhancement method based on a multi-scale fusion U-Net network structure. The method comprises the following steps: S1, constructing a training data set and a test sample set; s2, constructing a low-illumination enhancement network model; s3, jointly calculating a plurality of loss functions; s4, training a low-illumination enhancement network model; s5, performing conversion quantitative deployment on the trained low-illumination enhanced network model; s6, preprocessing the image data to be enhanced; s7, inputting the preprocessed image data into the low-illumination enhancement network model for image enhancement; and S8, outputting an enhanced normal illumination image by the low illumination enhancement network model. According to the method, on the premise of ensuring high accuracy, the parameter quantity of the model is reduced, the portability of the model is enhanced, and the low-illumination image enhancement technology has a good development prospect in the embedded field.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image enhancement, and in particular to a low-illumination image enhancement method based on a multi-scale fusion U-Net network structure. Background Art

[0002] In today's information society, digital images are widely used because they can efficiently transmit a large amount of information and are in line with the intuitive perception of the human eye. However, due to environmental factors (such as rainy days, foggy days, nights, etc.) and limitations of shooting equipment (such as underexposure, low light, etc.), images often have problems such as excessive noise, loss of details, and poor visual effects, resulting in poor image quality. Therefore, image enhancement technology, especially low-light image enhancement, has become an important research direction in the field of digital image processing.

[0003] In low-light environments, images taken often appear dark in color, unclear in details, and poor in visual effects, which seriously affects the transmission of information and poses challenges to image enhancement technology. In recent years, image enhancement models based on deep learning have made significant progress, but it is still difficult to achieve ideal visual effects under different lighting conditions. In addition, the deployment of these image enhancement models on edge devices is difficult and inefficient. Summary of the invention

[0004] The present invention provides a low-illumination image enhancement method based on a multi-scale fusion U-Net network structure, which can effectively solve the limitations of low-illumination image enhancement in existing methods and greatly improve image quality, and can solve at least one of the above technical problems.

[0005] In order to solve the above technical problems, the present invention adopts the following technical solutions:

[0006] A low-illumination image enhancement method based on a multi-scale fusion U-Net network structure comprises the following steps:

[0007] S1, construct training data set and test sample set;

[0008] S2. Construct a low-light enhancement network model, wherein the low-light enhancement network model includes:

[0009] An illumination coefficient estimation module, wherein the illumination coefficient estimation module is used to estimate the channel stretch coefficient corresponding to the RGB channel of the image pixel by pixel;

[0010] An encoder network structure, the encoder network structure comprising a residual downsampling module and an illumination enhancement module, the residual downsampling module is used to perform shallow feature extraction and downsampling operations, and the illumination enhancement module is used to improve the illumination of the image by using channel attention;

[0011] The lighting enhancement module includes a multi-scale illumination extraction block and a local residual layer;

[0012] The illumination extraction block includes three layers of feature extraction units, each of which is composed of a convolution layer, a normalization layer, and a ReLU activation function layer, and each of which concatenates the output of the feature extraction unit in the previous layer in the input feature map;

[0013] The lighting enhancement module comprises two cascaded lighting extraction blocks of different scales, the convolution kernel sizes of the two lighting extraction blocks are 3×3 and 5×5 respectively, the last layer of the lighting extraction block performs local residual on the lighting extraction block input and the lighting extraction block output, and the lighting enhancement module output is composed of a feature map of the outputs of the two lighting extraction blocks of different scales spliced ​​in the channel dimension;

[0014] A decoder network structure, wherein the decoder network structure includes a residual upsampling module, a skip connection module, and an attention mechanism module, wherein the feature maps of the corresponding encoders are fused in the decoder to enhance the image reconstruction quality and generate contextual attention weights to optimize the expression capability of the feature maps;

[0015] S3, jointly calculate multiple loss functions;

[0016] S4, training low-light enhancement network model;

[0017] S5. Convert and quantize the trained low-light enhancement network model;

[0018] S6, preprocessing the image data to be enhanced;

[0019] S7, inputting the preprocessed image data into the low illumination enhancement network model for image enhancement;

[0020] S8, the low illumination enhancement network model outputs the enhanced normal illumination image.

[0021] Furthermore, the illumination coefficient estimation module includes seven convolutional layers, each of which includes a 3×3 convolution kernel and a ReLU activation function;

[0022] The first to fourth convolutional layers are connected in sequence. The fifth convolutional layer inputs the feature map of the outputs of the third and fourth convolutional layers in the channel dimension. The sixth convolutional layer inputs the feature map of the outputs of the second and fifth convolutional layers in the channel dimension. The seventh convolutional layer outputs the illumination enhancement coefficient map.

[0023] Furthermore, the residual downsampling module combines the residual and double convolution design, reduces the dimension of the input feature map through the first layer of the convolution layer to extract features, realizes channel alignment through 1×1 convolution, and further enhances the feature expression through the second layer of the convolution layer. All convolution operations are followed by a normalization layer and a ReLU activation function layer, and finally the main path and input path features are fused through an addition operation to achieve efficient spatial dimension reduction and feature retention.

[0024] Further, the residual upsampling module includes a main path and a residual path;

[0025] The main path includes two convolutional layers, a normalization layer, and an activation function. The convolutional layer is used to extract local information of the input features, and the activation function is used to introduce nonlinear characteristics.

[0026] The residual path includes a 1×1 convolution layer and a normalization layer to match the number of channels, which is used to superimpose the input features with the output of the main path;

[0027] The outputs of the main path and the residual path are added, and the spatial resolution is enlarged through the deconvolution layer to achieve feature map upsampling.

[0028] Furthermore, the jump connection module includes two convolutional layers, each of which is followed by a normalization layer and a ReLU activation function layer, and the number of input and output channels of the convolutional layer remains consistent, the convolution kernel size is 3×3, and the padding is 1.

[0029] Furthermore, in S3, after each image reconstruction, a joint loss function needs to be calculated to determine the quality of image reconstruction and optimize network parameters;

[0030] The joint loss function is a weighted combination of reconstruction loss, perception loss, structural similarity loss, and color loss, which is:

[0031]

[0032] Among them, λ r , p , s and λ c are the weights of reconstruction loss, perceptual loss, structural similarity loss, and color loss, respectively. and L c They are reconstruction loss, perception loss, structural similarity loss and color loss respectively. for joint losses;

[0033] Reconstruction loss The calculation formula is:

[0034]

[0035] Among them, E′ n is the enhanced result through network training, Y is ′ n The corresponding expected normal lighting real value;

[0036] Extract E through VGG network ′ n The features of and Y will be defined based on the pre-trained VGG-16 network Perceived loss The calculation formula is:

[0037]

[0038] Among them, w ij 、h ij and c ij Describes the dimensions of each feature map in the VGG-16 network, φ ij Represents the feature map obtained by the jth convolutional layer of the i-th block in the VGG-16 network;

[0039] Structural Similarity Loss The calculation formula is:

[0040]

[0041]

[0042] Among them, SSIM is the image structure similarity, S and T represent the images to be compared, μ S and μ T is the average pixel value of the two images, and is the variance of the two images, σ ST is the covariance of the two images, c1 and c2 are two small constants used to prevent the denominator from being zero;

[0043] Color Loss The calculation formula is:

[0044]

[0045]

[0046] Among them, S p and T p denote the RGB color vector of pixel p in images S and T respectively. CA(·,·) denotes a three-dimensional vector with RGB color. The color loss between pixels is the angle between the two color vectors. The total color loss is the sum of the color losses of all pixels. The greater the color deviation between the two images, the greater the loss function.

[0047] Furthermore, in S4, the input of the low illumination enhancement network model is the low illumination image and the normal exposure image, and the output is the predicted reconstructed image. The training process further includes:

[0048] S41, randomly dividing the images to be trained into several batches, each batch containing the same number of images;

[0049] S42. Use the batches of images to train and optimize the low-light enhancement network model until the calculated joint loss function reaches a loss threshold or the number of iterations reaches a number threshold.

[0050] Furthermore, the S5 further includes:

[0051] S51, converting the trained low-illumination image enhancement network model into an ONNX format, and then converting the low-illumination image enhancement network model in the ONNX format into a format supported by a hardware device;

[0052] S52, quantizing the low-illumination image enhancement network model in a format supported by the hardware device;

[0053] S53, deploying the quantized low-illumination image enhancement network model on an edge device, and loading the quantized low-illumination image enhancement network model;

[0054] S54, input the preprocessed image into the quantized low-light image enhancement network model, obtain an output result from the quantized low-light image enhancement network model, and save the output result in an appropriate format and range through OpenCV.

[0055] Furthermore, the quantization process in S52 further includes:

[0056] S521, converting the floating point numbers in the low-light image enhancement network model into integer data to reduce the memory usage of the model and speed up the reasoning speed, and calculating the scaling factor and offset respectively:

[0057]

[0058]

[0059] S522, floating point data X f Perform quantization operation and convert to uint8 type data X q , the calculation formula is:

[0060]

[0061] Among them, X max and Xmin Respectively represent the maximum and minimum values ​​of floating-point numbers, X f Represents a floating point number, X q Indicates the quantized uint8 type data, the round function indicates the rounding operation, and the clamp function is used to ensure the quantization result X q In the interval [0, 255], clamp is defined as:

[0062]

[0063] Among them, the clamp function limits the randomly changing values ​​to a given interval, a and b are both expressed as constants, and x is expressed as a variable.

[0064] Furthermore, the preprocessing process in S6 further includes:

[0065] S61, reading the low-light image to be enhanced and converting it into RGB format;

[0066] S62, converting the image data into a NumPy array, and converting the data format into np.float32 to ensure the accuracy of subsequent processing;

[0067] S63. Convert the image data into NHWC format to ensure the consistency between the data format accepted by the low-illumination image enhancement network model and the required data format.

[0068] The beneficial effects of the present invention are embodied in:

[0069] 1. Enhanced low-light image processing capability. The present invention can effectively enhance low-light images during the training process by combining the channel attention mechanism (CAM) and the spatial attention mechanism (SAM) residual module. The channel attention mechanism can perform weighted correction on the image in the channel dimension to solve the color cast problem of low-light images, while the spatial attention mechanism constrains the regional deviation of the feature map in the spatial dimension, thereby effectively solving the problem of spatial inhomogeneity.

[0070] 2. The present invention constructs a lighting enhancement module, which uses a multi-scale lighting extraction block to extract lighting information of different scales in the image, and the last layer of the lighting extraction block applies local residual learning to improve the lighting effect and capture the details and brightness distribution in the image. The module restores the number of channels of the feature map through a compression layer of a 1×1 convolution kernel, and maintains the key information of brightness enhancement through a residual connection method, thereby improving the network's ability to model lighting distribution.

[0071] 3. The present invention adopts a multi-scale fusion mechanism, which can extract global information (such as brightness distribution) and local information (such as texture and details) at different resolutions. It performs well in low-illumination image enhancement, can effectively enhance the structure and texture of the image, and achieve a good balance between brightness, structure and details.

[0072] 4. The present invention adopts a relatively lightweight design scheme in network design and optimizes the model calculation amount. These advantages make this method have a wide range of application potentials on embedded devices, especially suitable for low-light image enhancement tasks that require low power consumption, high efficiency and real-time processing, and can effectively improve image quality and adapt to embedded platforms. BRIEF DESCRIPTION OF THE DRAWINGS

[0073] The drawings described herein are used to provide further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute improper limitations on the present application.

[0074] Figure 1 It is a schematic diagram of the overall process of the low-illumination image enhancement method according to an embodiment of the present invention.

[0075] Figure 2 Schematic diagram of a low-light enhancement network model according to an embodiment of the present invention.

[0076] Figure 3 Schematic diagram of the lighting enhancement module structure of an embodiment of the present invention.

[0077] Figure 4 It is a schematic diagram of a 3×3 illumination extraction block structure according to an embodiment of the present invention.

[0078] Figure 5 This is a set of comparison charts of the visualization effects of image enhancement processing in indoor scenes of the LOL dataset using various algorithms.

[0079] Figure 6 This is another set of comparison charts of the visualization effects of image enhancement processing in indoor scenes of the LOL dataset using multiple algorithms.

[0080] Figure 7 It is a structural block diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0081] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all of the embodiments. In the absence of conflict, the embodiments in this application and the features in the embodiments can be combined with each other. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0082] It should be noted that if there are descriptions involving "first", "second", etc. in the embodiments of the present invention, the descriptions of "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or suggesting their relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In addition, the meaning of "and / or" appearing in the full text includes three parallel schemes. Taking "A and / or B" as an example, it includes scheme A, or scheme B, or a scheme that satisfies both A and B. In addition, "multiple" refers to more than two.

[0083] See also Figure 1 The embodiment of the present invention provides a low-illumination image enhancement method based on a multi-scale fusion U-Net network structure, comprising the following steps:

[0084] S1, construct training data set and test sample set;

[0085] S2, build low-light enhancement network model;

[0086] S3, jointly calculate multiple loss functions;

[0087] S4, training low-light enhancement network model;

[0088] S5. Convert and quantize the trained low-light enhancement network model;

[0089] S6, preprocessing the image data to be enhanced;

[0090] S7, inputting the preprocessed image data into the low illumination enhancement network model for image enhancement;

[0091] S8, the low illumination enhancement network model outputs the enhanced normal illumination image.

[0092] In this embodiment, in S1, in order to train the low-light enhancement network model, it is necessary to construct a training data set and a test sample set. The training data set is a sequence set composed of multiple data pairs, and each data pair is composed of a low-light image and a normal-light image.

[0093] First, multiple low-light images in various scenes and normal-light images corresponding to the low-light images can be obtained as data pairs. The present invention collects a total of 2289 pairs of images from the public data sets LOL, LOL-v1 and LOL-v2 sources. The data content is low / normal-light image pairs. 90% of the 2289 pairs of images are training sets and 10% are test sets. The training set is 2074 pairs of images and the test set is 215 pairs of images. All of these images are converted into PNG format, and before the pictures are input into the low-light enhancement network model for training, the images are randomly rotated (0°, 90°, 180°, 270°) and flipped as data expansion.

[0094] See also Figure 2 In this embodiment, in S2, the low-light image is input into a multi-scale fusion U-Net network, and an enhanced image is output. The multi-scale fusion U-Net network includes:

[0095] An illumination coefficient estimation module, wherein the illumination coefficient estimation module is used to estimate the channel stretch coefficient corresponding to the RGB channel of the image pixel by pixel;

[0096] An encoder network structure, the encoder network structure comprising a residual downsampling module and an illumination enhancement module, the residual downsampling module is used to perform shallow feature extraction and downsampling operations, and the illumination enhancement module is used to improve the illumination of the image by using channel attention;

[0097] The decoder network structure includes a residual upsampling module, a jump connection module and an attention mechanism module, and the attention mechanism module includes an SSA module and a CBAM module. The feature maps of the corresponding encoders are fused in the decoder to enhance the image reconstruction quality, generate contextual attention weights to optimize the expressiveness of the feature maps, and further improve the importance of features through channel attention and spatial attention mechanisms.

[0098] See also Figure 3 In this embodiment, the illumination coefficient estimation module includes seven convolutional layers, each of which includes a 3×3 convolution kernel and a ReLU activation function;

[0099] The first to fourth convolutional layers are connected in sequence. The fifth convolutional layer inputs the feature map of the outputs of the third and fourth convolutional layers in the channel dimension. The sixth convolutional layer inputs the feature map of the outputs of the second and fifth convolutional layers in the channel dimension. The seventh convolutional layer outputs the illumination enhancement coefficient map.

[0100] The network structure of the illumination coefficient estimation module is shown in Table 1 below. The input image first passes through the convolution layers Conv1 to Conv4 to gradually extract features. The convolution layer uses a 3×3 convolution kernel, which is padded with 1. The ReLU activation function is used after the convolution layer. Through multiple convolution layers, higher-level features are gradually learned. The input of Conv5 is the feature map of the outputs of Conv3 and Conv4 spliced ​​in the channel dimension. The input of Conv6 is the feature map of the outputs of Conv2 and Conv5 spliced ​​in the channel dimension. The convolution layer uses a 3×3 convolution kernel, which is padded with 1. The ReLU activation function is used after the convolution layer. The input of the Conv_Out layer is the feature map of the outputs of Conv1 and Conv6 spliced ​​in the channel dimension. The convolution layer uses a 3×3 convolution kernel, which is padded with 1. Finally, the illumination enhancement coefficient map is output through Conv_Out, and this coefficient is multiplied element by element with the input image to complete the preliminary illumination enhancement of the image.

[0101] Table 1 Illumination coefficient estimation module network

[0102]

[0103] See also Figure 3-Figure 4 ,In this embodiment, the lighting enhancement module includes a multi-scale illumination extraction block and a local residual layer;

[0104] The illumination extraction block includes three layers of feature extraction units, each of which is composed of a convolution layer, a normalization layer, and a ReLU activation function layer, and each of which concatenates the output of the feature extraction unit in the previous layer in the input feature map;

[0105] The lighting enhancement module includes two cascaded illumination extraction blocks of different scales, the sizes of the convolution kernels of the two illumination extraction blocks are 3×3 and 5×5 respectively, and the last layer of the illumination extraction block performs local residual on the illumination extraction block input and the illumination extraction block output.

[0106] The network structure of the lighting enhancement module is shown in Table 2 below. The lighting enhancement module further restores the number of channels of the feature map through a compression layer of a 1×1 convolution kernel, and maintains the key information of brightness enhancement through a residual connection method, thereby improving the network's modeling ability for illumination distribution. The spliced ​​feature map is further processed through a 3×3 convolution layer, and then finally adjusted through a 3×3 convolution layer, and the output image is normalized by a Sigmoid activation function.

[0107] Table 2 Lighting enhancement module network

[0108]

[0109] In this embodiment, the residual downsampling module combines residual and double convolution designs, and the network structure is shown in Table 3 below. DownConv1 and DownConv2 represent downsampling convolution layers, and the first convolution layer DownConv1 is used to reduce the dimension of the input feature map and extract features; the Conv layer enhances the feature expression, and all convolution operations are followed by a normalization layer and a ReLU activation function layer; a 1×1 convolution kernel and a step size of 2 are used in the DownConv2 layer to directly map the resolution and number of channels of the input feature map to the target dimension as a residual term; finally, the main path and input path features are fused through an addition operation to achieve efficient spatial dimension reduction and feature retention.

[0110] Table 3 Residual downsampling module network

[0111]

[0112] In this embodiment, the network structure of the residual upsampling module is shown in Table 4 below. The residual upsampling module includes a main path and a residual path;

[0113] The main path includes two convolutional layers, a normalization layer, and an activation function. The convolutional layer is used to extract local information of the input features, and the activation function is used to introduce nonlinear characteristics.

[0114] The residual path includes a 1×1 convolution layer and a normalization layer to match the number of channels, which is used to superimpose the input features with the output of the main path;

[0115] The outputs of the main path and the residual path are added, and the spatial resolution is enlarged through the deconvolution layer to achieve feature map upsampling and reduce the gradient vanishing problem.

[0116] In the process of residual up and down sampling of the present invention, the second branch of the residual passes through a 1×1 convolution layer and a normalization layer for matching the number of channels, so as to realize the superposition of input features and main path output, optimize the model with relatively lightweight calculation, and reduce the cost of labor efficiency investment.

[0117] Table 4 Residual upsampling module network

[0118]

[0119] In this embodiment, the network structure of the jump connection module is shown in Table 5 below. The jump connection module includes two convolutional layers Conv1 and Conv2. Each convolutional layer is followed by a normalization layer and a ReLU activation function layer. The number of input and output channels of the convolutional layer remains consistent, the convolution kernel size is 3×3, and the padding is 1.

[0120] In the skip connection module, the input feature map and the feature map after convolution are spliced ​​in the channel dimension, which effectively retains the low-level feature information of the input and enhances the feature expression ability of the network. The spliced ​​feature map is compressed by a 1×1 convolution layer to ensure that the number of channels of the output feature map is consistent with that of the input feature map. Finally, the output is nonlinearly transformed by the ReLU activation function to introduce nonlinear characteristics and enhance the expression ability of the model.

[0121] Table 5. Skip connection module network

[0122]

[0123] In this embodiment, in S3, after each image reconstruction, a joint loss function needs to be calculated to determine the quality of image reconstruction and optimize network parameters;

[0124] The joint loss function is a weighted combination of reconstruction loss, perception loss, structural similarity loss, and color loss, which is:

[0125]

[0126] Among them, λ r , p , s and λ c are the weights of reconstruction loss, perceptual loss, structural similarity loss, and color loss, respectively. and L c They are reconstruction loss, perception loss, structural similarity loss and color loss respectively. It is a joint loss that uses four loss constraints to achieve noise removal, local detail recovery and color deviation correction in the reconstructed image.

[0127] The joint loss function plays a vital role in the image reconstruction process. By combining multiple loss functions, various aspects of image reconstruction can be comprehensively considered to improve the quality of the final image. The weighted combination of these loss functions can be adjusted according to actual needs during the training process to ensure that the network can achieve balanced and optimized results on different task objectives.

[0128] Reconstruction loss Directly measuring the reconstruction error at the pixel level of the image can effectively ensure that the generated image is as similar as possible to the original image visually and avoid large-scale visual distortion. The calculation formula is:

[0129]

[0130] Among them, E ′ n is the enhanced result through network training, Y is′ n The corresponding expected normal lighting real value;

[0131] Perceived loss By calculating the difference of images in the high-level semantic feature space, it helps the model capture texture, shape, and high-level semantic information, which ensures that the generated images are more perceptually similar to real images. ′ n The features of and Y will be defined based on the pre-trained VGG-16 network The calculation formula is:

[0132]

[0133] Among them, w ij 、h ij and c ij Describes the dimensions of each feature map in the VGG-16 network, φ ij Represents the feature map obtained by the jth convolutional layer of the i-th block in the VGG-16 network;

[0134] Structural Similarity Loss Paying attention to the overall structure, edges, contrast and other features of the image can prevent the generated image from becoming too blurred or losing its original structural features. The calculation formula is:

[0135]

[0136]

[0137] Among them, SSIM is the image structure similarity, S and T represent the images to be compared, μ S and μ T is the average pixel value of the two images, and is the variance of the two images, σ ST is the covariance of the two images, c1 and c2 are two small constants used to prevent the denominator from being zero;

[0138] Color Loss Ensure that the color distribution of the generated normal illumination image is consistent with the original image to avoid color cast or unnatural tones. The calculation formula is:

[0139]

[0140]

[0141] Among them, S p and T pdenote the RGB color vector of pixel p in images S and T respectively. CA(·,·) denotes a three-dimensional vector with RGB color. The color loss between pixels is the angle between the two color vectors. The total color loss is the sum of the color losses of all pixels. The greater the color deviation between the two images, the greater the loss function.

[0142] In this embodiment, in S4, the input of the low illumination enhancement network model is a low illumination image and a normal exposure image, and the output is a predicted reconstructed image. The training process further includes:

[0143] S41, randomly dividing the images to be trained into several batches, each batch containing the same number of images;

[0144] S42. Use the batches of images to train and optimize the multi-scale fusion U-Net network, that is, the low-light enhancement network model, until the calculated joint loss function reaches a loss threshold or the number of iterations reaches a number threshold.

[0145] As an embodiment of the present invention, the configuration of the hardware platform selected for the experiment is a CPU processor Intel (R) Xeon (R) CPU E5-2682 v4 @ 2.5GHz, memory 64G, graphics card Nvidia GeForce RTX 3090, video memory 24G; the software environment is Ubuntu 20.04, Python version is 3.8.0, deep learning framework Pytorch, version 1.11.0. CUDA version is 11.3.1; the software development environment of the experiment is PyCharm 2023.2.1 and Matlab R2019a.

[0146] The training batch of the low-light enhancement network model is 128, and the number of iterations is 300. The present invention uses a learning rate scheduler, the initial learning rate of the network is 0.0001, and the learning rate is adjusted at 100 and 200 rounds of training, and the adjustment rate is 0.1. The hyperparameter λ in the joint loss function in the experiment r , p , s and λ c The values ​​are 1, 1, 3 and 1 respectively. By constructing the training data set and the test sample set as described above, the integrated data set is input into the low-light enhancement network model for training.

[0147] In this embodiment, S5 further includes:

[0148] S51, converting the trained low-illumination image enhancement network model into an ONNX format, and then converting the low-illumination image enhancement network model in the ONNX format into a format supported by a hardware device;

[0149] S52, quantizing the low-illumination image enhancement network model in a format supported by the hardware device;

[0150] S53, deploying the quantized low-illumination image enhancement network model on an edge device, and loading the quantized low-illumination image enhancement network model;

[0151] S54, input the preprocessed image into the quantized low-light image enhancement network model, obtain an output result from the quantized low-light image enhancement network model, and save the output result in an appropriate format and range through OpenCV.

[0152] In the above step S52, the main quantization process is as follows:

[0153] Convert the floating point numbers in the low-light image enhancement network model to integer data to reduce the model's memory usage and speed up inference. Calculate the scaling factor and offset as:

[0154]

[0155]

[0156] For floating point data X f Perform quantization operation and convert to uint8 type data X q , the formula is:

[0157]

[0158] Among them, X max and X min Indicates the maximum and minimum values ​​of floating-point numbers, X f Represents a floating point number, X q Indicates the quantized uint8 type data, the round function indicates the rounding operation, and the clamp function is used to ensure the quantization result X q In the interval [0, 255], clamp is defined as:

[0159]

[0160] The clamp function is used to limit the randomly changing values ​​to a given interval, a and b represent constants, and x represents a variable. This process prevents overflow and distortion by limiting the data to a given interval, ensuring that the quantization results are suitable for hardware inference.

[0161] Through the above quantization steps, the floating-point model can be converted into an integer model that is more suitable for hardware accelerator processing, thereby optimizing the model's operating efficiency and memory usage and improving reasoning performance.

[0162] In this embodiment, the preprocessing process in S6 further includes:

[0163] S61, reading the low-light image to be enhanced and converting it into RGB format;

[0164] S62, converting the image data into a NumPy array, and converting the data format into np.float32 to ensure the accuracy of subsequent processing;

[0165] S63. Convert the image data into NHWC format to ensure the consistency between the data format accepted by the low-illumination image enhancement network model and the required data format.

[0166] See also Figure 5-Figure 6 , the present invention provides a plurality of sets of low-light data set comparison experiments, and shows the comparison results of different algorithms on the LOL data set. Among them, (a) is the input low-light image, (l) is the corresponding real image (normal exposure image), (b)-(j) are the results of other comparison methods, and (k) is the result of the method proposed in the present invention. (b) The image enhanced by the HE method has dense noise points and serious color cast. (c) The image enhanced by the AHE and (i) RUASNet methods is darker, indicating that the two methods cannot effectively enhance the image brightness when the input image illumination is low, and are accompanied by a small amount of noise. (d) The LIME method has deficiencies in local brightness optimization in low-light image processing, resulting in uneven processing effects in different areas, thereby affecting the visual consistency of the overall picture. (e) It can be seen from the figure that the MSR method has significantly improved the brightness of the low-light image, and the enhanced image has a better light and dark gradient effect, and the color saturation is also improved. However, the image clarity is insufficient, and in the high-brightness area of ​​the image, such as Figure 5 The contrast between the calligraphy and painting in the middle and back is insufficient. Figure 6 The background texture is not clear, and the over-enhancement phenomenon occurs, resulting in loss of image details. (f) The EFINet and (h) SCINet methods are not good at optimizing brightness, resulting in unclear light-dark contrast, making it difficult to identify the details of the dark areas of the image, and the picture is accompanied by a lot of noise, resulting in poor visual perception. (j) The EFINet method restores the color reproduction of the image to a large extent, but often causes loss of image details or the presence of random noise, and the picture is not smooth enough. (g) URetinexNet and (k) the method of the present invention are not good at optimizing brightness. Figure 5 The medium enhancement effect is better and the image clarity is higher, but Figure 6 In the figure, (g) the enhanced image processed by the URetinexNet method has serious color cast, and the wood grain texture contrast and color saturation of the background board are low. The image enhanced by the method of the present invention is more natural, with improved contrast and brightness, while restoring the image color and detail information.

[0167] From the comparison of the above experimental results, it can be seen that this method has more obvious advantages than the other nine methods in terms of subjective visual effects.

[0168] In the LOL test data set, image quality evaluation is performed on the enhanced images of the present invention and the enhanced images of the existing method, and the evaluation indicators are: peak signal-to-noise ratio PSNR and structural similarity SSIM.

[0169] Table 6 Comparison of quantitative data in the LOL test dataset

[0170]

[0171] As shown in Table 6 above, the method of the present invention is compared with the current mainstream low-light image enhancement methods HE, AHE, LIME, MSR, Zero-DEC, URetinexNet, SCINet, RUASNet, and EFINet in detail on the PSNR and SSIM evaluation indicators on the LOL dataset. The larger the value of the evaluation indicator, the higher the value of the SSIM, and the closer it is to 1, the better the effect of low-light enhancement is, and the more it can restore the image effect of normal light. It can be seen that the method of the present invention has a good performance advantage compared with the current mainstream low-light image enhancement methods, and the PSNR and SSIM values ​​both exceed the existing methods.

[0172] An embodiment of the present invention also provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the processor executes the steps of the low-illumination image enhancement method based on the multi-scale fusion U-Net network structure as described above.

[0173] See also Figure 7 An embodiment of the present invention also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the low-illumination image enhancement method based on the multi-scale fusion U-Net network structure as described above.

[0174] An embodiment of the present invention also provides a computer program product comprising instructions, which, when executed on a computer, enables the computer to execute the steps of the above-mentioned low-illumination image enhancement method based on a multi-scale fusion U-Net network structure.

[0175] It can be understood that the system, device and storage medium provided in the embodiments of the present invention correspond to the method provided in the embodiments of the present invention. The explanation, examples and beneficial effects of the relevant contents can refer to the corresponding parts in the above-mentioned low-illumination image enhancement method based on the multi-scale fusion U-Net network structure.

[0176] It should be noted that those skilled in the art can understand that all or part of the steps implemented in the embodiments of the present invention can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using hardware, it can be implemented in whole or in part in the form of purchasing standard parts or modified parts. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, server or data center. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium may be a magnetic medium (eg, a floppy disk, a hard disk, a magnetic tape), an optical medium (eg, a DVD), or a semiconductor medium (eg, a solid state disk (SSD)).

[0177] In summary, the low-illumination image enhancement algorithm proposed in the present invention can well maintain the color information of the image and improve the image illumination non-uniformly, which effectively improves the image enhancement effect compared with the existing methods.

[0178] It should be understood that the examples and implementation modes described herein are for illustrative purposes only and are not intended to limit the present invention. Those skilled in the art may make various modifications or changes based on the examples and implementation modes. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A low-illumination image enhancement method based on a multi-scale fusion U-Net network structure, characterized in that: The following steps are involved: S1, construct training data set and test sample set; S2. Construct a low-light enhancement network model, wherein the low-light enhancement network model includes: An illumination coefficient estimation module, wherein the illumination coefficient estimation module is used to estimate the channel stretch coefficient corresponding to the RGB channel of the image pixel by pixel; An encoder network structure, the encoder network structure comprising a residual downsampling module and an illumination enhancement module, the residual downsampling module is used to perform shallow feature extraction and downsampling operations, and the illumination enhancement module is used to improve the illumination of the image by using channel attention; The lighting enhancement module includes a multi-scale illumination extraction block and a local residual layer; The illumination extraction block includes three layers of feature extraction units, each of which is composed of a convolution layer, a normalization layer, and a ReLU activation function layer, and each of which concatenates the output of the feature extraction unit in the previous layer in the input feature map; The lighting enhancement module comprises two cascaded lighting extraction blocks of different scales, the convolution kernel sizes of the two lighting extraction blocks are 3×3 and 5×5 respectively, the last layer of the lighting extraction block performs local residual on the lighting extraction block input and the lighting extraction block output, and the lighting enhancement module output is composed of a feature map of the outputs of the two lighting extraction blocks of different scales spliced ​​in the channel dimension; A decoder network structure, wherein the decoder network structure includes a residual upsampling module, a skip connection module, and an attention mechanism module, wherein the feature maps of the corresponding encoders are fused in the decoder to enhance the image reconstruction quality and generate contextual attention weights to optimize the expression capability of the feature maps; S3, jointly calculate multiple loss functions; S4, training low-light enhancement network model; S5. Convert and quantize the trained low-light enhancement network model; S6, preprocessing the image data to be enhanced; S7, inputting the preprocessed image data into the low illumination enhancement network model for image enhancement; S8, the low illumination enhancement network model outputs the enhanced normal illumination image.

2. The low-illumination image enhancement method based on the multi-scale fusion U-Net network structure according to claim 1, characterized in that: The illumination coefficient estimation module includes seven convolutional layers, each of which includes a 3×3 convolution kernel and a ReLU activation function; The first to fourth convolutional layers are connected in sequence. The fifth convolutional layer inputs the feature map of the outputs of the third and fourth convolutional layers in the channel dimension. The sixth convolutional layer inputs the feature map of the outputs of the second and fifth convolutional layers in the channel dimension. The seventh convolutional layer outputs the illumination enhancement coefficient map.

3. The low-illumination image enhancement method based on a multi-scale fusion U-Net network structure as claimed in claim 1, characterized in that: The residual downsampling module combines the residual and double convolution design, reduces the dimension of the input feature map through the first layer of the convolution layer to extract features, realizes channel alignment through 1×1 convolution, and further enhances the feature expression through the second layer of the convolution layer. All convolution operations are followed by a normalization layer and a ReLU activation function layer, and finally the main path and input path features are fused through an addition operation to achieve efficient spatial dimension reduction and feature retention.

4. The low-illumination image enhancement method based on a multi-scale fusion U-Net network structure as claimed in claim 1, characterized in that: The residual upsampling module includes a main path and a residual path; The main path includes two convolutional layers, a normalization layer, and an activation function. The convolutional layer is used to extract local information of the input features, and the activation function is used to introduce nonlinear characteristics. The residual path includes a 1×1 convolution layer and a normalization layer to match the number of channels, which is used to superimpose the input features with the output of the main path; The outputs of the main path and the residual path are added, and the spatial resolution is enlarged through the deconvolution layer to achieve feature map upsampling.

5. The low-illumination image enhancement method based on a multi-scale fusion U-Net network structure according to claim 1, characterized in that: The jump connection module includes two convolutional layers, each of which is followed by a normalization layer and a ReLU activation function layer. The number of input and output channels of the convolutional layer is consistent, the convolution kernel size is 3×3, and the padding is 1.

6. The low-illumination image enhancement method based on a multi-scale fusion U-Net network structure as claimed in claim 1, characterized in that: In S3, after each image reconstruction, a joint loss function needs to be calculated to determine the quality of image reconstruction and optimize network parameters; The joint loss function is a weighted combination of reconstruction loss, perception loss, structural similarity loss, and color loss, which is: Among them, λ r , p , s and λ c are the weights of reconstruction loss, perceptual loss, structural similarity loss, and color loss, respectively. and L c They are reconstruction loss, perception loss, structural similarity loss and color loss respectively. for joint losses; Reconstruction loss The calculation formula is: Among them, E′ n is the enhanced result through network training, Y is the same as E′ n The corresponding expected normal lighting real value; Extract E′ through VGG network n The features of and Y will be defined based on the pre-trained VGG-16 network Perceived loss The calculation formula is: Among them, w ij 、h ij and c ij Describes the dimensions of each feature map in the VGG-16 network, φ ij Represents the feature map obtained by the jth convolutional layer of the i-th block in the VGG-16 network; Structural Similarity Loss The calculation formula is: Among them, SSIM is the image structure similarity, S and T represent the images to be compared, μ S and μ T is the average pixel value of the two images, and is the variance of the two images, σ ST is the covariance of the two images, c1 and c2 are two small constants used to prevent the denominator from being zero; Color Loss The calculation formula is: Among them, S p and T p denote the RGB color vector of pixel p in images S and T respectively. CA(·,·) denotes a three-dimensional vector with RGB color. The color loss between pixels is the angle between the two color vectors. The total color loss is the sum of the color losses of all pixels. The greater the color deviation between the two images, the greater the loss function.

7. The low-illumination image enhancement method based on a multi-scale fusion U-Net network structure as claimed in claim 6, characterized in that: In S4, the input of the low-light enhancement network model is a low-light image and a normal exposure image, and the output is a predicted reconstructed image. The training process further includes: S41, randomly dividing the images to be trained into several batches, each batch containing the same number of images; S42. Use the batches of images to train and optimize the low-light enhancement network model until the calculated joint loss function reaches a loss threshold or the number of iterations reaches a number threshold.

8. The low-illumination image enhancement method based on a multi-scale fusion U-Net network structure as claimed in claim 7, characterized in that: The S5 further comprises: S51, converting the trained low-illumination image enhancement network model into an ONNX format, and then converting the low-illumination image enhancement network model in the ONNX format into a format supported by a hardware device; S52, quantizing the low-illumination image enhancement network model in a format supported by the hardware device; S53, deploying the quantized low-illumination image enhancement network model on an edge device, and loading the quantized low-illumination image enhancement network model; S54, input the preprocessed image into the quantized low-light image enhancement network model, obtain an output result from the quantized low-light image enhancement network model, and save the output result in an appropriate format and range through OpenCV.

9. The low-illumination image enhancement method based on a multi-scale fusion U-Net network structure as claimed in claim 8, characterized in that: The quantization process in S52 further includes: S521, converting the floating point numbers in the low-light image enhancement network model into integer data to reduce the memory usage of the model and speed up the reasoning speed, and calculating the scaling factor and offset respectively: S522, floating point data X f Perform quantization operation and convert to uint8 type data X q , the calculation formula is: Among them, X max and X min Respectively represent the maximum and minimum values ​​of floating-point numbers, X f Represents a floating point number, X q Indicates the quantized uint8 type data, the round function indicates the rounding operation, and the clamp function is used to ensure the quantization result X q In the interval [0, 255], clamp is defined as: Among them, the clamp function limits the randomly changing values ​​to a given interval, a and b are both expressed as constants, and x is expressed as a variable.

10. The low-light image enhancement method based on a multi-scale fusion U-Net network structure according to claim 1, characterized in that: The pre-processing process in S6 further includes: S61, reading the low-light image to be enhanced and converting it into RGB format; S62, converting the image data into a NumPy array, and converting the data format into np.float32 to ensure the accuracy of subsequent processing; S63. Convert the image data into NHWC format to ensure the consistency between the data format accepted by the low-illumination image enhancement network model and the required data format.

Citation Information

Patent Citations

  • Low-illumination image enhancement method based on multi-scale stacked attention network

    CN114972107A

  • Low-illumination image enhancement method based on multilevel feature extraction fusion

    CN115393225A

  • Low-light image enhancement method based on U-shaped network and attention mechanism

    CN116596793A

  • A multi-scale cascaded hourglass depth map completion method guided by RGB images

    JP7610899B1

Cited By

  • Low-illumination image enhancement method, device, equipment and medium

    CN120852193A

  • Data preprocessing method and system for big data analysis

    CN121707840A

  • Image enhancement method based on stage quality perception

    CN121810510A

  • Method and device for removing rain imprint in image

    CN122335611A