Lightweight low-illumination image enhancement system and method
The lightweight low-light image enhancement system, constructed by a feature extraction module, a brightness and noise adjustment submodule, and a channel attention submodule, solves the problems of excessive model complexity and computational resource requirements on embedded devices, and achieves efficient image enhancement and real-time processing performance.
Patent Information
- Application Number
- CN202511004428.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-21
- Publication Date
- 2025-10-28
AI Technical Summary
Existing low-light image enhancement technologies face a contradiction between model complexity and excessive computational resource requirements on embedded devices, making it difficult to achieve lightweight models without sacrificing image enhancement effects.
A lightweight low-light image enhancement system is constructed using a feature extraction module, a brightness and noise adjustment submodule, and a channel attention submodule. Through deep convolution and channel fusion techniques, it accurately captures detailed features and suppresses noise, dynamically adjusts the weights of key features, and outputs high-quality enhanced images.
While reducing computational complexity and resource consumption, it significantly improves the enhancement effect and real-time processing performance of low-light images, making it suitable for embedded devices with limited resources.
Smart Images

Figure CN120852199A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular to a lightweight low-light image enhancement system and method. Background Technology
[0002] In the field of low-light image enhancement, the rapid development of deep learning technology has spurred a large number of algorithms to improve image quality. However, existing methods generally face the challenge of excessive model complexity and computational resource requirements, which fundamentally contradicts the lightweight and efficient requirements of real-time processing in embedded devices.
[0003] This contradiction has driven current research towards designing compact and efficient network architectures, such as employing techniques like depthwise separable convolutions, lightweight attention mechanisms, and model pruning to reduce computational burden. In some practical studies, zero-shot learning is achieved through a series of curve parameters in Zero-Reference Deep Curve Estimation (Zero-DCE), significantly reducing computational requirements; the multi-branch network in the Multi-Branch Low-Light Enhancement network (MBLLEN) focuses on extracting and fusing features at different scales, simplifying the model and improving processing speed; and the kernel-induced illumination decomposition and adjustment network in the KinD++ network architecture provides a lightweight and effective solution suitable for resource-constrained applications.
[0004] However, the aforementioned technologies do not take into account the practical application scenarios of embedded devices. Therefore, based on the current image enhancement techniques, how to achieve lightweight models without sacrificing image enhancement effects, and enable these models to run efficiently on resource-constrained embedded devices, remains a research hotspot. Summary of the Invention
[0005] To address the aforementioned technical problems, this application provides a lightweight low-light image enhancement system and method, which enables the deployment of a lightweight low-light image enhancement system in embedded devices, thereby improving image enhancement effects while reducing the complexity of the system during image enhancement. The technical solution is as follows:
[0006] Firstly, a lightweight low-light image enhancement system is provided for use in embedded systems, including:
[0007] The feature extraction module is used to extract detailed features from low-light images to obtain the first image;
[0008] The illumination channel attention module includes a luminance noise adjustment submodule and a channel attention submodule; wherein, the luminance noise adjustment submodule is used to suppress noise in the low-illuminance image and the first image, and fuse the noise-suppressed low-illuminance image and the first image to obtain a second image; the channel attention submodule is used to extract key features in the second image and adjust the weights of the key features to obtain a third image;
[0009] The image output module, connected to the channel attention submodule, is used to output the target enhancement image after standardizing the third image.
[0010] In one possible implementation, the feature extraction module includes a first convolutional layer, a first depthwise separable convolutional layer, and a second depthwise separable convolutional layer connected in series.
[0011] The first convolutional layer is used to extract detailed features from low-light images;
[0012] Both the first depthwise separable convolutional layer and the second depthwise separable convolutional layer are used to filter out some detailed features to obtain the first image.
[0013] In one possible implementation, the output of the first convolutional layer is further provided with a first activation function layer, which is used to perform a nonlinear transformation on the image output by the first convolutional layer.
[0014] In one possible implementation, the brightness noise adjustment submodule is configured as follows:
[0015] The low-light image and the first image are converted into grayscale images respectively to obtain their respective first grayscale images and second grayscale images;
[0016] The first grayscale image is divided to obtain multiple first grayscale sub-images;
[0017] The second grayscale image is divided to obtain multiple second grayscale sub-images;
[0018] The mean brightness and noise variance of each first grayscale sub-image and each second grayscale sub-image are calculated respectively. The mean brightness is used to reflect the brightness of each grayscale sub-image, and the noise variance is used to reflect the noise intensity of each grayscale sub-image.
[0019] Adjust the brightness of each grayscale sub-image based on the average brightness value of each grayscale sub-image;
[0020] The noise mask for each grayscale sub-image is determined based on the noise variance of each grayscale sub-image.
[0021] Based on the noise mask of each grayscale sub-image, noise suppression is performed on each grayscale sub-image after brightness adjustment;
[0022] The first grayscale sub-image after noise suppression and the second grayscale sub-image after noise suppression are fused to obtain the second image.
[0023] In one possible implementation, the channel attention submodule includes an average pooling layer, a second convolutional layer, and a weight processing layer connected in series.
[0024] The average pooling layer is used to convert the spatial information of the second image into channel information to obtain the key features of each channel.
[0025] The second convolutional layer is used to adjust the weights of each channel, and the kernel size of the second convolutional layer is calculated based on the number of input channels in the second image.
[0026] The weighting layer is used to multiply the key features of each channel with their corresponding weights and then add them together to obtain the third image.
[0027] In one possible implementation, the feature extraction module, the illumination channel attention module, and the image output module constitute an image enhancement model;
[0028] Training the image enhancement model includes:
[0029] Obtain a dataset comprising pairs of low-light images and normal-light images;
[0030] The image enhancement model is constructed using the feature extraction module, the illumination channel attention module, and the image output module.
[0031] The image enhancement model is trained using the dataset.
[0032] In one possible implementation, the process of training the image enhancement model using the dataset further includes:
[0033] For each training iteration, a loss function is used to calculate the loss value of the target enhanced image output by the image enhancement model;
[0034] When the loss value reaches a preset condition, the training of the image enhancement model is considered complete.
[0035] In one possible implementation, the loss function includes at least a pixel loss function, a gradient loss function, a structural similarity loss function, a color consistency loss function, and an exposure loss function;
[0036] The pixel loss function is used to calculate the pixel loss value of the target enhanced image;
[0037] The gradient loss function is used to calculate the gradient of the target enhanced image in the horizontal and vertical directions, and to calculate the mean of the gradient in each direction;
[0038] The structural similarity loss function is used to calculate the structural loss value of the target enhanced image;
[0039] The color consistency loss function is used to calculate the color loss value of the target enhanced image;
[0040] The exposure loss function is used to calculate the exposure loss value of the target enhancement image.
[0041] In one possible implementation, the preset condition is: the total loss value tends to stabilize and the individual loss values reach a minimum value;
[0042] The total loss value is determined by the pixel loss value, the gradient loss value, the structure loss value, the color loss value, and the exposure loss value.
[0043] Secondly, a lightweight low-light image enhancement method is provided, applied to any of the lightweight low-light image enhancement systems described above, comprising:
[0044] Acquire low-light images;
[0045] Extract the detailed features of the low-light image to obtain the first image;
[0046] Noise suppression is applied to the low-light image and the first image, and the noise-suppressed low-light image and the first image are then fused to obtain a second image;
[0047] Extract key features from the second image and adjust the weights of the key features to obtain the third image;
[0048] After standardizing the third image, the target enhanced image is output.
[0049] The technical solutions provided in this application can achieve the following technical effects:
[0050] (1) First, the system of this application consists only of a feature extraction module, an illumination channel attention module and an image output module. The system architecture is simple, which reduces the computational complexity and resource consumption when enhancing low-illuminance images and makes it easy to deploy in embedded devices.
[0051] (2) Secondly, the various modules in the system of this application cooperate with each other to achieve the purpose of enhancing low-light images. Specifically, this includes: firstly, a feature extraction module is constructed based on depthwise convolution, which performs spatial convolution on each input channel to accurately capture the detailed features within a single channel, and then combines channel fusion operations to retain the ability to express detailed feature information. Furthermore, a brightness noise adjustment submodule is used to suppress noise in the image, and then a channel attention submodule extracts the key features of each channel, multiplies the key features of each channel with their corresponding weights, and then adds them together to strengthen important key features. Finally, the image output module first restores the spatial resolution of the image after a series of processing steps, and then maps it to the target channel to obtain the target enhanced image. It can be seen that the system of this application can efficiently focus on important information such as detailed features and key features in low-light images, enhance the ability to express information about key features, improve adaptability and real-time processing performance on embedded devices, and thus optimize the image enhancement effect. Attached Figure Description
[0052] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below. In the drawings:
[0053] Figure 1 This is a structural diagram of a lightweight low-light image enhancement system provided in an embodiment of this application;
[0054] Figure 2 This is a structural diagram of the feature extraction module in the system embodiment of this application;
[0055] Figure 3 This is a structural diagram of the illuminance channel attention module in the system embodiment of this application;
[0056] Figure 4 This is a structural diagram of the channel attention submodule in the system embodiment of this application;
[0057] Figure 5 This is a flowchart of a lightweight low-light image enhancement method provided in an embodiment of this application;
[0058] Figure 6 This is a structural diagram of an electronic device provided in an embodiment of this application. Detailed Implementation
[0059] Exemplary embodiments of the present application will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this application will be thorough and complete, and will fully convey the scope of the present application to those skilled in the art.
[0060] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such use can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the term "comprising" and its variations should be interpreted as open-ended terms meaning "including but not limited to."
[0061] This application provides a lightweight low-light image enhancement system deployed in an embedded system. The embedded system consists of multiple embedded devices, each with limited available resources. Therefore, it is necessary to develop a lightweight image enhancement system with simplified processing logic.
[0062] like Figure 1 As shown, the lightweight low-light image enhancement system mainly includes a feature extraction module, an illumination channel attention module, and an image output module. The feature extraction module is mainly used to extract detailed features of the image, the illumination channel attention module is mainly used to remove interference information in the image and adjust the weights of key features, and the image output module standardizes the images processed by the feature extraction module and the illumination channel attention module before outputting them to ensure that the output images meet the user's requirements. Moreover, the entire system architecture for image processing is simple and easy to deploy in embedded devices.
[0063] In this embodiment, the system deployed in the embedded device is mainly used for enhancing low-light images. Low-light images refer to images taken in environments with insufficient light (such as at night, in low indoor light, or on cloudy days) that have low overall brightness, blurred details, significant noise, and limited dynamic range.
[0064] The following sections will provide a detailed description of each module in the system.
[0065] like Figure 2 As shown, the feature extraction module consists of, from top to bottom, a first convolutional layer, a first activation function layer, a first depthwise separable convolutional layer, and a second depthwise separable convolutional layer. Figure 2The layers are denoted as Conv1, Relu1, DSC1, and DSC2, respectively. The first convolutional layer, Conv1, uses a 3×3 convolutional kernel and has 32 output channels. It also incorporates reflection padding to reduce boundary effects in low-light images. The first activation function layer, Relu1, uses a rectified linear unit (ReLU), also known as the rectifier function, which is primarily used for nonlinear transformations. The first depthwise separable convolutional layer, DSC1, is identical to the second depthwise separable convolutional layer, DSC2, both consisting of depthwise convolution and pointwise convolution. This expansion increases the number of channels in the image output from the first convolutional layer, Conv1, from 32 to 64, and then to 128. The number of channels is represented by C in the attached diagram.
[0066] In practical applications, the low-light image is first input into the first convolutional layer Conv1, which extracts the detailed features of the low-light image. These detailed features typically encompass multi-layered visual and structural information, representing details lost due to insufficient lighting, such as brightness, texture, color, and edge details. After extracting the detailed features, the first convolutional layer Conv1 outputs 32 feature maps, meaning one feature map corresponds to each channel. Each feature map identifies the detailed features. Then, the first activation function layer ReLU1 performs a non-linear transformation on each feature map to enhance its expressive power of detailed features. Next, the first depthwise separable convolutional layer DSC1 maps the 32 feature maps to 64 feature maps, and the second depthwise separable convolutional layer DSC2 maps these 64 feature maps to 128 feature maps. Finally, a 1×1 convolution is used to combine the feature maps from different channels to output the first image. The 1×1 convolution is located at the output of the second depth separable convolutional layer DSC2, and can also be considered as a component of the second depth separable convolutional layer DSC2.
[0067] Therefore, it can be seen that the feature extraction module is based on depthwise convolution. By performing spatial convolution on each input channel separately, it accurately captures the detailed features within a single channel. Subsequently, combined with channel fusion operations, it significantly reduces the redundant computation of traditional convolution while preserving the ability to express detailed features, thus forming an efficient and compact feature extraction architecture.
[0068] like Figure 3 As shown, the illumination channel attention module includes a luminance noise adjustment submodule and a channel attention submodule. The luminance noise adjustment submodule is connected to the feature extraction module. In addition to receiving the first image output from the feature extraction module, the luminance noise adjustment submodule also receives a low-light image. In this embodiment, the composition of the luminance noise adjustment submodule is not limited, with the aim of suppressing noise in both the low-light image and the first image.
[0069] In this embodiment, the process of suppressing noise by the brightness noise adjustment submodule is shown in steps S101 to S107.
[0070] Step S101: Convert the low-light image and the first image into grayscale images respectively to obtain the corresponding first grayscale image and second grayscale image.
[0071] Step S102: Divide the first grayscale image into multiple first grayscale sub-images, and divide the second grayscale image into multiple second grayscale sub-images. In this embodiment, a 3×3 sliding window is used to divide the first grayscale image and the second grayscale image respectively.
[0072] Step S103: Calculate the mean brightness and noise variance of each first grayscale sub-image and each second grayscale sub-image respectively. The mean brightness is used to reflect the brightness of each grayscale sub-image, and the noise variance is used to reflect the noise intensity of each grayscale sub-image.
[0073] To illustrate the process of calculating the average brightness, let's take the calculation of the average brightness of one of the first grayscale sub-images as an example:
[0074] 1. Gray-level histogram calculation: Calculate the frequency of each gray-level value in the first gray-level sub-image and generate a normalized histogram. Where g∈[0, 255];
[0075] 2. Determine the segmentation threshold t using the Otsu thresholding method: Calculate the inter-class variance σ. 2 (t)=ω0(t)(μ 0 (t)-μ T ) 2 +ω1(t)(μ1(t)-μ T ) 2 Iterate through all possible thresholds t and select the one that makes σ 2 The largest t is used as the segmentation threshold, where ω0 and ω1 are the probabilities of the two classes, μ0 and μ1 are the means of the two classes, and μ T The global mean;
[0076] 3. Region division and quantization: When g>t, it is a high grayscale region; when t / 2<g≤t, it is a medium grayscale region; when g≤t / 2, it is a low grayscale region.
[0077] As can be seen, by calculating the average brightness as described above, we can determine the high grayscale area, medium grayscale area, and low grayscale area present in each first grayscale sub-image or each second grayscale sub-image.
[0078] To illustrate the process of calculating noise variance, let's take the calculation of the noise variance of one of the first grayscale sub-images as an example. The formula for calculating noise variance is:
[0079] in, x is the mean of the noise data for all first grayscale sub-images. i Let n represent the noise data of the i-th first grayscale sub-image, where n represents the number of first grayscale sub-images.
[0080] As can be seen, by calculating the noise variance as described above, the noise intensity of each first grayscale sub-image or each second grayscale sub-image can be determined.
[0081] Step S104: Adjust the brightness of each grayscale sub-image based on the average brightness of each grayscale sub-image. For each first grayscale sub-image or each second grayscale sub-image, based on the determined high grayscale area, medium grayscale area, and low grayscale area, apply an attenuation factor to the high grayscale area and a dynamic gain factor to the low grayscale area. Then, achieve brightness balance of each grayscale sub-image by pixel-by-pixel weighting, that is, adjust the high grayscale area and low grayscale area of each grayscale sub-image to the same brightness as the medium grayscale area.
[0082] Step S105: Determine the noise mask for each grayscale sub-image based on the noise variance of each grayscale sub-image. In one possible implementation, the noise variance is substituted into the following formula to obtain the noise mask for each grayscale sub-image: Where T is a preset value.
[0083] Step S106: Based on the noise mask of each grayscale sub-image, noise suppression is performed on each grayscale sub-image after brightness adjustment. For each first grayscale sub-image or each second grayscale sub-image after brightness adjustment, a 5x5 Gaussian blur is applied to regions with noise mask values greater than the first mask value to smooth the noise, and a 3x3 mean filter is applied to regions with noise mask values less than the second mask value to preserve details, thereby achieving noise suppression. In this embodiment, the first mask value is greater than the second mask value.
[0084] Step S107: Fuse the noise-suppressed first grayscale sub-image and the noise-suppressed second grayscale sub-image to obtain the second image. In this embodiment, a 1x1 convolutional layer is used to fine-tune the channel dimensions of the denoised first and second grayscale sub-images, aligning the number of channels in the two images before fusing them to obtain the second image. That is, the second image incorporates the key information of the original low-light image while retaining the detailed features of the first image, providing reliable image data for the subsequent channel attention submodule.
[0085] like Figure 4 As shown, the channel attention submodule consists of, from top to bottom, an average pooling layer, a dimensionality reduction layer, a second convolutional layer, a dimensionality increase layer, a second activation function layer, and a weight processing layer. Figure 4The layers are denoted by mean-pool, Squeeze, conv2, Unsqueeze, sigmoid, and Weighted Sum, respectively. The mean-pool layer primarily compresses the spatial information of the image output from the brightness noise adjustment submodule into a single channel description; that is, the feature information of each channel is aggregated into a scalar. The Squeeze layer reduces the dimensionality of the image output from the mean-pool layer to meet the requirements of the second convolutional layer. The second convolutional layer, covn2, uses 1×1 convolutions to establish relationships between channels and adjust their weights. The Unsqueeze layer restores the image output from covn2 to the same dimension as the second image. The sigmoid activation function normalizes the channel weights, ensuring they fall within the range of 0-1. The Weighted Sum layer multiplies the key features and corresponding weights of each channel and then sums the products to obtain the third image.
[0086] In this embodiment, the size of the second image is [B,C,H,W], where B is the batch size, C is the number of channels, and H and W are the height and width, respectively. The second image is first input into the mean-poo average pooling layer. The mean-poo average pooling layer compresses the spatial information of the second image into a single channel description. This channel description contains channel information, also known as key features. These key features are used to capture the global contextual information of the second image. Based on the obtained key features, the mean-poo average pooling layer outputs an image of size [B,C,1,1]. Before the image of size [B,C,1,1] enters the second convolutional layer covn2, a dimensionality reduction layer Squeeze is used to reduce the dimensionality of the image of size [B,C,1,1] to the image of size [B,C,1]. Simultaneously, the kernel size of the second convolutional layer covn2 is dynamically calculated using the following formula: Here, C refers to the number of channels. Next, the dynamically adjusted second convolutional layer, covn2, adjusts the weights of each channel based on the importance of their key features to enhance the information representation capability of these important features. The image output from covn2 is then passed through an upscaling layer, Unsqueeze, which restores the image to its original dimension, resulting in an image of size [B, C, 1, 1]. After the second activation function layer, sigmoid, normalizes the weights of each channel, the weighted sum layer multiplies the key features of each channel with their corresponding weights, and then sums the products to obtain the third image. This process strengthens the important key features in each channel, thereby improving the system's focus on key features in low-light images.
[0087] Therefore, it can be seen that the illumination channel attention module integrates the brightness noise adjustment submodule and the channel attention submodule. Through a series of steps set between the two, including grayscale conversion, grayscale image segmentation, adjusting the brightness in the grayscale image, noise suppression, global average pooling, dynamic convolution kernel generation, 1D convolution operation, Sigmoid activation, and weighted summation processing of key features, the system's ability to enhance key features in low-light images, noise suppression effect, and detail capture efficiency are significantly improved.
[0088] like Figure 1 As shown, the image output module includes a first upsampling layer, a second upsampling layer, a third convolutional layer, and a third function layer. The first and second upsampling layers have identical structures, each consisting of a 3×3 convolutional layer, a normalization layer, and a ReLU activation function layer connected in series (not shown in the figure). The first and second upsampling layers are connected to restore the spatial resolution of the third image. The third convolutional layer is a 3×3 convolutional layer used to map the spatially resolved third image back to the target number of channels. This target number of channels can be preset by the user or default to the same number of channels as the low-light image. The third function layer uses the Tanh activation function, which limits the output value to the range [-1, 1], i.e., limits the output image to the target number of channels; this image is also called the target enhanced image.
[0089] In this embodiment, the third image has 128 channels. After the third image is input into the first upsampling layer and the second upsampling layer, the first upsampling layer reduces the number of channels of the third image from 128 to 64, and then the second upsampling layer reduces it from 64 to 32, so that the third image is restored to the number of channels required by the user, or restored to the same size as the low-light image. Finally, the third convolutional layer maps the image with the target number of channels from 32 channels back to the output number of 3 channels, corresponding to the three channels of the RGB image, and generates the final target enhanced image through the Tanh activation function.
[0090] In summary, this application provides a lightweight low-light image enhancement system that integrates a lightweight feature extraction structure and an illumination channel attention module. The feature extraction module, built upon depthwise convolution, performs spatial convolution on each input channel individually, accurately capturing detailed features within a single channel. Combined with channel fusion operations, it significantly reduces redundant computations in traditional convolution while preserving the expressive power of detailed features, forming an efficient and compact feature extraction architecture. Further, after extracting detailed features, a first image is obtained. The first image and the low-light image are then input into the illumination channel attention module. First, a brightness noise adjustment submodule suppresses noise in both the first and low-light images, and the noise-suppressed images are fused to obtain a second image. Next, the channel attention submodule dynamically adjusts the convolution kernel size based on the number of channels to generate weights, extracting key features for each channel. The key features of each channel are multiplied by their corresponding weights and then summed to enhance important key features, resulting in a third image. Finally, the third image enters the image output module, which first restores the spatial resolution of the third image and then maps it to the target channel to obtain the target enhanced image. It is evident that while significantly reducing the number of parameters and computational load, the system can efficiently focus on important information such as detailed features and key features in low-light images, enhance the ability to express information about key features, improve adaptability and real-time processing performance on embedded devices, and thus optimize image enhancement effects.
[0091] It should be noted that the feature extraction module, illumination channel attention module, and image output module mentioned above are actually used as a whole image enhancement model. Therefore, before the model is put into use, it needs to be trained. The training process may include the following steps S201 to S209:
[0092] Step S201: Obtain the dataset, which includes pairs of low-light images and normal-light images;
[0093] Step S202: Build an image enhancement model using the feature extraction module, the illumination channel attention module, and the image output module;
[0094] Step S203: Set the initial number of training rounds i = 0, and set the maximum number of training rounds I;
[0095] Step S204: Use the dataset to train the image enhancement model for the i-th round to obtain the target enhanced image;
[0096] Step S205: Calculate the loss value of the target enhancement image using a loss function;
[0097] Step S206: Determine whether the loss value meets the preset condition. If yes, proceed to step S209; otherwise, proceed to step S207.
[0098] Step S207: Determine whether the current training round number i is greater than or equal to the maximum training round number I. If yes, proceed to step S209; otherwise, proceed to step S208.
[0099] Step S208: Increment the training round number i by 1, and return to step S204 to continue training;
[0100] Step S209: Confirm that the training of the image enhancement model is complete.
[0101] First, the dataset can be multiple images taken in advance in both low-light and standard-light environments, or it can be obtained from an existing low-light image database, such as a large number of images obtained from the LOL dataset. In this embodiment, the dataset is divided into a training subset and a validation subset, for example, setting 485 pairs as the training subset and 15 pairs as the validation subset.
[0102] In step S205, the loss functions involved include at least the pixel loss function, gradient loss function, structural similarity loss function, color consistency loss function, and exposure loss function. Specifically, the calculation formulas for each loss function are shown below.
[0103] The formula for calculating the pixel loss function is: Among them, I i and These represent the pixel values of the normal light image and the target enhancement image, respectively, where N is the total number of pixels.
[0104] The formula for calculating the gradient loss function is: Among them, I i,j Let represent the pixel value at position i,j. The gradients in the horizontal and vertical directions are calculated using this formula, and the average of the gradients in each direction is taken.
[0105] The formula for calculating the structural similarity loss function is: SSIM stands for Structural Similarity Index, which is obtained by calculating the similarity of brightness, contrast, and structure within a local window.
[0106] The formula for calculating the color consistency loss function is: Among them, M R M G M B represents the mean values of the R, G, and B channels of the target enhanced image, respectively, and ò is the stable term.
[0107] The formula for calculating the exposure loss function is: Among them, P i The local patch mean of the target image is enhanced, where T is the target exposure value and N is the total number of local patches.
[0108] Based on the loss functions mentioned above, after obtaining the target enhancement image, it is necessary to calculate the loss value of the target enhancement image using each of the aforementioned loss functions.
[0109] Furthermore, when calculating the loss values of the target augmented image in each dimension using various loss functions, it is also necessary to call the total loss function to calculate the total loss value of the target augmented image. The formula for calculating the total loss function is as follows:
[0110] L total =λ1L L1 +λ2L SSIM +λ3L smooth +λ4L color +λ5L exposure ,
[0111] Wherein, λ1, λ2, λ3, λ4, and λ5 are the preset weights corresponding to pixels, gradients, structural similarity, color consistency, and exposure dimensions, respectively, and are weight values set in advance.
[0112] In step S206, the preset condition is that the total loss value tends to stabilize and the individual loss values reach their minimum values. For example, when the number of training rounds reaches a certain number, the total loss value fluctuates within a certain range, while the loss values of each dimension are basically the minimum values in the past training rounds. At this time, it is said that the loss value has reached the preset condition.
[0113] In the above model training process, this embodiment adopts a gradient adaptive weight optimization strategy. In the early stage of training, weights are assigned to each loss value based on experience. After training for a period of time, the weights can be dynamically adjusted according to the gradient information of each loss value, so that the loss terms with higher optimization difficulty receive more attention and the optimization accuracy is improved.
[0114] In addition, in this embodiment, when the total loss value tends to stabilize, the individual loss values reach their minimum values, or the number of training rounds reaches the maximum number of training rounds, the training is terminated and the optimal model parameters are saved, resulting in an image enhancement model with strong overall performance.
[0115] It should also be noted that, based on the trained image enhancement model, after it is deployed in an embedded device, it executes a lightweight low-light image enhancement method. For example... Figure 5 As shown, the lightweight low-light image enhancement method may include the following steps S301 to S305:
[0116] Step S301: Acquire a low-light image;
[0117] Step S302: Extract detailed features from the low-light image using the feature extraction module to obtain the first image;
[0118] Step S303: The brightness noise adjustment submodule performs noise suppression on the low-light image and the first image, and then merges the noise-suppressed low-light image and the first image to obtain the second image;
[0119] Step S304: Extract key features from the second image using the channel attention submodule and adjust the weights of the key features to obtain the third image;
[0120] Step S305: The third image is normalized by the image output module and then the enhanced target image is output.
[0121] The lightweight low-light image enhancement method provided in this embodiment is executed by the lightweight low-light image enhancement system described in the above embodiment. Its implementation method and principle are the same. For details on the implementation of each module, please refer to the relevant descriptions in the above system embodiments, which will not be repeated here.
[0122] To facilitate the explanation of the enhancement effect achieved by the system in this embodiment using the above method, this embodiment uses 485 pairs of images from the dataset for training and 15 pairs of images for validation. Structural similarity (SSIM) and peak signal-to-noise ratio (PSNR) are used as objective evaluation metrics to measure the enhancement results. A higher SSIM value indicates better consistency between the enhancement result and human visual features; PSNR measures the ratio between effective information and noise in an image to reflect whether the image is distorted, and a higher PSNR value indicates better image quality.
[0123] In verifying the enhancement effect, this embodiment utilizes the deep learning framework PyTorch, specifically Python 3.9 and CUDA 11.6, and is implemented on the Ubuntu 20.04 operating system. The batch size is set to 16, the maximum number of training epochs is 200, the optimizer is Adam, the initial learning rate is set to 0.0001, and the input image size is set to 400 pixels in height and 600 pixels in width. Based on the same training data, image enhancement operations are performed using existing algorithms (such as RetinexNet, MBLLEN, Uformer, EnligtenGAN) and the method of this embodiment. The method of this embodiment is denoted by Ours, resulting in the following...
[0124] The results of the objective evaluation indicators shown in Table 1:
[0125]
[0126] Table 1
[0127] As can be seen from the four objective evaluation indicators in Table 1, the method of this embodiment has a better effect on enhancing low-light images than other algorithms, and the method of this embodiment has fewer parameters and less computation, making it suitable for embedded devices.
[0128] Furthermore, this embodiment also compares the average FPS of RetinexNet, MBLLEN, Uformer, EnligtenGAN and the method of this application on the validation set. The hardware used for inference testing is an AMD 5700G CPU, and the comparison results are shown in Table 2.
[0129]
[0130] Table 2
[0131] As can be seen from Table 2, the method in this embodiment outperforms other algorithms in terms of inference speed on the LOL dataset, significantly reducing inference time and making it suitable for scenarios requiring fast response.
[0132] It should be noted that the sequence numbers of the steps in the above embodiments do not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. In practical applications, all the above possible implementation methods can be arbitrarily combined in a combined manner to form possible embodiments of this application, which will not be described in detail here.
[0133] Based on the same inventive concept, embodiments of this application also provide an electronic device, including a processor and a memory, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute a lightweight low-light image enhancement method of any of the above embodiments.
[0134] In an exemplary embodiment, an electronic device is provided, such as Figure 6 As shown, Figure 6 The illustrated electronic device 600 includes a processor 601 and a memory 603. The processor 601 and the memory 603 are connected, for example, via a bus 602. Optionally, the electronic device 600 may also include a transceiver 604. It should be noted that in practical applications, the transceiver 604 is not limited to one type, and the structure of this electronic device 600 does not constitute a limitation on the embodiments of this application.
[0135] Processor 601 may be a CPU (Central Processing Unit), DSP (Digital Signal Processor), ASIC (Application Specific Integrated Circuit), FPGA (Field Programmable Gate Array), or other programmable logic device, transistor logic device, hardware component, or any combination thereof. It can implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 601 may also be a combination that implements computational functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc.
[0136] Bus 602 may include a pathway for transmitting information between the aforementioned components. Bus 602 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. Bus 602 can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 6 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.
[0137] The memory 603 may be a ROM (Read Only Memory) or other type of static storage device capable of storing static information and instructions, RAM (Random Access Memory) or other type of dynamic storage device capable of storing information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), CD-ROM (Compact Disc Read Only Memory) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto.
[0138] The memory 603 stores computer program code that executes the scheme of this application, and its execution is controlled by the processor 601. The processor 601 executes the computer program code stored in the memory 603 to implement the content shown in the foregoing method embodiments.
[0139] Among them, electronic devices include, but are not limited to: mobile terminals such as mobile phones, laptops, digital radio receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), and in-vehicle terminals (such as in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 6 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0140] Based on the same inventive concept, this application also provides a storage medium storing a computer program, wherein the computer program is configured to execute a lightweight low-light image enhancement method of any of the above embodiments at runtime.
[0141] Those skilled in the art will clearly understand that the specific working process of the systems, devices, and modules described above can be referred to the corresponding process in the foregoing method embodiments. For the sake of brevity, it will not be repeated here.
[0142] Those skilled in the art will understand that the technical solution of this application, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several program instructions to cause an electronic device (e.g., a personal computer, server, or network device) to execute all or part of the steps of the methods described in the embodiments of this application when running the program instructions. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, portable hard drive, read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
[0143] Alternatively, all or part of the steps of the foregoing method embodiments can be implemented by hardware (such as electronic devices like personal computers, servers, or network devices) associated with program instructions. The program instructions can be stored in a computer-readable storage medium. When the program instructions are executed by the processor of the electronic device, the electronic device executes all or part of the steps of the methods described in the embodiments of this application.
[0144] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that within the spirit and principles of this application, modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein; and these modifications or substitutions do not cause the corresponding technical solutions to leave the protection scope of this application.
Claims
1. A lightweight low-light image enhancement system, applied in an embedded system, characterized in that, include: The feature extraction module is used to extract detailed features from low-light images to obtain the first image; The illumination channel attention module includes a luminance noise adjustment submodule and a channel attention submodule; wherein, the luminance noise adjustment submodule is used to suppress noise in the low-illuminance image and the first image, and fuse the noise-suppressed low-illuminance image and the first image to obtain a second image; the channel attention submodule is used to extract key features in the second image and adjust the weights of the key features to obtain a third image; The image output module, connected to the channel attention submodule, is used to output the target enhancement image after standardizing the third image.
2. The system according to claim 1, characterized in that, The feature extraction module includes a first convolutional layer, a first depthwise separable convolutional layer, and a second depthwise separable convolutional layer connected in series. The first convolutional layer is used to extract detailed features from low-light images; Both the first depthwise separable convolutional layer and the second depthwise separable convolutional layer are used to filter out some detailed features to obtain the first image.
3. The system according to claim 2, characterized in that, The output of the first convolutional layer is further provided with a first activation function layer, which is used to perform nonlinear transformation on the image output by the first convolutional layer.
4. The system according to claim 1, characterized in that, The brightness noise adjustment submodule is configured as follows: The low-light image and the first image are converted into grayscale images respectively to obtain their respective first grayscale images and second grayscale images; The first grayscale image is divided to obtain multiple first grayscale sub-images; The second grayscale image is divided to obtain multiple second grayscale sub-images; The mean brightness and noise variance of each first grayscale sub-image and each second grayscale sub-image are calculated respectively. The mean brightness is used to reflect the brightness of each grayscale sub-image, and the noise variance is used to reflect the noise intensity of each grayscale sub-image. Adjust the brightness of each grayscale sub-image based on the average brightness value of each grayscale sub-image; The noise mask for each grayscale sub-image is determined based on the noise variance of each grayscale sub-image. Based on the noise mask of each grayscale sub-image, noise suppression is performed on each grayscale sub-image after brightness adjustment; The first grayscale sub-image after noise suppression and the second grayscale sub-image after noise suppression are fused to obtain the second image.
5. The system according to claim 1, characterized in that, The channel attention submodule includes an average pooling layer, a second convolutional layer, and a weight processing layer connected in series. The average pooling layer is used to convert the spatial information of the second image into channel information to obtain the key features of each channel. The second convolutional layer is used to adjust the weights of each channel, and the kernel size of the second convolutional layer is calculated based on the number of input channels in the second image. The weighting layer is used to multiply the key features of each channel with their corresponding weights and then add them together to obtain the third image.
6. The system according to claim 1, characterized in that, The feature extraction module, the illumination channel attention module, and the image output module together form an image enhancement model; Training the image enhancement model includes: Obtain a dataset comprising pairs of low-light images and normal-light images; The image enhancement model is constructed using the feature extraction module, the illumination channel attention module, and the image output module. The image enhancement model is trained using the dataset.
7. The system according to claim 6, characterized in that, The process of training the image enhancement model using the dataset also includes: For each training iteration, a loss function is used to calculate the loss value of the target enhanced image output by the image enhancement model; When the loss value reaches a preset condition, the training of the image enhancement model is considered complete.
8. The system according to claim 7, characterized in that, The loss function includes at least the pixel loss function, gradient loss function, structural similarity loss function, color consistency loss function, and exposure loss function; The pixel loss function is used to calculate the pixel loss value of the target enhanced image; The gradient loss function is used to calculate the gradient of the target enhanced image in the horizontal and vertical directions, and to calculate the mean of the gradient in each direction; The structural similarity loss function is used to calculate the structural loss value of the target enhanced image; The color consistency loss function is used to calculate the color loss value of the target enhanced image; The exposure loss function is used to calculate the exposure loss value of the target enhancement image.
9. The system according to claim 8, characterized in that, The preset conditions are: the total loss value tends to stabilize and the individual loss values reach their minimum values; The total loss value is determined by the pixel loss value, the gradient loss value, the structure loss value, the color loss value, and the exposure loss value.
10. A lightweight low-light image enhancement method, applied to the system as described in any one of claims 1-9, characterized in that, include: Acquire low-light images; Extract the detailed features of the low-light image to obtain the first image; Noise suppression is applied to the low-light image and the first image, and the noise-suppressed low-light image and the first image are then fused to obtain a second image; Extract key features from the second image and adjust the weights of the key features to obtain the third image; After standardizing the third image, the target enhanced image is output.
Citation Information
Cited By
Real-time image enhancement and intelligent exposure method and system for unmanned aerial vehicle inspection
CN121883783A