Real-time low-light image enhancement method, system, device and storage medium
By designing a low-light image enhancement network based on a lightweight pyramid depth model, the problem that the prior art is difficult to achieve real-time light enhancement on embedded platforms is solved, and efficient and real-time low-light image enhancement effect is achieved.
Patent Information
- Application Number
- CN202110829368.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-07-22
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2041-07-22
AI Technical Summary
Existing low-light image enhancement methods are difficult to achieve real-time lighting enhancement on embedded platforms with limited computing and storage resources, and the huge parameter space leads to large memory usage.
A real-time low-light image enhancement network based on lightweight pyramid depth model is designed, using a three-layer pyramid structure and depth separation dense convolution blocks, feature extraction and enhancement through sliding windows and maximum pooling downsampling, and network training is performed using L2 loss, perceived loss and SSIM loss.
Real-time low-light image enhancement on embedded platforms, with good visual effects and color fidelity, and short running time on resource-constrained platforms, such as the time taken only 0.335s on Nvidia Jetson Xavier NX.
Smart Images

Figure CN113538312B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a real-time low-light image enhancement method, system, device and storage medium based on a lightweight pyramid depth model, and relates to the technical field of image processing. Background Art
[0002] In real life, complex lighting conditions, such as night, rain, and haze, will affect image quality. Traditional enhancement methods are mainly divided into histogram equalization and Retinex theory. Histogram equalization is to stretch an image with concentrated grayscale to make its grayscale evenly distributed, thereby achieving the effect of enhancing global contrast. Kokufuta et al. solved the problem of local brightness enhancement by grayscale mapping the central pixel of the sliding window, but LHE requires a high computational cost and will cause over-enhancement of some parts of the image. The basis of Retinex theory is that the color of an object is determined by the object's ability to reflect light. The color of an object is not affected by the non-uniformity of illumination and is consistent. Jobson et al. proposed a Retinex algorithm based on a multiple iteration strategy. The pixel value of a single point depends on the result of a specific path surrounding it, and it approaches the ideal value after multiple iterations. The retinal theory can be used to adjust the brightness of the image without affecting other attributes of the image, so that other attribute information of the image can be avoided from being increased or decreased during the image enhancement operation. However, since the Gaussian filter does not have the characteristic of edge preservation, halo phenomenon is prone to occur in areas with higher brightness, affecting the quality of the image.
[0003] In recent years, learning-based enhancement methods have achieved remarkable success. Wei et al. combined Retinex theory with CNN network and proposed RetinexNet. Inspired by RetinexNet, the KIND algorithm uses the reflectance map of high-exposure images as label images, learns mapping through deep convolutional neural networks, and adds more illumination to bright areas. Inspired by bilateral grid processing and local affine color transformation, Gharbi et al. designed HDRNet, which trains convolutional neural networks to predict the coefficients of local affine models in bilateral space. Lv et al. proposed an end-to-end multi-branch enhancement network (MBLLEN). MBLLEN extracts effective feature representations through feature extraction modules, enhancement modules, and fusion modules, thereby improving image enhancement performance. Although the above methods have good restoration effects, due to the huge parameter space, they inevitably lead to large memory usage, making it difficult to run on embedded platforms with limited computing and storage resources, and the inference time is long, making it impossible to achieve real-time illumination enhancement. Summary of the invention
[0004] In view of the above problems, the object of the present invention is to provide a real-time low-light image enhancement method, system, device and medium based on a lightweight pyramid depth model that can achieve low-light enhancement on an embedded platform.
[0005] To achieve the above object, the present invention adopts the following technical solutions:
[0006] In a first aspect, the present invention provides a real-time low-light image enhancement method, comprising:
[0007] Get a low-light benchmark dataset;
[0008] Design a low-light enhancement network based on a lightweight pyramid deep model;
[0009] The low-light enhancement network is trained using a low-light benchmark dataset to obtain a low-light enhancement network model;
[0010] A low-light enhancement network model is used to process low-light images to achieve image illumination enhancement.
[0011] The real-time low-light image enhancement method further comprises: the low-light benchmark data set adopts a low-light / normal-light paired data set.
[0012] The real-time low-light image enhancement method further adopts a three-layer pyramid structure, the first layer includes a plurality of multi-scale depth-separable dense convolution blocks, and the second and third layers each include a plurality of depth-separable dense convolution blocks.
[0013] The real-time low-light image enhancement method further includes training a low-light enhancement network using a low-light benchmark data set to obtain a low-light enhancement network model, comprising:
[0014] The input low-light image is cropped using a non-overlapping sliding window of a set size, and the cropped image is downsampled twice using the maximum pooling method to obtain two downsampled images of different resolutions, which are used as the input of the second and third layers respectively.
[0015] The image is trained from the third level. After several depth-separable dense convolution blocks, the trained feature map is upsampled to the second level and concatenated with the input feature map of the second level.
[0016] After the input feature map of the second layer is concatenated with the output feature map of the third layer, after passing through several depth-separable dense convolution blocks, the trained feature map is upsampled to the first layer and concatenated with the input feature map of the first layer;
[0017] After the input feature map of the first layer is concatenated with the output feature map of the second layer, the final enhanced image is output after passing through several multi-scale depth-separable dense convolution blocks.
[0018] The real-time low-light image enhancement method further comprises the second layer and the third layer respectively including four depth-separable dense convolution blocks, and the depth-separable dense convolution blocks include a series module of two groups of depth-separable convolution blocks, InstanceNorm normalization function and LeakyReLU activation function; two adjacent depth-separable dense convolution blocks are connected by a jump connection.
[0019] The real-time low-light image enhancement method further comprises a first layer including four multi-scale depth-separable dense convolution blocks, wherein the multi-scale dense depth-separable convolution block is composed of a 3×3 depth-separable dense convolution block and a 5×5 depth-separable dense convolution block in parallel, and its output is the concat connection of the outputs of the two depth-separable dense convolution blocks.
[0020] The real-time low-light image enhancement method further requires that the feature distance between the enhanced image and the normal exposure image be constrained by a loss function during the training process, and the network parameters be iteratively corrected, including:
[0021] When the image is trained at the third level, the L2 loss function is used to reconstruct the illumination information of the image between the enhanced image and the normal exposure image;
[0022] When the image is trained on the second level and the first level, perceptual loss and SSIM loss are used to reconstruct the detail information.
[0023] In a second aspect, the present invention further provides a real-time low-light image enhancement system, the system comprising:
[0024] A data set acquisition unit, configured to acquire a low-light benchmark data set;
[0025] A low-light enhancement network design unit, configured to design a low-light enhancement network based on a lightweight pyramid structure;
[0026] A low-light enhancement network model training unit is configured to train the low-light enhancement network through a low-light benchmark data set to obtain a low-light enhancement network model;
[0027] The illumination enhancement unit is configured to process the low-light image using a low-light enhancement network model to achieve illumination enhancement of the image.
[0028] In a third aspect, the present invention further provides a processing device, which includes at least a processor and a memory, wherein a computer program is stored in the memory, and wherein the processor executes the computer program to implement the real-time low-light image enhancement method.
[0029] In a fourth aspect, the present invention further provides a computer storage medium having computer-readable instructions stored thereon, wherein the computer-readable instructions can be executed by a processor to implement the real-time low-light image enhancement method.
[0030] The present invention adopts the above technical solution, which has the following advantages:
[0031] 1. The present invention realizes image illumination enhancement with good visual effects and color fidelity by designing a low-light enhancement network with a three-layer pyramid structure;
[0032] 2. Based on the depthwise separable convolution, the present invention designs the depthwise separable dense convolution block and the multi-scale depthwise separable dense convolution block, which not only reduces the number of model parameters and the amount of calculation, but also improves the feature extraction performance, thereby improving the low-light enhancement effect;
[0033] 3. The present invention can achieve real-time low-light enhancement on multiple platforms such as embedded platforms and GPU computing platforms. Through verification and evaluation on the LOL dataset and the SCIE dataset, it only takes 35ms to enhance a low-light image on a 1080Ti and only 0.335s on an Nvidia Jetson Xavier NX;
[0034] In summary, the present invention can be widely used in image processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Various other advantages and benefits will become apparent to those of ordinary skill in the art by reading the detailed description of the preferred embodiments below. The accompanying drawings are only for the purpose of illustrating the preferred embodiments and are not to be considered as limiting the present invention. Throughout the accompanying drawings, the same reference numerals are used to represent the same components. In the accompanying drawings:
[0036] Figure 1 is a method flow chart of an embodiment of the present invention;
[0037] Figure 2 It is a block diagram of the low-light enhancement network structure of an embodiment of the present invention, wherein MDSCDB is a multi-scale depth-separable dense convolution block, and DSCDB is a depth-separable dense convolution block.
[0038] Figure 3is a structural diagram of a depth-separable dense convolution block according to an embodiment of the present invention, wherein MDSCDB is a multi-scale depth-separable dense convolution block, and DSCDB is a depth-separable dense convolution block;
[0039] Figure 4 2 is a comparison diagram of the low light enhancement algorithm of the embodiment of the present invention, FIG. (a) is a comparison diagram of the illumination enhancement results of the LOL image, and FIG. (b) is a comparison diagram of the illumination enhancement results of the SCIE image. DETAILED DESCRIPTION
[0040] The exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments described herein. On the contrary, these embodiments are provided in order to enable a more thorough understanding of the present invention and to fully convey the scope of the present invention to those skilled in the art.
[0041] It should be understood that the terms used herein are only for the purpose of describing specific example embodiments and are not intended to be limiting. Unless the context clearly indicates otherwise, the singular forms "one", "an" and "said" as used herein may also be meant to include plural forms. The terms "include", "comprise", "contain", and "have" are inclusive, and therefore specify the existence of stated features, steps, operations, elements and / or parts, but do not exclude the existence or addition of one or more other features, steps, operations, elements, parts, and / or combinations thereof. The method steps, processes, and operations described herein are not interpreted as necessarily requiring them to be performed in the specific order described or illustrated, unless the execution order is clearly indicated. It should also be understood that additional or alternative steps may be used.
[0042] For ease of description, spatially relative terms may be used herein to describe the relationship of one element or feature relative to another element or feature as shown in the figures, such as "inside", "outside", "inner side", "outer side", "below", "above", etc. Such spatially relative terms are intended to include different orientations of the device in use or operation in addition to the orientation depicted in the figures.
[0043] Embodiment 1
[0044] like Figure 1 As shown, the real-time low-light image enhancement method based on the lightweight pyramid depth model provided in this embodiment includes the following contents:
[0045] S1. Use the LOL dataset as the training dataset.
[0046] Specifically, the LOL dataset is a low-light enhanced benchmark dataset, which contains 500 pairs of normal / low-light images captured by real scenes, with an image size of 600×400.
[0047] S2. Designing a low-light enhanced network
[0048] The low-light enhancement network structure designed in this embodiment is as follows Figure 2 As shown in the figure. The network is a three-layer pyramid structure. The first layer consists of four multi-scale depth-separable dense convolution blocks MDSCDB, and the second and third layers are composed of four depth-separable dense convolution blocks DSCDB respectively. In addition, adjacent depth-separable dense convolution blocks DSCDB are connected by jump connections to prevent gradient disappearance. At the end of each layer, the visualization enhancement result is output through 3×3 vanilla convolution and tanh activation function. It should be noted that the final low-light enhancement result is the output of the first layer.
[0049] S3, such as Figure 2 , Figure 3 As shown in FIG. 1 , the low-light enhancement network is trained using a low-light benchmark dataset to obtain a low-light enhancement network model. The specific process includes:
[0050] S31. Use a non-overlapping sliding window of a set size, such as 96×96, to crop the input low-light image, and perform two maximum pooling downsamplings on the cropped image to downsample the image to 1 / 2 and 1 / 4 of the original resolution. This is an example and is not limited to this.
[0051] S32. After the input image is downsampled, it is distributed on three resolution levels, namely, the original resolution, 1 / 2 of the original resolution, and 1 / 4 of the original resolution.
[0052] The image is trained from the third level. After passing through four 3×3 depthwise separable dense convolution blocks DSCDB, the trained feature map is upsampled to the second level and concatenated with the input feature map of the second level.
[0053] In some implementations, the depthwise separable dense convolutional block DSCDB is Figure 3 As shown in Figure 1, it includes a series module of two sets of depth-wise separable convolution blocks, InstanceNorm normalization function and LeakyReLU activation function. Two adjacent depth-wise separable dense convolution blocks are connected by skip connections.
[0054] S33. After the input feature map of the second layer is concatenated with the output feature map of the third layer, it passes through four 3×3 depth-separable dense convolution blocks DSCDB, and then the trained feature map is upsampled to the first layer and concatenated with the input feature map of the first layer.
[0055] S34, after the input feature map of the first layer is concatenated with the output feature map of the second layer, it passes through four multi-scale depth-separable dense convolution blocks MDSCDB, and then passes through a 3×3 convolution and tanh activation function to output the final enhanced image. Among them, the multi-scale depth-separable dense convolution block MDSCDB is as follows Figure 3 As shown in Figure 1, it consists of a 3×3 depth-separable dense convolution block and a 5×5 depth-separable dense convolution block in parallel, and its output is the concat connection of the outputs of the two depth-separable dense convolution blocks.
[0056] Furthermore, in order to restore the illumination information of the image, the feature distance between the enhanced image and the normally exposed image needs to be constrained by multiple loss functions during the training process, and the network parameters need to be iteratively corrected, including:
[0057] When the image is trained at the third level, since the input image resolution is only 24×24, the feature map mainly contains illumination information, and there is little color and contour information. Therefore, the L2 loss function is used as the reconstruction loss for the global brightness between the enhanced image and the normal exposure image to restore the illumination information of the image. The reconstruction loss can be expressed as:
[0058]
[0059] Among them, H, W and C represent the height, width and number of channels of the training image respectively, P and T represent the enhanced image and the normally exposed image respectively, i, j represents the i-th row and j-th column, and c represents the number of channels of the feature map.
[0060] When the image is trained at the second level and the first level, due to the resolution of 48×48 and 96×96, the color and contour information in the image is more significant, which helps to reconstruct the details of the image. Therefore, between the enhanced image and the normal exposure image, the perceptual loss and SSIM loss are used to restore the details of the image, such as color and contour. The perceptual loss can be expressed as:
[0061]
[0062] Among them, H, W and C represent the height, width and number of channels of the feature map respectively. represents the feature map extracted after the mth convolutional layer in the VGG network, P and T represent the enhanced image and the normally exposed image respectively, i,j, represents the i-th row and j-th column, and c represents the number of channels of the feature map.
[0063] The SSIM loss can be expressed as:
[0064]
[0065] Among them, μ and σ represent the mean and variance respectively, C1 and C2 are constant values in SSIM loss, and their sizes can be set to 0.01 and 0.03 respectively.
[0066] S4. Use the low-light enhancement network model to process the low-light image to be processed to achieve image light enhancement. The specific process is as follows:
[0067] The input image is cropped to a size of 96×96 by a non-overlapping sliding window, and then convolution and pooling operations are performed twice, respectively, so that the input image is distributed in three levels of different resolutions.
[0068] The third layer inputs a small-size image after two downsamplings, and reconstructs the features through four depth-wise separable dense convolution blocks DSCDB. The reconstructed result is upsampled and returned to the second layer.
[0069] Similarly, after receiving the reconstructed features from the third level, the second level concatenates them with the downsampled features, reconstructs the features through four depth-wise separable dense convolutional blocks DSCDB, and then upsamples the reconstructed features from the second level and returns them to the first level.
[0070] The first level concatenates the reconstructed features returned by the second level with the original resolution features, reconstructs the features through four multi-scale depth-wise separable dense convolution blocks MDSCDB, and outputs the final enhanced image after 3×3 vanilla convolution and tanh activation function.
[0071] Some preferred embodiments of the present invention further include the step of transplanting the trained low-light enhancement model to a resource-constrained platform.
[0072] The experimental process of the real-time low-light image enhancement method based on the lightweight pyramid depth model of the present invention is described in detail below through specific embodiments.
[0073] 1. Experimental platform configuration
[0074] This embodiment is implemented under the pytorch framework. The operating platform is based on Intel i7-6850K CPU, Nvidia GTX1080Ti and Ubuntu 16.04 operating system for training and testing.
[0075] In the low-light enhancement network, this embodiment trains the above-mentioned low-light enhancement network model on the LOL data set. In order to prove the performance of the present invention on a resource-constrained platform, this embodiment is further tested on the embedded platform Nvidia-Jetson XavierNX. The Nvidia Jetson Xavier NX platform is equipped with a 6-core Nvidia Carmel ARM v8.2 CPU and a 384-core GPU, supports JetPack SDK, and supports mainstream artificial intelligence frameworks and algorithms, such as TensorFlow, PyTorch, Caffe / Caffe2, Keras, MXNet, etc.
[0076] 2. Training lightweight low-light enhancement network model
[0077] 2.1 Datasets and Evaluation Metrics
[0078] In the experiment, LOL was used as the training dataset. LOL contains 500 pairs of normal / low-light images captured by real scenes, with an image size of 600×400. The images were cropped into 96×96 image blocks by sliding windows, with a total of 25,260 images as the training dataset and 780 images as the validation and test sets. In addition, the SCIE dataset was used as the test set to evaluate the performance of the algorithm. The SCIE dataset includes 589 multiple exposure image sequences of indoor and outdoor scenes, each sequence has 3-18 low-contrast images with different exposure levels and corresponding high-quality high-contrast reference images, with a size between 3000×2000-6000×4000. The 10 images with the lowest exposure were selected as test images and resized to 600×400.
[0079] This embodiment uses PSNR (dB) and SSIM as image quality evaluation indicators of the low-light enhancement algorithm. Generally speaking, the higher the values of these two indicators, the better the image enhancement effect. In addition, Param, FLOPS and running time are used as lightweight indicators of the low-light enhancement method. The lower the values of these three indicators, the smaller the network model and the higher the model operation efficiency.
[0080] 2.2 Training parameter settings
[0081] The network of this embodiment is implemented in the Pytorch framework and trained using the ADAM optimizer on a 1080Ti GPU. The batch size is set to 16 for level 0 and 48 for level 1-2 according to the network level, and the epoch is set to 160 for level 0 and 40 for level 1-2 according to the network level. In addition, the learning rate is 10e-5, and each feature map channel is set to 32.
[0082] 2.3 Results Analysis
[0083] The enhancement network model of the present invention is compared with various enhancement algorithms, and all the results are evaluated by PSNR (dB), SSIM, Param (M), FLOPs (B) and Running time (s). Six representative algorithms, including Dong, RetinexNet, HDRNet, KinD, EnlightenGAN and MBLLEN, are used for illumination enhancement tests on LOL and SCIE datasets.
[0084] The experimental results of various low-light enhancement algorithms are shown in Table 1. Under each evaluation metric, the best result is indicated in italics, the second best result is indicated in bold, and “-” indicates that the result is unavailable.
[0085] In terms of low-light image enhancement, the PSNR and SSIM of the invention are slightly higher than EnlightenGAN and better than other algorithms. In terms of model lightweighting, the enhanced model of the invention has the lowest number of parameters and computational complexity. The running time of 35ms on GPU1080Ti is very close to 31ms of HDRNet, and the deduction time is lower than other algorithms. The running time deployed on the edge computing device Nvidia Xavier NX is only 0.335s, which is much better than all the comparison algorithms.
[0086] Figure 4 (a) and (b) show the illumination enhancement effect of LOL and SCIE dataset images respectively. The enhanced visual effect of the present invention is comparable to that of EnlightenGAN and MBLLEN algorithms. It is superior to RetinexNet, HDRNet, KinD and other algorithms in terms of brightness, clarity, noise and color reconstruction.
[0087] Experimental results show that the image enhancement method of the present invention achieves a balance between efficient operation and enhancement effect. It can not only obtain a good lighting enhancement effect, but also complete efficient enhancement processing on edge computing devices.
[0088] Table 1 Comparison of low light enhancement algorithms
[0089]
[0090] Embodiment 2
[0091] The above-mentioned embodiment 1 provides a real-time low-light image enhancement method based on a lightweight pyramid depth model. Correspondingly, this embodiment provides a real-time low-light image enhancement system. The real-time low-light image enhancement system provided in this embodiment can implement the real-time low-light image enhancement method of embodiment 1, and the system can be implemented by software, hardware, or a combination of software and hardware. For example, the system may include integrated or separate functional modules or functional units to execute the corresponding steps in each method of embodiment one. Since the real-time low-light image enhancement system of this embodiment is basically similar to the method embodiment, the process described in this embodiment is relatively simple, and the relevant parts can refer to the partial description of embodiment one. The real-time low-light image enhancement system of this embodiment is merely schematic.
[0092] This embodiment provides a real-time low-light image enhancement system, the system comprising:
[0093] A data set acquisition unit, configured to acquire a low-light benchmark data set;
[0094] A low-light enhancement network design unit, configured to design a low-light enhancement network based on a lightweight pyramid structure;
[0095] A low-light enhancement network model training unit is configured to train the low-light enhancement network through a low-light benchmark data set to obtain a low-light enhancement network model;
[0096] The illumination enhancement unit is configured to process the low-light image using a low-light enhancement network model to achieve illumination enhancement of the image.
[0097] Embodiment 3
[0098] This embodiment provides a processing device for implementing the real-time low-light image enhancement method based on a lightweight pyramid depth model provided in the first embodiment. The processing device may be a processing device for a client, such as a mobile phone, a laptop computer, a tablet computer, a desktop computer, etc., to execute the real-time low-light image enhancement method based on a lightweight pyramid depth model of the first embodiment.
[0099] The processing device includes a processor, a memory, a communication interface and a bus, and the processor, the memory and the communication interface are connected through the bus to complete mutual communication. The memory stores a computer program that can be run on the processor, and the processor executes the real-time low-light image enhancement method based on the lightweight pyramid depth model provided in the first embodiment when running the computer program.
[0100] Preferably, the memory may be a high-speed random access memory (RAM), and may also include a non-volatile memory, such as at least one disk memory.
[0101] Preferably, the processor may be a central processing unit (CPU), a digital signal processor (DSP) or other general-purpose processors of various types, which are not limited here.
[0102] Embodiment 4
[0103] The real-time low-light image enhancement method based on a lightweight pyramid depth model of the first embodiment is specifically implemented as a computer program product, which may include a computer-readable storage medium carrying computer-readable program instructions for executing the real-time low-light image enhancement method based on a lightweight pyramid depth model described in the first embodiment.
[0104] Computer readable storage media can be tangible devices that hold and store instructions used by instruction execution devices. Computer readable storage media can be, for example, but not limited to, electronic storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any combination thereof.
[0105] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein, and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A real-time low-light image enhancement method, characterized in that include: Get a low-light benchmark dataset; Design a low-light enhancement network based on a lightweight pyramid structure; The low-light enhancement network is trained using the low-light benchmark dataset to obtain a low-light enhancement network model, including: The input low-light image is cropped using a non-overlapping sliding window of a set size, and the cropped image is downsampled twice using the maximum pooling method to obtain two downsampled images of different resolutions, which are used as the input of the second and third layers respectively. The image is trained from the third level. After several depth-separable dense convolution blocks, the trained feature map is upsampled to the second level and concatenated with the input feature map of the second level. After the input feature map of the second layer is concatenated with the output feature map of the third layer, after passing through several depth-separable dense convolution blocks, the trained feature map is upsampled to the first layer and concatenated with the input feature map of the first layer; After the input feature map of the first layer is concatenated with the output feature map of the second layer, it passes through several multi-scale depth-separable dense convolution blocks to output the final enhanced image; The second layer and the third layer respectively include four depth-separable dense convolution blocks, each of which includes a series connection module of two groups of depth-separable convolution blocks, InstanceNorm normalization function and LeakyReLU activation function; two adjacent depth-separable dense convolution blocks are connected by a jump connection; The first layer includes four multi-scale depth-separable dense convolution blocks, each of which is composed of a 3×3 depth-separable dense convolution block and a 5×5 depth-separable dense convolution block in parallel, and its output is the concat connection of the outputs of the two depth-separable dense convolution blocks; During the training process, the feature distance between the enhanced image and the normal exposure image needs to be constrained by the loss function, and the network parameters need to be iteratively corrected, including: When the image is trained at the third level, the L2 loss function is used to reconstruct the illumination information of the image between the enhanced image and the normal exposure image; When the image is trained at the second level and the first level, perceptual loss and SSIM loss are used to reconstruct detail information; A low-light enhancement network model is used to process low-light images to achieve image illumination enhancement.
2. The real-time low-light image enhancement method according to claim 1, characterized in that: The low-light benchmark dataset uses a low-light / normal-light paired dataset.
3. The real-time low-light image enhancement method according to claim 1, characterized in that: The low-light enhancement network adopts a three-layer pyramid structure, the first layer includes a plurality of multi-scale depth-separable dense convolution blocks, and the second and third layers each include a plurality of depth-separable dense convolution blocks.
4. A real-time low-light image enhancement system, characterized in that: The system includes: A data set acquisition unit, configured to acquire a low-light benchmark data set; A low-light enhancement network design unit, configured to design a low-light enhancement network based on a lightweight pyramid structure; The low-light enhancement network model training unit is configured to train the low-light enhancement network through a low-light benchmark data set to obtain a low-light enhancement network model, including: The input low-light image is cropped using a non-overlapping sliding window of a set size, and the cropped image is downsampled twice using the maximum pooling method to obtain two downsampled images of different resolutions, which are used as the input of the second and third layers respectively. The image is trained from the third level. After several depth-separable dense convolution blocks, the trained feature map is upsampled to the second level and concatenated with the input feature map of the second level. After the input feature map of the second layer is concatenated with the output feature map of the third layer, after passing through several depth-separable dense convolution blocks, the trained feature map is upsampled to the first layer and concatenated with the input feature map of the first layer; After the input feature map of the first layer is concatenated with the output feature map of the second layer, it passes through several multi-scale depth-separable dense convolution blocks to output the final enhanced image; The second layer and the third layer respectively include four depth-separable dense convolution blocks, each of which includes a series connection module of two groups of depth-separable convolution blocks, InstanceNorm normalization function and LeakyReLU activation function; two adjacent depth-separable dense convolution blocks are connected by a jump connection; The first layer includes four multi-scale depth-separable dense convolution blocks, each of which is composed of a 3×3 depth-separable dense convolution block and a 5×5 depth-separable dense convolution block in parallel, and its output is the concat connection of the outputs of the two depth-separable dense convolution blocks; During the training process, the feature distance between the enhanced image and the normal exposure image needs to be constrained by the loss function, and the network parameters need to be iteratively corrected, including: When the image is trained at the third level, the L2 loss function is used to reconstruct the illumination information of the image between the enhanced image and the normal exposure image; When the image is trained at the second level and the first level, perceptual loss and SSIM loss are used to reconstruct detail information; The illumination enhancement unit is configured to process the low-light image using a low-light enhancement network model to achieve illumination enhancement of the image.
5. A processing device, comprising at least a processor and a memory, wherein a computer program is stored in the memory, wherein: When the processor runs the computer program, the computer program is executed to implement the real-time low-light image enhancement method according to any one of claims 1 to 3.
6. A computer storage medium, characterized in that: Computer-readable instructions are stored thereon, and the computer-readable instructions can be executed by a processor to implement the real-time low-light image enhancement method according to any one of claims 1 to 3.
Citation Information
Patent Citations
quick low-illumination target detection method based on convolutional neural network
CN113052210A