Low light enhancement network structure based on improved PPM structure and training method thereof
By introducing improved PPM structure and Unet structure into the low-light enhancement network, the network structure is simplified and computing power needs are reduced, and the shortcomings of existing low-light enhancement technologies in contrast, detail recovery and color recovery are solved, efficient low-light enhancement effects are achieved and the deployment process is simplified.
Patent Information
- Application Number
- CN202311547608.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-20
- Publication Date
- 2025-05-20
AI Technical Summary
The existing low-light enhancement technology has shortcomings in contrast, detail recovery and color recovery, and the network structure is complex, computing power is high, and deployment is difficult.
A low-light enhanced network structure based on improved PPM structure is proposed, using a one-layer Unet structure and an improved pyramid pooling module (PPM) to simplify the network structure, reduce the computing power requirements, and generate 8 enhanced weight maps through the PixelShuffle upsampling method.
It realizes that while ensuring low-light enhancement effects, it reduces network computing, simplifies structure, and facilitates terminal deployment. It can complete low-light enhancement of 1080P images in real time, and improves the visual effect of dark-light scenes.
Smart Images

Figure CN120020819A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of low-light enhancement in deep learning neural network vision tasks, and relates to a low-light enhancement network structure based on an improved PPM structure and a training method thereof. Background Art
[0002] With the development of technology, the consumption demand for intelligent devices is increasing day by day. Introducing vision technology into intelligent devices has become an essential weapon for major market competitions. Vision technology is inseparable from the imaging quality of images, and the imaging quality in extreme scenarios is particularly crucial. The low-light scene image enhancement technology is one of the links. In low-light environments, the signal-to-noise ratio of images is low; short-exposure images have a lot of noise and significant detail loss, while long-exposure images are prone to overexposure, being relatively blurry and unrealistic. How to obtain high-quality low-light images or low-light enhanced images is a technical problem currently faced. When traditional image enhancement methods cannot effectively improve the quality of low-light images, scholars have proposed different deep learning-based methods for enhancing low-light images, opening up a new idea for this field. However, deep learning-based low-light image enhancement methods still face some problems, such as contrast enhancement, detail restoration, and color restoration.
[0003] Low-light image enhancement (LLIE) aims to improve the perceptual quality of images acquired in low-light environments. The recent progress in this field is mainly dominated by deep learning methods (including different learning strategies, network architectures, loss functions, training data, etc.). Learning strategies include unsupervised, self-supervised, etc.; network architectures are also innovative, including those based on CNN, GAN, and those combining CNN with Retinex theory; loss functions include L1, L2, SSIM, MS_SSIM, perceptual loss, smooth loss, etc.; many scholars have also released some low-light datasets, such as LOL, SCIE, SID, etc.
[0004] However, traditional image enhancement methods, such as global histogram equalization and gamma transformation, have poor enhancement effects on low-light images. And some existing deep learning methods may have their own problems in terms of contrast, color, detail restoration, model structure complexity, dataset, etc.; moreover, most of the existing deep learning-based low-light enhancement networks have problems such as complex structures, high computing power, and difficult deployment. Therefore, there is an urgent need to propose a low-light enhancement network structure with low computing power, simple structure, and easy deployment on terminal chips while ensuring the low-light enhancement effect. Summary of the Invention
[0005] To solve the above problems, the purpose of this application is to provide a deep learning neural network structure based on the RAW domain for the low-light enhancement field. This network structure introduces an improved pyramid pooling module, with a simple structure, convenient for terminal deployment and extremely low computing power. Taking 1080P image input as an example, the number of network parameters is 12,192, and the Flops is about 5.32GFlops. It can complete real-time low-light enhancement of 1080P images on the Junzheng T41 series chips, improving the visual effect of the picture in low-light scenes.
[0006] Specifically, the present invention provides a low-light enhancement network structure based on an improved PPM structure, and the network structure includes:
[0007] The network only uses one layer of Unet structure for network context information interaction. The overall network structure of Unet is U-shaped and can fuse shallow features and deep features;
[0008] The network input converts the Bayer format image into RGGB. Compared with directly inputting the Bayer format, the size of all network features is reduced by 1 / 4; the network output feature size is (W / 2, H / 2, 32). Through the PixelShuffle upsampling method, the network output is converted from (W / 2, H / 2, 32) to (W, H, 8), that is, the result is 8 enhanced weight maps;
[0009] Improve the traditional pyramid pooling module PPM structure, and achieve the reduction of network computational complexity through the reduction and reorganization of multiple convolutions and upsamplings of the original structure; the improved PPM structure only contains one layer of convolution, and the computing power for processing 1080P images is 6.3MFLOPs; and while reducing the computing power, it ensures the interaction and fusion of global and local information;
[0010] The network output is a curve graph iterated N times. The quadratic function of each iteration refers to the Zero-DCE network, specifically as shown in formula (0);
[0011] LE n (x) = LE n-1 (x) + A n LE n-1 (x)(1 - LE n-1 (x)) Formula (0)
[0012] Where, A n is the weight map of the nth channel after the network output is upsampled by PixelShuffle, where n ≤ 8; LE n is the final low-light enhancement prediction map.
[0013] The specific implementation of the network is shown in Table 1:
[0014] Table 1
[0015]
[0016] In the table, W and H respectively represent the width and height of the input Bayer format image.
[0017] The specific implementation of the PPM structure is shown in Table 2 as follows:
[0018] Table 2
[0019]
[0020]
[0021] In the table, h and w are a set of values smaller than H and W.
[0022] In Table 2, according to the sizes of the input H and W, h = H / 16 and w = W / 16 can be set. That is, when the network input is a 1080P image, h = 67 and w = 120.
[0023] This application also includes a training method for a low-light enhancement network structure based on an improved PPM structure. The method includes:
[0024] S1. Construct a low-light enhancement network with the improved PPM structure as described above;
[0025] S2. Network training. The loss function used for network training is shown in the following formula (1):
[0026] loss = w 1 *MAE_loss + w 2 *TV_loss Formula (1)
[0027] Among them, MAE_loss represents the mean absolute error loss function, as shown in formula (2); TV_loss represents the gradient smoothing loss function, as shown in formula (3); w 1 、w 2 are the weights corresponding to the loss functions respectively, and can be set to w 1 = 20, w 2 = 5;
[0028]
[0029] Among them, Ω represents the position domain of all pixel points of the image, (i, j) is a position at the u-th row and j-th column on the image, x is the prediction result map obtained by network training of the input image, y is the GroundTruth image corresponding to the input image; M represents the total number of pixel points of the image,
[0030]
[0031] Where N is the number of iterations, and are the horizontal and vertical gradient operators respectively;
[0032] S3. Use the paired low-light training set to train and iterate the above network. The initial learning rate is set to le-4, and the decay coefficient for each epoch is 0.95.
[0033] The method is set to N = 8.
[0034] Therefore, the advantages of this application are as follows:
[0035] 1. The network structure is simple, easier to be deployed on the chip side, with fast model inference speed, and can realize low-light enhancement in real time. The total number of convolutional layers is small and the convolutional channels are small (channel = 16 or 32 or 48), greatly reducing the network computation amount; the introduced improved version of the PPM structure can fuse global and local image information. Compared with the original PPM structure, it reduces the number of convolutions and the number of upsamplings, and can eliminate redundant computations to a great extent;
[0036] 2. The network output is a curve graph iterated N times. The final prediction graph is obtained by performing curve enhancement iteration on the input graph multiple times. The weight coefficient can be set for the curve graph to control the final brightness effect;
[0037] 3. Select RAW format data for image enhancement. RAW domain data contains more color gamuts and a higher dynamic range. Therefore, the deep model based on RAW data can reconstruct clearer details, high contrast, and has better color information. At the same time, convert the single-channel bayer format data into RGGB four-channel data, and convert the image resolution that the network needs to process from (H, W, 1) to (H / 2, W / 2, 4). The feature size of the entire network is reduced by 1 / 4, accelerating the network inference speed. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] The drawings described herein are used to provide a further understanding of the present invention, constitute a part of this application, and do not limit the present invention.
[0039] Figure 1 is a schematic diagram of a low-light enhancement network structure involving the introduced improved PPM structure.
[0040] Figure 2 is a schematic diagram of a traditional PPM structure.
[0041] Figure 3 is a schematic diagram of the improved PPM structure involved in this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0042] To better understand the technical content and advantages of the present invention, the present invention will be further described in detail with reference to the accompanying drawings.
[0043] This application proposes a low-light enhancement network based on an improved PPM structure and its training method, and the main content is introduced as follows:
[0044] The low-light enhancement network based on the improved PPM structure, the schematic diagram of its structure is as Figure 1 shown;
[0045] (1) The network only uses one layer of Unet structure for network context information interaction. The overall network structure of Unet is U-shaped and can fuse shallow features and deep features. The specific network implementation is shown in Table 1:
[0046] Table 1
[0047]
[0048]
[0049] In the table, W and H respectively represent the width and height of the input Bayer format image. For example, for a 1080P image, the corresponding H and W in the Bayer format are 1080 and 1920 respectively.
[0050] (2) The network input converts the Bayer format image to RGGB. Compared with directly inputting the Bayer format, the size of all network features is reduced by 1 / 4. The network output feature size is (W / 2, H / 2, 32). Through the PixelShuffle upsampling method, the network output is converted from (W / 2, H / 2, 32) to (W, H, 8), that is, the result is 8 enhanced weight maps;
[0051] (3) Improve the traditional pyramid pooling module PPM structure, and streamline the network computation by reducing and reorganizing multiple convolutions and upsamplings of the traditional PPM structure. The improved PPM structure only contains one layer of convolution, and the computing power for processing 1080P images is about 6.3 MFLOPs. Similar parameter configurations of the traditional PPM structure require nearly four convolutional layers, and the computing power is about 134 MFLOPs. And while reducing the computing power, the interaction and fusion of global and local information are ensured. The traditional PPM structure is as shown in the appendix Figure 2 shown, and the improved PPM structure is as shown in the appendix Figure 3As shown; the improved PPM structure ensures the interaction and fusion of global and local information while greatly reducing the computing power; the specific structure implementation is shown in Table 2 (in the table, w and h are a set of values smaller than W and H, and w = W / 16, h = H / 16 can be taken; that is, when the network input is a 1080P image, h = 67 and w = 120; mainly to unify the output sizes of each Pool layer for easy concat. Since taking larger values of this set will correspondingly increase the computational amount of convolution Conv5, it is not advisable to take too large values).
[0052] Table 2
[0053]
[0054]
[0055] (4) The network output is a curve for N consecutive iterations. The quadratic function for each iteration refers to the Zero-DCE network and is specifically shown in formula (0);
[0056] LE n (x) = LE n-1 (x) + A n LE n-1 (x)(1 - LE n-1 (x)) Formula (0)
[0057] Among them, A n is the weight map of the nth channel (n ≤ 8) after upsampling the network output by PixelShuffle; LE n is the final low-light enhancement prediction map.
[0058] This application also relates to a training method for a low-light enhancement network structure based on the improved PPM structure, and the method includes:
[0059] S1. Construct the low-light enhancement network with the improved PPM structure as described above;
[0060] S2. Network training. The loss function used for network training is shown in the following formula (1):
[0061] loss = w 1 *MAE_loss + w 2 *TV_loss Formula (1)
[0062] Among them, MAE_loss represents the mean absolute error loss function (shown in formula (2)); TV_los represents the gradient smoothing loss function (shown in formula (3)); w 1 、w 2 are the weights of the corresponding loss functions respectively, and can be set to w 1 = 20, w2 = 5;
[0063]
[0064] where Ω represents the position domain of all pixel points of the image, (i, j) is a position on the i-th row and j-th column of the image, x is the prediction result image obtained by training the input image through the network, and y is the GroundTruth image corresponding to the input image; M represents the total number of pixel points of the image,
[0065]
[0066] where N is the number of iterations (set to N = 8 in this method), and are the horizontal and vertical gradient operators respectively.
[0067] S3. Use the paired low-light training set to perform training iterations on the above network. The initial learning rate is set to le-4, and the decay coefficient for each epoch is 0.95;
[0068] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, various changes and modifications can be made to the embodiments of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A low-light enhancement network structure based on an improved PPM structure, characterized in that: The network structure includes: The network uses only one layer of Unet structure for network context information interaction. The overall network structure of Unet is U-shaped and can integrate shallow features and deep features. The network input converts the Bayer format image into RGGB; the network output feature size is (W / 2, H / 2, 32), and the network output is converted from (W / 2, H / 2, 32) to (W, H, 8) through the PixelShuffle upsampling method, that is, the result is 8 enhanced weight maps; Improve the traditional pyramid pooling module PPM structure, and reduce the amount of network calculation by reducing and reorganizing multiple convolutions and upsampling of the original structure; the improved PPM structure contains only one layer of convolution, and the computing power for processing 1080P images is 6.3MFLOPs; and while reducing the computing power, it ensures the interactive fusion of global and local information; The network output is a curve graph of N iterations, and the quadratic function of each iteration refers to the Zero-DCE network, as shown in formula (0); LE n (x) = LE n-1 (x) + A n LE n-1 (x)(1 - LE n-1 (x)) Equation (0) Among them, A n That is, the nth channel weight map of the network output after PixelShuffle upsampling, where n≤8; LE n The final low-light enhancement prediction map.
2. According to claim 1, a low-light enhancement network structure based on an improved PPM structure is characterized in that: The specific implementation of the network is shown in Table 1: Table 1 In the table, W and H represent the width and height of the input Bayer format image respectively.
3. According to claim 2, a low-light enhancement network structure based on an improved PPM structure is characterized in that: The specific implementation of the PPM structure is shown in Table 2: Table 2 In the table, h and w are a set of values smaller than H and W.
4. The low-light enhancement network structure based on the improved PPM structure according to claim 3 is characterized in that: In Table 2, according to the input H and W sizes, h=H / 16 and w=W / 16 can be set, that is, when the network input is a 1080P image, h=67 and w=120.
5. A training method for a low-light enhancement network structure based on an improved PPM structure, characterized in that: The method comprises: S1. Constructing a low-light enhancement network of the improved PPM structure of any one of claims 1-4 above; S2. Network training. The loss function used in network training is shown in the following formula (1): loss=w1*MAE_loss+w2*TV_loss formula (1) Wherein, MAE_loss represents the mean absolute error loss function, as shown in formula (2); TV_loss represents the gradient smoothing loss function, as shown in formula (3); w1 and w2 are the weights of the corresponding loss functions, which can be set to w1=20 and w2=5 respectively; Among them, Ω represents the location domain of all pixels in the image, (i, j) is a position in the i-th row and j-th column on the image, x is the predicted result image obtained by network training of the input image, and y is the GroundTruth image corresponding to the input image; M represents the total number of pixels in the image, Where N is the number of iterations, and are the horizontal and vertical gradient operators respectively; S3. Use the paired low-light training set to train the above network iteratively, with the initial learning rate set to le-4 and the decay coefficient of each epoch to 0.
95.
6. The training method of the low-light enhancement network structure based on the improved PPM structure according to claim 5, characterized in that: The method is set as N=8.