Real-time low-illumination image enhancement method based on multi-lookup table collaborative network
Through the method based on the multi-lookup table collaborative network, the problem that low-illumination image enhancement in the prior art cannot process local and global information is solved, and high-quality real-time low-illumination image enhancement is achieved, with good robustness and color recovery effect.
Patent Information
- Application Number
- CN202510075399.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-06-10
AI Technical Summary
The existing low-illumination image enhancement method based on lookup table network cannot process local pixel information and global information, resulting in inaccurate dark noise amplification and color enhancement.
A real-time low-illumination image enhancement method based on multi-looking table collaborative network is proposed. By constructing one-dimensional, three-dimensional and four-dimensional multi-looking table collaborative enhancement network and global enhancement modules, the network receptive field is expanded and the image global information is fully utilized for adaptive enhancement.
It achieves good detail retention and brightness recovery, and has higher image color accuracy, stronger robustness, and real-time enhancement, suitable for edge device deployment.
Smart Images

Figure CN120125448A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and particularly relates to a real-time low-light image enhancement method based on a multi-lookup table collaborative network. Background Art
[0002] Low-light images can be defined as images taken under insufficient lighting conditions, which have serious degradation problems such as low visibility, low contrast, and noise, do not conform to the subjective visual effects of people, and will also affect the performance of downstream visual tasks, such as object detection and tracking, image and video segmentation, face recognition, autonomous driving, etc. The goal of low-light image enhancement is to improve image visibility and contrast, and at the same time restore various inherent distortions in the dark environment. It is an important and challenging task in the field of computer vision, has a very wide application prospect, and is one of the current research hotspots.
[0003] The image enhancement technology based on the lookup table network has developed rapidly in recent years, showing excellent enhancement performance. Compared with the enhancement methods based on large deep neural networks, it has obvious advantages in terms of inference speed and computational overhead. However, the currently proposed image enhancement methods based on the lookup table network have poor application effects in low-light scenarios. One is that the current methods only perform one-to-one mapping of single pixels based on a three-dimensional lookup table, unable to process local inter-pixel information, which will amplify the dark noise in low-light images. The other is that most of the current methods learn adaptive weights based on lightweight convolutional neural networks and improve adaptability by combining different lookup tables. Although this method can improve efficiency, it cannot fully utilize the global information of the image, limiting the enhancement performance of the lookup table network. Summary of the Invention
[0004] To solve the problems that existing methods based on lookup table networks cannot process local pixel information and cannot fully utilize the global information of images in the low-light image enhancement task, the present invention proposes a real-time low-light image enhancement method based on a multi-lookup table collaborative network for real-time image enhancement in low-light scenarios. Obtain a training set containing paired low-light and normal-light images; construct an enhancement network based on the collaboration of simulated one-dimensional, three-dimensional, and four-dimensional lookup tables, and a global enhancement module based on a visual state space model, where the lookup table network processes the information of multiple pixel points on a single channel simultaneously to expand the network receptive field, and the global enhancement module is used to extract the global information of the image, obtain global correction parameters for adaptive enhancement; construct a loss function, perform low-light image enhancement on the training set through the constructed network, and calculate the loss value to adjust the parameters of the network model; traverse the trained simulated network with multi-lookup table collaboration, and store the results in the corresponding lookup tables in the form of key-value pairs; perform fine-tuning training on the entire network on the training set to obtain a trained low-light image enhancement network; input the low-light image to be processed into the trained enhancement network, and an enhanced normal-light image can be obtained; the present invention achieves good detail retention and brightness restoration in the low-light image enhancement task, and can achieve real-time enhancement based on the lookup table structure.
[0005] The technical solution of the present invention is implemented as follows:
[0006] A real-time low-light image enhancement method based on a multi-lookup table collaborative network, the steps are as follows:
[0007] Step 1: Obtain a training set containing paired low-light and normal-light images;
[0008] Step 2: Construct a simulated network with multi-lookup table collaboration;
[0009] Step 3: Construct a global enhancement module;
[0010] Step 4: Construct a loss function, perform low-light image enhancement on the training set through the constructed network, and calculate the loss value to adjust the parameters of the network model;
[0011] Step 5: Traverse the trained simulated network with multi-lookup table collaboration, and store the results in the corresponding lookup tables in the form of key-value pairs;
[0012] Step 6: Perform fine-tuning training on the entire network on the training set to obtain a trained low-light image enhancement network;
[0013] Step 7: Input the low-light image to be processed into the trained low-light image enhancement network, and an enhanced normal-light image can be obtained.
[0014] Preferably, the specific steps for obtaining the training set including paired low-illumination and normal-illumination images in step one are as follows:
[0015] S11: Obtain paired low-illumination images T l and normal-illumination images T n , and divide them into a training set, a validation set, and a test set according to a ratio;
[0016] S12: Perform data preprocessing on the paired images in the training set, crop them into image patches of uniform size, and randomly perform data augmentation operations such as rotation and flipping;
[0017] Preferably, the specific steps for constructing the simulation network with multiple look-up tables in cooperation in step two are as follows:
[0018] S21: Construct a simulation network of the look-up table to simulate the operation process of the look-up table in the network. The processing process of the look-up table is divided into three stages: one-dimensional look-up table, four-dimensional look-up table, and three-dimensional look-up table;
[0019] S22: In the first stage, use a one-dimensional look-up table for non-interactive indexing. On each RGB channel of the image I, use a scanning area of 3×3 composed of nine one-dimensional look-up tables to scan the entire channel, and average the results of the nine one-dimensional look-up tables to obtain the corresponding output pixel value, obtaining the intermediate image I s1 , and the calculation formula is as follows: where LUT i represents the indexing process of the one-dimensional look-up table corresponding to the i-th position, and I (i) represents the pixel value at the i-th position in the 3×3 range;
[0020] S23: The one-dimensional look-up table is implemented in the simulation network of the cooperative look-up table through the following network: a linear layer with an input channel number of 1 and an output channel number of C, plus a linear layer with an input channel number of C and an output channel number of 1, and a tanh activation layer;
[0021] S24: In the second stage, use a four-dimensional look-up table for interactive indexing. On each RGB channel of the intermediate image I s1 indexed by the one-dimensional look-up table in the first stage, use three four-dimensional look-up tables LUT c , LUT d , LUT y with complementary indexing regions to form a 3×3 scanning area to scan the entire channel;
[0022] S25: Rotate the input 3×3 image patch by 0°, 90°, 180°, and 270° respectively with the i 0 pixel position as the center and use them as four inputs respectively, and then reverse-rotate the output results by the corresponding angles. The calculation formula is as follows; Among them represents the forward processing process of the four-dimensional lookup table. I and I' respectively represent the input and output pixel values, and R j and respectively represent the operations of rotating the image j times by 90° and performing k times of inverse rotation by 90°;
[0023] S26: Multiply the index results of all three four-dimensional lookup tables corresponding to the four inputs by the learnable scale factors and then average them to obtain the corresponding pixel values, resulting in the intermediate image I s2 , and the calculation formula is as follows: I s2 =(s c ·LUT c [I 0 , I 1 , I 3 , I 4 +s d ·LUT d [I 0 , I 2 , I 6 , I 8 +s y ·LUT y [I 0 , I 4 ,I 5 , I 7 ) / 3, where LUT c , LUT d , LUT y respectively represent the index processes of the corresponding four-dimensional lookup tables, and s c , s d , s y respectively represent the learnable scale factors corresponding to the three four-dimensional lookup tables;
[0024] S27: In the third stage, use a three-dimensional lookup table to perform single-point pixel mapping on the intermediate image I s2 to obtain the enhanced image enhanced(I 1 ), and the calculation formula is as follows: enhanced(I 1 ) = LUT p [I s2 , where LUT p represents the index processing process of the three RGB channels of the image by the three-dimensional lookup table.
[0025] Preferably, the specific steps for implementing the three index-region complementary four-dimensional lookup tables LUT c , LUT d , LUT y in the step S24 are as follows:
[0026] S241: Label the pixels in the 3×3 scanning area from left to right and from top to bottom as {i 0 , i 1 , i 2 , i 3 , i 4 , i 5 , i 6 , i 7 , i 8};
[0027] S242: Among them, the four-dimensional lookup table LUT c is responsible for indexing the pixel information at the four positions of {i 0 , i 1 , i 3 , i 4}. Specifically, it is implemented through the following network in the simulation network of the collaborative lookup table: a convolutional layer with a convolutional kernel size of 2×2, five dense convolutional layers, with 64 channels in each layer, and a tanh activation layer;
[0028] S243: Among them, the four-dimensional lookup table LUT d is responsible for indexing the pixel information at the four positions of {i 0 , i 2 , i 6 , i 8}. Specifically, it is implemented through the following network in the simulation network of the collaborative lookup table: a dilated convolutional layer with a convolutional kernel size of 2×2 and a dilation rate of 2, five dense convolutional layers, with 64 channels in each layer, and a tanh activation layer;
[0029] S244: The four-dimensional lookup table LUT y is responsible for indexing the pixel information at the four positions of {i 0 , i 4 , i 5 , i 7}. Specifically, it is implemented through the following network in the simulation network of the collaborative lookup table: a convolutional layer with a convolutional kernel size of 1×4, five dense convolutional layers, with 64 channels in each layer, and a tanh activation layer.
[0030] Preferably, the specific steps for constructing the global enhancement module in step three are as follows:
[0031] S31: Construct a global enhancement module based on a parallel Vision State-Space Module (VSS), which consists of the following components: a 3×3 depthwise separable convolutional layer, a layer normalization layer, a Parallel VSS Layer (PVSS) composed of four parallel vision state-space models. The number of channels C is evenly divided into four parts, and the number of input channels for each parallel branch is C / 4. A residual connection is used for each parallel channel with a scaling factor; the outputs of each branch in the parallel VSS layer are concatenated in the channel dimension; then there is a group normalization layer, an average pooling layer, and a linear layer with C input channels and 10 output channels. Finally, the output vector of size (10,1) is reshaped into two vectors of sizes (3,3) and (1) respectively, and the identity matrix and 1 are added respectively;
[0032] S32: Concatenate the input image I and the image enhanced(I 1 ) obtained in step S27 in the channel dimension as the input I m ;
[0033] S33: Input I m into the global enhancement module, and two global enhancement parameters are output. The vector of size (3,3) is used as a color adjustment matrix parameter W, and the vector of size (1) is used as a gamma correction parameter γ;
[0034] S34: Use the color adjustment matrix parameter W and the gamma correction parameter γ to perform global enhancement calculation G(·) on the image enhanced(I 1 ) obtained in step S27. The calculation formula is: c i , c j ∈ {R, G, B}, where ε = 1 - e 8 to ensure numerical stability. After calculation, the final enhanced image enhanced(I 2 ) is obtained.
[0035] Preferably, the vision state-space model in step S31 mainly includes two branches. The first branch is successively a linear layer, a depthwise separable convolutional layer, a SiLU activation layer, a 2D-SSM selective scan module, and a layer normalization layer. The second branch is successively a linear layer and a SiLU activation layer; the output results of the first and second branches are multiplied element-wise and input into the last linear layer.
[0036] Preferably, in the process of calculating the loss value of the constructed network to adjust the parameters of the network model in step four, since the lookup table needs to be stored using the int8 integer data type to reduce space, and the simulation network training of the collaborative lookup table requires using floating-point numbers to calculate gradients. Therefore, during the training process of the entire network, the forward process of the network is quantized into integers, and the calculated gradients are retained as floating-point numbers in the backward process.
[0037] Preferably, the specific steps of traversing the trained simulation network with multiple lookup tables in collaboration and storing the results in the corresponding lookup tables in the form of key-value pairs in step five are as follows:
[0038] S51: Traverse the input and output values of each lookup table corresponding module in the simulation network with multiple lookup tables in collaboration, and cache the results in the corresponding lookup tables in the form of key-value pairs;
[0039] S52: The one-dimensional lookup table uses the complete [0 - 255] interval to store key-value pairs and performs direct indexing during the inference process;
[0040] S53: The four-dimensional lookup table uses a sampling interval of [0 - 16] to store key-value pairs and uses the four-dimensional tetrahedron interpolation algorithm for interpolation indexing during the inference process;
[0041] S54: The three-dimensional lookup table uses a sampling interval of [0 - 32] to store key-value pairs and uses the three-dimensional tetrahedron interpolation algorithm for interpolation indexing during the inference process.
[0042] Preferably, the specific steps of fine-tuning the entire network on the training set in step six are as follows:
[0043] S61: Replace the simulation module of the corresponding lookup table in the above-mentioned simulation network with multiple lookup tables in collaboration with the lookup table constructed in step five, and then combine it with the global enhancement module to form the final network with multiple lookup tables in collaboration. During the subsequent inference process of the network, directly obtain the prediction value by indexing the lookup table;
[0044] S62: Treat the parameters in all lookup tables as trainable parameters, and perform fine-tuning training of the adaptation interpolation algorithm on the entire network on the training set;
[0045] S63: After testing with the training and validation sets, obtain the trained low-light image enhancement network.
[0046] Advantages of the present invention: Compared with the 3D LUTs method proposed by Zeng et al. and the 4D LUTs method proposed by Liu et al., which are image enhancement methods based on look-up tables, the present invention solves the problems that the look-up table cannot process local pixel information and cannot make full use of the global information of the image. In the low-light image enhancement task, it has stronger robustness to noise, and at the same time has higher color accuracy after enhancement. Compared with the method based on a complex deep neural network, the present invention retains the advantage of high execution efficiency of the look-up table network, is more lightweight, can achieve real-time low-light image enhancement, and is easy to deploy on devices, and has a wide application prospect in edge devices. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0048] Figure 1 is the overall flowchart of the present invention;
[0049] Figure 2 is the model framework diagram of the multi-look-up table collaborative network of the present invention;
[0050] Figure 3 is the partial branch diagram of the one-dimensional look-up table construction process in the embodiment of the present invention;
[0051] Figure 4 is the partial branch diagram of the four-dimensional look-up table construction process in the embodiment of the present invention;
[0052] Figure 5 is the partial branch diagram of the global enhancement module in the embodiment of the present invention;
[0053] Figure 6 is the comparison diagram of low-light image enhancement results in the specific embodiment of the present invention, where (a) is the input low-light image, (b) is the corresponding normal-illumination label image, (c) is the enhancement result of 3D LUTs, (d) is the enhancement result of 4D LUTs, (e) is the enhancement result of AdaInt, (f) is the enhancement result of Zero-DCE, (g) is the enhancement result of SNR-Aware, and (h) is the enhancement result of the algorithm of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0054] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0055] Figure 1 is the overall flowchart of the present invention. Obtain a training data set; construct a simulation network with collaborative multi-lookup tables and a global enhancement module; after training the network through the constructed loss function, traverse the output of the network, store the results in the corresponding lookup tables in the form of key-value pairs, and then perform fine-tuning; obtain the final low-light image enhancement network. The specific implementation is as follows:
[0056] Step 1. Obtain a training set containing paired low-light and normal-light images, including the following steps:
[0057] S11: Obtain paired low-light images T l and normal-light images T n . In this embodiment, the publicly available LOL data set commonly used in low-light image enhancement methods is adopted, which includes three subsets, LOL-v1, LOL-v2-real, and LOL-v2-syn, of different scenes. The LOL-v1 subset contains a total of 500 pairs of paired low-light and normal-light images, with an image size of 600×400 pixels; the LOL-v2-real subset contains a total of 789 pairs of paired low-light and normal-light images, with an image size of 600×400 pixels; the LOL-v2-syn subset contains a total of 1000 pairs of paired low-light and normal-light images, with an image size of 384×384 pixels. The above data sets are respectively divided into a training set, a validation set, and a test set according to a ratio of 8:1:1;
[0058] S12: Perform data preprocessing on the paired images in the training set, uniformly crop them into image patches of 128×128, and perform data augmentation operations such as rotation and flipping according to a ratio of 50%.
[0059] Step 2. Construct a simulation network with collaborative multi-lookup tables, as Figure 2 shown, including the following steps:
[0060] S21: Construct a simulation network of the lookup table to simulate the operation process of the lookup table in the network. The processing process of the lookup table is divided into three stages: a one-dimensional lookup table, a four-dimensional lookup table, and a three-dimensional lookup table, as Figure 2 shown;
[0061] S22: In the first stage, use a one-dimensional lookup table for non-interactive indexing, as Figure 3As shown, on each RGB channel of the image I, a scanning area of 3×3 composed of nine one-dimensional lookup tables is used to scan the entire channel. The results of the nine one-dimensional lookup tables are averaged to obtain the corresponding output pixel value, and the intermediate image I is obtained. s1 , and the calculation formula is as follows: where LUT i represents the indexing process of the one-dimensional lookup table corresponding to the i-th position, and I (i) represents the pixel value at the i-th position in the 3×3 range;
[0062] S23: The one-dimensional lookup table is implemented through the following network in the simulation network of the lookup table: a linear layer with an input channel number of 1 and an output channel number of 64, plus a linear layer with an input channel number of 64 and an output channel number of 1, and a tanh activation layer;
[0063] S24: In the second stage, a four-dimensional lookup table is used for interactive indexing. As Figure 4 shown, on each RGB channel of the intermediate image I obtained by indexing with the one-dimensional lookup table in the first stage s1 , three four-dimensional lookup tables LUT c , LUT d , and LUT y with complementary indexing regions are used to form a 3×3 scanning area to scan the entire channel. The specific method is as follows:
[0064] S241: The pixels in the 3×3 scanning area are marked as {i 0 , i 1 , i 2 , i 3 , i 4 , i 5 , i 6 , i 7 , i 8} from left to right and top to bottom;
[0065] S242: Among them, the four-dimensional lookup table LUT c is responsible for indexing the pixel information at the four positions of {i 0 , i 1 , i 3 , i 4 . Specifically, it is implemented through the following network in the simulation network of the lookup table: a convolutional layer with a convolutional kernel size of 2×2, five dense convolutional layers, each with 64 channels, and a tanh activation layer;
[0066] S243: Among them, the four-dimensional lookup table LUT d is responsible for indexing {i 0 , i 2 , i 6, i 8 The pixel information of four positions is specifically implemented through the following network in the analog network of the lookup table: a dilated convolutional layer with a convolutional kernel size of 2×2 and a dilation rate of 2, five dense convolutional layers, each with 64 channels, and a tanh activation layer;
[0067] S244: Among them, the four-dimensional lookup table LUT y is responsible for indexing {i 0 , i 4 , i 5 , i 7} The pixel information of four positions; specifically implemented through the following network in the analog network of the lookup table: a convolutional layer with a convolutional kernel size of 1×4, five dense convolutional layers, each with 64 channels, and a tanh activation layer;
[0068] S25: The input 3×3 image block is rotated 0°, 90°, 180°, and 270° respectively with the i 0 pixel position as the center and used as four inputs respectively, and then the output results are rotated inversely by the corresponding angles. The calculation formula is as follows; Among them represents the forward processing process of the four-dimensional lookup table, I and I′ represent the input and output pixel values respectively, R j and represent the operations of rotating the image j times by 90° and rotating it inversely j times by 90° respectively;
[0069] S26: Multiply the index results of all three corresponding four-dimensional lookup tables in the four inputs by the learnable scale factor and then average them to obtain the corresponding pixel value, obtaining the intermediate image I s2 , and the calculation formula is as follows: I s2 =(s c ·LUT c [I 0 , I 1 , I 3 , I 4 +s d ·LUT d [I 0 , I 2 , I 6 , I 8 +s y ·LUT y [I 0 , I 4 , I 5 , I 7 ) / 3, where LUT c , LUT d , LUT y represent the indexing processes of the corresponding four-dimensional lookup tables respectively, sc , s d , s y respectively represent the learnable scale factors corresponding to three four-dimensional lookup tables;
[0070] S27: In the third stage, a three-dimensional lookup table is used to perform single-point pixel mapping on the intermediate image I s2 to obtain the enhanced image enhanced(I 1 ), and the calculation formula is as follows: enhanced(I 1 ) = LUT p [I s2 , where LUT p represents the index processing process of the three-dimensional lookup table for the three RGB channels of the image.
[0071] Step 3. Construct a global enhancement module, as Figure 5 shown, including the following steps:
[0072] S31: Construct a global enhancement module based on the parallel vision state-space model (Vision State-Space Module, VSS), as Figure 5 (a) shown, which consists of the following components: a 3×3 depthwise separable convolutional layer, a linear normalization layer, a parallel VSS layer (Parallel VSS Layer, PVSS) composed of four parallel vision state-space models, the number of channels C is evenly divided into four parts, and the number of input channels for each parallel branch is C / 4, where each parallel channel uses a scale factor for residual connection; the outputs of each branch in the parallel VSS layer are concatenated in the channel dimension; then there is a group normalization layer, an average pooling layer, and a linear layer with the number of input channels C and the number of output channels 10. Finally, the output vector of size (10,1) is reshaped into two vectors of sizes (3, 3) and (1) respectively, and the identity matrix and 1 are added respectively; in this embodiment, the number of channels C is set to 64;
[0073] S311: The vision state-space model in the step S31 mainly includes two branches, as Figure 5 (b) shown, the first branch is successively a linear layer, a depthwise separable convolutional layer, a SiLU activation layer, a 2D-SSM selective scan module, as Figure 5 (c) shown, and a linear normalization layer; the second branch is successively a linear layer and a SiLU activation layer; the output results of the first and second branches are multiplied pointwise and input into the last linear layer; the number of all channels is 16.
[0074] S32: The input image I and the image enhanced(I obtained in step S271 ) After connection in the channel dimension, it serves as the input I of the global enhancement module m ;
[0075] S33: Input I m into the global enhancement module, and two global enhancement parameters are output. The vector of size (3, 3) output is used as a color adjustment matrix parameter W, and the vector of size (1) output is used as the gamma correction parameter γ;
[0076] S34: Use the color adjustment matrix parameter W and the gamma correction parameter γ to perform global enhancement calculation G(·) on the image enhanced(I 1 ) obtained in step S27. The calculation formula is: c i , c j ∈ {R, G, B}, where ε = 1 - e 8 to ensure numerical stability. After calculation, the final enhanced image enhanced(I 2 ) is obtained.
[0077] Step Four: Construct a loss function, perform low-light image enhancement on the training set through the constructed network, and calculate the loss value to adjust the parameters of the network model, including the following steps:
[0078] S41: Use the combination of L1 smooth loss L s and perceptual loss L p as the total loss function L t of the network. The calculation formula is: L t = αL s + (1 - α)L p , where α is a hyperparameter and is set to 0.8;
[0079] S42: The calculation formula of the above smooth loss L s is as follows: where x is the difference between the normalized pixel pre-test and the label image pixel value;
[0080] S43: The calculation formula of the above perceptual loss L p is as follows: where represents the predicted image, i represents the label image, F i represents the feature map of the i-th layer after inputting the image into the VGG network, and N represents the number of layers of the VGG network;
[0081] S44: During the training process of the entire network, the forward process of the network is quantized to integers, and the calculated gradient is retained as a floating point number in the backward process, as Figure 4 shown;
[0082] S45: Iteratively optimize the loss function on the training set to obtain a trained simulation network with collaborative multiple lookup tables and a global enhancement module.
[0083] Step Five: Traverse the trained simulation network with collaborative multiple lookup tables and store the results in the corresponding lookup tables in the form of key-value pairs, including the following steps:
[0084] S51: Traverse the input and output values of the modules corresponding to each lookup table in the simulation network with collaborative multiple lookup tables, and cache the results in the corresponding lookup tables in the form of key-value pairs;
[0085] S52: The one-dimensional lookup table uses the complete [0 - 255] interval to store key-value pairs and performs direct indexing during the inference process;
[0086] S53: The four-dimensional lookup table uses a sampling interval of [0 - 16] to store key-value pairs and performs interpolation indexing using the four-dimensional tetrahedron interpolation algorithm during the inference process;
[0087] S54: The three-dimensional lookup table uses a sampling interval of [0 - 32] to store key-value pairs and performs interpolation indexing using the three-dimensional tetrahedron interpolation algorithm during the inference process.
[0088] Step Six: Fine-tune the entire network on the training set. The specific steps are as follows:
[0089] S61: Replace the simulation modules of the corresponding lookup tables in the above simulation network with collaborative multiple lookup tables with the lookup tables constructed in Step Five, and then combine them with the global enhancement module to form the final network with collaborative multiple lookup tables. During the inference process of the subsequent network, directly obtain the predicted values by indexing the lookup tables;
[0090] S62: Treat the parameters in all lookup tables as trainable parameters. Based on the loss function described in Step S41, treat the parameters in all lookup tables as trainable parameters and perform fine-tuning training of the entire network on the training set for the adaptive interpolation algorithm;
[0091] S63: After testing with the training and validation sets, obtain a trained low-light image enhancement network.
[0092] Step Seven: Input the low-light image to be processed into the trained low-light image enhancement network to obtain the enhanced normal-light image.
[0093] The present invention is based on a multi-lookup table collaborative network for low-light image enhancement, which solves the problems that existing lookup table-based methods cannot handle the noise in low-light images and the inaccurate color enhancement. By collaboratively using different types of lookup tables, the receptive field is expanded and the network robustness is improved. The present invention combines a lightweight global enhancement module, further improving the adaptability of the lookup table network and the accuracy of color restoration, and having a very high execution efficiency. The training method for constructing lookup tables by simulating the network proposed by the present invention can stably and quickly construct the required lookup tables. The present invention can achieve high-quality real-time low-light image enhancement.
[0094] In an embodiment of the present invention, the LOL dataset is selected for experiments, which includes three subsets: LOL-v1, LOL-v2-real, and LOL-v2-syn. The algorithm of the present invention is used to perform low-light image enhancement on the images, and is compared with five methods: 3D LUTs, 4DLUTs, AdaInt, Zero-DCE, and SNR-Aware. The comparison results are as Figure 6 shown.
[0095] Figure 6 (a) is the input low-light image, Figure 6 (b) is the corresponding normal-light label image, Figure 6 (c) is the enhancement result of 3DLUTs, Figure 6 (d) is the enhancement result of 4D LUTs, Figure 6 (e) is the enhancement result of AdaInt, Figure 6 (f) is the enhancement result of Zero-DCE, Figure 6 (g) is the enhancement result of SNR-Aware, Figure 6 (h) is the enhancement result of the algorithm of the present invention. It can be observed that when enhancing the extremely dark scene images in the LOL-v1 and LOL-v2-real test sets, the subjective visual performance of the algorithm of the present invention far exceeds that of Zero-DEC and other real-time low-light image enhancement methods based on lookup tables. After the enhancement of 3D LUTs and 4D LUTs, it can be clearly seen that the noise on some planes is amplified, while the present invention can well suppress the noise after enhancement. At the same time, the present invention has better enhancement effects on color and brightness. When enhancing the images in the LOL-v2-syn test set, overexposure occurs in the bright regions of the images enhanced by methods such as SNR-Aware, 3D LUTs, 4D LUTs, and AdaInt. The color and brightness information of the images enhanced by the present invention is more accurate. Based on the comprehensive comparison results, the algorithm of the present invention has the best visual effect in processing low-light images.
[0096] In order to effectively and objectively quantify the evaluation of the method proposed in the present invention for processing low - illumination images, two evaluation metrics, namely peak signal - to - noise ratio (PSNR) and structural similarity (SSIM), as well as the execution efficiency of the algorithm are used to evaluate the experimental results.
[0097] PSNR is the most commonly used objective observation method for evaluating image quality. The larger the PSNR value between two images, the better the denoising effect and the more similar the images. The calculation formula for the peak signal - to - noise ratio metric is as follows:
[0098]
[0099]
[0100] Where z is the number of bits per pixel, MSE is the mean square error between the labeled image and the enhanced image, m is the height of the image, n is the width of the image, R(p, q) represents the gray value of the pixel at the p - th row and q - th column of the labeled image, and F(p, q) represents the gray value of the pixel at the p - th row and q - th column of the enhanced image.
[0101] The SSIM metric is often used to measure the structural similarity between two images. It is independent of brightness and contrast and mainly focuses on the structural information of the images. The larger the SSIM value, the more complete the structural information of the enhanced image is retained. The calculation formula for the structural similarity metric is as follows:
[0102]
[0103] Where i and j represent the labeled image and the enhanced image respectively, C 1 , C 2 are constants, μ, σ 2 and σ ij represent the mean value, variance, and covariance respectively.
[0104] The execution efficiency of the algorithm is evaluated by comparing the storage occupancy, computational complexity, and inference speed of the algorithm. A lower storage occupancy, lower computational complexity, and faster inference speed indicate a better algorithm.
[0105] In an embodiment of the present invention, the PSNR and SSIM metrics are used to evaluate the effects of the method proposed in the present invention and the comparative methods on the enhancement of low - illumination images. The results are shown in Table 1.
[0106] In an embodiment of the present invention, the algorithm execution efficiencies of the method proposed in the present invention and the comparative methods are compared. The results are shown in Table 2.
[0107] Table 1 Quantitative comparison results of PSNR and SSIM values of test - set images
[0108]
[0109] Table 2 Quantitative comparison results of the execution efficiency of the method of the present invention and the comparative methods
[0110]
[0111] As can be seen from Table 1, after being processed by various methods, the method of the present invention has achieved the optimal or sub-optimal results in terms of PSNR and SSIM indicators. Moreover, compared with the lookup table-based methods such as 3D LUTs and 4D LUTs, the method of the present invention has made great improvements in the above indicators. Among them, there is an improvement amplitude of more than 0.1 in the SSIM indicator of the LOL-v1 test set and an improvement amplitude of more than 2.0 dB in the PSNR indicator, indicating that the method of the present invention can better handle the noise in low-light images, retain the structural information of the images, and at the same time can more accurately restore the color, brightness and contrast of the images after enhancement.
[0112] As can be seen from Table 2, compared with other methods, the method of the present invention has lower parameter quantity and floating-point operation quantity FLOPs, and shorter running time. Therefore, it performs better in terms of storage occupancy, computational complexity and inference speed, indicating that the method of the present invention has higher execution efficiency and can achieve real-time low-light image enhancement.
[0113] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A real-time low-light image enhancement method based on a multi-lookup table collaborative network, characterized in that: Construct an enhancement network based on simulating the coordination of multiple lookup tables in one, three and four dimensions, and a global enhancement module based on the visual state space model; Construct a loss function training network, simulate the network output by traversing the lookup table, cache it in the corresponding lookup table in the form of key-value pairs, and complete the construction of the lookup table; perform adaptive fine-tuning training on the constructed lookup table based on the training set to obtain a trained low-light image enhancement network. The steps are as follows: Step 1: Obtain a training set containing paired low-illumination and normal-illumination images; Step 2: Construct a simulation network with multiple lookup tables working together; Step 3: Build a global enhancement module; Step 4: construct a loss function, perform low-light image enhancement on the training set through the constructed network, and calculate the loss value to adjust the parameters of the network model; Step 5: traverse the trained multi-lookup table collaborative simulation network and store the results in the form of key-value pairs into the corresponding lookup table; Step 6: Fine-tune the entire network on the training set to obtain a trained low-light image enhancement network; Step 7: Input the low-light image to be processed into the trained low-light image enhancement network to obtain the enhanced normal-light image.
2. The real-time low-light image enhancement method based on a multi-lookup table collaborative network according to claim 1, characterized in that: The specific steps of constructing a simulation network with multiple lookup tables in step 2 are as follows: S21: constructing a simulation network of a lookup table to simulate the operation process of the lookup table in the network. The processing process of the lookup table is divided into three stages: a one-dimensional lookup table, a four-dimensional lookup table, and a three-dimensional lookup table; S22: In the first stage, a one-dimensional lookup table is used for non-interactive indexing. On each RGB channel of image I, a scanning area with a range of 3×3 consisting of nine one-dimensional lookup tables is used to scan the entire channel. The results of the nine one-dimensional lookup tables are averaged to obtain the corresponding output pixel value to obtain the intermediate image I. s1 , the calculation formula is as follows: Among them, LUT i Represents the indexing process of the one-dimensional lookup table corresponding to the i-th position, I (i) Represents the pixel value at the i-th position in the 3×3 range; S23: The one-dimensional lookup table is implemented in the simulation network of the collaborative lookup table through the following network: a linear layer with 1 input channel and C output channels, plus a linear layer with C input channels and 1 output channels, and a tanh function activation layer; S24: The second stage uses a four-dimensional lookup table for interactive indexing, and the intermediate image I obtained by indexing the one-dimensional lookup table in the first stage s1 For each RGB channel, a four-dimensional lookup table LUT with three complementary index areas is used. c , LUT d , LUT y Form a 3×3 scanning area to scan the entire channel; S25: The input 3×3 image block is rotated by 0°, 90°, 180°, and 270° respectively with the i0 pixel position as the center, and then used as four inputs respectively, and then the output result is reversely rotated by the corresponding angles. The calculation formula is as follows; in represents the forward processing of the four-dimensional lookup table, I and I′ represent the input and output pixel values respectively, R j and Respectively represent the operation of rotating the image by 90° j times and performing the reverse rotation by 90° j times; S26: Multiply the index results of all three four-dimensional lookup tables corresponding to the four inputs by the learnable scaling factor and average them to obtain the corresponding pixel values to obtain the intermediate image I s2 , the calculation formula is as follows: s2 =(s c LUT c [I0, I1, I3, I4]+s d LUT d [I0,I2,I6,I8]+s y LUT y [I0, I4, I5, I7]) / 3, where LUT c , LUT d , LUT y Respectively represent the indexing process of the corresponding four-dimensional lookup table, s c 、s d 、s y They represent the learnable scaling factors corresponding to the three four-dimensional lookup tables respectively; S27: The third stage uses a three-dimensional lookup table to convert the intermediate image I s2 Perform single-point pixel mapping to obtain the enhanced image enhanced(I1), the calculation formula is as follows: enhanced(I1) = LUT p [I s2 ], where LUT p It represents the index processing process of the three-dimensional lookup table on the RGB channels of the image.
3. The real-time low-light image enhancement method based on a multi-lookup table collaborative network according to claim 2, characterized in that: The four-dimensional lookup table LUT of the three index regions complementary in step S24 c , LUT d , LUT y The specific steps to achieve this are: S241: Mark the pixels of the 3×3 scan area from left to right and from top to bottom as {i0, i1, i2, i3, i4, i5, i6, i7, i8} respectively; S242: Four-dimensional lookup table LUT c Responsible for indexing the pixel information of the four positions {i0, i1, i3, i4}, which is specifically implemented in the simulation network of the collaborative lookup table through the following network: a convolution layer with a convolution kernel size of 2×2, five dense convolution layers, each with 64 channels, and a tanh activation layer; S243: Four-dimensional lookup table LUT d Responsible for indexing the pixel information of the four positions {i0, i2, i6, i8}, which is specifically implemented in the simulated network of the collaborative lookup table through the following network: a dilated convolution layer with a convolution kernel size of 2×2 and a dilation rate of 2, five dense convolution layers, each with 64 channels, and a tanh activation layer; S244: 4D Lookup Table LUT y Responsible for indexing the pixel information of the four positions {i0, i4, i5, i7}; specifically, it is implemented through the following network in the simulation network of the collaborative lookup table: a convolution layer with a convolution kernel size of 1×4, five dense convolution layers, each with 64 channels, and a tanh activation layer.
4. The real-time low-light image enhancement method based on a multi-lookup table collaborative network according to claim 1, characterized in that: The specific steps of constructing the global enhancement module in step 3 are: S31: Construct a global enhancement module based on a parallel vision state-space module (VSS), which consists of the following components: a 3×3 depthwise separable convolutional layer, a linear normalization layer, a parallel VSS layer (PVSS) composed of four parallel vision state-space models, the number of channels C is evenly divided into four parts, the number of input channels of each parallel branch is C / 4, and each parallel channel is residually connected using a scaling factor: the output of each branch in the parallel VSS layer is connected in the channel dimension; followed by a group normalization layer, an average pooling layer and a linear layer with C input channels and 10 output channels. Finally, the output vector of size (10,1) is reshaped into two vectors of size (3,3) and (1), and the unit matrix and 1 are added respectively; The visual state space model in step S31 mainly includes two branches, the first branch is a linear layer, a depth-separable convolution layer, a SiLU activation layer, a 2D-SSM selection scanning module and a linear normalization layer, and the second branch is a linear layer and a SiLU activation layer; the output results of the first and second branches are dot-multiplied and input into the last linear layer; S32: Concatenate the input image I and the image enhanced (I1) obtained in step S27 in the channel dimension as the input I of the global enhancement module m ; S33: I m Input into the global enhancement module, and output two global enhancement parameters. The output vector of size (3, 3) is used as a color adjustment matrix parameter W, and the output vector of size (1) is used as a gamma correction parameter γ. S34: Perform global enhancement calculation G(·) on the image enhanced(I1) obtained in step S27 using the color adjustment matrix parameter W and the gamma correction parameter γ. The calculation formula is: where ε = 1-e 8 To ensure numerical stability, the final enhanced image enhanced (I2) is obtained after calculation.
5. The real-time low-light image enhancement method based on a multi-lookup table collaborative network according to claim 1, characterized in that: The specific steps of traversing the trained multi-lookup table collaborative simulation network in step 5 and storing the results in the form of key-value pairs in the corresponding lookup table are: S51: traverse the input and output values of the corresponding module of each lookup table in the multi-lookup table collaborative simulation network, and cache the results in the form of key-value pairs into the corresponding lookup table; S52: One-dimensional lookup table uses the full [0-255] interval to store key-value pairs, with direct indexing during inference; S53: The four-dimensional lookup table uses a sampling interval of [0-16] to store key-value pairs, and uses a four-dimensional tetrahedron interpolation algorithm for interpolation indexing during the inference process; S54: The three-dimensional lookup table uses a sampling interval of [0-32] to store key-value pairs, and uses a three-dimensional tetrahedron interpolation algorithm for interpolation indexing during the inference process.
6. The real-time low-light image enhancement method based on a multi-lookup table collaborative network according to claim 1, characterized in that: The specific steps of fine-tuning the entire network on the training set in step 6 are: S61: replacing the simulation module of the corresponding lookup table in the above multi-lookup table collaborative simulation network with the lookup table constructed in step 5, and then combining with the global enhancement module to form the final multi-lookup table collaborative network, and directly obtaining the prediction value by indexing the lookup table in the subsequent network reasoning process; S62: All parameters in the lookup table are considered as trainable parameters, and the entire network is fine-tuned using the adaptive interpolation algorithm on the training set; S63: After training and testing the validation set, a trained low-light image enhancement network is obtained.