A joint optimization method for low-light image enhancement and instance segmentation

Through the joint optimization method of low-light image enhancement and instance segmentation based on Retinex theory, the problems of segmentation accuracy and visual effect under low-light conditions are solved by using illumination map and reflectance map estimators and semantic feature fusion, and efficient instance segmentation and enhancement effects are achieved.

CN119313906BActive Publication Date: 2025-10-03TIANJIN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411422708.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-12
Publication Date
2025-10-03
Estimated Expiration
2044-10-12

AI Technical Summary

Technical Problem

Existing low-light image instance segmentation methods find it difficult to simultaneously improve visual effects and segmentation accuracy during the enhancement process, and fail to effectively utilize the semantic information generated during the segmentation process.

Method used

A joint optimization method of low-light image enhancement and instance segmentation based on Retinex theory is adopted. Through the illumination map and reflectance map estimator decomposition module, combined with the instance segmentation network CondInst and the illumination adjustment module, the enhancement process is optimized by using semantic feature fusion and instance-guided color histogram loss function.

Benefits of technology

The instance segmentation accuracy in low-light conditions is improved, the color cast of the enhanced image is reduced, and enhanced images with better visual effects are obtained.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119313906B_ABST
    Figure CN119313906B_ABST
Patent Text Reader

Abstract

The present invention relates to a joint optimization method for low-light image enhancement and instance segmentation based on Retinex theory, comprising the following steps: simultaneously inputting a low-light image and a normal-light image of the same scene into a decomposition module, ensuring that the obtained reflectance map and illumination map conform to the Retinex model through a cross-validated loss function; inputting the reflectance map into an instance segmentation module, dynamically generating a convolution kernel of a segmentation branch through a conditioning mechanism, intercepting intermediate features containing instance-level semantic information before convolution of the segmentation branch and inputting them into an illumination adjustment module; inputting the intermediate features containing instance-level semantic information intercepted in the segmentation module and an illumination map estimated by an illumination estimator into an illumination adjustment module for adjustment; the illumination adjustment module is divided into a semantic feature fusion submodule and an illumination enhancement submodule, and an instance-guided color histogram loss function is introduced for supervision to obtain an enhanced image with less color deviation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of deep learning algorithms and computer vision technology, and relates to a joint optimization algorithm for low-light image enhancement and instance segmentation based on Retinex theory, covering methods such as convolutional neural networks and attention mechanisms. Background Art

[0002] Low-light images are those captured under conditions of insufficient exposure or extremely low contrast. These images not only degrade the visual perception of the human eye but also lead to reduced accuracy in high-level tasks such as object detection and instance segmentation. Instance segmentation, a fundamental challenge in computer vision, aims to identify and segment the pixels corresponding to individual object instances in an image. This technology has been widely used in fields such as autonomous driving, medical systems, and agricultural analysis. However, in low-light scenes, these methods often struggle to achieve satisfactory results due to the obscuration of useful information and the introduction of noise.

[0003] At present, the research on instance segmentation methods for low-light images is still in its early stages. Most methods rely on low-light image enhancement (LLIE) as a preprocessing step and then use general instance segmentation methods for segmentation. These methods can be divided into two categories based on the specific model structure: (1) The two tasks are regarded as independent processes and the output of LLIE is directly used as the input of instance segmentation. Such methods are often complex in structure and cannot significantly improve the performance of instance segmentation. This is because low-light enhancement methods usually include steps such as noise reduction and brightness adjustment to improve visual quality, but this may lead to overexposure and loss of details, thereby losing information that is crucial for subsequent high-level visual tasks. For example, LIIS regards low-light enhancement and instance segmentation as two separate tasks, adds an additional post-processing step for denoising in the enhancement part, and adds a wavelet feature fusion module to the segmentation part to improve the performance of instance segmentation. Although the accuracy of segmentation can be improved, it will also sacrifice efficiency and face the risk of introducing erroneous prior knowledge into the artificially constructed loss function. (2) The two tasks are connected through a cascade network and processed as a whole. Such methods can usually achieve improved instance segmentation performance, but the visual effect of the images generated by the LLIE method is usually poor. This is because enhanced images for machine vision are often detrimental to human vision. For example, FeatEnhancer integrates a low-light enhancement network before the object detection task, which is highly correlated with instance segmentation. Visualizing the intermediate features between the enhanced and segmented parts reveals that the resulting image has a significant color shift compared to the true image under normal light. Furthermore, both architectures fail to effectively utilize the rich semantic information generated during the segmentation process, which is beneficial for improving the performance of LLIE methods.

[0004] In summary, how to make full use of the complementary information in the two tasks and improve the accuracy of the instance segmentation task while ensuring that the output image of the LLIE method is beneficial to human vision has become an urgent problem to be solved. Summary of the Invention

[0005] To address the challenges of instance segmentation in low-light conditions, this paper proposes a joint optimization method for low-light image enhancement and instance segmentation based on Retinex theory. The technical solution is as follows:

[0006] A joint optimization method for low-light image enhancement and instance segmentation based on Retinex theory includes the following steps:

[0007] Step 1: The low-light image and the normal-light image of the same scene are simultaneously input into the decomposition module. The decomposition module includes an illumination map estimator and a reflectance map estimator, which are used to estimate the illumination map and reflectance map respectively. The illumination map estimator generates a single output feature layer, while the reflectance map estimator generates three output feature layers corresponding to the three RGB channels. The cross-validation loss function is used to ensure that the obtained reflectance map and illumination map conform to the Retinex model. The loss function of the decomposition model is divided into two parts: reconstruction loss and reflectance loss. Figure 1 Causative losses;

[0008] Step 2: Select the instance segmentation network CondInst as the network of the segmentation module, input the reflectance map into the instance segmentation module, dynamically generate the convolution kernel of the segmentation branch through the CondInst conditioning mechanism, intercept the intermediate features containing instance-level semantic information before the convolution of the segmentation branch and send them to the illumination adjustment module;

[0009] In step 3, the intermediate features containing instance-level semantic information intercepted in the segmentation module and the illumination map estimated by the illumination estimator are sent to the illumination adjustment module for adjustment. The illumination adjustment module is divided into two parts: a semantic feature fusion submodule and an illumination enhancement submodule. First, the semantic feature fusion submodule constructed based on channel attention is used to fuse the intermediate features containing instance-level semantic information intercepted in the segmentation module with the illumination map to obtain fused features. Then, the fused features are enhanced by the illumination enhancement submodule to obtain an enhanced illumination map. Finally, the enhanced illumination map is element-wise multiplied with the reflectance map estimated by the reflectance map estimator, and an instance-guided color histogram loss function is introduced for supervision to obtain an enhanced image with less color deviation. The instance-guided color histogram loss function is calculated as follows: by counting the contribution value of each point to the specified pixel intensity in the three channels of the image and adding them together, three intensity histograms are constructed to characterize the intensity distribution of each channel. Then, the indexes of all instances in the target image are obtained by mask annotation, and the intensity histogram of each instance between the enhanced image and the true value are compared to obtain the instance-guided color histogram loss.

[0010] Furthermore, in step 1, the illumination map estimator applies a sigmoid activation function at the end of the module to normalize the output to the range [0, 1]. The reflectance map estimator uses a U-Net architecture.

[0011] Furthermore, in step 1, the reconstruction loss function is:

[0012] L DEC =MSE(I,L·R)+MSE(R,I / L)+Smooth(L)

[0013] Where I is the input image, i.e., a low-light image or a normal-light image, L is the illumination map, and R is the reflectance map. Smooth(·) represents the loss function that constrains the illumination map to be piecewise smooth; MSE(·) represents the mean squared error loss; MSE(I,L·R) represents the reconstruction loss, which ensures that the element-wise product of the illumination map and the reflectance map is consistent with the input image; and MSE(R,I / L) further constrains the reflectance map.

[0014] Furthermore, the specific mathematical expression of Smooth(·) is as follows:

[0015]

[0016] Where i and j in the subscript represent the i-th pixel and j-th pixel in the image, L represents the illumination map, N is the total number of pixels, and N h(i) Represents all pixels in the h×h area near i; calculate w based on the color matching of I in YUV space i,j ,σ r is the standard deviation, c represents the image channel; |·| represents the absolute value operation;

[0017] Furthermore, the consistency loss of the reflection map is defined as follows:

[0018] L CR =SSIM(L N ·R L ,I N )+MSE(L N ·R L ,I N )+SSIM(R L ,R N )+||R L -R N ||

[0019] Among them, R L is the reflectance map of the low-light image, R N is the reflection image of the normal light image, I Nis the normal-light image corresponding to the low-light image; SSIM(·) represents the structural similarity loss function between the two images; ||·|| represents the operation of taking the L1 norm.

[0020] Furthermore, the decomposition module is trained separately first, and the parameters of the decomposition module are optimized through gradient backpropagation. When the loss no longer decreases and tends to be stable, the weights are saved, and the illumination map and reflection map obtained at this time are used as part of the input of the subsequent segmentation module and illumination adjustment module, and the weights of the module are continued to be updated in subsequent stages.

[0021] Furthermore, the method of step 2 is as follows:

[0022] 1) Obtain the intermediate features of the segmentation module containing instance-level semantic information:

[0023] The instance segmentation network CondInst is selected as the network of the segmentation module, and the reflectance map is used as input to train the segmentation module. The intermediate features of its segmentation branch containing instance-level semantic information are intercepted as the input of the illumination adjustment module.

[0024] 2) Get the segmentation results:

[0025] The convolution kernels for different categories generated by the convolution kernel prediction branch of CondInst predict the outgoing features of the segmentation branch; DiceLoss is used as the mask loss to measure the difference between the predicted instance mask and the true mask, and the gradient descent method is used to update the weights of the segmentation module and the reflection map estimator.

[0026] Furthermore, in step 3, the method for constructing the semantic feature fusion submodule based on channel attention is as follows:

[0027] Taking the mask features and illumination map as input, the transposed attention mechanism is used to fuse them into the output features. The volume fusion process is expressed as follows:

[0028]

[0029] Among them, F M represents the mask feature, I L Represents the low-light image illumination map; first, F M and I L Perform convolution and upsampling operations to align them in size and get F' M and I' L ; Then, F M 'and I L 'Reshape into a single channel feature by rearranging elements and To input into the attention mechanism; then, calculate the attention map: and Multiply by the transpose of and divide by Where C represents the number of channels, and then after the Softmax function, we get a C×C attention map with a value range of [0,1]. Multiply the transpose of and add the result to I' L Add them together and send them into the feedforward network (FFN) as input to finally obtain the fused feature F f .

[0030] Furthermore, in step 3, the illumination enhancement submodule selects the same structure as the illumination map estimator.

[0031] Furthermore, the method of constructing the instance-guided color histogram loss function is as follows:

[0032] The grayscale histogram of each channel is obtained by kernel density estimation, and the expression for calculating the frequency of each subinterval of the histogram is:

[0033]

[0034] Among them, P C represents the instance set in the Cth channel, P Cn is the pixel set of the nth instance in the Cth channel; the enhanced image and the normal light image whose histogram needs to be calculated are divided into K sub-intervals, the length of each interval is L = 1 / K, and k ranges from 1 to K; the center value of the kth interval is represented by μ k =(k-0.5)×1 / K; for the Cth channel of the i-th instance, is the frequency of the kth interval; by selecting the Sigmoid function as the kernel function, the sum of the contribution of each pixel in the region to the gray value intensity is calculated to obtain the histogram frequency; where x j represents the jth pixel in the patch, σ is the scaling factor;

[0035] The instance-guided color histogram loss function is obtained by calculating the L1 loss, and the formula is:

[0036]

[0037] Here, H(I') and H(I) represent the grayscale histograms of the enhanced image and the normal light image, respectively. This loss function is used to measure the difference in grayscale distribution between the corresponding instances of the enhanced image and the normal light image to guide model training and reduce the color cast in the foreground area.

[0038] The essential features and beneficial effects of the present invention are:

[0039] 1. A “decomposition-first, segmentation-later” strategy is proposed to alleviate the negative impact of poor illumination in low-light scenes, thereby improving the performance of instance segmentation under such conditions.

[0040] 2. The semantic feature fusion module integrates instance-level semantic information into the enhancement process. The semantic information guides the enhancement model to make differentiated enhancements for regions with different semantics, thereby obtaining enhanced images with better visual effects.

[0041] 3. An instance-guided color histogram loss function is proposed to reduce the color cast in the foreground area by strengthening the supervision of color consistency between the corresponding instances of the enhanced image and the normal light image. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 The overall structure of the network is given;

[0043] Figure 2 The structure of the semantic feature fusion module is given. DETAILED DESCRIPTION

[0044] The basic scheme of the present invention will be described below with reference to the accompanying drawings.

[0045] The present invention performs processing according to the following steps, which can successively obtain a reflection map with clear details, obtain a high-precision segmentation map through the reflection map, and finally further process the semantic features of the segmentation module and the features after fusion of the illumination map to obtain a high-quality enhanced image.

[0046] In step 1, the low-light image and the normal-light image are simultaneously input into the decomposition module. The reflection map and illumination map are obtained through the reflection map estimator and the illumination map estimator. The cross-validation loss function is used to ensure that the reflection map details are clear and the illumination map gradient is smooth.

[0047] The specific steps are as follows:

[0048] 1) Dataset preparation:

[0049] We use 1561 pairs of normal-light and low-light image pairs from the LIS dataset, covering eight categories including bicycles, cars, and motorcycles, as training sets, and 669 pairs as validation sets.

[0050] 2) Design decomposition module:

[0051] The decomposition module consists of an illumination map estimator and a reflectance map estimator. It serves as the first part of the instance-aware enhancement flow and the illumination-independent instance segmentation flow, respectively, to estimate the illumination map and reflectance map. The illumination map estimator's structure consists of five convolutional layers, with a sigmoid activation function applied at the end of the module to normalize the output to the range [0, 1]. The reflectance map estimator adopts a U-Net architecture, consisting of a series of convolutional layers that first downsamples the input, then upsamples it, and combines skip connections. According to Retinex theory, the three channels of an RGB image share the same illumination. Therefore, the illumination map estimator generates a single output feature layer, while the reflectance map estimator generates three output feature layers. By introducing a cross-validation loss function, we ensure that the obtained reflectance and illumination maps conform to the Retinex model.

[0052] 3) Construct decomposition loss function:

[0053] The loss function of the decomposition model can be divided into two parts: reconstruction loss and reflection loss Figure 1 Causative loss.

[0054] Reconstruction loss: Since the network is designed based on Retinex theory, the output needs to be constrained to conform to the Retinex physical model:

[0055] L DEC =MSE(I,L·R)+MSE(R,I / L)+Smooth(L)

[0056] Where I is the input image, L is the illumination map, and R is the reflectance map. Smooth(·) represents the loss function that constrains the illumination map to be piecewise smooth. MSE(·) represents the mean squared error loss. MSE(I,L·R) represents the reconstruction loss, which ensures that the element-wise product of the illumination map and the reflectance map is consistent with the input image. MSE(R,I / L) further constrains the reflectance map to make it more reasonable. During training, the gradient calculation of L is stopped to maintain training stability. The specific mathematical expression of Smooth(·) is as follows:

[0057]

[0058] Where i represents the i-th pixel in the image, N is the total number of pixels, and N h(i) represents all pixels in the h×h area near i, w {i,j} Based on the color matching calculation of I in YUV space, the standard deviation σ r Set to 10, c represents the image channel, and |·| represents the absolute value operation.

[0059] reflection Figure 1 Consistency loss: The consistency loss of the reflection map is defined as follows:

[0060] LCR =SSIM(L N ·R L ,I N )+MSE(L N ·R L ,I N )+SSIM(R L ,R N )+||R L -R N ||

[0061] Among them, R L is the reflectance map of the low-light image, R N is the reflectance map of the normal light image. Since the low light image and the normal light image depict the same object in the same scene, they should have the same reflectance characteristics. N is the normal light image corresponding to the low light image. The value of SSIM(·) is 1 minus the structural similarity between the two items in the brackets. L1 loss and SSIM loss are in R L and R N The weights between are selected based on experimental results to ensure that a more detailed reflection map is generated. ||·|| represents the operation of taking the L1 norm. By calculating L N ·R L with I N The structural similarity and mean square error loss realize the cross-validation between the low-light reflection map and the normal-light image illumination map, ensuring the rationality of the low-light reflection map.

[0062] 4) Training to obtain illumination map and reflection map

[0063] Because the enhancement model and the segmentation model have different fitting difficulties, the decomposition module is trained separately first, and its parameters are optimized through gradient backpropagation. When the loss stops decreasing and tends to stabilize, the weights are saved. The illumination map and reflectance map obtained at this time are used as part of the input to the subsequent segmentation module and illumination adjustment module, and the weights of this module are further updated in subsequent stages.

[0064] In step 2, the reflection map is input into the instance segmentation module, and the convolution kernel of the segmentation branch is dynamically generated through the conditioning mechanism. At the same time, the intermediate features containing instance-level semantic information before the convolution of the segmentation branch are taken out and sent to the illumination adjustment module.

[0065] The specific steps are as follows:

[0066] 1) Obtain the intermediate features of the segmentation module containing instance-level semantic information:

[0067] The classic instance segmentation network CondInst is selected as the network of the segmentation module, and the reflectance map is used as input to train the module. The intermediate features of its segmentation branch containing instance-level semantic information are intercepted as the input of the illumination adjustment module.

[0068] 2) Get the segmentation results:

[0069] The convolution kernel prediction branch of CondInst generates convolution kernels for different categories to predict the outgoing features of the segmentation branch. DiceLoss is used as the mask loss to measure the difference between the predicted instance mask and the true mask, and gradient descent is used to update the weights of the segmentation module and the reflectance map estimator.

[0070] In step 3, the intermediate features containing instance-level semantic information intercepted in the segmentation module and the illumination map estimated by the illumination estimator are fed into the illumination adjustment module for adjustment. Its specific structure is divided into two parts: a semantic feature fusion submodule and an illumination enhancement submodule. First, the semantic feature fusion submodule constructed based on channel attention is used to fuse the intermediate features containing instance-level semantic information intercepted in the segmentation module with the illumination map to obtain the fused features. The fused features are then enhanced by the illumination enhancement submodule to obtain an enhanced illumination map. Finally, the enhanced illumination map is element-wise multiplied with the reflectance map estimated by the reflectance map estimator, and an instance-guided color histogram loss function is introduced for supervision to obtain an enhanced image with less color deviation.

[0071] The specific steps are as follows:

[0072] 1) Construct semantic feature fusion submodule:

[0073] In order to integrate semantic features into the illumination adjustment process and enable instance-level semantic features to guide enhancement, an attention-based semantic feature fusion submodule is constructed, which takes mask features and illumination maps as input and fuses them into output features using the transposed attention mechanism.

[0074] The specific fusion process can be expressed as follows:

[0075]

[0076] Among them, F M represents the mask feature, I L Represents the low-light image illumination map. First, F M and I L Perform convolution and upsampling operations to align them in size and get F' M and I' L Next, F M 'and I L 'Reshape into a single channel feature by rearranging elements and Then, we calculate the attention map by and Multiply by the transpose of and divide by (where C represents the number of channels, which is set to 24), and then passes through the Softmax function to obtain an attention map of size C×C, with a value range between [0,1]. Next, the attention map is compared with Multiply the transpose of and add the result to I' L Add them together and send them into the feedforward network (FFN) as input to finally obtain the fused feature F f The whole process effectively combines the mask features and the low-light image features through the attention mechanism, thus enhancing the fusion effect.

[0077] 2) Construct the overall loss function of the illumination enhancement submodule and the adjustment module, and introduce the instance-guided color histogram loss function:

[0078] To improve network efficiency, the same structure as the illumination map estimator is selected as the structure of the illumination enhancement submodule.

[0079] Introducing instance-guided color histogram loss for low-light enhancement (L ICH ) to address the color cast of instances in the enhanced image. By summing the contribution of each point to the intensity of a given pixel in each of the three image channels, we construct three intensity histograms to represent the intensity distribution of each channel. Mask annotation is then used to obtain the indices of all instances in the target image, and the intensity histogram of each instance is compared between the enhanced image and the ground truth.

[0080] The grayscale histogram of each channel is obtained by kernel density estimation, and the specific expression for calculating the frequency of each subinterval of the histogram is:

[0081]

[0082] Among them, P C represents the instance set in the Cth channel, P Cn is the set of pixels of the nth instance in the Cth channel. We obtain these regions through the annotation of the dataset. The image is divided into K sub-intervals, each of which has a length of L = 1 / K, where k ranges from 1 to K. The center value of the kth interval is denoted as μ k =(k-0.5)×1 / K. For the Cth channel of the i-th instance, is the frequency of the kth interval. The histogram frequency is obtained by selecting the Sigmoid function as the kernel function and calculating the sum of the contributions of each pixel in the region to the grayscale intensity. In this expression, x j represents the j-th pixel in the patch, and σ is the scaling factor, which is set to 400.

[0083] The final loss function is obtained by calculating the L1 loss, the formula is:

[0084]

[0085] Where H(I') and H(I) represent the grayscale histograms of the processed and original images, respectively. This loss function measures the difference in grayscale distribution between corresponding instances in the processed and original images to guide model training and reduce color casts in foreground areas.

[0086] 3) Obtain enhanced results

[0087] The fused features serve as the input to the illumination enhancement submodule, which is supervised by the aforementioned loss function. Gradient descent is used to update the weights of the illumination enhancement submodule and the feature fusion module. When the loss function stops decreasing and approaches a plateau, an enhanced illumination map is obtained. Based on Retinex theory, the enhanced illumination map is element-wise multiplied with the reflectance map output by the reflectance map estimator to produce the enhanced image.

[0088] The overall framework based on the above steps is as follows Figure 1 As shown, the specific structure of the feature fusion module used in step 2 is as follows Figure 2 shown.

[0089] The embodiments of the present invention are as follows:

[0090] The first step is data preparation and evaluation protocol.

[0091] We trained our proposed joint optimization algorithm for low-light image enhancement and instance segmentation using 1,561 pairs of normal-light and low-light images from the LIS dataset, covering eight categories, including bicycles, cars, and motorcycles. We validated our algorithm using 669 low-light images from the LIS dataset. We used COCO mean average precision (mAP) to evaluate the effectiveness of instance segmentation in low-light images, facilitating comparison with the classic two-stage segmentation approach of first enhancing and then segmenting. To verify the effectiveness of our enhancement method, we report the peak signal-to-noise ratio (PSNR) and structural similarity index (SSIM) to assess the quality of the enhanced images.

[0092] The second step is parameter setting.

[0093] The network parameters involved include: (1) the input image is resized to [800, 1216] in the training phase, and the input image resolution is 800×1200 in the inference phase; (2) data enhancement is performed using operations such as random rotation and flipping; (3) pre-trained ResNet-50 is used as the detector backbone; (4) the batch size is 2; (5) SGD optimizer; the training is divided into two stages. In the first stage, the initial learning rate is set to 0.0001, and 200 epochs are trained. At this time, only the decomposition module is trained; in the second stage, the learning rate is initialized to 0.001, and 36 epochs are trained. At this time, the decomposition module, segmentation module and illumination adjustment module are trained together. The experiment is carried out using an RTX 3090 graphics card.

[0094] The third step is model training.

[0095] Due to the different difficulty of fitting the enhancement model and the segmentation model, the training is divided into two stages. In the first stage, pairs of low-light images and normal-light images are input into the decomposition model to obtain their respective illumination maps and reflectance maps, and the decomposition model parameters are optimized according to the decomposition loss function. In the second stage, the decomposition model, segmentation model and illumination adjustment model are trained together, and the reflectance map obtained by decomposition is used as the input of the instance segmentation module. The reflection map estimation module is further optimized based on the segmentation results. The intermediate features of the segmentation process and the illumination map obtained by decomposition are the same. Figure 1 They are both used as inputs to the feature fusion module to obtain the final adjusted illumination map, and the enhancement-related losses are used to optimize the parameters of the feature fusion module and the illumination enhancement module.

[0096] The fourth step is inference and prediction.

[0097] The algorithm's performance was verified on a public LIS instance segmentation dataset and a public LOL low-light enhancement dataset. During the inference and prediction phase, a low-light image was input into the network. The low-light image was first decomposed into an illumination map and a reflectance map. The reflectance map was fed into the segmentation module to generate segmentation results and extract instance-level semantic features. The semantic features and illumination map were fed into the illumination adjustment module, where they were fused by the semantic feature fusion model. The fused features were then fed into the illumination enhancement model to produce an adjusted illumination map. Finally, the adjusted illumination map was element-wise multiplied with the segmentation map to produce the final enhanced image.

[0098] Under the LIS dataset, a single RTX3090 graphics card is used for training:

[0099] (1) Due to the introduction of an additional instance-guided color histogram loss function, the color distribution of the enhanced result is more strictly supervised, and the enhanced result has less color deviation compared to ordinary enhancement methods;

[0100] (2) The structural similarity loss and L1 loss are introduced to design a decomposition loss function, which results in a reflection map with clear details, and thus makes the enhanced result with clear edges and rich details;

[0101] (3) By adopting the segmentation strategy of decomposition first and then segmentation, the segmentation accuracy is significantly improved compared with the conventional method of enhancement first and then segmentation. It also has a plug-and-play feature, and replacing different segmentation networks can obtain more accurate segmented images.

Claims

1. A joint optimization method for low-light image enhancement and instance segmentation based on Retinex theory, comprising the following steps: In step 1, a low-light image and a normal-light image of the same scene are simultaneously fed into a decomposition module. The decomposition module includes an illumination map estimator and a reflectance map estimator, which are used to estimate the illumination map and reflectance map, respectively. The illumination map estimator generates a single output feature layer, while the reflectance map estimator generates three output feature layers corresponding to the three RGB channels. A cross-validated loss function is used to ensure that the obtained reflectance and illumination maps conform to the Retinex model. The decomposition model's loss function is divided into two parts: reconstruction loss and reflectance map consistency loss. Step 2: Select the instance segmentation network CondInst as the network of the segmentation module, input the reflectance map into the instance segmentation module, dynamically generate the convolution kernel of the segmentation branch through the CondInst conditioning mechanism, intercept the intermediate features containing instance-level semantic information before the convolution of the segmentation branch and send them to the illumination adjustment module; In step 3, the intermediate features containing instance-level semantic information intercepted in the segmentation module and the illumination map estimated by the illumination estimator are sent to the illumination adjustment module for adjustment. The illumination adjustment module is divided into two parts: a semantic feature fusion submodule and an illumination enhancement submodule. First, the semantic feature fusion submodule constructed based on channel attention is used to fuse the intermediate features containing instance-level semantic information intercepted in the segmentation module with the illumination map to obtain fused features. Then, the fused features are enhanced by the illumination enhancement submodule to obtain an enhanced illumination map. Finally, the enhanced illumination map is element-wise multiplied with the reflectance map estimated by the reflectance map estimator, and an instance-guided color histogram loss function is introduced for supervision to obtain an enhanced image with less color deviation. The instance-guided color histogram loss function is calculated as follows: by counting the contribution value of each point to the specified pixel intensity in the three channels of the image and adding them together, three intensity histograms are constructed to characterize the intensity distribution of each channel. Then, the indexes of all instances in the target image are obtained by mask annotation, and the intensity histogram of each instance between the enhanced image and the true value are compared to obtain the instance-guided color histogram loss.

2. The joint optimization method for low-light image enhancement and instance segmentation according to claim 1, characterized in that: In step 1, the illumination map estimator applies a sigmoid activation function at the end of the module to normalize the output to the range of [0, 1]; the reflectance map estimator adopts the U-Net architecture.

3. The joint optimization method for low-light image enhancement and instance segmentation according to claim 1, characterized in that: In step 1, the reconstruction loss function is: L DEC =MSE(I,L·R)+MSE(R,I / L)+Smooth(L) Where I is the input image, i.e., a low-light image or a normal-light image, L is the illumination map, and R is the reflectance map. Smooth(·) represents the loss function that constrains the illumination map to be piecewise smooth; MSE(·) represents the mean squared error loss; MSE(I,L·R) represents the reconstruction loss, which ensures that the element-wise product of the illumination map and the reflectance map is consistent with the input image; and MSE(R,I / L) further constrains the reflectance map.

4. The joint optimization method for low-light image enhancement and instance segmentation according to claim 3, characterized in that: The specific mathematical expression of Smooth(·) is as follows: Where i and j in the subscript represent the i-th pixel and j-th pixel in the image, L represents the illumination map, N is the total number of pixels, and N h(i) Represents all pixels in the h×h area near i; calculate w based on the color matching of I in YUV space i,j ,σ r is the standard deviation, c represents the image channel; |·| represents the absolute value operation.

5. The joint optimization method for low-light image enhancement and instance segmentation according to claim 3, characterized in that: The consistency loss of the reflection map is defined as follows: L CR =SSIM(L N ·R L ,I N )+MSE(L N ·R L ,I N )+SSIM(R L ,R N )+||R L -R N || Among them, R L is the reflectance map of the low-light image, R N is the reflection image of the normal light image, I N is the normal-light image corresponding to the low-light image; SSIM(·) represents the structural similarity loss function between the two images; ||·|| represents the operation of taking the L1 norm.

6. The joint optimization method for low-light image enhancement and instance segmentation according to claim 3, characterized in that: First, the decomposition module is trained separately, and the parameters of the decomposition module are optimized through gradient backpropagation. When the loss stops decreasing and tends to be stable, the weights are saved, and the illumination map and reflectance map obtained at this time are used as part of the input of the subsequent segmentation module and illumination adjustment module, and the weights of the module are continued to be updated in subsequent stages.

7. The joint optimization method for low-light image enhancement and instance segmentation according to claim 1, characterized in that: The method for step 2 is as follows: 1) Obtain the intermediate features of the segmentation module containing instance-level semantic information: The instance segmentation network CondInst is selected as the network of the segmentation module, and the reflectance map is used as input to train the segmentation module. The intermediate features of its segmentation branch containing instance-level semantic information are intercepted as the input of the illumination adjustment module. 2) Get the segmentation results: The convolution kernels for different categories generated by the convolution kernel prediction branch of CondInst predict the outgoing features of the segmentation branch; DiceLoss is used as the mask loss to measure the difference between the predicted instance mask and the true mask, and the gradient descent method is used to update the weights of the segmentation module and the reflection map estimator.

8. The joint optimization method for low-light image enhancement and instance segmentation according to claim 1, characterized in that: In step 3, the method for constructing the semantic feature fusion submodule based on channel attention is as follows: Taking the mask features and illumination map as input, the transposed attention mechanism is used to fuse them into the output features. The volume fusion process is expressed as follows: Among them, F M represents the mask feature, I L Represents the low-light image illumination map; first, F M and I L Perform convolution and upsampling operations to align them in size and get F' M and I' L ; Then, F M 'and I L 'Reshape into a single channel feature by rearranging elements and To input into the attention mechanism; then, calculate the attention map: and Multiply by the transpose of and divide by Where C represents the number of channels, and then after the Softmax function, we get a C×C attention map with a value range of [0,1]. Multiply the transpose of and add the result to I' L Add them together and send them into the feedforward network (FFN) as input to finally obtain the fused feature F f .

9. The joint optimization method for low-light image enhancement and instance segmentation according to claim 1, characterized in that: In step 3, the illumination enhancement submodule selects the same structure as the illumination map estimator.

10. The joint optimization method for low-light image enhancement and instance segmentation according to claim 1, characterized in that: The method to construct the instance-guided color histogram loss function is as follows: The grayscale histogram of each channel is obtained by kernel density estimation, and the expression for calculating the frequency of each subinterval of the histogram is: Among them, P C represents the instance set in the Cth channel, P Cn is the pixel set of the nth instance in the Cth channel; the enhanced image and the normal light image whose histogram needs to be calculated are divided into K sub-intervals, the length of each interval is L = 1 / K, and k ranges from 1 to K; the center value of the kth interval is represented by μ k =(k-0.5)×1 / K; for the Cth channel of the i-th instance, is the frequency of the kth interval; by selecting the Sigmoid function as the kernel function, the sum of the contribution of each pixel in the region to the gray value intensity is calculated to obtain the histogram frequency; where x j represents the jth pixel in the patch, σ is the scaling factor; The instance-guided color histogram loss function is obtained by calculating the L1 loss, and the formula is: Here, H(I') and H(I) represent the grayscale histograms of the enhanced image and the normal light image, respectively. This loss function is used to measure the difference in grayscale distribution between the corresponding instances of the enhanced image and the normal light image to guide model training and reduce the color cast in the foreground area.

Citation Information

Patent Citations

  • Low-illumination underwater image enhancement method based on multi-scale detail enhancement

    CN112561804A

  • Low-light image enhancement method combining attention mechanism and Retinex model

    CN114266707A