TEDS low-illumination image enhancement method based on improved Retinexformer
By improving the Retinexformer model, combining the image preliminary enhancement module and the image denoising enhancement module, the problem of low-light image enhancement in TEDS data sets is solved, and effective enhancement of low-light images and improved accuracy of target detection is achieved.
Patent Information
- Application Number
- CN202411968185.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-06-06
AI Technical Summary
The existing Retinexformer model cannot be directly used for low-light image enhancement in the TEDS dataset, because there is no normal sample data corresponding to the low-light samples in the TEDS dataset that corresponds to the pixel-level corresponding relationships, and the model cannot be directly constructed using the loss of the L1 or L2 category during training.
Improve the Retinexformer model, and low-light image enhancement is performed through the image preliminary enhancement module and the image denoising enhancement module. The image preliminary enhancement module uses the convolution module to extract lighting features and enhance them by splicing low-light pictures and target lighting information. The image denoising enhancement module consists of the IGCAB module, which uses the cross attention mechanism and lighting information to perform image denoising and enhancement.
Through the improved Retinexformer model, the low-light images in the TEDS dataset can be effectively enhanced, making the enhanced images easier to detect the correct target by the target detection model, and improves the accuracy of EMU operation fault detection.
Smart Images

Figure CN120107094A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of computer vision image processing, and in particular relates to a TEDS low-light image enhancement method based on an improved Retinexformer. Background Art
[0002] The Trouble of moving EMU Detection System (TEDS) currently in use is equipped with a combination of area array and line array cameras to capture images of the bottom and side parts of the EMU. It also integrates automatic image recognition and automatic abnormality analysis functions, playing an important role in discovering and eliminating abnormalities in EMU parts. In the actual application of the current TEDS system, low-light images caused by environmental factors have obvious light and dark changes compared to normal-light images, and produce noise that affects the detection visual task. At present, when using target detection models such as Yolo trained with normal-light images to perform visual tasks on low-light images, the detection accuracy will decline significantly. Low-light images will cause target detection models such as Yolo to be unable to detect the target, which seriously affects the TEDS system's detection of EMU running faults. Therefore, it is necessary to first perform low-light enhancement operations on the acquired images, and the enhanced images are more likely to be detected by the target detection model as the correct target.
[0003] At present, low-light enhancement algorithms are divided into two categories according to different methods: enhancement methods and generative methods. The enhancement method is to enhance the image in a specific way to achieve the low-light enhanced image. This method is usually based on the retinal cortex theory (Reinex) theory. The leading method in the current field is the Retinexformer model. The generative method uses the regeneration of the image to regenerate the enhanced image. The leading methods in the current field are mainly designed based on generative adversarial models. However, since the images obtained after low-light enhancement of the TEDS dataset need to ensure that reliable enhanced normal sample data is provided for the detection task, the generative method cannot be used. In addition, there is no normal sample data in the TEDS dataset that forms pixel-level correspondence with the low-light samples. When training the model, it is impossible to directly use the L1 or L2 category loss to construct the pixel-level correspondence relationship, which makes the existing Retinexformer and other models cannot be used directly and need to be improved. Summary of the invention
[0004] The technical problem to be solved by the present invention is to provide a TEDS low-light image enhancement method based on an improved Retinexformer, which is used to enhance image data acquired from TEDS motor vehicle image data in a low-light environment.
[0005] In order to solve the above technical problems, the present invention provides a TEDS low-light image enhancement method based on an improved Retinexformer: collecting TEDS low-light pictures, and then combining the TEDS low-light pictures and target illumination information I T The data is imported into tensor form and normalized separately. The normalized TEDS low-light image and target illumination information I T As the input of the offline trained low-light enhancement model, the low-light enhancement model includes the output of the image preliminary enhancement module and the image denoising enhancement module. The image preliminary enhancement module outputs the preliminary enhanced image and the illumination feature map as the input of the image denoising enhancement module. The image denoising enhancement module outputs the enhanced image I E .
[0006] As an improvement of the TEDS low-light image enhancement method based on the improved Retinexformer of the present invention:
[0007] The image preliminary enhancement module includes an image stitching module, which stitches the TEDS low-light image and the target illumination information I T The channels are concatenated and then passed through two convolution modules to obtain the illumination feature I F , and then pass through a convolution module to obtain the enhanced illumination map I M ; The enhanced illumination map I M Multiply it with the TEDS low-light image and then add it to get the initial enhanced image I L .
[0008] As a further improvement of the TEDS low-light image enhancement method based on the improved Retinexformer of the present invention:
[0009] The image denoising and enhancement module is mainly composed of a group of IGCAB modules:
[0010] (1) Input preliminary enhanced image I L First through convolution, and then with the illumination feature I F Input an IGCAB module to get the first layer feature encoding processing result and
[0011] (2) The features and After downsampling using convolution, the result is input into two IGCAB modules to obtain the second-layer feature encoding processing result.
[0012] (3) The features and After downsampling using convolution, the result is input into two IGCAB modules to obtain the third-layer feature encoding processing result. and
[0013] (4) The features After upsampling using the deconvolution module and feature Splicing, and then use convolution to adjust the channel, and combine the result of channel adjustment with the feature Input two IGCAB modules together to obtain the second layer image feature decoding processing result
[0014] (5) The features After upsampling using the deconvolution module and feature Splicing, and then use convolution to adjust the channel, and combine the result of channel adjustment with the feature Input two IGCAB modules together to get the first layer image feature decoding processing result
[0015] (6) Features Use convolution to process the channels and get the enhanced image I E .
[0016] As a further improvement of the TEDS low-light image enhancement method based on the improved Retinexformer of the present invention:
[0017] The IGCAB module is:
[0018] First, the input image features and lighting characteristics Perform the first cross-attention mechanism to calculate the output illumination features in, After two different linear layers, Only one independent linear layer is passed through;
[0019] Then the illumination feature With image features Perform a second attention mechanism to calculate the output image features in, After two different linear layers, Only one linear layer is passed through.
[0020] As a further improvement of the TEDS low-light image enhancement method based on the improved Retinexformer of the present invention:
[0021] The offline training of the low-light enhancement model includes pre-training and enhancement training:
[0022] (1) Pre-training
[0023] The weight parameter of the low light enhancement model is randomly initialized to W 0 , and then use the public datasets LOLv1 and LOLv2 in turn to perform inner loop training and outer loop training on the model to obtain the model weight parameter W;
[0024] (2) Enhanced training
[0025] Use the weight parameter W to initialize the low-light enhancement model; collect real train pictures as training sets and test sets, and divide the samples in the training sets and test sets into normal light samples and low light samples after cropping and scaling, and then obtain the target light information I of the normal light samples T and image illumination information I N , the normal illumination sample and the target illumination information I T Synthesize into TEDS synthetic sample Then the TEDS synthesis sample and image illumination information I N Splicing is performed on the channel, the splicing result is input into the model for forward propagation, the model output is compared with the normal illumination sample S for loss calculation, the loss is back-propagated, and the model parameters are updated using the Adam optimizer.
[0026] As a further improvement of the TEDS low-light image enhancement method based on the improved Retinexformer of the present invention:
[0027] The training process of the data set LOLv1 is as follows:
[0028] The dataset LOLv1 includes the input image input image S LOLV1 and label L LOLV1 , use Gaussian filtering to filter the label L LOLV1 Get lighting information I LOLV1 , S LOLV1 with I LOLV1 After concatenation on the channel, the model is input for forward propagation, and the model output is consistent with the label L LOLV1 Calculate the loss, backpropagate the loss, and get the weight parameter Using the weight parameter After initializing the model again, perform an outer loop to obtain the model weight parameters
[0029]
[0030] The training process of the data set LOLv2 is as follows:
[0031] The dataset LOLv2 includes the input image input image S LOLV2 and label L LOLV2 , use Gaussian filtering to filter the label L LOLV2 Get lighting information I LOLV2 , S LOLV2 with I LOLV2 After concatenation on the channel, the model is input for forward propagation, and the model output is consistent with the label L LOLV2 Calculate the loss, backpropagate the loss, and get the weight parameter Using the weight parameter After initializing the model again, perform an outer loop to obtain the model weight parameter W 2 :
[0032]
[0033] As a further improvement of the TEDS low-light image enhancement method based on the improved Retinexformer of the present invention:
[0034] The normal light samples and low light samples are obtained as follows:
[0035] Calculate the average pixel value I of the train image avg , will I avg Samples with >50 are classified as normal light samples, and the rest are classified as low light samples.
[0036]
[0037] Where X is the input image, W and H are the width and height of the image respectively.
[0038] As a further improvement of the TEDS low-light image enhancement method based on the improved Retinexformer of the present invention:
[0039] The target illumination information I T The acquisition method is:
[0040]
[0041] Where j represents the number of iterations, represents the target illumination information obtained in the jth iteration. When j = 0, is an all-zero matrix, I j Represents the jth image X j Image lighting information:
[0042] I j =G(X j , s), s=80 (3).
[0043] As a further improvement of the TEDS low-light image enhancement method based on the improved Retinexformer of the present invention:
[0044] Synthesis of the TEDS synthesis sample The method is:
[0045]
[0046] Where S is the normal illumination sample, is the TEDS synthetic sample, μ N and is the distribution mean and variance of the average value of the illuminated pixels of the normal illumination sample, μ L and is the mean and variance of the distribution of the average light pixel values of the low light sample.
[0047] As a further improvement of the TEDS low-light image enhancement method based on the improved Retinexformer of the present invention:
[0048] The loss function of the low-light enhancement model is:
[0049] loss(x,y)=L TV (x)+L SSIM (x,y)+L ILLU (x,y) (13)
[0050] in,
[0051] Where x i,j Represents the value of the i-th row and j-th column of the input matrix x, H is the image height, and W is the image width;
[0052] L SSIM (x,y)=1-SSIM(x,y) (10)
[0053]
[0054] Where μ x ,μ y Represents the average value of x and y respectively, s x ,s y represents the variance of x and y, s xy represents the covariance of x and y;
[0055] c 1 =(K 1 L) 2 ,c 2 =(K 2 L) 2 ,,L is the gray level of the image;
[0056] L ILLU (x,y)=L1(G(x,σ),G(y,σ)), σ=80 (12).
[0057] The beneficial effects of the present invention are mainly reflected in:
[0058] 1. The present invention uses the TEDS data set and combines the illumination information of the target image in the data set as the prior information of the model for input, so that the model output result can be better close to the image domain where the target image is located, so that the enhanced image can better complete the detection task. At the same time, the EMA method is used to summarize the normal illumination sample data, and the illumination information of the normal illumination sample is obtained in a relatively simple way.
[0059] 2. The loss function used in the model training of the present invention is modified to TV Loss, SSIM Loss and Illumination Loss, so as to better achieve the purpose of keeping the image content and structure unchanged before and after enhancement, so that the enhanced image has higher credibility.
[0060] 3. The model of the present invention replaces the IGAB module with an IGCAB module designed based on a two-way cross-attention mechanism, and adds the target image illumination information as the prior information of the model. The IGCAB module decodes the image information according to the target illumination information. Since the illumination environment of TEDS image data is relatively stable, adding the target image illumination information as the prior information of the model can make the enhanced image closer to the illumination conditions of the image required by the target detection network, so as to increase the detection accuracy of the target detection network model. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] The specific implementation modes of the present invention are further described in detail below with reference to the accompanying drawings.
[0062] Figure 1 Schematic diagram of the offline training process of the low-light enhancement model of the present invention;
[0063] Figure 2 It is a structural schematic diagram of the low-light enhancement model based on the improved Retinexformer of the present invention;
[0064] Figure 3 for Figure 2 Schematic diagram of the structure of the IGCAB module;
[0065] Figure 4 This is an example of the comparison of target detection effects before and after low-light enhancement of TEDS data;
[0066] Figure 5 This is an example of the comparison of semantic segmentation effects before and after low-light enhancement of TEDS data. DETAILED DESCRIPTION
[0067] The present invention is further described below in conjunction with specific embodiments, but the protection scope of the present invention is not limited thereto:
[0068] Example 1: TEDS low-light image enhancement method based on improved Retinexformer, such as Figure 1 As shown in the figure, it aims to solve the problem that the images obtained from TEDS train graphics under low light conditions perform poorly in computer vision detection tasks, such as failing to correctly detect the specified target. To solve this problem, this paper improves the Retinexformer model and performs offline training. The offline training includes two stages. First, the improved Retinexformer is pre-trained on the LOL dataset, and then enhanced training is performed on the TEDS dataset constructed by field collection. The obtained low-light enhancement model is used to enhance the low-light images of the TEDS dataset, so that the enhanced images can achieve better detection results under the already trained target detection models (such as the Yolov7 network and the Deeplabv3+ network).
[0069] S1. EMU image acquisition and dataset construction
[0070] The high-speed imaging camera installed near the rails directly photographs the train and obtains real train images with a size of W×1024×3, where W is the image width, which is usually greater than 10000. After cropping, train images of the same size are obtained, which are 1024×1024×3, and then scaled to 512×512×3, thereby constructing a train image dataset.
[0071] S2. Sort and distinguish the normal illumination samples and low illumination samples of the EMU image dataset, obtain the target illumination information of the normal illumination samples, and synthesize TEDS synthetic samples for model training.
[0072] S2.1. Calculate the average pixel value of illumination distribution of each sample in the collected train image dataset
[0073] The Gaussian filter approximation method is used to obtain the illumination information of each sample in the train image dataset. The filter parameter s is 80, and the average illumination pixel value I of each sample is calculated based on the illumination information. avg The calculation formula is as follows:
[0074]
[0075] In the above formula, X is the input image, W and H are the width and height of the image respectively, and G represents the Gaussian filter approximation method.
[0076] S2.2. Distinguishing between normal illumination samples and low illumination samples in the train image dataset
[0077] Average value of illuminated pixels I avg The judgment threshold is 50, and the train data set is I avg Samples with values >50 are classified as normal light samples, and the rest are classified as low light samples.
[0078] S2.3. Obtain the target illumination information of the normal illumination samples of the train image dataset and calculate the target illumination information of each normal illumination sample.
[0079] In order to improve the algorithm calculation performance, the EMA (Exponential Moving Average, EMA) method is used to iterate all the normal illumination samples obtained in step S2.2, and the target illumination information I of the normal illumination samples of the dynamic vehicle image dataset is approximately calculated. T , the iterative process is:
[0080]
[0081] Where j represents the number of iterations, represents the target illumination information obtained in the jth iteration. When j = 0, is an all-zero matrix, used as the initial value. j represents the image illumination information obtained from the jth image. Assuming that the jth image is X j , X j The process of calculating image illumination information using Gaussian filtering approximation can be expressed as:
[0082] I j =G(X j , s), s=80 (3)
[0083] S2.4. Obtain TEDS synthetic samples for enhanced model training. Since the TEDS dataset does not have pixel-level corresponding input and labels, it is necessary to synthesize based on the existing normal illumination samples. The synthesis method is shown in formula (4):
[0084]
[0085] In the above formula, S is the normal illumination sample image, is a TEDS synthetic sample. N and is the statistical distribution mean and variance of the average value of the normal illumination sample illumination pixel, μ L and The TEDS data set obtained by the present invention is statistically analyzed to obtain the mean value and variance of the average value of the illumination pixel of the low-light sample. N =84.49, variance The mean μ of low-light samples L =28.71, variance
[0086] The obtained TEDS synthetic samples were divided into a training set and a test set in a ratio of 8:2 and used as the enhanced training for the model of the present invention.
[0087] S3. Design a low-light enhancement model based on the improved Retinexformer model
[0088] Retinexformer uses the self-attention-based IGAB (Illumination-Guided AttentionBlock, IGAB) module to build according to the UNet structure. After the image is input, it is downsampled through two encoding layers, and the number of IGAB modules used in each layer is 1 and 2 respectively, to obtain the underlying features of the image; the underlying features of the image are upsampled after two IGABs and enter the decoding layer, and the number of IGAB modules used in each layer is 2 and 2 respectively. Skip connections are used between the encoding layer and the decoding layer for feature transfer.
[0089] The low-light enhancement model based on the improved Retinexformer model of the present invention (hereinafter referred to as the low-light enhancement model) is as follows: Figure 2 As shown in the figure, the IGAB module in the Retinexformer model is replaced with the IGCAB (Illumination-Guided Cross-Attention Block, IGCAB) module designed based on a two-way cross attention mechanism, and the target image illumination information obtained in step S2.3 is added as the prior information of the model. The IGCAB module decodes the image information according to the target illumination information. Since the illumination environment of TEDS image data is relatively stable, adding the target image illumination information as the prior information of the model can make the enhanced image closer to the illumination conditions of the image required by the target detection network, so as to increase the detection accuracy of the target detection network model.
[0090] S3.1. Build a preliminary image enhancement module
[0091] The image preliminary enhancement module includes an image stitching module, three convolution modules and an image operation module. It takes the original information of the image sample and the target illumination information as input, and outputs a preliminary enhanced image and an illumination feature map.
[0092] (1) Build a feature map stitching module to stitch the original information of the image sample and the target illumination information obtained in step S2.3 on the channel to obtain a 6-channel illumination feature map.
[0093] (2) Build an input convolution module and transform the 6-channel illumination feature map into a high-dimensional feature vector through a convolution module with a convolution kernel size of 1×1.
[0094] (3) Build a feature extraction convolution module and input the high-dimensional feature vector into the feature extraction convolution module with a convolution kernel size of 5×5 to obtain the illumination feature I F (Illumination Feature).
[0095] (4) Build the output convolution module to transform the illumination feature I F Input the convolution module with a convolution kernel size of 1×1 again to obtain the enhanced illumination map I M (Illumination Map).
[0096] (5) Build an image operation module, multiply the enhanced illumination map with the input original image, and then add them together to obtain the initial enhanced image I L (Light-up image).
[0097] The above process (1)-(5) can be described in mathematical form as follows:
[0098] I F , I M =f(I T ,X) (5)
[0099] I L =I M ×X+X (6)
[0100] where X∈R H′W′3 is the input image, I T ∈R H′W′3 is the target illumination information, f(x, y) is the full convolutional neural network, I F Illumination Feature, I M Represents the enhanced illumination map (Illumination Map), I L Represents the preliminary enhanced image (Light-up image).
[0101] S3.2. Build an image denoising and enhancement module, which is mainly composed of 9 IGCAB modules. The preliminary enhanced image and illumination features output by the image preliminary enhancement module are used as input to output the final low-light enhanced image. The process is as follows:
[0102] I E =d(I F , I L ) (7)
[0103] In the above formula, I E ∈R H′W′3 , represents the enhanced image obtained. d(I F , I L ) represents the model decoding process. The decoding process represents the entire decoding part model structure which is composed of 9 IGCABs as the backbone and a small number of convolutional layers.
[0104] S3.2.1. Build IGCAB module
[0105] The IGCAB module is built based on the dual-path cross-attention mechanism. The module has two inputs and two outputs are obtained through cross-attention calculation (such as Figure 3 ). In the present invention, the image feature (Image Feature Input) and the illumination feature (Illumination Feature Input) are used as the input of the IGCAB module, and the processed image feature output (Image Feature Output) and the illumination feature output (Illumination Feature Output) are obtained. This process will make full use of the illumination information to perform illumination restoration on the image. For the convenience of describing the process, the input image feature (Image Feature Input) is recorded as The input illumination feature is The output image feature (Image Feature Output) is The output illumination feature (Illumination FeatureOutput) is The IGCAB module uses a double cross attention mechanism for calculation, and its calculation process is as follows:
[0106] enter and Perform the first cross-attention mechanism calculation. After two different linear layers, we get the result V 1 and K 1 , Then after an independent linear layer, the result Q is obtained1 .Q 1 With K 1 After matrix multiplication, Softmax processing is performed, and the result is the same as V 1 Matrix product is performed again, and normalization and multi-layer perceptron MLP processing are performed. The result is the output illumination feature
[0107] Use the obtained lighting features With image features Perform the second attention mechanism calculation. After two different linear layers, we get the result V 2 and K 2 , Then after an independent linear layer, the result Q is obtained 2 .Q 2 With K 2 After matrix multiplication, Softmax processing is performed, and the result is the same as V 2 Perform matrix product again, normalize and process with multi-layer perceptron MLP, and the result is the output image feature. The process can be written as:
[0108]
[0109] S3.2.2. Build the image denoising and enhancement module according to the UNet model structure: the depth is 3, that is, a total of three layers of feature processing. The model is divided into two parts: encoding and decoding. Encoding is mainly for downsampling features, while decoding is for upsampling. The model initially enhances the input image I L Illumination Characteristics I F Downsampling encoding is performed, but only upsampling decoding is performed on the encoded image features. The model processing process is as follows (the top semicircle of the variable pointing upwards indicates the encoding result, and the semicircle pointing downwards indicates the decoding result):
[0110] (1) First, use 3×3 convolution to initially enhance the input image I L The data is preprocessed, and then the preprocessed feature image and illumination feature I F Input an IGCAB module to get the first layer feature encoding processing result and
[0111] (2) The first layer of feature encoding processing and Use 4×4 convolution for downsampling with a convolution step of 2. The result is then input into two IGCAB modules to obtain the second-layer feature encoding processing result. and
[0112] (3) The second layer of feature encoding processing and The same 4×4 convolution is used for downsampling, and the convolution step size is 2. The result is input into two IGCAB modules to obtain the third-layer feature encoding processing result. and
[0113] (4) The third layer image coding feature processing results Use a 2×2 deconvolution module for upsampling with a step size of 2. The upsampling result is consistent with the image feature result processed by the second layer. The concatenation is performed and the channel is adjusted using 1×1 convolution. The channel-adjusted result is combined with the illumination feature result processed by the second layer. Input two IGCAB modules together to obtain the second layer image feature decoding processing result
[0114] (5) Decoding the second layer image features Use a 2×2 deconvolution module for upsampling with a step size of 2. The upsampling result is consistent with the image feature result processed in the first layer. Splicing is performed and channel adjustment is performed using 1×1 convolution. The channel-adjusted result is combined with the illumination feature result processed by the first layer. Input two IGCAB modules together to get the first layer image feature decoding processing result
[0115] (6) Decoding the first layer image features Use 3×3 convolution for channel processing to get the enhanced image I E .
[0116] S4. Pre-training using the public dataset LOL
[0117] LOL stands for Low-Light dataset. It uses the meta-learning method to pre-train the model, so that the model has the initial ability to enhance image illumination, improve the sensitivity to illumination enhancement task data, and reduce the dependence on TEDS samples. LOL consists of low-light and normal-light image pairs. LOLv1 comes from Deepretinex decomposition for low-light enhancement (Chen Wei, Wenjing Wang, WenhanYang, and Jiaying Liu. BMVC, 2018), and LOLv2 comes from Sparse gradient regularized deepretinex network for robust low-light image enhancement (Wenhan Yang, WenjingWang, Haofeng Huang, Shiqi Wang, and Jiaying Liu. TIP, 2021.)
[0118] The pre-training process is divided into an inner loop and an outer loop. The process is as follows:
[0119] S4.1. Obtain the public datasets LOLv1 and LOLv2, and build a low-light enhancement model according to step S3, and randomly initialize the model parameters. The model parameters after random initialization are recorded as W 0 .
[0120] S4.2. Set the loss function used for model training. The loss function used in training is composed of the sum of TV Loss, SSIM Loss and Illumination Loss. Assume that x is the model output and y is the label. The specific process is as follows:
[0121] (1) TV Loss (Total Variation Loss) is the total variation loss, whose main function is to reduce image noise, denoted by L TV (x). The full-resolution loss actually sums the gradients in the pixel domain, and the loss can be written as:
[0122]
[0123] Where x i,j Represents the value of the i-th row and j-th column of the input matrix x, H is the image height, and W is the image width.
[0124] (2) SSIM Loss (Structural Similarity Loss) is the structural similarity loss, expressed as L SSIM (x,y) is:
[0125] L SSIM (x,y)=1-SSIM(x,y) (10)
[0126] It is used to measure the structural similarity between x and y to reduce the gap between the generated image and the target image. The structural similarity calculation can be written as:
[0127]
[0128] Where μ x ,μ y Represents the average value of x and y respectively. x ,s y represents the variance of x and y, s xy represents the covariance of x and y.
[0129] c 1 =(K 1 L) 2 ,c 2 =(K 2 L) 2 , used to stabilize the SSIM calculation results, L is the grayscale level of the image, the default is K 1 =0.01,K 1 = 0.03. If Pytorch is used for model training, ssim can be directly called from the pytorch_msssim library to calculate the SSIM(x,y) indicator.
[0130] (3) Illumination Loss is the illumination loss, expressed as L ILLU (x,y) is:
[0131] L ILLU (x, y) = L1 (G (x, σ), G (y, σ)), σ = 80 (12)
[0132] Used to determine the similarity of illumination distribution between x and y. Where G() represents Gaussian filtering.
[0133] (4) The total model loss is the sum of the above three, that is,
[0134] loss(x,y)=L TV (x)+L SSIM (x,y)+L ILLU (x,y) (13)
[0135] S4.3. Use the public dataset LOLv1 to perform inner loop training on the model. The data of the public dataset includes the input image S LOLV1 (Low Light) and Label L LOLV1 (normal lighting image), the variable subscript indicates the class name of the dataset. Set the optimizer and hyperparameters for model training. The optimizer used for model training is Adam, the learning rate is set to 2e-4, and the momentum value is set to 0.9. The model training process is divided into the following steps:
[0136] (1) Use Python to import LOLv1 data as tensor data and divide all data by 255 for normalization.
[0137] (2) The model input data is divided into two parts: input image S LOLV1 and lighting information I LOLV1 I LOLV1 Gaussian filtering G() can be used to filter the label L LOLV1 To obtain an approximate value, the calculation method is as follows:
[0138] I LOLV1 =G(L LOLV1 , s), s=80 (14)
[0139] S LOLV1 with I LOLV1 Splicing is performed on the channel, and the splicing result is input into the model for forward propagation. The model output is consistent with the label L LOLV1 Calculate the loss. Back-propagate the loss and use the Adam optimizer to update the model parameters. Repeat this process 10 times to train the model and get the weights.
[0140] S4.4. Use the model weights obtained in the previous step After initializing the model weights, perform an outer loop to update the model parameters and obtain the model parameters W 1 , the update process is as follows:
[0141]
[0142] S4.5. Use W 1 Initialize the model weights and repeat S4.3 to train the model in the inner loop using the public dataset LOLv2. The obtained model weights are named
[0143] S4.6. Repeat step S4.4 using the public dataset LOLv2 and use the model weights obtained in the previous step After initializing the model weights, perform an outer loop to update the model parameters and obtain the model parameters W 2The weight W is obtained 2 Pre-training parameters for subsequent model training.
[0144]
[0145] S5. Use TEDS synthetic samples to enhance the low-light enhancement model training.
[0146] The TEDS synthetic sample generated in step 2 The input model is forward propagated, the model output result is compared with the normal illumination image S for loss calculation, and the calculated loss is back-propagated to update the model parameters. The training process of the model of the present invention is as follows: Figure 1 shown.
[0147] S5.1. The loss function of training is the same as that of pre-training, that is, formula (13).
[0148] S5.2. Use W obtained from step 4 training 2 The model weights are imported as the enhanced training parameters of the model. The optimizer and hyperparameters of the model enhanced training are set. The optimizer used in the enhanced training is Adam, the learning rate is set to 2e-4, and the momentum value is set to 0.9. The model enhanced training process is:
[0149] (1) Use Python to import the TEDS synthetic sample obtained in step S2.4 and the target illumination information obtained in step S2.3 into data in the form of tensors, and divide all data by 255 for normalization.
[0150] (2) During enhanced training, the input data is divided into two parts: TEDS synthetic samples and the image illumination information I of the normal illumination sample S N , where the image illumination information calculation method is:
[0151] I N =G(S, s), s = 80 (16)
[0152] The input TEDS synthesis sample Image illumination information I with normal illumination sample S N Splicing is performed on the channel, the splicing result is input into the model for forward propagation, and the model output is compared with the normal illumination sample S for loss calculation.
[0153] The loss is back-propagated and the model parameters are updated using the Adam optimizer.
[0154] S6. Use the low-light enhancement model trained by enhanced training to perform TEDS low-light image enhancement, and verify the effect of low-light enhancement on downstream target detection tasks (Yolov7 and Deeplabv3+).
[0155] S6.1. Use the low-light image enhancement model trained in step S5 to enhance the collected low-light images. The user only needs to input the sample to be enhanced and the target illumination information used in the second stage of enhancement training into the model, and the model output result is the low-light enhanced sample. The specific steps are:
[0156] (1) The model input data consists of two parts: the sample to be enhanced and the target illumination information I obtained in S2.3 T Using Python will require enhanced sample and target lighting information I T Import the data in the form of tensors and divide the data by 255 for normalization.
[0157] (2) The normalized TEDS low-light image tensor and target illumination information I T The tensors are concatenated on the channels, and the concatenated results are input into the low-light enhancement model trained in step S5. The output of the model is the enhanced image.
[0158] S6.2. Verify the effectiveness of the low-light image enhancement model on the TEDS EMU dataset
[0159] The low-light enhancement model of the present invention can effectively enhance the low-light images in the TEDS dataset under existing conditions to improve the detection effect of downstream target detection tasks. In this experiment, the RetinexFormer model, the RetinexFormer+IGCAB model (indicates that the IGAB module in the Retinexformer model is replaced with an IGCAB module designed based on a two-way cross attention mechanism), and the low-light model of the present invention (RetinexFormer+IGCAB model+public dataset pre-training step) are respectively trained offline on the same TEDS dataset (the training set and test set of TEDS synthetic samples obtained in step 2 of Example 1), and the offline trained RetinexFormer model, RetinexFormer+IGCAB model, and the low-light enhancement model of the present invention are respectively used to enhance the low-light samples in the TEDS dataset to obtain enhanced sample set 1, enhanced sample set 2, and enhanced sample set 3.
[0160] The enhanced sample set 1, enhanced sample set 2, enhanced sample set 3 and enhanced sample set 4 are respectively input into the Yolov7 and Deeplabv3 models to obtain the detection results shown in Table 1.
[0161] Table 1 TEDS low light comparison test data
[0162]
[0163] As can be seen from Table 1, it is obvious that the accuracy of the target detection model using the low-light samples after image enhancement of the present invention has been significantly improved, which proves the effectiveness of the present invention. At the same time, the two-stage training method of pre-training + enhanced training adopted by the present invention has better results than training on the TEDS dataset alone.
[0164] The visualization results are as follows Figure 4 (Device Identification) and Figure 5 (Device segmentation) as shown.
[0165] Finally, it should be noted that the above examples are only some specific embodiments of the present invention. Obviously, the present invention is not limited to the above embodiments, and there are many variations. All variations that can be directly derived or associated with the content disclosed by a person skilled in the art should be considered as the protection scope of the present invention.
Claims
1. TEDS low-light image enhancement method based on improved Retinexformer, characterized by: Collect TEDS low-light images, and then combine the TEDS low-light images and target illumination information T The data is imported into tensor form and normalized separately. The normalized TEDS low-light image and target illumination information I T As the input of the offline trained low-light enhancement model, the low-light enhancement model includes the output of the image preliminary enhancement module and the image denoising enhancement module. The image preliminary enhancement module outputs the preliminary enhanced image and the illumination feature map as the input of the image denoising enhancement module. The image denoising enhancement module outputs the enhanced image I E .
2. The TEDS low-light image enhancement method based on improved Retinexformer according to claim 1, characterized in that: The image preliminary enhancement module includes an image stitching module, which stitches the TEDS low-light image and the target illumination information I T The channels are concatenated and then passed through two convolution modules to obtain the illumination feature I F , and then pass through a convolution module to obtain the enhanced illumination map I M ; The enhanced illumination map I M Multiply it with the TEDS low-light image and then add it to get the initial enhanced image I L .
3. The TEDS low-light image enhancement method based on improved Retinexformer according to claim 2, characterized in that: The image denoising and enhancement module is mainly composed of a group of IGCAB modules: (1) Input preliminary enhanced image I L First through convolution, and then with the illumination feature I F Input an IGCAB module to get the first layer feature encoding processing result and (2) The features and After downsampling using convolution, the result is input into two IGCAB modules to obtain the second-layer feature encoding processing result. and (3) The features and After downsampling using convolution, the result is input into two IGCAB modules to obtain the third-layer feature encoding processing result. and (4) The features After upsampling using the deconvolution module and feature Splicing, and then use convolution to adjust the channel, and combine the result of channel adjustment with the feature Input two IGCAB modules together to obtain the second layer image feature decoding processing result (5) The features After upsampling using the deconvolution module and feature Splicing, and then use convolution to adjust the channel, and combine the result of channel adjustment with the feature Input two IGCAB modules together to get the first layer image feature decoding processing result (6) Features Use convolution to process the channels and get the enhanced image I E .
4. The TEDS low-light image enhancement method based on improved Retinexformer according to claim 3 is characterized in that: The IGCAB module is: First, the input image features and lighting characteristics Perform the first cross-attention mechanism to calculate the output illumination features in, After two different linear layers, Only one independent linear layer is passed through; Then the illumination feature With image features Perform a second attention mechanism to calculate the output image features in, After two different linear layers, Only one linear layer is passed through.
5. The TEDS low-light image enhancement method based on improved Retinexformer according to claim 4, characterized in that: The offline training of the low-light enhancement model includes pre-training and enhancement training: (1) Pre-training The weight parameter of the low-light enhancement model is randomly initialized to W0, and then the public datasets LOLv1 and LOLv2 are used in turn to perform inner loop training and outer loop training on the model to obtain the model weight parameter W; (2) Enhanced training Use the weight parameter W to initialize the low-light enhancement model; collect real train pictures as training sets and test sets, and divide the samples in the training sets and test sets into normal light samples and low light samples after cropping and scaling, and then obtain the target light information I of the normal light samples T and image illumination information I N , the normal illumination sample and the target illumination information I T Synthesize into TEDS synthetic sample Then the TEDS synthesis sample and image illumination information I N Splicing is performed on the channel, the splicing result is input into the model for forward propagation, the model output is compared with the normal illumination sample S for loss calculation, the loss is back-propagated, and the model parameters are updated using the Adam optimizer.
6. The TEDS low-light image enhancement method based on improved Retinexformer according to claim 5, characterized in that: The training process of the data set LOLv1 is as follows: The dataset LOLv1 includes the input image input image S LOLV1 and label L LOLV1 , use Gaussian filtering to filter the label L LOLV1 Get lighting information I LOLV1 , S LOLV1 with I LOLV1 After concatenation on the channel, the model is input for forward propagation, and the model output is consistent with the label L LOLV1 Calculate the loss, backpropagate the loss, and get the weight parameter Using the weight parameter After initializing the model again, perform an outer loop to obtain the model weight parameters The training process of the data set LOLv2 is as follows: The dataset LOLv2 includes the input image input image S LOLV2 and label L LOLV2 , use Gaussian filtering to filter the label L LOLV2 Get lighting information I LOLV2 , S LOLV2 with I LOLV2 After concatenation on the channel, the model is input for forward propagation, and the model output is consistent with the label L LOLV2 Calculate the loss, backpropagate the loss, and get the weight parameter Using the weight parameter After initializing the model again, perform an outer loop to obtain the model weight parameter W2:
7. The TEDS low-light image enhancement method based on improved Retinexformer according to claim 6, characterized in that: The normal light samples and low light samples are obtained as follows: Calculate the average pixel value I of the train image avg , will I avg Samples with >50 are classified as normal light samples, and the rest are classified as low light samples. Where X is the input image, W and H are the width and height of the image respectively.
8. The TEDS low-light image enhancement method based on improved Retinexformer according to claim 7, characterized in that: The target illumination information I T The acquisition method is: Where j represents the number of iterations, represents the target illumination information obtained in the jth iteration. When j = 0, is an all-zero matrix, I j Represents the jth image X j Image lighting information: I j =G(X j ,s),s=80 (3)。 9. The TEDS low-light image enhancement method based on improved Retinexformer according to claim 8, characterized in that: Synthesis of the TEDS synthesis sample The method is: Where S is the normal illumination sample, is the TEDS synthetic sample, μ N and is the distribution mean and variance of the average value of the illuminated pixels of the normal illumination sample, μ L and is the mean and variance of the distribution of the average light pixel values of the low light sample.
10. The TEDS low-light image enhancement method based on improved Retinexformer according to claim 9, characterized in that: The loss function of the low-light enhancement model is: loss(x,y)=L TV (x)+L SSIM (x,y)+L ILLU (x,y) (13) in, Where x i,j Represents the value of the i-th row and j-th column of the input matrix x, H is the image height, and W is the image width; L SSIM (x,y)=1-SSIM(x,y) (10) Where μ x ,μ y Represents the average value of x and y respectively, s x ,s y represents the variance of x and y, s xy represents the covariance of x and y; c1 = (K1L) 2 ,c2=(K2L) 2 ,,L is the gray level of the image; L ILLU (x,y)=L1(G(x,σ),G(y,σ)),σ=80 (12)。