A Transformer-based multi-scale optimization low-light image enhancement method
Through a multi-scale optimization low-illumination image enhancement method based on Transformer, combined with multi-task recovery and enhancer network, attention module and color correction module, the problems of brightness, color difference, noise and contrast in low-illumination image enhancement are solved, and the overall improvement of image quality and the improvement of downstream task performance are achieved.
Patent Information
- Application Number
- CN202210828621.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-13
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2042-07-13
AI Technical Summary
Existing low-illumination image enhancement methods are difficult to comprehensively solve problems such as brightness, chromatic aberration, noise and contrast, resulting in image details loss and noise interference, affecting visual perception and downstream task performance.
A multi-scale optimization low-illumination image enhancement method based on Transformer was designed, and by building multi-task recovery and enhancer networks and multi-scale optimization networks, combining attention modules and color correction modules, enhancement for different scales and problems.
This method can significantly improve the overall and detailed quality of low-illumination images, comprehensively solve problems such as brightness, chromatic aberration, noise and contrast, improve visual perception experience and improve downstream tasks performance.
Smart Images

Figure CN115205147B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of image and video processing and computer vision, and in particular to a Transformer-based multi-scale optimized low-illumination image enhancement method. Background Art
[0002] Images taken under insufficient lighting, backlight or non-uniform lighting conditions lose detail information, color information, etc., have low contrast, and result in missing useful information. On the other hand, there is more noise interference in dark areas, which disrupts the image information. The low-light image enhancement process often faces additional problems such as insufficient brightness, image noise, color deviation, and artifacts, which often affect human visual perception experience and reduce the detection performance of downstream tasks such as target detection, face recognition, and semantic segmentation. Therefore, in recent years, low-light image enhancement methods have been continuously developed to solve the above problems.
[0003] Traditional methods are mainly divided into two categories: one is the histogram equalization method, which uses histogram equalization to expand the dynamic range of the image. It has a small number of learnable parameters and lacks effective modeling of image structure information; the other is the low-light image enhancement method based on Retinex, which decomposes the image into a reflection component that determines the inherent properties of the image and an illumination component that represents the characteristics of the ambient light. Since the reflection component is directly used as the enhanced image, the noise problem is difficult to solve, which will cause loss of details and color damage, and cannot properly enhance images with relatively high dynamic range. In recent years, deep learning has developed rapidly, and many deep learning-based methods have been proposed. Some algorithms that focus on color difference restoration ignore noise processing, resulting in obvious noise in the enhancement results, disrupting the observable information of the image and having a strong interference effect on downstream tasks; some deep learning methods based on Retinex theory begin to decompose the noise component and remove the noise, but often ignore detail retention, and remove useful information of the image while removing noise. However, low-light image enhancement, as an upstream task, needs to be applied to multiple downstream tasks such as detection and recognition. It is necessary to comprehensively consider image problems and improve image quality from multiple aspects. On the other hand, the quality of human visual perception is also affected by comprehensive image factors, rather than a single factor.
[0004] Most existing methods can only propose solutions to one or two problems, but cannot comprehensively pay attention to more problems, resulting in the one-sidedness of the methods. It is difficult to fundamentally improve human visual perception and effectively apply it to downstream tasks. The present invention designs a multi-scale optimized low-light image enhancement method based on Transformer. First, according to the severity of the low-light image in various problems, a multi-task recovery and enhancement sub-network based on Transformer is constructed to perform preliminary comprehensive enhancement on the brightness, color difference, noise, contrast, etc. of the low-light image. Then, a multi-scale optimized low-light image enhancement network based on Transformer is designed to enhance the image at different scales to ensure that the overall image and image details are taken into account. In addition, an attention module and a color correction module are set in each branch to further enhance the color difference and contrast problems in a targeted manner. Summary of the invention
[0005] The purpose of the present invention is to provide a Transformer-based multi-scale optimized low-light image enhancement method, which combines a multi-scale network structure and a comprehensive enhancement of multiple targets, and is conducive to significantly improving the performance of low-light image enhancement.
[0006] To achieve the above object, the technical solution of the present invention is: a multi-scale optimized low-light image enhancement method based on Transformer, comprising the following steps:
[0007] Step A: preprocess the data by first pairing the data, then perform data segmentation and data enhancement.
[0008] Step B: Build a Transformer-based multi-task restoration and enhancement sub-network, which consists of the head, body parts, and tail, pre-trained using a multi-task dataset and fine-tuned using a synthetic low-light image dataset;
[0009] Step C, design a Transformer-based multi-scale optimized low-light image enhancement network, which consists of a multi-scale shallow feature extraction sub-network, a multi-scale feature fusion sub-network, a Transformer-based low-light image feature enhancement sub-network and a multi-scale low-light image enhancement sub-network;
[0010] Step D: Design a loss function based on the network structure to guide the parameter optimization of the network model;
[0011] Step E: training the network with the image block data, optimizing the network model parameters, and converging to a balanced state;
[0012] Step F: Cut the low-light image into blocks, and then input the image blocks into the trained image enhancement network respectively. After obtaining the corresponding enhanced image blocks, the image blocks are spliced into an enhanced complete image.
[0013] Compared with the prior art, the present invention has the following beneficial effects: The present invention constructs a Transformer-based multi-task restoration and enhancement network according to the severity of various problems in low-light images, which can comprehensively enhance the brightness, color difference, noise, contrast, etc. in low-light images. A Transformer-based multi-scale optimized low-light image enhancement network is designed, which can effectively pay attention to both the overall image and image details, and an attention module and a color correction module are set in each branch to further enhance the color difference and contrast problems in a targeted manner. Unlike other methods that only solve one or two problems, the present invention can comprehensively solve the degradation problems in low-light images. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 It is a flow chart for realizing the method of the present invention.
[0015] Figure 2 It is a structural diagram of the Transformer-based multi-task recovery and enhancement sub-network in an embodiment of the present invention.
[0016] Figure 3 It is a structural diagram of a Transformer-based multi-scale optimized low-light image enhancement network in an embodiment of the present invention.
[0017] Figure 4 is a structural diagram of the attention module in an embodiment of the present invention.
[0018] Figure 5 is a structural diagram of a color correction module in an embodiment of the present invention. DETAILED DESCRIPTION
[0019] The technical solution of the present invention is described in detail below in conjunction with the accompanying drawings.
[0020] The present invention provides a Transformer-based multi-scale optimization low-light image enhancement method. Figure 1 As shown, the following steps are included:
[0021] Step A: Preprocess the data by first pairing the data, then perform data segmentation and data enhancement.
[0022] Step B: Construct a Transformer-based multi-task restoration and enhancement sub-network consisting of the head, body parts, and tail, pre-trained using a multi-task dataset and fine-tuned using a synthetic low-light image dataset;
[0023] Step C: Design a Transformer-based multi-scale optimized low-light image enhancement network, which consists of a multi-scale shallow feature extraction sub-network, a multi-scale feature fusion sub-network, a Transformer-based low-light image feature enhancement sub-network, and a multi-scale low-light image enhancement sub-network;
[0024] Step D: Design the loss function according to the network structure to guide the parameter optimization of the network model;
[0025] Step E: Use image block data to train the network, optimize network model parameters, and converge to a balanced state;
[0026] Step F: The low-light image is cut into blocks, and then the image blocks are respectively input into the trained image enhancement network. After obtaining the corresponding enhanced image blocks, the image blocks are spliced into an enhanced complete image.
[0027] Furthermore, the step A comprises the following steps:
[0028] Step A1: Pairing the real low-light image with the corresponding label image;
[0029] Step A2: All real low-light images are randomly cut into blocks three times in different ways, and the paired label images of the real low-light image I are also cut into blocks in the same way. The random cutting process is as follows: Given a real low-light image I, the image size is h×w×3, and the size of the cut image blocks is q×q×3, where h and w represent the height and width of the low-light image I, respectively, and q represents the height and width of the cut image block.
[0030] Step A3: Perform the same random horizontal flipping, vertical flipping, and rotation operations on the paired image blocks for data enhancement to obtain a real low-light image dataset.
[0031] Step A4: Preprocess the noisy image and the labeled image, the underwater image and the labeled image, the low-resolution image and the labeled image, and the synthetic low-light image and the labeled image in the same manner as steps A1 to A3, respectively, to obtain a noisy image dataset, an underwater image dataset, a low-resolution image dataset, and a synthetic low-light image dataset.
[0032] Furthermore, the step B comprises the following steps:
[0033] Step B1: Construct a Transformer-based multi-task recovery and enhancement sub-network, which consists of a head, a body, and a tail. The network input passes through the head, the body, and the tail in sequence to obtain the output result. The parameters of the body are shared between tasks, but the parameters of the head and the tail are not shared.
[0034] Step B2: Use the multi-task dataset to pre-train the network model in step B1, and use the network objective loss function to optimize the model parameters. Here, L1 loss is used, and the calculation method is as follows:
[0035]
[0036] in, represents the corrupted image of the ith task, Indicates that the input is The results obtained by Transformer-based multi-task recovery and enhancement sub-network, I gt Indicates a damaged image The corresponding label image,||·|| 1 represents L1 loss, and n represents n tasks, including image denoising, underwater image enhancement, image super-resolution, and synthetic low-light image enhancement.
[0037] During the training process, each batch randomly selects a training from n-1 tasks except the synthetic low-light enhancement task. Each batch contains the same number of pairs of damaged images and labeled images randomly selected from the corresponding data set constructed in step A4. The damaged images are input into the network described in step B1 to obtain the result images. According to the target loss function of the network, the gradient of each parameter in the network is calculated using the back propagation method, and the parameters of the network are updated using the stochastic gradient descent method until the number of iterations reaches the threshold, the trained model is saved, and the network training process is completed.
[0038] Step B3: Use the synthetic low-light image dataset to fine-tune the network model parameters obtained in step B2. Load the network model pre-trained in B2 as a pre-trained model and use it as the initial network parameters for the subsequent fine-tuning of network parameters. In each batch, the same number of synthetic low-light images are randomly selected from the synthetic low-light image dataset and input into the network described in B1 to obtain enhanced images. According to the target loss function of the network, the back propagation method is used to calculate the gradient of each parameter in the network, and the parameters of the network are updated using the stochastic gradient descent method until the number of iterations reaches the threshold, the trained model is saved, and the network training process is completed. Finally, the model parameters of the Transformer-based multi-task recovery and enhancement sub-network are obtained, which will be used as a pre-trained model in subsequent steps.
[0039] Furthermore, the step B1 comprises the following steps:
[0040] Step B11: Construct the head structure. The head structure of each task consists of a convolutional layer and two residual structures. The damaged image of the i-th task After the head structure, the output feature map is obtained The calculation process is as follows:
[0041] Res i (x) = Cinv i (ReLU(Conv i (x)))+x,i=1,2,…n
[0042]
[0043] Among them, i represents the i-th task, n represents a total of n tasks, and Res i (·) represents the residual structure corresponding to the i-th task, x is the input of the residual structure, Conv i It represents the convolution layer corresponding to the i-th task, and ReLU is the activation function.
[0044] Step B12: Construct the body part structure. The body part adopts the encoder-decoder structure of Vision Transformer. The parameters of the body parts corresponding to all tasks are shared and the structure is the same. Therefore, the feature maps obtained by passing the head of different task images are processed in the same way. Unified as F corrupted First, the feature map obtained from the head Break it down into a series of blocks P is the size of the block, and a learnable position code is added to each block. After that, it is sent to the Transformer encoder. Similarly, after the Transformer decoder, a series of block outputs are obtained. Then output N blocks Reorganize into feature map
[0045] Step B13: Construct the tail structure. The tail adopts a multi-tail structure to adapt to different tasks. The tail of the super-resolution task consists of a convolution layer and an upsampling layer. The tails of other tasks are composed of a convolution layer. The input is the feature map F obtained by the body part. o , the output is the result image The calculation is as follows:
[0046]
[0047] Among them, i represents the i-th task, n represents a total of n tasks, the n-th task is the super-resolution task, Conv i It represents the convolution layer corresponding to the i-th task, and Upsample represents the upsampling layer.
[0048] Step B14: Construct a Transformer-based multi-task restoration and enhancement sub-network by integrating the head, body, and tail. The network consists of the head, body, and tail. The network input passes through the head, body, and tail in turn to obtain the output result. The parameters of the body are shared between multiple tasks, but the parameters of the head and tail are not shared. Specifically, the damaged image mentioned in step B11 is Input the head, and get the output feature map from step B11 Will Input body part, and get output feature map F from step B12 o ; F o Input the tail, and get the output image from step B13
[0049] Furthermore, the step C comprises the following steps:
[0050] Step C1: Design a multi-scale shallow feature extraction sub-network. The input of the network is the low-light image block I in the low-light image dataset obtained in step A3. in After three stages of feature extraction by ResNet, three features of different scales are obtained Add the features obtained from the original scale A total of four shallow features are obtained, where the first to fourth scales increase in size, and the calculation method is as follows:
[0051]
[0052]
[0053] Among them, I in represents the input low-light image block, represents the shallow features of the output low-light image at the i-th scale, f 4-i Represents the operation of the 4-ith stage of ResNet.
[0054] Step C2: Design a multi-scale feature fusion sub-network. The network consists of four branches. The input feature maps of the four branches are the output feature maps of step C1. The larger the branch number, the larger the scale. The input features of each branch are first transformed to the same number of channels through a 1×1 convolution layer, which is recorded as feature F. i , i = 1, 2, 3, 4, and then perform feature fusion. The feature maps of each branch after feature fusion are composed of the feature maps F of each branch before feature fusion. iAfter scale transformation, the result is added. This is because the scales of the branches before and after fusion are different, so during the transformation process, it is necessary to perform upsampling or convolution with a step size of 2 to transform to the scale of a certain branch before adding. The feature map obtained by the jth branch after feature fusion is The fusion process is as follows:
[0055]
[0056]
[0057]
[0058] Among them, Conv1 represents a 1×1 convolution with a step size of 1, a(F i ,j) represents F i The feature map obtained after converting from the i-th scale to the j-th scale is: Represents the upsampling layer, which increases the size of the feature map by 2 j-i times, (Conv3) i-j represents ij 3×3 convolutions with a stride of 2.
[0059] Step C3: Design a Transformer-based low-light image feature enhancement subnetwork, the input of which is the output feature map of the first branch in step C2. The output feature map is obtained by adding the input feature map and the feature map obtained after the input feature map passes through the Transformer-based multi-task recovery and enhancement sub-network in step B. The calculation method is as follows:
[0060]
[0061] in, Indicates that the feature map The feature map obtained after feeding the body part of the Transformer-based multi-task recovery and enhancement sub-network obtained in step B.
[0062] Step C4: Design a multi-scale low-light image enhancement sub-network. The network consists of three branches, each of which contains an attention module and a color correction module. The attention module consists of two parallel branches: a channel attention module and a spatial attention module.
[0063] Step C5: Design a Transformer-based multi-scale optimized low-light image enhancement network, which is composed of a multi-scale shallow feature extraction sub-network, a multi-scale feature fusion sub-network, a Transformer-based low-light image feature enhancement sub-network, and a multi-scale low-light image enhancement sub-network. Input low-light image block I inFirst, enter the multi-scale shallow feature extraction subnetwork in step C1 to obtain shallow features of four scales The first to fourth scales are increased in sequence; then the multi-scale feature fusion subnetwork in step C2 is entered to obtain the features of the four scales after fusion. Then Input the low-light image feature enhancement subnetwork based on Transformer in step C3 to obtain the output enhanced image feature t of the first scale out ;Bundle and t out Input the multi-scale low-light image enhancement subnetwork in step C4, where and t out Input the first branch of the multi-scale low-light image enhancement sub-network; and Input the second and third branches respectively to obtain image blocks of different scales after enhancement of each branch Among them, the enhanced image block R obtained by the third branch is the final enhancement result of the entire network.
[0064] Furthermore, the step C4 comprises the following steps:
[0065] Step C41: Design the channel attention module, assuming the input feature map is F att , obtained by convolution layer, layer standardization, activation function, convolution layer, layer standardization After that, it is divided into two branches. One branch performs global average pooling of the space to obtain a vector of scale 1×1×C and then combines it with F ch1 Multiply, then go through convolution, activation function, and convolution to get The other branch performs the global maximum pooling of the space to obtain a vector of size 1×1×C and then adds it to F ch1 Multiply, then go through convolution, activation function, and convolution to get F ch11 With F ch12 Add them together and pass the Sigmoid activation function to get the weight of each pixel in the feature map with a scale of H×W×C. Use this weight and F ch After multiplication, we get the output feature map O of the channel attention module ch The calculation process is as follows:
[0066] F ch1 =LN(Conv3(ReLU(LN(conv3(F att )))))
[0067] F ch =F ch1
[0068] Fch11 =Conv1(ReLU(Conv1(F ch1 ×AvgPooling s (F ch1 ))))
[0069] F ch12 =Conv1(ReLU(Conv1(F ch1 ×MaxPooling s (F ch1 ))))
[0070] O ch =F ch ×Sigmoid(F ch11 +F ch12 )
[0071] Among them, O ch The feature map F represents the input att After the channel attention module, the output feature map is obtained, F ch1 It represents the feature map after convolution layer, layer normalization, activation function, convolution layer, and layer normalization. Conv3 represents a 3×3 convolution with a step size of 1. LN represents layer normalization. ReLU is the activation function. F ch11 is the feature map obtained by the global average pooling branch of the space, F ch12 The feature map obtained by the global maximum pooling branch of the representation space. Conv1 is a 1×1 convolution for dimensionality increase and decrease. AvgPooling s is global average pooling, MaxPooling s is the global maximum pooling, and Sigmoid represents the Sigmoid activation function.
[0072] Step C42: Design the spatial attention module. The input of the spatial attention module is the same as that of the channel attention module, so the input feature map of the spatial attention module is also F att , the input feature map is obtained by convolution layer, batch normalization, activation function, convolution layer, batch normalization After that, it is divided into two branches. One branch performs average pooling in the channel dimension to obtain a vector of scale H×W×1, which is then combined with F sp1 Multiply to get The other branch performs the maximum pooling of the channel dimension to obtain a vector of size H×W×1, which is then combined with Multiply to get F sp11 With F sp12After concatenation of the channel dimensions, a 1x1 convolution is used to reduce the number of channels to C. The Sigmoid activation function is then used to obtain the weight of each pixel in the feature map with a scale of H×W×C. This weight is used together with F sp After multiplication, we get the output of the spatial attention module. The calculation process is as follows:
[0073] F sp1 =BNConv3(ReLU(BN(Conv3(F att )))))
[0074] F sp =F sp1
[0075] F sp11 =F sp1 ×AvgPooling c (F sp1 )
[0076] F sp12 =F sp1 ×MaxPooling c (F sp1 )
[0077] O sp =F sp ×Sigmoid(Conv1(Concat(F sp11 , F sp12 )))
[0078] Among them, O sp The feature map F represents the input att After the spatial attention module, the output feature map is obtained, F sp1 It represents the feature map after convolution layer, batch normalization, activation function, convolution layer, and batch normalization. Conv3 represents a 3×3 convolution with a step size of 1. BN represents batch normalization. ReLU is the activation function. F sp11 is the feature map obtained by the average pooling branch of the channel, F sp12 Represents the feature map obtained by the maximum pooling branch of the channel, AvgPooling c is the average pooling of the channel dimension, MaxPooling c It is the maximum pooling of channel dimension, and Sigmoid represents the Sigmoid activation function.
[0079] Step C43: Design an attention module, integrating the channel attention module and the spatial attention module. The attention module consists of two parallel modules: the channel attention module and the spatial attention module. Specifically, let the input feature map of the attention module be F att, the output O of the channel attention module obtained in step C41 ch And the output O of the spatial attention module obtained in step C42 sp It is concatenated through the channel dimension, and then convolved with F after a 3×3 convolution with a step size of 1. att Add together to get the output feature map O of the attention module att . It is expressed as follows:
[0080] O att =F att +Conv3(Concat(O ch , O sp ))
[0081] Step C44: Design a color correction module, assuming that the input of the color correction module is F color The input features first pass through the average pooling layer, and then pass through 1×1 convolution, activation function, 1×1 convolution, activation function, 1×1 convolution for dimensionality reduction and dimensionality increase, and then pass through twice the nearest neighbor upsampling and F color The concatenation is done along the channel dimension and then sent to the SE module. The number of channels is adjusted by a 1×1 convolution with a step size of 1 and the color-corrected feature map O is output. color The calculation method is as follows:
[0082] F c =Upsample(Conv1(ReLU(Conv1(ReLU(Conv1(AvgPooling(F color )))))))
[0083] SE(y)=y×Sigmoid(Conv1(ReLU(Conv1(AvgPooling(y)))))
[0084] O color =Conv1(SE(concat(F c , F color )))
[0085] Among them, F c and O color They represent the intermediate feature map and output feature map of the color correction module respectively, SE(·) represents the SE module, y represents the input of the SE module, Conv1 represents a 1×1 convolution with a step size of 1, ReLU represents the activation function, and Sigmoid represents the Sigmoid activation function.
[0086] Step C45: Design a multi-scale low-light image enhancement subnetwork. The subnetwork consists of three branches. Each branch has the same network structure, which includes a 3*3 convolution layer, an attention module, a 3*3 convolution layer, a ReLU activation function, a color correction module and two 3*3 convolution layers. The scales of the first to third branches increase in sequence. The input of the first branch of the network is the output feature map t obtained in step C3. out Output feature map obtained after deconvolution and step C2 The concatenated feature map is passed through a 3*3 convolutional layer, an attention module, a 3*3 convolutional layer, a ReLU activation function, a color correction module, and a 3*3 convolutional layer to output the feature map. The input of the second branch of the network is the output feature of the first branch Output feature map obtained after deconvolution and step C2 The concatenated feature map is passed through a 3*3 convolutional layer, an attention module, a 3*3 convolutional layer, a ReLU activation function, a color correction module, and a 3*3 convolutional layer to output the feature map. The input of the third branch of the network is the output feature of the second branch Output feature map obtained after deconvolution and step C2 The concatenated feature map is passed through a 3*3 convolutional layer, an attention module, a 3*3 convolutional layer, a ReLU activation function, a color correction module, and a 3*3 convolutional layer to output the feature map. Output feature map of each branch After a 3×3 convolution with a step size of 1, it is combined with the low-light image block I input in step C1 in The image block is scaled to the same scale as the branch. Add together to get the enhanced image block of the corresponding scale of the branch Used for subsequent calculation of loss function. The enhanced image block obtained by the third branch is the final enhancement result R of the entire network. The calculation method is as follows:
[0087]
[0088]
[0089]
[0090]
[0091]
[0092] in, represents the feature map output by the i-th branch, deconv represents the deconvolution operation, and Ff i represents the feature map of the i-th branch before passing through the color correction module in step C44, Conv3 represents a 3×3 convolution with a step size of 1, lrelu represents the LeakyRelu activation function, Atten represents the attention module in step C43, Concat represents concatenation by channel dimension, and Color represents the color correction module in step C44. represents the enhanced image block of the i-th scale obtained by the i-th branch, Resize represents the image block that is scaled from the input low-light image block to the image block with the same scale as the branch i. i It means scaling the image block to the scale corresponding to branch i, and R represents the enhanced image block finally output by the network.
[0093] Further, the step D is implemented as follows:
[0094] Design the target loss function of the network. The total target loss function of the network is as follows:
[0095]
[0096] Among them, R is the final output result of the Transformer-based multi-scale optimized low-light image enhancement network designed in step C, G is the corresponding label image block, ||·|| 1 Represents L1 loss, Resize 2 Resize means scaling the image block to the size corresponding to the scale of the second branch of the multi-scale low-light image enhancement subnetwork in step C4. 1 It means scaling the image block to the size corresponding to the scale of the first branch of the multi-scale low-light image enhancement subnetwork in step C4.
[0097] Further, the step E is implemented as follows:
[0098] The paired low-light image blocks and the corresponding label image blocks are randomly divided into several batches, each batch contains N pairs of image blocks, and the low-light image blocks are input into the Transformer-based multi-scale optimized low-light image enhancement network model designed in step C. The loss designed in step D is calculated for the final output enhanced image blocks and the image blocks output by each branch and the corresponding scale label image blocks, and the gradient of the parameters in the network is calculated using the back propagation method according to the loss, and the network parameters are updated using the stochastic gradient descent method.
[0099] Furthermore, the step F comprises the following steps:
[0100] Step F1: Cut the low-light image into blocks in height and width directions with a step size of q / 2 starting from the upper left corner. Each image block is cut into the same size as the image block during training, i.e., q×q×c.
[0101] Step F2: sequentially input the image blocks cut out in step F1 into the trained image enhancement network, and output the enhanced image blocks;
[0102] Step F3: splicing the enhanced image blocks in step F2 according to the order of cutting in step F1 to obtain the final low-illumination image enhancement result.
[0103] The above are preferred embodiments of the present invention. Any changes made according to the technical solution of the present invention, as long as the resulting functions do not exceed the scope of the technical solution of the present invention, belong to the protection scope of the present invention.
Claims
1. A Transformer-based multi-scale optimization low-light image enhancement method. It is characterized in that The steps include: Step A: preprocess the data by first pairing the data, then perform data segmentation and data enhancement. Step B: Build a Transformer-based multi-task restoration and enhancement sub-network, which consists of the head, body parts, and tail, pre-trained using a multi-task dataset and fine-tuned using a synthetic low-light image dataset; Step C, design a Transformer-based multi-scale optimized low-light image enhancement network, which consists of a multi-scale shallow feature extraction sub-network, a multi-scale feature fusion sub-network, a Transformer-based low-light image feature enhancement sub-network and a multi-scale low-light image enhancement sub-network; Step D: Design a loss function based on the network structure to guide the parameter optimization of the network model; Step E: training the network with the image block data, optimizing the network model parameters, and converging to a balanced state; Step F, cutting the low-light image into blocks, and then inputting the image blocks into the trained image enhancement network respectively, and after obtaining the corresponding enhanced image blocks, splicing the image blocks into an enhanced complete image; Step C is implemented as follows: Step C1, design a multi-scale shallow feature extraction sub-network, the input of the network is the real low-light image block in the data set obtained in step A, after the feature extraction of the three stages of ResNet, the features of three different scales are added to the features obtained at the original scale, and the features obtained at the original scale are added to obtain shallow features of four scales in total, where the first scale to the fourth scale are successively larger; Step C2, design a multi-scale feature fusion subnetwork, the network consists of four branches, the input feature maps of the four branches are the output feature maps of step C1, the input features of each branch are first transformed to the same number of channels through a 1×1 convolution layer, and then feature fusion is performed, the feature maps of each branch after feature fusion are obtained by adding the feature maps of each branch before feature fusion after scale transformation, because the scales of the branches before and after fusion are different, so in the transformation process, it is necessary to perform an upsampling operation or a convolution operation with a step size of 2 to transform to the scale of a certain branch and then perform an addition operation; Step C3, designing a Transformer-based low-light image feature enhancement subnetwork, the input of which is the output feature map of the first branch in step C2, and the output feature map is obtained by adding the input feature map and the feature map obtained after the input feature map passes through the Transformer-based multi-task restoration and enhancement subnetwork in step B; Step C4: Design a multi-scale low-light image enhancement subnetwork. The network consists of three branches. Each branch includes an attention module and a color correction module. The attention module consists of two parallel branches: a channel attention module and a spatial attention module.
2. According to claim 1, a Transformer-based multi-scale optimization low-light image enhancement method, It is characterized in that The specific implementation steps of step A are as follows: Step A1: pairing the real low-light image with the corresponding label image; Step A2: All real low-light images are subjected to three different random block processings, and the paired label images of the real low-light image I are also subjected to the same block processing method; the random block processing process is as follows: given a real low-light image I, the image size is h×w×3, and the size of the cut image blocks is q×q×3, where h and w represent the height and width of the real low-light image I, respectively, and q represents the height and width of the cut image block; Step A3, performing the same random horizontal flipping, vertical flipping, and rotation operations on the paired image blocks for data enhancement to obtain a real low-light image dataset; Step A4: Preprocess the noisy image and the labeled image, the underwater image and the labeled image, the low-resolution image and the labeled image, and the synthetic low-illumination image and the labeled image in the same manner as steps A1 to A3, respectively, to obtain a noisy image dataset, an underwater image dataset, a low-resolution image dataset, and a synthetic low-illumination image dataset.
3. According to claim 2, a Transformer-based multi-scale optimization low-light image enhancement method, It is characterized in that The specific implementation steps of step B are as follows: Step B1: Construct a Transformer-based multi-task recovery and enhancement sub-network, which consists of a head, a body, and a tail. The network input passes through the head, the body, and the tail in sequence to obtain the output result. The parameters of the body are shared between tasks, but the parameters of the head and the tail are not shared. Step B2: Use the multi-task dataset to pre-train the Transformer-based multi-task recovery and enhancement sub-network built in step B1, and use the network objective loss function to optimize the model parameters. Here, L1 loss is used, and the calculation method is as follows: in, represents the corrupted image of the i-th task, Indicates that the input is The results obtained by Transformer-based multi-task recovery and enhancement sub-network, I gt Indicates a damaged image The corresponding label image,||·|| 1 represents L1 loss, n represents n tasks, including image denoising, underwater image enhancement, image super-resolution, and synthetic low-light image enhancement; During the training process, each batch randomly selects a training from n-1 tasks except the synthetic low-light image enhancement task. Each batch contains the same number of pairs of damaged images and labeled images randomly extracted from the corresponding data set constructed in step A4. The damaged image is input into the Transformer-based multi-task restoration and enhancement sub-network constructed in step B1 to obtain the result image. According to the target loss function of the network, the back propagation method is used to calculate the gradient of each parameter in the network, and the stochastic gradient descent method is used to update the parameters of the network until the number of iterations reaches a threshold. The trained model is saved to complete the network training process. Step B3, use the synthetic low-light image dataset to fine-tune the network model parameters obtained in step B2; load the network model pre-trained in B2 as a pre-trained model, and use it as the initial network parameters for the subsequent fine-tuning of network parameters; randomly extract the same number of synthetic low-light images from the synthetic low-light image dataset in each batch, and input them into the Transformer-based multi-task restoration and enhancement sub-network constructed in B1 to obtain enhanced images; according to the network's target loss function, use the back propagation method to calculate the gradient of each parameter in the network, and use the stochastic gradient descent method to update the network parameters until the number of iterations reaches the threshold, save the trained model, and complete the network training process; finally, obtain the model parameters of the Transformer-based multi-task restoration and enhancement sub-network, which will be used as a pre-trained model in subsequent steps.
4. According to claim 3, a Transformer-based multi-scale optimization low-light image enhancement method, It is characterized in that The specific implementation steps of step B1 are as follows: Step B11: Construct the head structure. The head structure of each task consists of a convolutional layer and two residual structures. The damaged image of the i-th task After the head structure, the output feature map is obtained The calculation process is as follows: Res i (x)=Conv i (ReLU(Conv i (x)))+x,i=1,2,…n Among them, i represents the i-th task, n represents a total of n tasks, and Res i (·) represents the residual structure corresponding to the i-th task, x is the input of the residual structure, Conv i represents the convolutional layer corresponding to the i-th task, and ReLU is the activation function; Step B12: Construct the body part structure. The body part adopts the encoder-decoder structure of Vision Transformer. The parameters of the body parts corresponding to all tasks are shared and the structure is the same. Therefore, the feature maps obtained by passing the head of different task images are processed in the same way. Unified as F corrupted ; First, take the feature map obtained from the head Break it down into a series of blocks P is the size of the block, and a learnable position code is added to each block. After that, it is sent to the Transformer encoder; similarly, after the Transformer decoder, a series of block outputs are obtained Then output N blocks Reorganize into feature map Step B13: Construct the tail structure. The tail adopts a multi-tail structure to adapt to different tasks. The tail of the super-resolution task consists of a convolution layer and an upsampling layer. The tails of other tasks are composed of a convolution layer. The input is the feature map F obtained by the body part. o , the output is the result image The calculation is as follows: Among them, i represents the i-th task, n represents a total of n tasks, the n-th task is the super-resolution task, Conv i represents the convolution layer corresponding to the i-th task, and Upsample represents the upsampling layer; Step B14: Construct a Transformer-based multi-task restoration and enhancement sub-network by integrating the head, body, and tail. The network input passes through the head, body, and tail in turn to obtain the output result. The parameters of the body are shared between multiple tasks, but the parameters of the head and tail are not shared. That is, the damaged image mentioned in step B11 is Input the head, and get the output feature map from step B11 Will Input body part, and get output feature map F from step B12 o ; F o Input the tail, and get the output image from step B13 5. According to claim 2, a Transformer-based multi-scale optimization low-light image enhancement method, It is characterized in that The specific implementation steps of step C are as follows: Step C1: Design a multi-scale shallow feature extraction sub-network, the input of which is the real low-light image block I in the real low-light image dataset obtained in step A3. in After the three stages of feature extraction of ResNet, three features of different scales are obtained Add the features obtained from the original scale A total of four shallow features are obtained, where the first to fourth scales increase in size, and the calculation method is as follows: Among them, I in represents the input real low-light image block, represents the shallow features of the output real low-light image at the i-th scale, f 4-i Represents the operation of the 4-ith stage of ResNet; Step C2: Design a multi-scale feature fusion sub-network. The network consists of four branches. The input feature maps of the four branches are the shallow features of the real low-light image output in step C1. The larger the branch number, the larger the scale. The input features of each branch are first transformed to the same number of channels through a 1×1 convolution layer, which is recorded as feature F. i , i = 1, 2, 3, 4, and then perform feature fusion. The feature maps of each branch after feature fusion are composed of the feature maps F of each branch before feature fusion. i After scale transformation, it is added. This is because the scales of the branches before and after fusion are different, so in the transformation process, it is necessary to undergo upsampling or convolution with a step size of 2 to transform to the scale of a certain branch and then perform the addition operation; the feature map obtained by the jth branch after feature fusion is The fusion process is as follows: Among them, Conv1 represents a 1×1 convolution with a step size of 1, a(F i ,j) represents F i The feature map obtained after converting from the i-th scale to the j-th scale is: Represents the upsampling layer, which increases the size of the feature map by 2 j-i times, (Conv3) i-j represents ij 3×3 convolutions with a stride of 2; Step C3: Design a Transformer-based low-light image feature enhancement subnetwork, the input of which is the output feature map of the first branch in step C2. The output feature map is obtained by adding the input feature map and the feature map obtained after the input feature map passes through the Transformer-based multi-task recovery and enhancement sub-network in step B. The calculation method is as follows: in, Indicates that the feature map The feature map obtained after feeding the body part of the Transformer-based multi-task recovery and enhancement sub-network obtained in step B; Step C4, design a multi-scale low-light image enhancement sub-network, the network consists of three branches, each branch includes an attention module and a color correction module, and the attention module consists of two parallel branches: a channel attention module and a spatial attention module; Step C5, design a Transformer-based multi-scale optimized low-light image enhancement network, which is composed of a multi-scale shallow feature extraction sub-network, a multi-scale feature fusion sub-network, a Transformer-based low-light image feature enhancement sub-network, and a multi-scale low-light image enhancement sub-network; input low-light image block I in First, enter the multi-scale shallow feature extraction subnetwork in step C1 to obtain shallow features of four scales The first to fourth scales are increased in sequence; then the multi-scale feature fusion subnetwork in step C2 is entered to obtain the features of the four scales after fusion. Then Input the low-light image feature enhancement subnetwork based on Transformer in step C3 to obtain the output enhanced image feature t of the first scale out ;Bundle and t out Input the multi-scale low-light image enhancement subnetwork in step C4, where and t out Input the first branch of the multi-scale low-light image enhancement sub-network; and Input the second and third branches respectively to obtain image blocks of different scales after enhancement of each branch Among them, the enhanced image block R obtained by the third branch is the final enhancement result of the entire network.
6. According to claim 5, a Transformer-based multi-scale optimization low-light image enhancement method, It is characterized in that The specific implementation steps of step C4 are as follows: Step C41, design the channel attention module, assuming the input feature map is F att , obtained by convolution layer, layer standardization, activation function, convolution layer, layer standardization After that, it is divided into two branches. One branch performs global average pooling of the space to obtain a vector of scale 1×1×C and then combines it with F ch1 Multiply, then go through convolution, activation function, and convolution to get The other branch performs the global maximum pooling of the space to obtain a vector of size 1×1×C and then adds it to F ch1 Multiply, then go through convolution, activation function, and convolution to get F ch11 With F ch12 Add them together and pass the Sigmoid activation function to get the weight of each pixel in the feature map with a scale of H×W×C. Use this weight and F ch After multiplication, we get the output feature map O of the channel attention module ch ; The calculation process is as follows: F ch1 =LN(Conv3(ReLU(LN(Conv3(F att ))))) F ch =F ch1 F ch11 =Conv1(ReLU(Conv1(F ch1 ×AvgPooling s (F ch1 )))) F ch12 =Conv1(ReLU(Conv1(F ch1 ×MaxPooling s (F ch1 )))) O ch =F ch ×Sigmoid(F ch11 +F ch12 ) Among them, O ch The feature map F represents the input att After the channel attention module, the output feature map is obtained, F ch1 It represents the feature map after convolution layer, layer normalization, activation function, convolution layer, and layer normalization. Conv3 represents a 3×3 convolution with a step size of 1. LN represents layer normalization. ReLU is the activation function. F ch11 is the feature map obtained by the global average pooling branch of the space, F ch12 The feature map obtained by the global maximum pooling branch of the representation space. Conv1 is a 1×1 convolution for dimensionality increase and decrease. AvgPooling s is global average pooling, MaxPooling s is the global maximum pooling, Sigmoid represents the Sigmoid activation function; Step C42: Design a spatial attention module. The input of the spatial attention module is the same as that of the channel attention module, so the input feature map of the spatial attention module is also F att , the input feature map is obtained by convolution layer, batch normalization, activation function, convolution layer, batch normalization After that, it is divided into two branches. One branch performs average pooling in the channel dimension to obtain a vector of scale H×W×1, which is then combined with F sp1 Multiply to get The other branch performs the maximum pooling of the channel dimension to obtain a vector of size H×W×1, which is then combined with F sp1 Multiply to get F sp11 With F sp12 After concatenation of the channel dimensions, a 1x1 convolution is used to reduce the number of channels to C. The Sigmoid activation function is then used to obtain the weight of each pixel in the feature map with a scale of H×W×C. This weight is used together with F sp After multiplication, we get the output of the spatial attention module; the calculation process is as follows: <h2 style=";text-align:left;direction:ltr">F<h2 style=";text-align:left;direction:ltr"> sp1 <h2 style=";text-align:left;direction:ltr"> =BN(Conv3(ReLU(BN(Conv3(F<h2 style=";text-align:left;direction:ltr"> att <h2 style=";text-align:left;direction:ltr"> ))))) F sp =F sp1 F sp11 =F sp1 ×AvgPooling c (F sp1 ) F sp12 =F sp1 ×MaxPoolingc(F sp1 ) O sp =F sp ×Sigmoid(Conv1(Concat(F sp11 ,F sp12 ))) Among them, O sp The feature map F represents the input att After the spatial attention module, the output feature map is obtained, F sp1 It represents the feature map after convolution layer, batch normalization, activation function, convolution layer, and batch normalization. Conv3 represents a 3×3 convolution with a step size of 1. BN represents batch normalization. ReLU is the activation function. F sp11 is the feature map obtained by the average pooling branch of the channel, F sp12 Represents the feature map obtained by the maximum pooling branch of the channel, AvgPooling c is the average pooling of the channel dimension, MaxPooling c is the maximum pooling of the channel dimension, and Sigmoid represents the Sigmoid activation function; Step C43, design an attention module, integrating the channel attention module and the spatial attention module; the attention module consists of two parallel modules: the channel attention module and the spatial attention module; specifically, let the input feature map of the attention module be F att , the output O of the channel attention module obtained in step C41 ch And the output O of the spatial attention module obtained in step C42 sp It is concatenated through the channel dimension, and then convolved with F after a 3×3 convolution with a step size of 1. att Add together to get the output feature map O of the attention module att ; means as follows: THE att =F att +Conv3(Concat(The ch ,THE sp )) Step C44: Design a color correction module, assuming that the input of the color correction module is F color The input features first pass through the average pooling layer, and then pass through 1×1 convolution, activation function, 1×1 convolution, activation function, 1×1 convolution for dimensionality reduction and dimensionality increase, and then pass through twice the nearest neighbor upsampling and F color The concatenation is done along the channel dimension and then sent to the SE module. The number of channels is adjusted by a 1×1 convolution with a step size of 1 and the color-corrected feature map O is output. color ; The calculation method is as follows: F c =Upsample(Conv1(ReLU(Conv1(ReLU(Conv1(AvgPooling(F color )))))) SE(y)=y×Sigmoid(Conv1(ReLU(Conv1(AvgPooling(y))))) A color =Conv1(SE(Concat(F c ,F color ))) Among them, F c and O color They represent the intermediate feature map and output feature map of the color correction module, respectively. SE(·) represents the SE module, y represents the input of the SE module, Conv1 represents a 1×1 convolution with a step size of 1, ReLU represents the activation function, and Sigmoid represents the Sigmoid activation function. Step C45, design a multi-scale low-light image enhancement subnetwork, which consists of three branches. Each branch has the same network structure, which includes a 3*3 convolutional layer, an attention module, a 3*3 convolutional layer, a ReLU activation function, a color correction module and two 3*3 convolutional layers in sequence. The scales of the first to third branches increase in sequence; the input of the first branch of the network is the output feature map t obtained in step C3 out Output feature map obtained after deconvolution and step C2 The concatenated feature map is passed through a 3*3 convolutional layer, an attention module, a 3*3 convolutional layer, a ReLU activation function, a color correction module, and a 3*3 convolutional layer to output the feature map. The input of the second branch of the network is the output feature of the first branch Output feature map obtained after deconvolution and step C2 The concatenated feature map is passed through a 3*3 convolutional layer, an attention module, a 3*3 convolutional layer, a ReLU activation function, a color correction module, and a 3*3 convolutional layer to output the feature map. The input of the third branch of the network is the output feature of the second branch Output feature map obtained after deconvolution and step C2 The concatenated feature map is passed through a 3*3 convolutional layer, an attention module, a 3*3 convolutional layer, a ReLU activation function, a color correction module, and a 3*3 convolutional layer to output the feature map. Output feature map of each branch After a 3×3 convolution with a step size of 1, it is combined with the low-light image block I input in step C1 in The image block is scaled to the same scale as the branch. Add together to get the enhanced image block of the corresponding scale of the branch Used for subsequent calculation of loss function; enhanced image block obtained by the third branch is the final enhancement result R of the entire network; the calculation method is as follows: in, represents the feature map output by the i-th branch, deconv represents the deconvolution operation, and Ff i represents the feature map of the i-th branch before the color correction module in step C44, Conv3 represents a 3×3 convolution with a step size of 1, lrelu represents the LeakyRelu activation function, Atten represents the attention module in step C43, Concat represents concatenation by channel dimension, and Color represents the color correction module in step C44; represents the enhanced image block of the i-th scale obtained by the i-th branch, Resize represents the image block that is scaled from the input low-light image block to the image block with the same scale as the branch i. i It means scaling the image block to the scale corresponding to branch i, and R represents the enhanced image block finally output by the network.
7. According to claim 1, a Transformer-based multi-scale optimization low-light image enhancement method, It is characterized in that The specific implementation of step D is: Design the target loss function of the network. The total target loss function of the network is as follows: Among them, R is the final output result of the Transformer-based multi-scale optimized low-light image enhancement network designed in step C, G is the corresponding label image block, ||.|| 1 Represents L1 loss, Resize 2 Resize means scaling the image block to the size corresponding to the scale of the second branch of the multi-scale low-light image enhancement subnetwork in step C4. 1 It means scaling the image block to the size corresponding to the scale of the first branch of the multi-scale low-light image enhancement subnetwork in step C4.
8. According to claim 1, a Transformer-based multi-scale optimization low-light image enhancement method, It is characterized in that The specific implementation of step E is as follows: The paired low-light image blocks and the corresponding label image blocks are randomly divided into several batches, each batch contains N pairs of image blocks, and the low-light image blocks are input into the Transformer-based multi-scale optimized low-light image enhancement network model designed in step C. The loss designed in step D is calculated for the final output enhanced image blocks and the image blocks output by each branch and the corresponding scale label image blocks, and the gradient of the parameters in the network is calculated using the back propagation method according to the loss, and the network parameters are updated using the stochastic gradient descent method.
9. According to claim 4, a Transformer-based multi-scale optimization low-light image enhancement method, It is characterized in that The specific implementation steps of step F are as follows: Step F1, cut the low-light image into blocks in height and width directions according to a step size of q / 2 starting from the upper left, and cut each image block into the same size as the image block during training, that is, q×q×C; Step F2, sequentially input the image blocks cut out in step F1 into the trained image enhancement network, and output the enhanced image blocks; Step F3: splice the enhanced image blocks in step F2 according to the order of cutting in step F1 to obtain the final low-illumination image enhancement result.
Citation Information
Patent Citations
Photographing method based on low-illumination image enhancement algorithm of brightness attention mechanism
CN111915526A
Low-illumination image enhancement method based on attention guidance and multi-scale feature fusion
CN114596233A