Dark and Blurry Image Enhancement Method Based on Task Decoupling
By designing a task-decoupled dark-light blur image enhancement network, including color transformation, fusion enhancement and high-frequency information reconstruction branches, the simultaneous problems of dark light and blur in dark-light blur images are solved, and the efficient enhancement and clarity improvement of the image is achieved.
Patent Information
- Application Number
- CN202310853029.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-12
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2043-07-12
AI Technical Summary
The prior art is difficult to effectively solve the problems of dark light and blur in dark blur images, resulting in poor visibility and perceptibility of the image and difficult to apply to computer vision tasks.
Using a task-based decoupling method, a network including color transformation branches, fusion enhancement branches and high-frequency information reconstruction branches are designed to gradually improve the brightness and clarity of the image through supervised learning.
It effectively improves the perception and visibility of dark light blur images, outputs high-quality clear images, and solves the dual challenges of dark light and blur problems.
Smart Images

Figure CN116797491B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of image processing and computer vision, and in particular to a method for enhancing low-light blurred images based on task decoupling. Background Art
[0002] In recent years, the progress and development of science and technology have significantly promoted the improvement of society and people's lives. With the miniaturization of the volume of image acquisition devices and the high efficiency of acquisition capabilities, images and image systems are closely related to people's daily lives and production development. Image-based processing systems have a wide range of applications in life scenarios. The image system can facilitate users to record and observe intuitively in real time, with high convenience. However, limited by the environment and usage status of the acquisition device, the obtained images lack ideal observability, which is reflected in phenomena such as motion blur and poor illuminance of the images. Especially when the lighting conditions are poor, the situation of poor illuminance is very common. At this time, the images taken often have multiple dark regions, and at the same time, object movement and camera shake may cause image blur, which brings difficulties to human eye reading or machine vision processing. Therefore, designing a method for enhancing low-light blurred images has important theoretical and application significance.
[0003] Low-light images can be generated due to insufficient exposure or environmental conditions. Low-light images usually contain regions with poor visibility; blurred images can be generated due to the movement of objects or photographic equipment during image shooting, which is a common problem in image acquisition. Both low-light images and blurred images cause difficulties in human perception, and at the same time, it makes these images difficult to be applied to advanced visual tasks of computers. The object and scene information in low-light blurred images is greatly masked. Specifically, the pixel distribution differences in low-light blurred images are small, the pixel values approach 0, and the color and texture information are difficult to detect. At the same time, the blur factor will also cause great damage to the texture structure and contour, especially the impact on detail information is very serious.
[0004] There have been some studies on solving a single problem in low-light blurred images. For example, studying low-light image enhancement or image deblurring problems. In the scenario of a single problem, the loss of color and detailed contours is limited to a certain characteristic. For example, if there is no blur problem in a low-light image, only the pixel values need to be changed accordingly, and there is no need to repair the structure and detail information of the image; similarly, if there is no low-light problem in a blurred image, only the texture and details need to be enhanced and restored, and there will be no significant changes in the overall tone, brightness, etc. of the image during this process. However, when the low-light and blur problems occur simultaneously, effective methods are still needed to solve these problems simultaneously. Summary of the Invention
[0005] In view of this, the purpose of the present invention is to provide a method for enhancing low-light blurred images based on task decoupling, which can improve the perception and visibility of low-light blurred images.
[0006] To achieve the above object, the present invention adopts the following technical solutions: A low-light and blurred image enhancement method based on task decoupling, comprising the following steps:
[0007] Step S1, construct a training dataset, preprocess each image of the low-light and blurred images to obtain a training dataset;
[0008] Step S2, design a low-light and blurred image enhancement network with task decoupling, which is composed of a color transformation branch, a fusion enhancement branch, and a high-frequency information reconstruction branch;
[0009] Step S3, design a loss function for training the network designed in step S2;
[0010] Step S4, use the training dataset to train the low-light and blurred image enhancement network based on task decoupling;
[0011] Step S5, input the low-light and blurred image to be measured into the designed network, and use the trained network to predict and generate a final result with better visual perception.
[0012] In a preferred embodiment, the step S1 includes the following steps:
[0013] Step S11, scale each image in the dataset to an image of the same size with dimensions H×W;
[0014] Step S12, perform normalization processing on the training images; given an image I train , calculate the normalized image using the following formula:
[0015]
[0016] where I train is an image with 8-bit color depth and size H×W, and I bit_max is an image with size H×W and all pixel values being 255.
[0017] In a preferred embodiment, the step S2 includes the following steps:
[0018] Step S21, design a color transformation branch, which is composed of a color transformation preprocessing network, a color transformation backbone network, and a color transformation head network; the color transformation branch takes the normalized low-light and blurred image X with dimensions H×W as input, and finally outputs a color transformation result with dimensions
[0019] Step S22: Design a fusion enhancement branch, which consists of a fusion enhancement preprocessing network, a fusion enhancement backbone network, and a fusion enhancement head network; the input of the fusion enhancement branch is: the normalized low-light blurred image X with size H×W, the output features of the color transformation branch, and the output features of the high-frequency information reconstruction branch, and finally outputs a fusion enhancement result with size H×W
[0020] Step S23: Design a high-frequency information reconstruction branch, which consists of a high-frequency information reconstruction preprocessing network, a high-frequency information reconstruction backbone network, and a high-frequency information reconstruction head network; the high-frequency information reconstruction branch takes the normalized low-light blurred image X with size H×W as the input, and finally outputs a high-frequency information reconstruction result with size H×W
[0021] In a preferred embodiment, the step S21 specifically includes the following steps:
[0022] Step S211: Design a color transformation preprocessing network, which includes 1 PixelUnShuffle layer and 1 convolutional layer; among them, the PixelUnShuffle layer performs pixel recombination on the low-light blurred image X, and changes the size of the input low-light blurred image X to The number of channels is changed to 16 times the original number of channels. The convolutional layer is a convolution with a convolution kernel size of 3×3, a stride of 1, and a padding of 1; the output of the color transformation preprocessing network is denoted as F 0 col ;
[0023] Step S212: Design a color transformation backbone network, which includes several color transformation modules with the same structure. Taking the i-th module as an example, i≥1, for the feature F input to the color transformation module i-1 col , first, the feature distillation and channel modulation unit FDCMU predicts the weight of the color transformation, and the output result is denoted as K i , similarly, the bias of the color transformation is also obtained by FDCMU, denoted as b i ; at the same time, based on F i-1 col predict the channel mapping, F i-1 col After passing through FDCMU, global average pooling GAP operation and multi-layer perceptron MLP in this way, finally obtain a vector with the same channel dimension as F i-1 col denoted as g i ; the output F of the color transformation module i col is obtained by the following transformation:
[0024] F i col = (F i-1 col * K i + b i ) exp(g i ).
[0025] where * is the element-wise multiplication of matrices, + is the element-wise addition of matrices, and exp(·) is the exponential operation with the broadcasting mechanism;
[0026] Step S213, design the Feature Distillation and Channel Modulation Unit (FDCMU); for a given input feature f D , the process of obtaining the output feature f D out by the FDCMU is described as follows:
[0027] f D 1a = PReLU(Conv 1×1 (DWConv 5×5 (f D ))).
[0028] f D 1b = DWConv 3×3 (Conv 1×1 (f D )).
[0029] f D 2 = PReLU(Conv 3×3 (PReLU(f D 1b ))).
[0030] f D 2b = PReLU(Conv 1×1 (f D 1b )).
[0031] f D 3 = Conv 1×1 (Concatenate(f D 1a , f D 1b , f D 2 , f D 2b )).
[0032]
[0033] Among them, Conv 1×1 is a convolution with a kernel size of 1×1, a stride of 1, and a padding of 0. Conv 3×3 is a convolution with a kernel size of 3×3, a stride of 1, and a padding of 1. DWConv 3×3 is a depthwise separable convolution with a kernel size of 3×3, a stride of 1, and a padding of 1. DWConv 5×5 is a depthwise separable convolution with a kernel size of 5×5, a stride of 1, and a padding of 2. PReLU(·) is the activation function; Concatenate(·) is the feature concatenation along the channels; w c and b c are learnable parameter vectors. ⊙ and are element-wise multiplication and element-wise addition with a broadcasting mechanism;
[0034] Step S214: Design a color transformation head network; the color transformation head network receives the output from the color transformation backbone network as its input and finally outputs a color transformation result with a size of ; The color transformation head network is composed of a convolutional layer, a ReLU activation function, and a convolutional layer stacked in sequence. The convolutional layer is a convolution with a kernel size of 3×3, a stride of 1, and a padding of 1.
[0035] In a preferred embodiment, step S22 includes the following steps:
[0036] Step S221: Design a fusion enhancement preprocessing network. The fusion enhancement preprocessing network consists of 1 convolutional layer. The convolutional layer is a convolution with a kernel size of 3×3, a stride of 1, and a padding of 1. The output feature is denoted as F 0 fuse ;
[0037] Step S222: Design a fusion enhancement backbone network. The fusion enhancement backbone network contains several fusion enhancement modules with the same structure. Taking the i-th module as an example, i≥1, the input received by the fusion enhancement module is: the feature F i-1 fuse from the previous fusion enhancement module. When i = 1, this input is the output of the fusion enhancement preprocessing network, the feature F i col from the color transformation module, and the feature F i high from the high-frequency information reconstruction module; First, the feature F i col from the color transformation module passes through the channel expansion and upsampling module CEUM to transform and output the feature F i col'; then, feature F i-1 fuse Extracts and fuses multi-scale features by the multi-scale feature fusion unit MFFU, and then adds them element-wise to feature F i high and F i col ' respectively, and obtains F after being processed by two convolutional layers i fuse ; The convolutional layer is a convolution with a kernel size of 3×3, a stride of 1, and a padding of 1
[0038] Step S223, design the channel expansion and upsampling module CEUM; for the given input feature f E , the process of obtaining the output feature f E out by CEUM is described as:
[0039] f E 1 = Conv 1×1 (f E ).
[0040] f E 2a = FDCMU(f E 1 ).
[0041] f E 2b = FDCMU(f E 1 ).
[0042] f E 3 = Conv 3×3 (Conv 1×1 (Concatenate(f E 2a , f E 2b ))).
[0043] f E out = Conv 1×1 (Conv 3×3 (PixelShuffle(f E 3 ))).
[0044] where Conv 1×1 is a convolution with a kernel size of 1×1, a stride of 1, and a padding of 0, Conv 3×3is a convolution with a convolution kernel size of 3×3, a stride of 1, and a padding of 1. FDCMU is the feature distillation and channel modulation unit described in step S213. Concatenate(·) is the feature concatenation along the channels. PixelShuffle(·) is the PixelShuffle layer, whose function is to transform the input feature of size into H×W, and the number of channels is transformed into 1 / 16 times the original number of channels;
[0045] Step S224, design a multi-scale feature fusion unit MFFU; for a given input feature f M , the process of obtaining the output feature f M out is described as:
[0046] f M 1 = Conv 3×3 (f M ).
[0047] f M 2 = ResBlocks(f M 1 ).
[0048] f M 2a = Up ×2 (ResBlocks(Down ×2 (f M 1 ))).
[0049] f M 2b = Up ×4 (ResBlocks(Down ×4 (f M 1 ))).
[0050] f M out = Conv 3×3 (Conv 1×1 (Concatenate(f M 1 , f M 2 , f M 2a , f M 2b ))).
[0051] where Conv 1×1is a convolution with a convolution kernel size of 1×1, a stride of 1, and a padding of 0, Conv 3×3 is a convolution with a convolution kernel size of 3×3, a stride of 1, and a padding of 1; Down ×4 (·) is a spatial 4-fold downsampling performed by the bilinear interpolation method, Up ×4 (·) is a spatial 4-fold upsampling performed by the bilinear interpolation method; Down ×2 (·) is a spatial 2-fold downsampling performed by the bilinear interpolation method, Up ×2 (·) is a spatial 2-fold upsampling performed by the bilinear interpolation method; ResBlocks is composed of a number of residual blocks (ResBlock). In this embodiment, the number of residual blocks is 15; the residual block is stacked by a convolutional layer, a PReLU activation function, and a convolutional layer, and the input and output are added by skipping. Among them, the convolutional layer is a convolution with a convolution kernel size of 3×3, a stride of 1, and a padding of 1;
[0052] Step S225, design a fusion enhancement head network; the fusion enhancement head network receives the output from the fusion enhancement backbone network as its input and finally outputs a fusion enhancement result of size H×W The fusion enhancement head network is sequentially stacked by a convolutional layer, a ReLU activation function, and a convolutional layer, where the convolutional layer is a convolution with a convolution kernel size of 3×3, a stride of 1, and a padding of 1.
[0053] In a preferred embodiment, the step S23 specifically includes the following steps:
[0054] Step S231, design a high-frequency information reconstruction preprocessing network. The high-frequency information reconstruction preprocessing network is composed of 1 convolutional layer. The convolutional layer is a convolution with a convolution kernel size of 3×3, a stride of 1, and a padding of 1; the output feature is denoted as F 0 high ;
[0055] Step S232, design a high-frequency information reconstruction backbone network. The high-frequency information reconstruction backbone network contains a number of high-frequency information reconstruction modules with the same structure. Taking the i-th module as an example, i≥1, for the feature F input to the high-frequency information reconstruction module i -1 high , the feature is processed by the multi-scale feature fusion unit (MFFU) and the gated Fourier transform unit GFU described in step S224 in sequence; the output of the high-frequency information reconstruction module is denoted as F i high ;
[0056] Step S233, design a gated Fourier transform unit GFU; for a given input feature f G , the output feature f is obtained by GFUG out The process is described as follows:
[0057] f G 1a = Conv 3×3 (PReLU(Conv 3×3 (f G )) + f G .
[0058] f G 1b = PReLU(Conv 1×1 (f G ).
[0059] f G 2 = IFFT(PReLU(Conv 1×1 (FFT(f G 1b )))) + f G 1b .
[0060] f G out = Conv 1×1 (Sigmoid(f G 1b ) * f G 2 .
[0061] Among them, Conv 1×1 is a convolution with a kernel size of 1×1, a stride of 1, and a padding of 0. Conv 3×3 is a convolution with a kernel size of 3×3, a stride of 1, and a padding of 1. PReLU(·) is an activation function; FFT(·) is the Fourier transform taking the real part, IFFT(·) is the inverse Fourier transform taking the real part, and Sigmoid(·) is an activation function; * is the element-wise multiplication of matrices;
[0062] Step S234, design a high-frequency information reconstruction head network; the high-frequency information reconstruction head network receives the output from the high-frequency information reconstruction backbone network as its input and outputs a high-frequency information reconstruction result of size H×W
[0063] In a preferred embodiment, in step S3, design a loss function for training the network designed in step S2; it includes the following steps:
[0064] Step S31, design the overall optimization objective of the network; the optimization objective is to minimize the total loss function
[0065]
[0066] Among them, represents the color transformation loss function, and λ col represents the weight of the color transformation loss function; represents the semantic perception loss function, and λ sem represents the weight of the semantic perception loss function; represents the pixel loss function, and λ pix represents the weight of the pixel loss function; represents the high-frequency information reconstruction loss function, and λ high represents the weight of the high-frequency information reconstruction loss function;
[0067] Step S32: Design the color transformation loss function; The calculation formula of is as follows:
[0068]
[0069] Among them, is the output result of the color transformation branch, with a size of Y ↓4 is the result of the reference image being reduced to through bilinear interpolation operation. AvgPool(·) is the average pooling operation with a kernel size of 4, and ||·|| 1 is the absolute value operation;
[0070] Step S33: Design the semantic perception loss function; The calculation formula of is as follows:
[0071]
[0072] Among them is the color transformation result output by the color transformation branch, Y ↓4 is the image of the reference image reduced by 4 times through bilinear interpolation, is the output result of the fusion enhancement branch, Y is the reference image; Φ(·) represents the operation of extracting the features of the Conv2-2, Conv3-2, and Conv4-2 layers using the VGG-16 model pre-trained on ImageNet; ||·|| 1 is the absolute value operation;
[0073] Step S34: Design the light pixel loss function; The calculation formula of is as follows:
[0074]
[0075] Among them, is the output result of the fusion enhancement branch, and Y is the reference image; ||·|| 1 is the absolute value operation;
[0076] Step S35: Design the high-frequency information reconstruction loss function; The calculation formula is as follows:
[0077]
[0078] Among them, is the output result of the high-frequency information reconstruction branch, Y is the reference image, HighFreq(·) is the operation to obtain the high-frequency information part, and the high-frequency information part is obtained by transforming the original image to the frequency domain by the discrete cosine transform (DCT), and then performing the inverse transform after obtaining the high-frequency; ||·|| 1 is the absolute value operation.
[0079] In a preferred embodiment, the step S4 includes the following steps:
[0080] Step S41: Select a random training image X from the dataset constructed in step S1;
[0081] Step S42: Training image encoding and enhancement; Input the image X, and pass it through the task-decoupled low-light and blurred image enhancement network, and calculate the output color transformation result fusion enhancement result and high-frequency information reconstruction result Calculate the loss of the total loss function in step S31
[0082] Step S43: Use the backpropagation method to calculate the gradients of the parameters in the task-decoupled low-light and blurred image enhancement network, and update the parameters using the Adam optimization method;
[0083] Step S44: The above steps are one iteration of the training process, and the entire training process requires 10 5 iterations, and in each iteration process, multiple image pairs are randomly sampled as a batch for training.
[0084] In a preferred embodiment, in the step S5, the to-be-detected low-light and blurred image is input into the designed network, and the output result of the trained network is used, where the fusion enhancement result is the final result.
[0085] Compared with the prior art, the present invention has the following beneficial effects: The existing low-light image enhancement methods or image deblurring methods cannot effectively solve the enhancement of low-light blurred images. The present invention proposes a low-light blurred image enhancement method based on task decoupling, which decouples the enhancement of low-light blurred images into color transformation and high-frequency information reconstruction by designing a parallel network, and designs corresponding supervised learning tasks; the designed network can effectively improve the brightness of the input image, eliminate the blurring effect, and output a clear image with normal illumination of high quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0086] Figure 1 It is a flowchart of the method according to an embodiment of the present invention.
[0087] Figure 2 It is a low-light blurred image enhancement network based on task decoupling according to an embodiment of the present invention.
[0088] Figure 3 It is a feature distillation and channel modulation unit (FDCMU) according to an embodiment of the present invention.
[0089] Figure 4 It is a channel expansion and upsampling module (CEUM) according to an embodiment of the present invention.
[0090] Figure 5 It is a multi-scale feature fusion unit (MFFU) according to an embodiment of the present invention.
[0091] Figure 6 It is a gated Fourier transform unit (GFU) according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0092] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0093] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this application belongs.
[0094] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application; as used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0095] The present invention provides a low-light blurred image enhancement method based on task decoupling, as Figures 1-6 shown, including the following steps:
[0096] Step S1: Construct a training dataset by preprocessing each image in the low-light and blurred images to obtain a training dataset;
[0097] Step S2: Design a task-decoupled low-light and blurred image enhancement network, which consists of a color transformation branch, a fusion enhancement branch, and a high-frequency information reconstruction branch;
[0098] Step S3: Design a loss function for training the network designed in Step S2;
[0099] Step S4: Use the training dataset to train the task-decoupled low-light and blurred image enhancement network;
[0100] Step S5: Input the low-light and blurred image to be tested into the designed network, and use the trained network to predict and generate a final result with better visual perception.
[0101] Furthermore, Step S1 includes the following steps:
[0102] Step S11: Scale each image in the dataset to an image of the same size with dimensions H×W.
[0103] Step S12: Normalize the training images. Given an image I train , the formula for calculating the normalized image is as follows:
[0104]
[0105] where I train is an image with 8-bit color depth and size H×W, and I bit_max is an image with size H×W and all pixel values being 255.
[0106] Furthermore, Step S2 includes the following steps, as Figure 2 shown:
[0107] Step S21: Design a color transformation branch, which consists of a color transformation preprocessing network, a color transformation backbone network, and a color transformation head network. The color transformation branch takes the normalized low-light and blurred image X with size H×W as input and finally outputs a color transformation result with size
[0108] Step S22: Design a fusion enhancement branch, which consists of a fusion enhancement preprocessing network, a fusion enhancement backbone network, and a fusion enhancement head network. The input of the fusion enhancement branch is: the normalized low-light blurred image X with size H×W, the output features of the color transformation branch, and the output features of the high-frequency information reconstruction branch. The final output is the fusion enhancement result with size H×W
[0109] Step S23: Design a high-frequency information reconstruction branch, which consists of a high-frequency information reconstruction preprocessing network, a high-frequency information reconstruction backbone network, and a high-frequency information reconstruction head network. The high-frequency information reconstruction branch takes the normalized low-light blurred image X with size H×W as the input, and the final output is the high-frequency information reconstruction result with size H×W
[0110] Furthermore, step S21 includes the following steps:
[0111] Step S211: Design a color transformation preprocessing network, which contains 1 PixelUnShuffle layer and 1 convolutional layer. The PixelUnShuffle layer reorganizes the pixels of the low-light blurred image X, and changes the size of the input low-light blurred image X to The number of channels is changed to 16 times the original number of channels. The convolutional layer is a convolution with a kernel size of 3×3, a stride of 1, and a padding of 1. The output of the color transformation preprocessing network is denoted as F 0 col .
[0112] Step S212: Design a color transformation backbone network, which contains several color transformation modules with the same structure. Taking the i-th (i≥1) module as an example, for the input feature F of the color transformation module i-1 col , first, the Feature Distillation and Channel Modulation Unit (FDCMU) predicts the weight of the color transformation, and the output result is denoted as K i , similarly, the bias of the color transformation is also obtained by FDCMU and denoted as b i . At the same time, based on F i-1 col predict the channel mapping, F i-1 col After passing through FDCMU, global average pooling (GAP) operation and multi-layer perceptron (MLP) in this way, finally, a vector with the same channel dimension as F i-1 col is denoted as g i . The output of the color transformation module is Fi col Obtained by the following transformation:
[0113] F i col = (F i-1 col * K i + b i ) exp(g i ).
[0114] Where * is the element-wise multiplication of matrices, + is the element-wise addition of matrices, and exp(·) is the exponential operation with a broadcasting mechanism.
[0115] Step S213, design a Feature Distillation and Channel Modulation Unit (FDCMU), as Figure 3 shown. For a given input feature f D , the process of obtaining the output feature f D out by the FDCMU can be described as:
[0116] f D 1a = PReLU(Conv 1×1 (DWConv 5×5 (f D ))).
[0117] f D 1b = DWConv 3×3 (Conv 1×1 (f D )).
[0118] f D 2 = PReLU(Conv 3×3 (PReLU(f D 1b ))).
[0119] f D 2b = PReLU(Conv 1×1 (f D 1b )).
[0120] f D 3 = Conv 1×1 (Concatenate(f D 1a , f D 1b , f D 2 , fD 2b )).
[0121]
[0122] Among them, Conv 1×1 is a convolution with a convolution kernel size of 1×1, a stride of 1, and a padding of 0. Conv 3×3 is a convolution with a convolution kernel size of 3×3, a stride of 1, and a padding of 1. DWConv 3×3 is a depthwise separable convolution with a convolution kernel size of 3×3, a stride of 1, and a padding of 1. DWConv 5×5 is a depthwise separable convolution with a convolution kernel size of 5×5, a stride of 1, and a padding of 2. PReLU(·) is an activation function. Concatenate(·) is feature concatenation along the channel. w c and b c are learnable parameter vectors. ⊙ and are element-wise multiplication and element-wise addition with a broadcasting mechanism.
[0123] Step S214: Design a color transformation head network. The color transformation head network receives the output from the color transformation backbone network as its input and finally outputs a color transformation result of size The color transformation head network is composed of a convolutional layer, a ReLU activation function, and a convolutional layer stacked in sequence. The convolutional layer is a convolution with a convolution kernel size of 3×3, a stride of 1, and a padding of 1.
[0124] Furthermore, step S22 includes the following steps:
[0125] Step S221: Design a fusion enhancement preprocessing network. The fusion enhancement preprocessing network consists of 1 convolutional layer. The convolutional layer is a convolution with a convolution kernel size of 3×3, a stride of 1, and a padding of 1. The output feature is denoted as F 0 fuse .
[0126] Step S222: Design a fusion enhancement backbone network. The fusion enhancement backbone network contains several fusion enhancement modules with the same structure. Taking the i-th (i≥1) module as an example, the input received by the fusion enhancement module is: the feature F i-1 fuse from the previous fusion enhancement module (when i = 1, this input is the output of the fusion enhancement preprocessing network), the feature F i col from the color transformation module, and the feature F i high from the high-frequency information reconstruction module. First, the feature F i col The output feature F is transformed through the Channel Expansion and Upsampling Module (CEUM). i col '. Then, the feature F i-1 fuse extracts and fuses multi-scale features by the Multi-scale Feature Fusion Unit (MFFU), and then adds them element-wise to the feature F i high and F i col ' respectively. After being processed by two convolutional layers, F i fuse is obtained. The convolutional layer is a convolution with a kernel size of 3×3, a stride of 1, and a padding of 1.
[0127] Step S223: Design the Channel Expansion and Upsampling Module (CEUM), as Figure 4 shown. For a given input feature f E , the process of obtaining the output feature f E out by CEUM can be described as follows:
[0128] f E 1 = Conv 1×1 (f E ).
[0129] f E 2a = FDCMU(f E 1 ).
[0130] f E 2b = FDCMU(f E 1 ).
[0131] f E 3 = Conv 3×3 (Conv 1×1 (Concatenate(f E 2a , f E 2b ))).
[0132] f E out = Conv 1×1 (Conv 3×3(PixelShuffle(f E 3 ))).
[0133] Among them, Conv 1×1 is a convolution with a convolution kernel size of 1×1, a stride of 1, and a padding of 0. Conv 3×3 is a convolution with a convolution kernel size of 3×3, a stride of 1, and a padding of 1. FDCMU is the feature distillation and channel modulation unit described in step S213. Concatenate(·) is the feature concatenation along the channel. PixelShuffle(·) is the PixelShuffle layer, and its function is to transform the input feature with a size of into H×W through pixel rearrangement, and the number of channels is then changed to 1 / 16 times the original number of channels.
[0134] Step S224: Design a multi-scale feature fusion unit (MFFU), as shown in Figure 5 . For a given input feature f M , the process of obtaining the output feature f M out from CEUM can be described as follows:
[0135] f M 1 =Conv 3×3 (f M ).
[0136] f M 2 =ResBlocks(f M 1 ).
[0137] f M 2a =Up ×2 (ResBlocks(Down ×2 (f M 1 ))).
[0138] f M 2b =Up ×4 (ResBlocks(Down ×4 (f M 1 ))).
[0139] f M out =Conv 3×3 (Conv 1×1 (Concatenate(f M 1,f M 2 ,f M 2a ,f M 2b ))).
[0140] Among them, Conv 1×1 is a convolution with a convolution kernel size of 1×1, a stride of 1, and a padding of 0. Conv 3×3 is a convolution with a convolution kernel size of 3×3, a stride of 1, and a padding of 1. Down ×4 (·) is a spatial 4-fold downsampling performed by the bilinear interpolation method. Up ×4 (·) is a spatial 4-fold upsampling performed by the bilinear interpolation method; Down ×2 (·) is a spatial 2-fold downsampling performed by the bilinear interpolation method. Up ×2 (·) is a spatial 2-fold upsampling performed by the bilinear interpolation method. ResBlocks is composed of several residual blocks (ResBlock). In this embodiment, the number of residual blocks is 15. The residual block is stacked by a convolutional layer, a PReLU activation function, and a convolutional layer, and the input and output are added by skip connection. Among them, the convolutional layer is a convolution with a convolution kernel size of 3×3, a stride of 1, and a padding of 1.
[0141] Step S225: Design a fusion enhancement head network. The fusion enhancement head network receives the output from the fusion enhancement backbone network as its input and finally outputs a fusion enhancement result of size H×W The fusion enhancement head network is sequentially stacked by a convolutional layer, a ReLU activation function, and a convolutional layer. Among them, the convolutional layer is a convolution with a convolution kernel size of 3×3, a stride of 1, and a padding of 1.
[0142] Further, step S23 includes the following steps:
[0143] Step S231: Design a high-frequency information reconstruction preprocessing network. The high-frequency information reconstruction preprocessing network is composed of 1 convolutional layer. The convolutional layer is a convolution with a convolution kernel size of 3×3, a stride of 1, and a padding of 1. The output feature is denoted as F 0 high 。
[0144] Step S232: Design a high-frequency information reconstruction backbone network. The high-frequency information reconstruction backbone network contains several high-frequency information reconstruction modules with the same structure. Taking the i-th (i≥1) module as an example, for the feature F input to the high-frequency information reconstruction module i -1 high, the features are processed by the multi-scale feature fusion unit (MFFU) and the gated Fourier transform unit (Gated FFTUnit (GFU)) described in step S224 in sequence. The output of the high-frequency information reconstruction module is denoted as F i high .
[0145] Step S224, design a gated Fourier transform unit (GFU), as Figure 6 shown. For a given input feature f G , the process of obtaining the output feature f G out by the GFU can be described as:
[0146] f G 1a = Conv 3×3 (PReLU(Conv 3×3 (f G )) + f G .
[0147] f G 1b = PReLU(Conv 1×1 (f G )).
[0148] f G 2 = IFFT(PReLU(Conv 1×1 (FFT(f G 1b )))) + f G 1b .
[0149] f G out = C onv 1×1 (Sigmoid(f G 1b ) * f G 2 ).
[0150] Among them, Conv 1×1 is a convolution with a kernel size of 1×1, a stride of 1, and a padding of 0. Conv 3×3 is a convolution with a kernel size of 3×3, a stride of 1, and a padding of 1. PReLU(·) is an activation function. FFT(·) is the Fourier transform taking the real part, IFFT(·) is the inverse Fourier transform taking the real part, and Sigmoid(·) is an activation function. * is the element-wise multiplication of matrices.
[0151] Step S225: Design a high-frequency information reconstruction head network. The high-frequency information reconstruction head network takes the output from the high-frequency information reconstruction backbone network as its input and outputs a high-frequency information reconstruction result with a size of H×W.
[0152] Furthermore, step S3 includes the following steps:
[0153] Step S31: Design the overall optimization objective of the entire network. The optimization objective is to minimize the total loss function.
[0154]
[0155] where represents the color transformation loss function, and λ col represents the weight of the color transformation loss function; represents the semantic perception loss function, and λ sem represents the weight of the semantic perception loss function; represents the pixel loss function, and λ pix represents the weight of the pixel loss function; represents the high-frequency information reconstruction loss function, and λ high represents the weight of the high-frequency information reconstruction loss function.
[0156] Step S32: Design the color transformation loss function. The calculation formula of
[0157]
[0158] is as follows: is the output result of the color transformation branch, with a size of Y ↓4 is the result of the reference image being reduced to through bilinear interpolation operation. AvgPool is an average pooling operation with a kernel size of 4, and ||·|| 1 is the absolute value operation.
[0159] Step S33: Design the semantic perception loss function. The calculation formula of
[0160]
[0161] is as follows: is the color transformation result output by the color transformation branch, Y ↓4 is the image of the reference image reduced by a factor of 4 through bilinear interpolation, is the output result of the fusion enhancement branch, and Y is the reference image. Φ(·) represents the operation of extracting the features of the Conv2-2, Conv3-2, and Conv4-2 layers using the VGG-16 model pre-trained on ImageNet. ||·|| 1 is the absolute value operation.
[0162] Step S34: Design the optical pixel loss function. The calculation formula is as follows:
[0163]
[0164] Among them, is the output result of the fusion enhancement branch, and Y is the reference image. ||·|| 1 is the absolute value operation.
[0165] Step S35: Design the high-frequency information reconstruction loss function. The calculation formula is as follows:
[0166]
[0167] Among them, is the output result of the high-frequency information reconstruction branch, Y is the reference image, HighFreq(·) is the operation to obtain the high-frequency information part, and the high-frequency information part is obtained by performing a discrete cosine transform (DCT) on the original image to the frequency domain, and then performing an inverse transform after obtaining the high frequency. ||·|| 1 is the absolute value operation.
[0168] Furthermore, step S4 includes the following steps:
[0169] Step S41: In the dataset constructed in step S1, select a random training image X.
[0170] Step S42: Training image encoding and enhancement. Input the image X, and pass it through the task-decoupled low-light and blurred image enhancement network, and calculate the output color transformation result fusion enhancement result and high-frequency information reconstruction result Calculate the loss of the total loss function in step S31
[0171] Step S43: Use the backpropagation method to calculate the gradients of the parameters in the task-decoupled low-light and blurred image enhancement network, and update the parameters using the Adam optimization method.
[0172] Step S44: The above steps are one iteration of the training process, and the entire training process requires 10 5 iterations, and in each iteration process, multiple image pairs are randomly sampled as a batch for training.
[0173] Further, step S5 includes the following steps:
[0174] Step S5: Input the low-light blurred image to be measured into the designed network, and use the output result of the trained network, where the fusion enhancement result is the final result.
[0175] The present invention proposes a method for enhancing low-light blurred images based on task decoupling. Based on the independent characteristics of the low-light and blur problems of images, the problem of low-light blurred images is creatively decoupled into color transformation and high-frequency information restoration, and corresponding tasks are designed, that is, both the color transformation task and the high-frequency information restoration task are supervised by corresponding reference images. The designed network adopts a parallel structure, learns color transformation with downsampled small images, can weaken the influence brought by noise and blur, and at the same time, the high-frequency information restoration task enables the network to learn the restoration of details independent of color transformation, which is beneficial to solving the dual problems of low-light and deblurring.
[0176] The above are the preferred embodiments of the present invention. All changes made according to the technical solution of the present invention, when the functions and effects produced do not exceed the scope of the technical solution of the present invention, shall fall within the protection scope of the present invention.
Claims
1. Dark and Blurry Image Enhancement Method Based on Task Decoupling Characterized in that It includes the following steps: Step S1: Construct a training dataset, and preprocess each image of the dark and blurry images to obtain a training dataset; Step S2: Design a dark and blurry image enhancement network with task decoupling, which is composed of a color transformation branch, a fusion enhancement branch, and a high-frequency information reconstruction branch; Step S3: Design a loss function for training the network designed in Step S2; Step S4: Use the training dataset to train the dark and blurry image enhancement network based on task decoupling; Step S5: Input the dark and blurry image to be measured into the designed network, and use the trained network to predict and generate a final result with visual perception; The said Step S2 includes the following steps: Step S21: Design a color transformation branch, which consists of a color transformation preprocessing network, a color transformation backbone network, and a color transformation head network; the color transformation branch takes the normalized low-light blurred image X with a size of H×W as input and finally outputs a color transformation result with a size of Step S22: Design a fusion enhancement branch, which consists of a fusion enhancement preprocessing network, a fusion enhancement backbone network, and a fusion enhancement head network; the input of the fusion enhancement branch is: the normalized low-light blurred image X with a size of H×W, the output features of the color transformation branch, and the output features of the high-frequency information reconstruction branch, and finally outputs a fusion enhancement result with a size of H×W Step S23: Design a high-frequency information reconstruction branch, which consists of a high-frequency information reconstruction preprocessing network, a high-frequency information reconstruction backbone network, and a high-frequency information reconstruction head network; the high-frequency information reconstruction branch takes the normalized low-light and blurred image X with a size of H×W as input and finally outputs a high-frequency information reconstruction result with a size of H×W In the said Step S3, designing a loss function for training the network designed in Step S2; includes the following steps: Step S31: Design the overall optimization objective of the entire network; the optimization objective is to minimize the total loss function Among them, represents the color transformation loss function, and λ col represents the weight of the color transformation loss function; represents the semantic perception loss function, and λ sem represents the weight of the semantic perception loss function; represents the pixel loss function, and λ pix represents the weight of the pixel loss function; represents the high-frequency information reconstruction loss function, and λ high represents the weight of the high-frequency information reconstruction loss function; Step S32: Design a color transformation loss function; The calculation formula is as follows: Among them, is the output result of the color transformation branch, with a size of Y ↓4 is the result of the reference image being reduced to through bilinear interpolation operation. AvgPool(·) is an average pooling operation with a kernel size of 4, and ||·|| 1 is an absolute value operation; Step S33: Design a semantic perception loss function; The calculation formula is as follows: Among them is the color transformation result output by the color transformation branch, Y ↓4 is the image obtained by reducing the reference image by a factor of 4 through bilinear interpolation, is the output result of the fusion enhancement branch, Y is the reference image; Φ(·) represents the operation of extracting the features of the Conv2-2, Conv3-2, and Conv4-2 layers using the VGG-16 model pre-trained on ImageNet; ||·|| 1 is the absolute value operation; Step S34: Design a light pixel loss function; The calculation formula thereof is as follows: Among them, is the output result of the fusion enhancement branch, and Y is the reference image; ||·|| 1 is the absolute value operation; Step S35: Design a high-frequency information reconstruction loss function; The calculation formula is as follows: wherein, is the output result of the high-frequency information reconstruction branch, Y is the reference image, HighFreq(·) is the operation to obtain the high-frequency information part, and the high-frequency information part is obtained by transforming the original image to the frequency domain through the discrete cosine transform (DCT), and then performing the inverse transform after obtaining the high-frequency components; ||·|| 1 is the absolute value operation.
2. The dark and blurry image enhancement method based on task decoupling according to Claim 1 Characterized in that The said Step S1 includes the following steps: Step S11: Scale each image in the dataset to an image with the same size of H×W; Step S12: Normalize the training images; given an image I train , calculate the normalized image using the following formula: where I train is an image of size H×W with an 8-bit color depth, and I bit_max is an image of size H×W with all pixel values equal to 255.
3. The dark and blurry image enhancement method based on task decoupling according to Claim 1 Characterized in that The said Step S21 specifically includes the following steps: Step S211: Design a color transformation preprocessing network, which includes 1 PixelUnShuffle layer and 1 convolutional layer; among them, the PixelUnShuffle layer performs pixel recombination on the low-light blurred image X, and changes the size of the input low-light blurred image X to The number of channels is changed to 16 times the original number of channels. The convolutional layer is a convolution with a kernel size of 3×3, a stride of 1, and a padding of 1. The output of the color transformation preprocessing network is denoted as F 0 col ; Step S212: Design a color transformation backbone network. The color transformation backbone network includes several color transformation modules with the same structure. Taking the i-th module as an example, where i≥1, for the feature F input to the color transformation module i-1 col , first, the feature distillation and channel modulation unit FDCMU predicts the weight of color transformation, and the output result is denoted as K i . The bias of color transformation is also obtained by FDCMU and denoted as b i ; meanwhile, based on F i-1 col predicts the channel mapping. F i-1 col After passing through FDCMU, global average pooling GAP operation and multi-layer perceptron MLP, finally, a vector with the same channel dimension as F i-1 col is obtained and denoted as g i ; the output F of the color transformation module i col is obtained by the following transformation: F i col =(F i-1 col *K i +b i ).exp(g i ). where * is the element-wise multiplication of matrices, + is the element-wise addition of matrices, and exp(·) is the exponential operation with a broadcasting mechanism; Step S213, design the Feature Distillation and Channel Modulation Unit (FDCMU); for a given input feature f D , the process of obtaining the output feature f D out by the FDCMU is described as follows: f D 1a = PReLU(Conv 1×1 (DWConv 5×5 (f D ))). f D 1b = DWConv 3×3 (Conv 1×1 (f D )) f D 2 = PReLU(Conv 3×3 (PReLU(f D 1b ))). f D 2b = PReLU(Conv 1×1 (f D 1b )) f D 3 = Conv 1×1 (Concatenate(f D 1a , f D 1b , f D 2 , f D 2b )) Among them, Conv 1×1 is a convolution with a convolution kernel size of 1×1, a stride of 1, and a padding of 0. Conv 3×3 is a convolution with a convolution kernel size of 3×3, a stride of 1, and a padding of 1. DWConv 3×3 is a depthwise separable convolution with a convolution kernel size of 3×3, a stride of 1, and a padding of 1. DWConv 5×5 is a depthwise separable convolution with a convolution kernel size of 5×5, a stride of 1, and a padding of 2. PReLU(·) is an activation function; Concatenate(·) is a feature concatenation along the channels; w c and b c are learnable parameter vectors. ⊙ and are element-wise multiplication and element-wise addition with a broadcasting mechanism; Step S214: Design a color transformation head network; the color transformation head network receives the output from the color transformation backbone network as its input and finally outputs a color transformation result of size The color transformation head network is composed of convolutional layers, ReLU activation functions, and convolutional layers stacked in sequence. Among them, the convolutional layer is a convolution with a kernel size of 3×3, a stride of 1, and a padding of 1. 4. The dark and blurry image enhancement method based on task decoupling according to Claim 1 Characterized in that The said Step S22 includes the following steps: Step S221: Design a fusion enhancement preprocessing network, which consists of 1 convolutional layer. The convolutional layer is a convolution with a kernel size of 3×3, a stride of 1, and a padding of 1; the output feature is denoted as F 0 fuse ; Step S222: Design a fusion-enhanced backbone network. The fusion-enhanced backbone network contains several fusion-enhanced modules with the same structure. Taking the i-th module as an example, where i ≥ 1, the input received by the fusion-enhanced module is: the feature F from the previous fusion-enhanced module i-1 fuse , when i = 1, this input is the output of the fusion-enhanced preprocessing network, the feature F from the color transformation module i col , and the feature F from the high-frequency information reconstruction module i high ; First, the feature F from the color transformation module i col is transformed through the channel expansion and upsampling module CEUM to output the feature F i col′ ; Then, the feature F i-1 fuse extracts and fuses multi-scale features by the multi-scale feature fusion unit MFFU, and then adds them element-wise to the features F i high and F i col′ respectively, and after being processed by two convolutional layers, F i fuse is obtained; The convolutional layer is a convolution with a kernel size of 3×3, a stride of 1, and a padding of 1 Step S223, design the channel expansion and upsampling module CEUM; for a given input feature f E , the process of obtaining the output feature f E out is described as follows: f E 1 = Conv 1×1 (f E ). f E 2a = FDCMU(f E 1 ). f E 2b = FDCMU(f E 1 ). f E 3 = Conv 3×3 (Conv 1×1 (Concatenate(f E 2a , f E 2b ))). f E out = Conv 1×1 (Conv 3×3 (PixelShuffle(f E 3 ))). Among them, Conv 1×1 is a convolution with a convolution kernel size of 1×1, a stride of 1, and a padding of 0. Conv 3×3 is a convolution with a convolution kernel size of 3×3, a stride of 1, and a padding of 1. FDCMU is the feature distillation and channel modulation unit described in step S213; Concatenate(·) is the feature concatenation along the channels; PixelShuffle(·) is the PixelShuffle layer, and its function is to transform the input feature with a size of into H×W through pixel rearrangement, and the number of channels is then transformed into 1 / 16 times the original number of channels. Step S224, design a multi-scale feature fusion unit MFFU; for a given input feature f M , the process of obtaining the output feature f M out from CEUM is described as follows: f M 1 = Conv 3×3 (f M ). f M 2 = ResBlocks(f M 1 ). f M 2a = Up ×2 (ResBlocks(Down ×2 (f M 1 ))). f M 2b = Up ×4 (ResBlocks(Down ×4 (f M 1 ))). f M out = Conv 3×3 (Conv 1×1 (Concatenate(f M 1 , f M 2 , f M 2a , f M 2b ))). Among them, Conv 1×1 is a convolution with a convolution kernel size of 1×1, a stride of 1, and a padding of 0. Conv 3×3 is a convolution with a convolution kernel size of 3×3, a stride of 1, and a padding of 1; Down ×4 (·) is a spatial 4-fold downsampling performed by the bilinear interpolation method. Up ×4 (·) is a spatial 4-fold upsampling performed by the bilinear interpolation method; Down ×2 (·) is a spatial 2-fold downsampling performed by the bilinear interpolation method. Up ×2 (·) is a spatial 2-fold upsampling performed by the bilinear interpolation method; ResBlocks is composed of a number of residual blocks (ResBlock). In this embodiment, the number of residual blocks is 15; the residual block is stacked by a convolutional layer, a PReLU activation function, and a convolutional layer, and the input and output are added by skip connection. Among them, the convolutional layer is a convolution with a convolution kernel size of 3×3, a stride of 1, and a padding of 1; Step S225: Design a fusion-enhanced head network; the fusion-enhanced head network receives the output from the fusion-enhanced backbone network as its input and finally outputs a fusion-enhanced result of size H×W. The fusion-enhanced head network is composed of a convolutional layer, a ReLU activation function, and a convolutional layer stacked in sequence, where the convolutional layer is a convolution with a kernel size of 3×3, a stride of 1, and a padding of 1.
5. The dark and blurry image enhancement method based on task decoupling according to Claim 4 Characterized in that The said Step S23 specifically includes the following steps: Step S231: Design a preprocessing network for high-frequency information reconstruction. The preprocessing network for high-frequency information reconstruction consists of 1 convolutional layer. The convolutional layer is a convolution with a kernel size of 3×3, a stride of 1, and a padding of 1. The output feature is denoted as F 0 high ; Step S232: Design a high-frequency information reconstruction backbone network. The high-frequency information reconstruction backbone network includes several high-frequency information reconstruction modules with the same structure. Taking the i-th module as an example, where i≥1, for the feature F input to the high-frequency information reconstruction module i-1 high , the feature is processed sequentially by the multi-scale feature fusion unit (MFFU) and the gated Fourier transform unit GFU described in step S224; the output of the high-frequency information reconstruction module is denoted as F i high ; Step S233, design the gated Fourier transform unit GFU; for a given input feature f G , the process of obtaining the output feature f G out is described as follows: f G 1a = Conv 3×3 (PReLU(Conv 3×3 (f G )) + f G ). f G 1b = PReLU(Conv 1×1 (f G )). f G 2 = IFFT(PReLU(Conv 1×1 (FFT(f G 1b )))) + f G 1b . f G out = Conv 1×1 (Sigmoid(f G 1b ) * f G 2 ). Among them, Conv 1×1 is a convolution with a convolution kernel size of 1×1, a stride of 1, and a padding of 0. Conv 3×3 is a convolution with a convolution kernel size of 3×3, a stride of 1, and a padding of 1. PReLU(·) is an activation function; FFT(·) is the Fourier transform taking the real part, IFFT(·) is the inverse Fourier transform taking the real part, and Sigmoid(·) is an activation function; * is the element-wise multiplication of matrices; Step S234: Design a high-frequency information reconstruction head network; the high-frequency information reconstruction head network receives the output from the high-frequency information reconstruction backbone network as its input and outputs a high-frequency information reconstruction result with a size of H×W 6. The dark and blurry image enhancement method based on task decoupling according to Claim 1 Characterized in that The said Step S4 includes the following steps: Step S41: In the dataset constructed in Step S1, select a random training image X; Step S42, training image encoding and enhancement; input image X passes through the task-decoupled low-light blurred image enhancement network, and the output color transformation result Fusion enhancement result And high-frequency information reconstruction result Calculate the total loss function loss in step S31 Step S43: Use the backpropagation method to calculate the gradients of each parameter in the dark and blurry image enhancement network with task decoupling, and update the parameters using the Adam optimization method; Step S44: The above steps are one iteration of the training process, and the entire training process requires 10 5 iterations. During each iteration, multiple image pairs are randomly sampled as a batch for training.
7. The dark and blurry image enhancement method based on task decoupling according to Claim 1 Characterized in that In the step S5, the to-be-tested low-light blurred image is input into the designed network, and the output result of the trained network is utilized, wherein the fusion enhancement result is the final result.
Citation Information
Patent Citations
Self-adaptive enhancement method based on dark-light color image
CN108389163A
A low-light image enhancement method and apparatus
CN109087269A