A multi-branch low-illumination image enhancement method based on frequency domain division

By adopting a multi-branch method based on frequency domain frequency division in low-illuminance image enhancement technology, processing branches for high-frequency and low-frequency information are designed, and the UNet structure of multi-scale features coordinated attention module is used to solve the problems of detail and noise processing in the prior art, and a more efficient low-illuminance image enhancement effect is achieved.

CN116363011BActive Publication Date: 2025-05-13FUZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310356172.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-06
Publication Date
2025-05-13
Estimated Expiration
2043-04-06

AI Technical Summary

Technical Problem

Existing low-illumination image enhancement technology is difficult to effectively distinguish and process details and noise in the image, which easily leads to details loss and noise amplification, and lacks targeted processing of global and local information.

Method used

The multi-branch low-illumination image enhancement method based on frequency domain frequency division is adopted. By designing the reference branch, high-frequency branch and low-frequency branch, it includes a shallow image enhancement module and a UNet structure based on multi-scale feature collaborative attention module, and targeted enhancement of the high-frequency information, low-frequency information and overall information of the image, and constrain the learning of high-frequency and low-frequency information by the loss function.

Benefits of technology

The performance of low-illumination image enhancement is significantly improved, and the details and noise in the image can be better separated and processed, noise amplified and detailed loss is prevented, and effective processing of global and local information is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116363011B_ABST
    Figure CN116363011B_ABST
Patent Text Reader

Abstract

The present invention relates to a multi-branch low-illumination image enhancement method based on frequency domain division. The method comprises: preprocessing the input image, including data pairing, random data cropping, and data enhancement processing, to obtain a training data set; designing a multi-branch low-illumination image enhancement network based on frequency domain division, the network consisting of three branches: a reference branch, a high-frequency branch, and a low-frequency branch, each branch including a shallow image enhancement module and a UNet structure based on a multi-scale feature collaborative attention module; designing a loss function for guiding the optimization of the network parameters designed in step B; using the training data set obtained in step A to train the multi-branch low-illumination image enhancement network based on frequency domain division in step B, to obtain a trained multi-branch low-illumination image enhancement model based on frequency domain division; inputting the low-illumination image to be tested into the trained multi-branch low-illumination image enhancement model based on frequency domain division, and predicting and generating a normal illumination image. The present invention can enhance low-illumination images and solve the problems of detail loss, high noise, and insufficient brightness of low-illumination images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the field of image and video processing and computer vision technology, in particular to a multi-branch low-illumination image enhancement method based on frequency domain division. Background Art

[0002] Low-light image enhancement is an important research direction in computer vision. Its purpose is to improve the brightness and contrast of images, remove noise, correct colors, etc., so that images are clearer and brighter, which is convenient for human observation and machine recognition. Low-light image enhancement technology plays an important role in many fields: it can improve the quality of night-time pictures; improve the quality of medical images, thereby improving the diagnosis and treatment level of doctors; improve the recognition ability of unmanned vehicles for nighttime road conditions, which helps to improve driving safety; improve the picture quality of surveillance cameras, thereby improving the monitoring effect. Therefore, it is very important to improve the level of low-light image enhancement technology.

[0003] Early histogram equalization methods improve the contrast and brightness of images by adjusting the grayscale value distribution of images to a uniform distribution. Model-based methods enhance images by establishing physical or statistical models of images. Multi-image fusion-based methods use multiple images with different exposure times or angles to fuse enhanced images. These traditional methods usually enhance low-light images by increasing brightness, enhancing contrast, etc., but this easily causes image details to be lost, and traditional methods are difficult to deal with some complex scenes. Therefore, in recent years, more and more deep learning-based methods have been proposed. Based on the global / local self-attention mechanism, local and global information in the image are learned and enhanced separately. The residual learning-based method learns the mapping relationship between input and output. Based on the generative adversarial network, the generator and discriminator are trained to generate realistic images. However, the details, noise and other information in low-light images are not completely unrelated. Enhancing details can also amplify noise at the same time, while removing noise may cause detail loss. In addition, the processing of details, texture, noise information and brightness, color, and scene often needs to be carried out separately.

[0004] Existing methods cannot distinguish between details and noise, and process details and noise in the same way, which easily leads to detail loss and noise amplification. In addition, some methods jointly process global information such as brightness, color, and scene, and local information such as details, texture, and noise, which lacks specificity. Summary of the invention

[0005] In view of this, the purpose of the present invention is to provide a multi-branch low-illumination image enhancement method based on frequency domain division, which specifically enhances the high-frequency information, low-frequency information and high- and low-frequency overall information of the image. Each branch first designs a shallow image enhancement module to preliminarily enhance the low-illumination image and improve its visibility in subsequent processing steps. Then, a UNet structure based on a multi-scale feature collaborative attention module is designed to further enhance the image. Finally, the processing results of the three branches are fused to obtain the final enhanced result. This method specifically enhances the high-frequency information and low-frequency information of the image, which is beneficial to significantly improve the performance of low-illumination image enhancement.

[0006] To achieve the above object, the present invention adopts the following technical solution: a multi-branch low-illumination image enhancement method based on frequency domain division, comprising the following steps:

[0007] Step A: preprocess the input image, including data pairing, random data cropping, and data enhancement, to obtain a training data set;

[0008] Step B, design a multi-branch low-light image enhancement network based on frequency domain division, which consists of three branches: a baseline branch, a high-frequency branch, and a low-frequency branch. Each branch contains a shallow image enhancement module and a UNet structure based on a multi-scale feature collaborative attention module;

[0009] Step C, designing a loss function for guiding the optimization of the network parameters designed in step B;

[0010] Step D: using the training data set obtained in step A to train the multi-branch low-illumination image enhancement network based on frequency domain division in step B, to obtain a trained multi-branch low-illumination image enhancement model based on frequency domain division;

[0011] Step E: input the low-illuminance image to be tested into the trained multi-branch low-illuminance image enhancement model based on frequency domain division to predict and generate a normal-illuminance image.

[0012] In a preferred embodiment, the specific implementation steps of step A are as follows:

[0013] Step A1: pairing the low-light image with its corresponding label image;

[0014] Step A2: randomly select a cropping area of ​​p×p×3 from each low-light image of size H×W×3 for cropping, and the corresponding label image is also cropped in the same way, where H and W are the height and width of the low-light image, respectively, and p is the height and width of the cropping area;

[0015] Step A3: Randomly perform horizontal flipping, vertical flipping, and rotation operations on the training paired images for data enhancement.

[0016] In a preferred embodiment, the specific implementation steps of step B are as follows:

[0017] Step B1, designing a shallow image enhancement module to preliminarily enhance low-light images;

[0018] Step B2: Design a UNet structure based on a multi-scale feature collaborative attention module to further enhance features;

[0019] Step B3, design a multi-branch low-light image enhancement network based on frequency domain division, including three branches: baseline branch, high-frequency branch and low-frequency branch. Each branch contains a shallow image enhancement module and a UNet structure based on a multi-scale feature collaborative attention module.

[0020] In a preferred embodiment, the specific implementation steps of step B1 are as follows:

[0021] Design a shallow image enhancement module with low-light images as input Perform three average poolings in the channel dimension to obtain three feature maps of size H×W×1. After splicing the three feature maps along the channel dimension, they pass through 3×3 convolution, batch normalization layer, ReLU activation function, 3×3 convolution, and Sigmoid activation function in sequence, and then multiply them pixel by pixel with the low-light image I to obtain the weight of each position, and then add them pixel by pixel with the low-light image I to obtain the output image. The specific formula is as follows:

[0022]

[0023] Where I′ is the output image of the shallow image enhancement module, I is the input low-light image, Sigmoid is the Sigmoid activation function, Conv3 is the 3×3 convolution, ReLU is the ReLU activation function, BN is batch normalization, Concat is the concatenation operation along the channel dimension, AvgPool is the average pooling of the channel dimension, and AvgPool(I) is 3 The average pooling of the channel dimension is performed 3 times. It is a pixel-by-pixel multiplication operation.

[0024] In a preferred embodiment, the specific implementation steps of step B2 are as follows:

[0025] Step B21, design a multi-scale feature collaborative attention module, which consists of a channel collaborative attention submodule and a spatial collaborative attention submodule;

[0026] Step B22, design a UNet structure based on a multi-scale feature collaborative attention module, which consists of three encoders E1, E2, E3, three decoders D1, D2, D3 and six multi-scale feature collaborative attention modules; each encoder and decoder consists of 3×3 convolution, ReLU activation function, 3×3 convolution, ReLU activation function; let the input be the feature map f, after passing through encoder E1, the feature map is obtained Then enter the multi-scale feature collaborative attention module in step B21 to obtain the feature map After the feature map F′1 is max-pooled, it passes through encoder E2 to obtain the feature map Then enter the multi-scale feature collaborative attention module in step B21 to obtain the feature map After the maximum pooling of feature map F′2, it passes through encoder E3 to obtain feature map Then enter the multi-scale feature collaborative attention module in step B21 to obtain the feature map After the feature map F′3 passes through the decoder D1, the feature map is obtained Then enter the multi-scale feature collaborative attention module in step B21 to obtain the feature map The feature map F′4 is deconvolved and added to the feature map F2, and then passes through the decoder D2 to obtain the feature map Then enter the multi-scale feature collaborative attention module in step B21 to obtain the feature map The feature map F′5 is deconvolved and added to the feature map F1, and then passes through the decoder D3 to obtain the feature map Then enter the multi-scale feature collaborative attention module in step B21 to obtain the output feature map The specific formula is as follows:

[0027] F1=E1(f)

[0028] F′1=att(F1)

[0029] F2=E2(Maxpool(F′1))

[0030] F′2=att(F2)

[0031] F3=E3(Maxpool(F′2))

[0032] F′3=att(F3)

[0033] F4=D1(F′3)

[0034] F′4=att(F4)

[0035] F5=D2(deconv(F′4)+F2)

[0036] F′5=att(F5)

[0037] F6=D3(deconv(F′5)+F1)

[0038] f′=att(F6)

[0039] Among them, f is the input feature map of the UNet structure based on the multi-scale feature collaborative attention module, f′ is the output feature map of the UNet structure based on the multi-scale feature collaborative attention module, E1, E2, E3 are encoders, D1, D2, D3 are decoders, att is the multi-scale feature collaborative attention module, Maxpool is the maximum pooling layer, deconv is the deconvolution operation, F1, F2, F3 are the output feature maps of encoders E1, E2, E3 respectively, F4, F5, F6 are the output feature maps of decoders D1, D2, D3 respectively, and F′ i ,i=1,2,3…6 are the output feature maps of the multi-scale feature collaborative attention module at different stages.

[0040] In a preferred embodiment, the specific implementation steps of step B21 are as follows:

[0041] Step B211: Design a channel collaborative attention submodule with three feature maps F as input 11 、F 12 、F 13 After the three feature maps are averaged in the spatial dimension, three vectors of size 1×1×Cf are obtained. After splicing along the channel dimension, a vector of size 1×1×3Cf is obtained. Then, 1×1 convolution, batch normalization, ReLU activation function, 1×1 convolution, and Sigmoid activation function are successively passed. The vector of size 1×1×3Cf is then divided into three vectors of size 1×1×Cf along the channel dimension. The input feature map F 11 、F 12 、F 13 After multiplying these three vectors of scale 1×1×Cf respectively, the output feature map F′ of the channel collaborative attention submodule is obtained 11 , F′ 12 , F′ 13 ;

[0042] Step B212: Design a spatial collaborative attention submodule, with the input being the three feature maps F′ obtained in step B221 11 , F′ 12 , F′ 13, after performing average pooling on the three feature maps in the channel dimension, three feature maps of size Hf×Wf×1 are obtained. After splicing along the channel dimension, a feature map of size Hf×Wf×3 is obtained. Then, 1×1 convolution, batch normalization, ReLU activation function, 1×1 convolution, and Sigmoid activation function are successively performed. The feature map of size HF×Wf×3 is then divided into three vectors of size Hf×Wf×1 along the channel dimension, and the input feature map F′ is converted into 11 , F′ 12 , F′ 13 The output feature map F″ of the spatial collaborative attention submodule is obtained by multiplying these three vectors of scale Hf×Wf×1 respectively. 11 , F″ 12 , F″ 13 ;

[0043] Step B213, design a multi-scale feature collaborative attention module, the input feature map is the encoder E in step B21 i ,i=1,2,3,decoder D i , output feature map F of i=1,2,3 i ,i=1,2,…6, here it is uniformly denoted as F, After F passes through 3×3 convolution, 5×5 convolution, and 7×7 convolution, we get three feature maps. F 11 、F 12 、F 13 Input channels coordinate attention to obtain feature map F′ 11 , F′ 12 , F′ 13 , F′ 11 , F′ 12 , F′ 13 Input spatial collaborative attention to obtain feature map F″ 11 , F″ 12 , F″ 13 , the feature map F″ 11 , F″ 12 , F″ 13 The output feature map F′ of the multi-scale feature collaborative attention module is obtained by adding them together; the specific formula is as follows:

[0044] F 11 =Conv3(F)

[0045] F 12 =Conv5(F)

[0046] F 13 =Conv7(F)

[0047] F′ 11,F′ 12 ,F′ 13 =att c (F 11 ,F 12 ,F 13 )

[0048] F" 11 ,F" 12 ,F" 13 =att s (F 11 ,F 12 ,F 13 )

[0049] F′=F" 11 +F" 12 +F" 13

[0050] Among them, Conv3, Conv5, and Conv7 are 3×3 convolution, 5×5 convolution, and 7×7 convolution respectively, F is the input feature map, F′ is the output feature map of the multi-scale feature collaborative attention module, and att c is the channel collaborative attention submodule, att s It is the spatial collaborative attention submodule.

[0051] In a preferred embodiment, the specific implementation steps of step B3 are as follows:

[0052] Step B3, design a multi-branch low-light image enhancement network based on frequency domain division, which consists of three branches: reference branch, high-frequency branch and low-frequency branch; the input image is a low-light image I, and the three branches respectively pass the input low-light image I through the shallow image enhancement module of step B1 to obtain preliminary enhanced images I′1, I′2, and I′3. After the images I′1, I′2, and I′3 are subjected to 3×3 convolution, the feature map is obtained. After the feature maps f1, f2, and f3 enter the UNet structure based on the multi-scale feature collaborative attention module designed in step B2 respectively, the feature maps are obtained. After 3×3 convolution, the feature maps f′1, f′2, and f′3 are added to the input low-light image I to obtain the output images of the reference branch, high-frequency branch, and low-frequency branch, respectively. At the same time, the feature maps f′1, f′2, and f′3 are concatenated along the channel dimension, and after a 3×3 convolution, they are added to the input low-light image I to obtain the final enhanced image I o ; The specific formula is as follows:

[0053] I′1=s1(I)

[0054] I′2=s2(I)

[0055] I′3=s3(I)

[0056] f1=Conv3(I′1)

[0057] f2=Conv3(I′2)

[0058] f3=Conv3(I′3)

[0059] f′1=UNet1(f1)

[0060] f′2=UNet2(f2)

[0061] f′3=UNet3(f3)

[0062]

[0063]

[0064]

[0065] I o =Conv3(Concat(f′1,f′2,f ′ ))+I

[0066] Where I is the input low-light image, I′1, I′2, I′3 are the output images of the shallow image enhancement modules on the benchmark branch, high-frequency branch and low-frequency branch, respectively, f1, f2, f3 are the intermediate feature maps on the benchmark branch, high-frequency branch and low-frequency branch, f′1, f′2, f′3 are the output feature maps of the UNet structure based on the multi-scale feature collaborative attention module on the benchmark branch, high-frequency branch and low-frequency branch, s1, s2, s3 are the shallow image enhancement modules on the benchmark branch, high-frequency branch and low-frequency branch, UNet1, UNet2, UNet3 are the UNet structures based on the multi-scale feature collaborative attention module on the benchmark branch, high-frequency branch and low-frequency branch, Conv3 is a 3×3 convolution, They are the initial enhanced image, high-frequency information enhanced image, and low-frequency information enhanced image output by the baseline branch, high-frequency branch, and low-frequency branch, respectively. Concat is splicing along the channel dimension. o It is the output image of the multi-branch low-light image enhancement network based on frequency domain division.

[0067] In a preferred embodiment, the specific implementation of step C is as follows:

[0068] Step C: Design the loss function, which consists of L2 loss and VGG perceptual loss.

[0069] The standard loss function is as follows:

[0070]

[0071] Where Φ(·) represents the operation of extracting Conv4-1 layer features using the VGG-16 classification model pre-trained on the ImageNet dataset; are the initial enhanced image, high-frequency information enhanced image, and low-frequency information enhanced image output by the baseline branch, high-frequency branch, and low-frequency branch, respectively. o is the output image of the multi-branch low-light image enhancement network based on frequency domain division, G is the label image corresponding to the low-light image I, G1 and G2 are the high-frequency information label image and low-frequency information label image extracted using discrete cosine transform, respectively, ||.||1 represents the L1 loss, ||.|| 2 represents L2 loss.

[0072] In a preferred embodiment, the specific implementation steps of step D are as follows:

[0073] The low-light images are divided into batches N, and the training data set in step A is divided into K batches; the low-light image I is input into the multi-branch low-light image enhancement network based on frequency domain division in step B, and the initial enhanced image output by the reference branch, high-frequency branch and low-frequency branch is obtained. High-frequency information enhances images Low frequency information enhanced image And the final enhanced image I o ; Using the loss function designed in step D, calculate the initial enhanced image output by the baseline branch, high-frequency branch, and low-frequency branch High-frequency information enhances images Low frequency information enhanced image And the final enhanced image I o The Adam optimization method is used to update the network parameters until convergence, and a multi-branch low-light image model based on frequency domain division is obtained.

[0074] In a preferred embodiment, the specific implementation steps of step E are as follows:

[0075] The low-light image to be tested is input into the trained multi-branch low-light image enhancement model based on frequency domain division, and the initial enhanced image output by the reference branch, high-frequency branch and low-frequency branch is predicted. High-frequency information enhances images Low frequency information enhanced image And the final enhanced image I o , take the final enhanced image I o is the generated normal illumination image.

[0076] Compared with the prior art, the present invention has the following beneficial effects: the present invention processes the high-frequency information and low-frequency information of the image separately, which can better separate high-frequency information such as detail texture and low-frequency information such as color brightness, and process them in a targeted manner. Moreover, the noise and detail information can be well distinguished through the setting of filters to prevent noise amplification and detail loss. A multi-branch low-light image enhancement network based on frequency domain division is designed. On the three branches, a shallow image enhancement module is first designed to preliminarily enhance the low-light image, which is conducive to enhancing the visibility of subsequent processing. Then, the features are further enhanced through the UNet structure based on the multi-scale feature collaborative attention module. The loss function constrains the image to learn high-frequency and low-frequency information. Unlike other methods that generally learn global information such as color and scene and local information such as details and noise, this method can better separate global and local information, and distinguish details from noise, so as to enhance in a targeted manner. BRIEF DESCRIPTION OF THE DRAWINGS

[0077] Figure 1 It is a flow chart of the implementation of the method in the preferred embodiment of the present invention.

[0078] Figure 2 It is a structural diagram of a multi-branch low-illumination image enhancement network based on frequency domain division in a preferred embodiment of the present invention.

[0079] Figure 3 It is a structural diagram of the UNet structure based on the multi-scale feature collaborative attention module in the preferred embodiment of the present invention.

[0080] Figure 4 It is a structural diagram of the multi-scale feature collaborative attention module in the preferred embodiment of the present invention. DETAILED DESCRIPTION

[0081] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0082] It should be noted that the following detailed descriptions are illustrative and are intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present application belongs.

[0083] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application; as used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or their combinations.

[0084] The present invention provides a multi-branch low-illumination image enhancement method based on frequency domain division, such as Figure 1-4 As shown, the following steps are included:

[0085] Step A: preprocess the input image, including data pairing, random data cropping, and data enhancement, to obtain a training data set;

[0086] Step B, design a multi-branch low-light image enhancement network based on frequency domain division, which consists of three branches: a baseline branch, a high-frequency branch, and a low-frequency branch. Each branch contains a shallow image enhancement module and a UNet structure based on a multi-scale feature collaborative attention module;

[0087] Step C, designing a loss function for guiding the optimization of the network parameters designed in step B;

[0088] Step D: using the training data set obtained in step A to train the multi-branch low-illumination image enhancement network based on frequency domain division in step B, to obtain a trained multi-branch low-illumination image enhancement model based on frequency domain division;

[0089] Step E: input the low-illuminance image to be tested into the trained multi-branch low-illuminance image enhancement model based on frequency domain division to predict and generate a normal-illuminance image.

[0090] Furthermore, the step A comprises the following steps:

[0091] Step A1: pairing the low-light image with its corresponding label image;

[0092] Step A2: randomly select a cropping area of ​​p×p×3 from each low-light image of size H×W×3 for cropping, and the corresponding label image is also cropped in the same way, where H and W are the height and width of the low-light image, respectively, and p is the height and width of the cropping area;

[0093] Step A3: Randomly perform horizontal flipping, vertical flipping, and rotation operations on the training paired images for data enhancement.

[0094] Furthermore, the step B comprises the following steps:

[0095] Step B1, designing a shallow image enhancement module to preliminarily enhance low-light images;

[0096] Step B2: Design a UNet structure based on a multi-scale feature collaborative attention module to further enhance features;

[0097] Step B3, design a multi-branch low-light image enhancement network based on frequency domain division, including three branches: baseline branch, high-frequency branch and low-frequency branch. Each branch contains a shallow image enhancement module and a UNet structure based on a multi-scale feature collaborative attention module.

[0098] Furthermore, the step B1 comprises the following steps:

[0099] Step B1: Design a shallow image enhancement module, with low-light image as input Perform three average poolings in the channel dimension to obtain three feature maps of size H×W×1. After splicing the three feature maps along the channel dimension, they pass through 3×3 convolution, batch normalization layer, ReLU activation function, 3×3 convolution, and Sigmoid activation function in sequence, and then multiply them pixel by pixel with the low-light image I to obtain the weight of each position, and then add them pixel by pixel with the low-light image I to obtain the output image. The specific formula is as follows:

[0100]

[0101] Where I′ is the output image of the shallow image enhancement module, I is the input low-light image, Sigmoid is the Sigmoid activation function, Conv3 is the 3×3 convolution, ReLU is the ReLU activation function, BN is batch normalization, Concat is the concatenation operation along the channel dimension, AvgPool is the average pooling of the channel dimension, and AvgPool(I) is 3 The average pooling of the channel dimension is performed 3 times. It is a pixel-by-pixel multiplication operation.

[0102] Furthermore, the step B2 comprises the following steps:

[0103] Step B21, design a multi-scale feature collaborative attention module, which consists of a channel collaborative attention submodule and a spatial collaborative attention submodule;

[0104] Step B22, design a UNet structure based on a multi-scale feature collaborative attention module, which consists of three encoders E1, E2, E3, three decoders D1, D2, D3 and six multi-scale feature collaborative attention modules. Each encoder and decoder consists of 3×3 convolution, ReLU activation function, 3×3 convolution, ReLU activation function. Let the input be the feature map f, after passing through encoder E1, the feature map is obtained Then enter the multi-scale feature collaborative attention module in step B21 to obtain the feature map After the feature map F′1 is max-pooled, it passes through encoder E2 to obtain the feature map Then enter the multi-scale feature collaborative attention module in step B21 to obtain the feature map After the maximum pooling of feature map F′2, it passes through encoder E3 to obtain feature map Then enter the multi-scale feature collaborative attention module in step B21 to obtain the feature map After the feature map F′3 passes through the decoder D1, the feature map is obtained Then enter the multi-scale feature collaborative attention module in step B21 to obtain the feature map The feature map F′4 is deconvolved and added to the feature map F2, and then passes through the decoder D2 to obtain the feature map Then enter the multi-scale feature collaborative attention module in step B21 to obtain the feature map The feature map F′5 is deconvolved and added to the feature map F1, and then passes through the decoder D3 to obtain the feature map Then enter the multi-scale feature collaborative attention module in step B21 to obtain the output feature map The specific formula is as follows:

[0105] F1=E1(f)

[0106] F′1=att(F1)

[0107] F2=E2(Maxpool(F′1))

[0108] F′2=att(F2)

[0109] F3=E3(Maxpool(F′2))

[0110] F′3=att(F3)

[0111] F4=D1(F′3)

[0112] F′4=att(F4)

[0113] F5=D2(deconv(F′4)+F2)

[0114] F′5=att(F5)

[0115] F6=D3(deconv(F′5)+F1)

[0116] f′=att(F6)

[0117] Among them, f is the input feature map of the UNet structure based on the multi-scale feature collaborative attention module, f′ is the output feature map of the UNet structure based on the multi-scale feature collaborative attention module, E1, E2, E3 are encoders, D1, D2, D3 are decoders, att is the multi-scale feature collaborative attention module, Maxpool is the maximum pooling layer, deconv is the deconvolution operation, F1, F2, F3 are the output feature maps of encoders E1, E2, E3 respectively, F4, F5, F6 are the output feature maps of decoders D1, D2, D3 respectively, and F′ i ,i=1,2,3…6 are the output feature maps of the multi-scale feature collaborative attention module at different stages.

[0118] Furthermore, the step B21 includes the following steps:

[0119] Step B211: Design a channel collaborative attention submodule with three feature maps F as input 11 、F 12 、F 13 After the three feature maps are averaged in the spatial dimension, three vectors of size 1×1×Cf are obtained. After splicing along the channel dimension, a vector of size 1×1×3Cf is obtained. Then, 1×1 convolution, batch normalization, ReLU activation function, 1×1 convolution, and Sigmoid activation function are successively passed. The vector of size 1×1×3CF is then divided into three vectors of size 1×1×CF along the channel dimension. The input feature map F 11 、F 12 、F 13 After multiplying these three vectors of scale 1×1×CF respectively, the output feature map F′ of the channel collaborative attention submodule is obtained 11 , F′ 12 , F′ 13 .

[0120] Step B212: Design a spatial collaborative attention submodule, with the input being the three feature maps F′ obtained in step B221 11 , F′ 12 , F′ 13 , after performing average pooling on the three feature maps in the channel dimension, three feature maps of size Hf×Wf×1 are obtained. After splicing along the channel dimension, a feature map of size Hf×Wf×3 is obtained. Then, 1×1 convolution, batch normalization, ReLU activation function, 1×1 convolution, and Sigmoid activation function are successively performed. The feature map of size Hf×Wf×3 is then divided into three vectors of size Hf×Wf×1 along the channel dimension. The input feature map F′ 11 , F′ 12 , F′ 13The output feature map F″ of the spatial collaborative attention submodule is obtained by multiplying these three vectors of scale Hf×Wf×1 respectively. 11 , F″ 12 , F″ 13 .

[0121] Step B213, design a multi-scale feature collaborative attention module, the input feature map is the encoder E in step B21 i ,i=1,2,3,decoder D i , output feature map F of i=1,2,3 i ,i=1,2,…6, here it is uniformly denoted as F, After F passes through 3×3 convolution, 5×5 convolution, and 7×7 convolution, we get three feature maps. F 11 、F 12 、F 13 Input channels coordinate attention to obtain feature map F′ 11 , F′ 12 , F′ 13 , F′ 11 , F′ 12 , F′ 13 Input spatial collaborative attention to obtain feature map F″ 11 , F″ 12 , F″ 13 , the feature map F″ 11 , F″ 12 , F″ 13 The output feature map F′ of the multi-scale feature collaborative attention module is obtained by adding them together. The specific formula is as follows:

[0122] F 11 =Conv3(F)

[0123] F 12 =Conv5(F)

[0124] F 13 =Conv7(F)

[0125] F′ 11 ,F′ 12 ,F′ 13 =att c (F 11 ,F 12 ,F 13 )

[0126] F″ 11 ,F″ 12 ,F″ 13 =att s (F 11 ,F12 ,F 13 )

[0127] F′=F″ 11 +F″ 12 +F″ 13

[0128] Among them, Conv3, Conv5, and Conv7 are 3×3 convolution, 5×5 convolution, and 7×7 convolution respectively, F is the input feature map, F′ is the output feature map of the multi-scale feature collaborative attention module, and att c is the channel collaborative attention submodule, att s It is the spatial collaborative attention submodule.

[0129] Further, the step B3 is implemented as follows:

[0130] Step B3: Design a multi-branch low-light image enhancement network based on frequency domain division, which consists of three branches: reference branch, high-frequency branch and low-frequency branch. The input image is a low-light image I. The three branches respectively pass the input low-light image I through the shallow image enhancement module of step B1 to obtain preliminary enhanced images I′1, I′2, and I′3. After 3×3 convolution, the images I′1, I′2, and I′3 are subjected to feature maps. After the feature maps f1, f2, and f3 enter the UNet structure based on the multi-scale feature collaborative attention module designed in step B2 respectively, the feature maps are obtained. After 3×3 convolution, the feature maps f′1, f′2, and f′3 are added to the input low-light image I to obtain the output images of the reference branch, high-frequency branch, and low-frequency branch, respectively. At the same time, the feature maps f′1, f′2, and f′3 are concatenated along the channel dimension, and after a 3×3 convolution, they are added to the input low-light image I to obtain the final enhanced image I o The specific formula is as follows:

[0131] I′1=s1(I)

[0132] I′2=s2(I)

[0133] I′3=s3(I)

[0134] f1=Conv3(I′1)

[0135] f2=Conv3(I′2)

[0136] f3=Conv3(I′3)

[0137] f′1=UNet1(f1)

[0138] f′2=Unet2(f2)

[0139] f′3=UNet3(f3)

[0140]

[0141]

[0142]

[0143] I o =Conv3(Concat(f′1,f′2,f′3))+I

[0144] Where I is the input low-light image, I′1, I′2, I′3 are the output images of the shallow image enhancement modules on the benchmark branch, high-frequency branch and low-frequency branch, respectively, f1, f2, f3 are the intermediate feature maps on the benchmark branch, high-frequency branch and low-frequency branch, f′1, f′2, f′3 are the output feature maps of the UNet structure based on the multi-scale feature collaborative attention module on the benchmark branch, high-frequency branch and low-frequency branch, s1, s2, s3 are the shallow image enhancement modules on the benchmark branch, high-frequency branch and low-frequency branch, UNet1, UNet2, UNet3 are the UNet structures based on the multi-scale feature collaborative attention module on the benchmark branch, high-frequency branch and low-frequency branch, Conv3 is a 3×3 convolution, They are the initial enhanced image, high-frequency information enhanced image, and low-frequency information enhanced image output by the baseline branch, high-frequency branch, and low-frequency branch, respectively. Concat is splicing along the channel dimension. o It is the output image of the multi-branch low-light image enhancement network based on frequency domain division.

[0145] Further, step C is implemented as follows:

[0146] Step C: Design the loss function, which consists of L2 loss and VGG perceptual loss.

[0147] The standard loss function is as follows:

[0148]

[0149] Here, Φ(·) represents the operation of extracting Conv4-1 layer features using the VGG-16 classification model pre-trained on the ImageNet dataset. are the initial enhanced image, high-frequency information enhanced image, and low-frequency information enhanced image output by the baseline branch, high-frequency branch, and low-frequency branch, respectively. ois the output image of the multi-branch low-light image enhancement network based on frequency domain division, G is the label image corresponding to the low-light image I, G1 and G2 are the high-frequency information label image and low-frequency information label image extracted using discrete cosine transform, respectively, ||.||1 represents the L1 loss, ||.|| 2 represents L2 loss.

[0150] Further, the step D is implemented as follows:

[0151] Step D: divide the low-light image into batches N, and divide the training data set in step A into K batches; input the low-light image I into the multi-branch low-light image enhancement network based on frequency domain division in step B, and obtain the initial enhanced image output by the reference branch, high-frequency branch and low-frequency branch High-frequency information enhances images Low frequency information enhanced image And the final enhanced image I o ; Using the loss function designed in step D, calculate the initial enhanced image output by the baseline branch, high-frequency branch, and low-frequency branch High-frequency information enhances images Low frequency information enhanced image And the final enhanced image I o The Adam optimization method is used to update the network parameters until convergence, and a multi-branch low-light image model based on frequency domain division is obtained.

[0152] Further, the step E is implemented as follows:

[0153] Step E: Input the low-light image to be tested into the trained multi-branch low-light image enhancement model based on frequency domain division, and predict the initial enhanced image output by the reference branch, high-frequency branch and low-frequency branch. High-frequency information enhances images Low frequency information enhanced image And the final enhanced image I o , take the final enhanced image I o is the generated normal illumination image.

[0154] The above are preferred embodiments of the present invention. Any changes made according to the technical solution of the present invention, as long as the resulting functions do not exceed the scope of the technical solution of the present invention, belong to the protection scope of the present invention.

Claims

1. A multi-branch low-light image enhancement method based on frequency domain division, characterized in that: The steps include: Step A: preprocess the input image, including data pairing, random data cropping, and data enhancement processing, to obtain a training data set; Step B, design a multi-branch low-light image enhancement network based on frequency domain division, which consists of three branches: a baseline branch, a high-frequency branch, and a low-frequency branch. Each branch contains a shallow image enhancement module and a UNet structure based on a multi-scale feature collaborative attention module; Step C, designing a loss function for guiding the optimization of the network parameters designed in step B; Step D: using the training data set obtained in step A to train the multi-branch low-illumination image enhancement network based on frequency domain division in step B, to obtain a trained multi-branch low-illumination image enhancement model based on frequency domain division; Step E: inputting the low-illuminance image to be tested into the trained multi-branch low-illuminance image enhancement model based on frequency domain division to predict and generate a normal-illuminance image; The specific implementation steps of step B are as follows: Step B1, designing a shallow image enhancement module to preliminarily enhance low-light images; Step B2: Design a UNet structure based on a multi-scale feature collaborative attention module to further enhance features; Step B3, design a multi-branch low-light image enhancement network based on frequency domain division, including three branches: a baseline branch, a high-frequency branch, and a low-frequency branch. Each branch contains a shallow image enhancement module and a UNet structure based on a multi-scale feature collaborative attention module; The specific implementation steps of step B3 are as follows: A multi-branch low-light image enhancement network based on frequency domain division is designed, which consists of three branches: reference branch, high-frequency branch and low-frequency branch. The input image is a low-light image I. The three branches respectively pass the input low-light image I through the shallow image enhancement module of step B1 to obtain preliminary enhanced images I'1, I'2, and I'3. After 3×3 convolution, the images I'1, I'2, and I'3 obtain feature maps After the feature maps f1, f2, and f3 enter the UNet structure based on the multi-scale feature collaborative attention module designed in step B2 respectively, the feature maps are obtained. After 3×3 convolution, the feature maps f'1, f'2, and f'3 are added to the input low-light image I to obtain the output images of the reference branch, high-frequency branch, and low-frequency branch, respectively. At the same time, the feature maps f'1, f'2, and f'3 are concatenated along the channel dimension, and after a 3×3 convolution, they are added to the input low-light image I to obtain the final enhanced image I o ; The specific formula is as follows: I′1=s1(I) I′2=s2(I) I′3=s3(I) f1=Conv3(I′1) f2=Conv3(I′2) f3=Conv3(I′3) f′1=UNet1(f1) f′2=UNet2(f2) f′3=UNet3(f3) I o =Conv3(Concat(f′1,f′2,f′3))+I Where I is the input low-light image, I'1, I'2, I'3 are the output images of the shallow image enhancement modules on the baseline branch, high-frequency branch and low-frequency branch, respectively, f1, f2, f3 are the intermediate feature maps on the baseline branch, high-frequency branch and low-frequency branch, f'1, f'2, f'3 are the output feature maps of the UNet structure based on the multi-scale feature collaborative attention module on the baseline branch, high-frequency branch and low-frequency branch, s1, s2, s3 are the shallow image enhancement modules on the baseline branch, high-frequency branch and low-frequency branch, UNet1, UNet2, UNet3 are the UNet structures based on the multi-scale feature collaborative attention module on the baseline branch, high-frequency branch and low-frequency branch, Conv3 is a 3×3 convolution, They are the initial enhanced image, high-frequency information enhanced image, and low-frequency information enhanced image output by the baseline branch, high-frequency branch, and low-frequency branch, respectively. Concat is splicing along the channel dimension. o It is the output image of the multi-branch low-light image enhancement network based on frequency domain division.

2. The multi-branch low-light image enhancement method based on frequency domain division according to claim 1, characterized in that: The specific implementation steps of step A are as follows: Step A1: pairing the low-light image with its corresponding label image; Step A2: randomly select a cropping area of ​​p×p×3 from each low-light image of size H×W×3 for cropping, and the corresponding label image is also cropped in the same way, where H and W are the height and width of the low-light image, respectively, and p is the height and width of the cropping area; Step A3: Randomly perform horizontal flipping, vertical flipping, and rotation operations on the training paired images for data enhancement.

3. The multi-branch low-illumination image enhancement method based on frequency domain division according to claim 1, characterized in that: The specific implementation steps of step B1 are as follows: Design a shallow image enhancement module with low-light images as input Perform three average poolings in the channel dimension to obtain three feature maps of size H×W×1. After splicing the three feature maps along the channel dimension, they pass through 3×3 convolution, batch normalization layer, ReLU activation function, 3×3 convolution, and Sigmoid activation function in sequence, and then multiply them pixel by pixel with the low-light image I to obtain the weight of each position, and then add them pixel by pixel with the low-light image I to obtain the output image. The specific formula is as follows: Where I' is the output image of the shallow image enhancement module, I is the input low-light image, Sigmoid is the Sigmoid activation function, Conv3 is the 3×3 convolution, ReLU is the ReLU activation function, BN is batch normalization, Concat is the concatenation operation along the channel dimension, AvgPool is the average pooling of the channel dimension, and AvgPool(I) is 3 The average pooling of the channel dimension is performed 3 times. It is a pixel-by-pixel multiplication operation.

4. The multi-branch low-light image enhancement method based on frequency domain division according to claim 1, characterized in that: The specific implementation steps of step B2 are as follows: Step B21, design a multi-scale feature collaborative attention module, which consists of a channel collaborative attention submodule and a spatial collaborative attention submodule; Step B22, design a UNet structure based on a multi-scale feature collaborative attention module, which consists of three encoders E1, E2, E3, three decoders D1, D2, D3 and six multi-scale feature collaborative attention modules; each encoder and decoder consists of 3×3 convolution, ReLU activation function, 3×3 convolution, ReLU activation function; let the input be the feature map f, after passing through encoder E1, the feature map is obtained Then enter the multi-scale feature collaborative attention module in step B21 to obtain the feature map After the feature map F'1 is max-pooled, it passes through encoder E2 to obtain the feature map Then enter the multi-scale feature collaborative attention module in step B21 to obtain the feature map The feature map F'2 is max-pooled and passes through encoder E3 to obtain the feature map Then enter the multi-scale feature collaborative attention module in step B21 to obtain the feature map After the feature map F'3 passes through the decoder D1, the feature map is obtained Then enter the multi-scale feature collaborative attention module in step B21 to obtain the feature map The feature map F'4 is deconvolved and added to the feature map F2, and then passes through the decoder D2 to obtain the feature map Then enter the multi-scale feature collaborative attention module in step B21 to obtain the feature map The feature map F'5 is deconvolved and added to the feature map F1, and then passes through the decoder D3 to obtain the feature map Then enter the multi-scale feature collaborative attention module in step B21 to obtain the output feature map The specific formula is as follows: F1=E1(f) F′1=att(F1) F2=E2(Maxpool(F′1)) F′2=att(F2) F3=E3(Maxpool(F′2)) F′3=att(F3) F4=D1(F′3) F′4=att(F4) F5=D2(deconv(F′4)+F2) F′5=att(F5) F6=D3(deconv(F′5)+F1) f'=att(F6) Among them, f is the input feature map of the UNet structure based on the multi-scale feature collaborative attention module, f' is the output feature map of the UNet structure based on the multi-scale feature collaborative attention module, E1, E2, E3 are encoders, D1, D2, D3 are decoders, att is the multi-scale feature collaborative attention module, Maxpool is the maximum pooling layer, deconv is the deconvolution operation, F1, F2, F3 are the output feature maps of encoders E1, E2, E3 respectively, F4, F5, F6 are the output feature maps of decoders D1, D2, D3 respectively, and F' i ,i=1,2,3…6 are the output feature maps of the multi-scale feature collaborative attention module at different stages.

5. The multi-branch low-light image enhancement method based on frequency domain division according to claim 4, characterized in that: The specific implementation steps of step B21 are as follows: Step B211: Design a channel collaborative attention submodule with three feature maps F as input 11 、F 12 、F 13 After the three feature maps are averaged in the spatial dimension, three vectors of size 1×1×Cf are obtained. After splicing along the channel dimension, a vector of size 1×1×3Cf is obtained. Then, 1×1 convolution, batch normalization, ReLU activation function, 1×1 convolution, and Sigmoid activation function are successively passed. The vector of size 1×1×3Cf is then divided into three vectors of size 1×1×Cf along the channel dimension. The input feature map F 11 、F 12 、F 13 After multiplying these three vectors of scale 1×1×Cf respectively, the output feature map F' of the channel collaborative attention submodule is obtained 11 、F' 12 , F' 13 ; Step B212: Design a spatial collaborative attention submodule, with the input being the three feature maps F' obtained in step B221 11 , F' 12 , F' 13 , after performing average pooling on the three feature maps in the channel dimension, three feature maps of size Hf×Wf×1 are obtained. After splicing along the channel dimension, a feature map of size Hf×Wf×3 is obtained. Then, 1×1 convolution, batch normalization, ReLU activation function, 1×1 convolution, and Sigmoid activation function are successively performed. The feature map of size Hf×Wf×3 is then divided into three vectors of size Hf×Wf×1 along the channel dimension, and the input feature map F' 11 , F' 12 , F' 13 The output feature map F″ of the spatial collaborative attention submodule is obtained by multiplying these three vectors of scale Hf×Wf×1 respectively. 11 , F″ 12 , F″ 13 ; Step B213, design a multi-scale feature collaborative attention module, the input feature map is the encoder E in step B21 i ,i=1,2,3,decoder D i , output feature map F of i=1,2,3 i ,i=1,2,…6, here it is uniformly denoted as F, After F passes through 3×3 convolution, 5×5 convolution, and 7×7 convolution, we get three feature maps F 11 、F 12 、F 13 Input channel coordinated attention to obtain feature map F' 11 , F' 12 , F' 13 , F' 11 , F' 12 , F' 13 Input spatial collaborative attention to obtain feature map F″ 11 , F″ 12 , F″ 13 , the feature map F″ 11 , F″ 12 , F″ 13 The output feature map F' of the multi-scale feature collaborative attention module is obtained by adding them together; the specific formula is as follows: F 11 =Conv3(F) F 12 =Conv5(F) F 13 =Conv7(F) F′ 11 ,F′ 12 ,F′ 13 =to c (F 11 ,F 12 ,F 13 ) F″ 11 "F" 12 "F" 13 =to s (F 11 ,F 12 ,F 13 ) F′=F″ 11 +F″ 12 +F″ 13 Among them, Conv3, Conv5, and Conv7 are 3×3 convolution, 5×5 convolution, and 7×7 convolution respectively. F is the input feature map, and F' is the output feature map of the multi-scale feature collaborative attention module. c is the channel collaborative attention submodule, att s It is the spatial collaborative attention submodule.

6. The multi-branch low-light image enhancement method based on frequency domain division according to claim 1, characterized in that: The specific implementation of step C is: Step C: Design the loss function, which consists of L2 loss and VGG perception loss. The total target loss function of the network is as follows: Where Φ(·) represents the operation of extracting Conv4-1 layer features using the VGG-16 classification model pre-trained on the ImageNet dataset; They are the initial enhanced image, high-frequency information enhanced image, and low-frequency information enhanced image output by the baseline branch, high-frequency branch, and low-frequency branch, respectively. o is the output image of the multi-branch low-light image enhancement network based on frequency domain division, G is the label image corresponding to the low-light image I, G1 and G2 are the high-frequency information label image and low-frequency information label image extracted using discrete cosine transform, respectively, ||.||1 represents the L1 loss, ||.|| 2 represents L2 loss.

7. The multi-branch low-illumination image enhancement method based on frequency domain division according to claim 1, characterized in that: The specific implementation steps of step D are as follows: The low-light images are divided into batches N, and the training data set in step A is divided into K batches; the low-light image I is input into the multi-branch low-light image enhancement network based on frequency domain division in step B, and the initial enhanced image output by the reference branch, high-frequency branch and low-frequency branch is obtained. High-frequency information enhances images Low frequency information enhanced image And the final enhanced image I o ; Using the loss function designed in step D, calculate the initial enhanced image output by the baseline branch, high-frequency branch, and low-frequency branch High-frequency information enhances images Low frequency information enhanced image And the final enhanced image I o The Adam optimization method is used to update the network parameters until convergence, and a multi-branch low-light image model based on frequency domain division is obtained.

8. The multi-branch low-illumination image enhancement method based on frequency domain division according to claim 1, characterized in that: The specific implementation steps of step E are as follows: The low-light image to be tested is input into the trained multi-branch low-light image enhancement model based on frequency domain division, and the initial enhanced image output by the reference branch, high-frequency branch and low-frequency branch is predicted. High-frequency information enhances images Low frequency information enhanced image And the final enhanced image I o , take the final enhanced image I o is the generated normal illumination image.

Citation Information

Patent Citations

  • Low-illumination image enhancement method based on improved depth separable generative adversarial network

    CN111915525A

  • Low-illumination image enhancement method based on multi-scale stacked attention network

    CN114972107A