Full-Resolution Low-Light Image Enhancement Method for Aggregating Context and Enhancing Details
Through the coordinated processing of the full resolution detail extraction and frequency airspace context information attention module, the problem of disco-coordination of details and context feature processing in low-illumination image enhancement is solved, and better image enhancement effect is achieved, improving image quality and performance of downstream tasks.
Patent Information
- Application Number
- CN202211600774.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-12
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-12-12
AI Technical Summary
Existing low-illumination image enhancement technology is difficult to effectively coordinate the details and contextual features, resulting in poor image enhancement effects, especially in low-illumination conditions, loss of details, color distortion and insufficient brightness.
A full-resolution low-illumination image enhancement method with aggregation context and enhanced details is designed. Through the full-resolution detail extraction module, frequency airspace context information attention module and feature aggregation module, the details and context features of low-illumination images are processed in a coordinated manner, and data preprocessing, network training and loss function optimization techniques are adopted.
It significantly improves the enhancement performance of low-illumination images, restores detailed information, improves the color and brightness of the images, improves image quality, and supports more efficient downstream tasks such as face recognition and night monitoring.
Smart Images

Figure CN115880177B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of image processing and computer vision, and particularly relates to a full-resolution low-light image enhancement method for aggregating context and enhancing details. Background Art
[0002] Low-light image enhancement is an important branch of image enhancement. Due to environmental reasons such as insufficient lighting, non-uniform lighting, backlighting, and the influence of scene conditions such as being easily disturbed during the camera imaging process, low-light images exhibit degraded conditions such as low brightness, high noise, and loss of color and detail information. This requires low-light image enhancement algorithms to process them. On the basis of retaining useful detail information, useful information such as color is restored, and noise is removed, so as to obtain a normal illumination image that meets the human visual perception experience or is more suitable for downstream task analysis and processing.
[0003] Low-light image enhancement technology has a wide range of application prospects. It can improve the visibility of night inspections or monitoring. Cameras are set up in public places for better monitoring and recording. Images captured by night surveillance cameras are mostly dark and unclear in low-light conditions and cannot provide strong evidence for some events. After being processed by low-light image enhancement technology, the clearer the video, the stronger the support for event judgment and decision-making. Low-light image enhancement technology can also improve the quality of human visual perception. The ambient light in the captured pictures is not satisfactory, but it can quickly reach a satisfactory level through the processing of this technology without the need to reshoot with a set. At the same time, low-light image enhancement technology can also improve the performance of downstream tasks. Many downstream technologies, such as face recognition and human key point recognition, have relatively high requirements for the quality of the input image. Whether the algorithm can correctly identify the positions of the face and body depends on whether the input image is clear enough. It is extremely challenging to identify the face or body contour when the input picture is dark or blurred, while low-light image enhancement technology can improve the image quality, thus significantly improving the recognition and detection accuracy of the algorithm.
[0004] Early methods were mainly based on histogram equalization and the Retinex theory. Histogram equalization enhances images by expanding the dynamic range of image pixels. This method can improve contrast well, but due to the lack of consideration for local areas, it is prone to overexposure and underexposure of images. The Retinex theory believes that an image can be described as the product of a reflection component R and an illumination component I, which requires prior knowledge. Poor prior knowledge will result in unrealistic enhancement with serious color differences and noise amplification. In recent years, many deep learning methods have been proposed. Some methods combine the Retinex theory with convolutional neural networks and directly classify the reflection as the enhanced image, resulting in loss of details and color deviation; some methods directly transfer the mainstream network architectures in other fields to low-light tasks, lacking consideration for low-light characteristics; and some methods independently solve some aspects of the problems existing in low-light images, such as insufficient illumination, high noise, and loss of color and detail information, ignoring the correlation between problems. However, low-light images contain low-level information such as details and noise, and high-level information such as color and scene. The two types of features are not completely unrelated. The processing of the scene and illumination helps the recovery of details, and the recovery of details in turn promotes the restoration of the overall scene.
[0005] Existing methods treat low-level information such as details and noise and high-level information such as color and scene equally and independently. However, low-light image enhancement itself is a relatively delicate image processing task that requires more attention to the processing of details and then the integration of the processing of high-level information such as color and scene. Moreover, connections should be established between the two types of features and they should be enhanced together, rather than being processed independently. Summary of the Invention
[0006] Aiming at the defects and deficiencies of the existing technology, the purpose of the present invention is to provide a full-resolution low-light image enhancement method that aggregates context and enhances details. This method aggregates detail features and context features for collaborative enhancement, which is beneficial to significantly improving the performance of low-light image enhancement.
[0007] The present invention designs a full-resolution low-light image enhancement method that aggregates context and enhances details. First, a full-resolution detail extraction module is designed to extract detail features. Then, a frequency-domain context information attention module is designed to extract context features such as color and scene in the frequency domain and spatial domain and use the attention module to learn the importance of features in the frequency domain and spatial domain. Finally, a feature aggregation and enhancement module is designed to aggregate the detail features and context features and then enhance them collaboratively.
[0008] It includes: performing data preprocessing, including data pairing, data random cropping, and data augmentation processing, to obtain a training data set; designing a full-resolution low-light image enhancement network that aggregates context and enhances details, which consists of a full-resolution detail extraction module, a frequency domain context information attention module, and a feature aggregation and enhancement module; designing a loss function to guide the parameter optimization of the network designed in step B; using the training data set obtained in step A to train the full-resolution low-light image enhancement network that aggregates context and enhances details in step B, converging to the Nash equilibrium, to obtain a trained full-resolution low-light image enhancement model that aggregates context and enhances details; inputting the low-light image to be tested into the trained full-resolution low-light image enhancement model that aggregates context and enhances details, and outputting the enhanced normal-light image. The present invention can enhance low-light images and solve problems such as missing details, color distortion, and insufficient brightness in low-light images.
[0009] The technical solution adopted by the present invention to solve its technical problems is:
[0010] A full-resolution low-light image enhancement method that aggregates context and enhances details, characterized in that:
[0011] Step A: Perform data preprocessing, including data pairing, data random cropping, and data augmentation processing, to obtain a training data set;
[0012] Step B: Design a full-resolution low-light image enhancement network that aggregates context and enhances details, including: a full-resolution detail extraction module, a frequency domain context information attention module, and a feature aggregation and enhancement module;
[0013] Step C: Design a loss function for guiding the parameter optimization of the network designed in step B;
[0014] Step D: Use the training data set obtained in step A to train the full-resolution low-light image enhancement network that aggregates context and enhances details in step B, converging to the Nash equilibrium, to obtain a trained full-resolution low-light image enhancement model that aggregates context and enhances details;
[0015] Step E: Input the low-light image to be tested into the trained full-resolution low-light image enhancement model that aggregates context and enhances details, and output the enhanced normal-light image.
[0016] Further, the specific implementation steps of step A are as follows:
[0017] Step A1: Pair the low-light image and the corresponding label image;
[0018] Step A2: Randomly crop each low-light image with a size of h×w×3 into an image of size p×p×3, and perform the same random cropping method on its corresponding label image, where h and w are the height and width of the low-light image and the label image, and p is the height and width of the cropped image;
[0019] Step A3: Randomly select 1 out of the following 8 enhancement methods to perform data enhancement on the training paired images: keep the original image, vertical flip, rotate 90 degrees, vertical flip after rotating 90 degrees, rotate 180 degrees, vertical flip after rotating 180 degrees, rotate 270 degrees, vertical flip after rotating 270 degrees.
[0020] Furthermore, the specific implementation steps of Step B are as follows:
[0021] Step B1: Construct a full-resolution detail extraction module, which consists of a shallow feature extraction sub-module, an attention sub-module based on CBAM, and a frequency domain transformation sub-module, and use the designed network to extract detail features;
[0022] Step B2: Design a frequency-spatial domain context information attention module, which consists of a multi-scale feature extraction sub-module and a frequency-spatial domain feature fusion sub-module, and use the designed network to extract context features;
[0023] Step B3: Design a feature aggregation and enhancement module, which consists of a feature aggregation convolutional block and a collaborative enhancement sub-module, aggregate the detail features extracted in Step B1 and the context features extracted in Step B2, and enhance the two types of features together;
[0024] Step B4: Design a full-resolution low-light image enhancement network that aggregates context and enhances details, including a full-resolution detail extraction module, a frequency-spatial domain context information attention module, and a feature aggregation and enhancement module.
[0025] Furthermore, the specific implementation steps of Step B1 are as follows:
[0026] Step B11: Design a shallow feature extraction sub-module. The input is the low-light image I. After obtaining the initial feature map F through a 3×3 convolution, it enters three branches. The first branch contains 1 3×3 convolution, the second branch contains 2 serial 3×3 convolutions, and the third branch contains 3 serial 3×3 convolutions. Concatenate the processing results F, F, F along the channel dimension, and then pass through a 3×3 convolution to obtain the feature map F output by the shallow feature extraction sub-module; The specific formula is expressed as follows: ori After that, it enters three branches. The first branch contains 1 3×3 convolution, the second branch contains 2 serial 3×3 convolutions, and the third branch contains 3 serial 3×3 convolutions. Concatenate the processing results F, F, F B1 、F B2 、F B3 along the channel dimension, and then pass through a 3×3 convolution to obtain the feature map F output by the shallow feature extraction sub-module; low ; The specific formula is as follows:
[0027] F ori =Conv3(I)
[0028] F B1 = Conv3(F ori )
[0029] F B2 = Conv3(Conv3(F ori ))
[0030] F B3 = Conv3(Conv3(Conv3(F ori )))
[0031] F low = Conv3(Concat(F B1 , F B2 , F B3 ))
[0032] where Conv3 is a 3×3 convolution and Concat is an operation of concatenation along the channel dimension;
[0033] Step B12: Construct an attention sub-module based on CBAM, which consists of sequential attention Att c in the channel dimension and attention Att s in the spatial dimension. The input feature map is the feature map F low obtained in Step B11, and the output feature map of the attention sub-module based on CBAM is F spa ; The specific formula is as follows:
[0034] F spa = Att s (Att c (F low ))
[0035] where Att c is the attention in the channel dimension and Att s is the attention in the spatial dimension;
[0036] Step B13: Design a frequency domain transformation sub-module. The input feature map is the feature map F spa obtained in Step B12. After converting the spatial domain to the frequency domain using the Fourier transform function, passing through a 3×3 convolution, a normalization layer, and a ReLU activation function in sequence, and then converting the frequency domain back to the spatial domain using the inverse Fourier transform function, the output feature map F fre is obtained; The specific formula is as follows:
[0037] F fre = idft(ReLU(BN(Conv3(dft(F spa )))))
[0038] Among them, DFT is the Fourier transform, IDFT is the inverse Fourier transform, ReLU is the ReLU activation function, BN is the batch normalization layer, and Conv3 is a 3×3 convolution;
[0039] Step B14: Construct a full-resolution detail extraction module, which consists of a shallow feature extraction sub-module, an attention sub-module based on CBAM, and a frequency domain transformation sub-module; let the input be the low-illumination image I processed in step A, and it passes through the shallow feature extraction sub-module, the attention sub-module, and the frequency domain transformation sub-module in sequence to obtain the feature map F low 、F spa 、F fre 。
[0040] Furthermore, the specific implementation steps of step B2 are as follows:
[0041] Step B21: Design a multi-scale feature extraction sub-module, and denote the input feature map as F, Let H, W, and C be the height, width, and number of channels of the feature F respectively. After passing through an average pooling layer with a kernel size of 2×2 and a stride of 2, it is successively reduced in dimension through a 1×1 convolution, a ReLU activation function, a 1×1 convolution, and a ReLU activation function to obtain an intermediate feature map Then it is divided into two branches. After the upper branch is further reduced in dimension through a 1×1 convolution, the output of the upper branch is obtained through an upsampling layer a is the number of channels after dimension reduction; the other branch passes through an average pooling layer with a kernel size of 2×2 and a stride of 2, and is successively reduced in dimension through a 1×1 convolution, a ReLU activation function, a 1×1 convolution, and a ReLU activation function to obtain an intermediate feature map The intermediate feature map F 121 Then it successively passes through an upsampling layer, a 1×1 convolution, a ReLU activation function, and an upsampling layer to obtain the output of the lower branch Add F 11 and F 12 After adding them, concatenate them with F in the channel dimension, and then adjust the channels through a 1×1 convolution after passing through the SE module to obtain the feature map output by the multi-scale feature extraction sub-module The specific formula is expressed as follows:
[0042] F1 = ReLU(Conv1(ReLU(Conv1(Avgpooling(F)))))
[0043] F 11 = Upsampling(Conv1(F1))
[0044] F 121 = ReLU(Conv1(ReLU(Conv1(Avgpooling(F1)))))
[0045] F 12 = Upsampling(ReLU(Conv1(Upsampling(F 121 ))))
[0046] F m = Conv1(SE(Concat(F 11 + F 12 , F)))
[0047] Where ReLU is the activation function, Conv1 is a 1×1 convolution, SE(·) is the SE module, Avgpooling is the average pooling layer with a kernel size of 2×2 and a stride of 2, Upsampling is the nearest neighbor upsampling layer with a factor of 2, and Concat is the concatenation operation along the channel dimension;
[0048] Step B22: Design the frequency-domain feature fusion sub-module, which consists of a serial connection of channel attention and spatial attention;
[0049] Step B23: Design the frequency-domain context information attention module, which consists of three multi-scale feature extraction sub-modules and the frequency-domain feature fusion sub-module; the inputs of the three multi-scale feature extraction sub-modules are the three feature maps F low 、F spa 、F fre obtained in Step B1 respectively. After being processed by the multi-scale feature extraction sub-modules designed in Step B21 respectively, the feature maps F low_m 、F spa_m 、F fre_m with context information are obtained. Then, after passing through the frequency-domain feature fusion sub-module designed in Step B22, the feature map F f output by the frequency-domain context information attention module is obtained.
[0050] Furthermore, the specific implementation steps of Step B22 are as follows:
[0051] Step B221: Design the channel attention, with the input being the feature map obtained in Step B23 The three feature maps are respectively subjected to global average pooling in the spatial dimension to obtain three vectors with a scale of 1×1×C, and then the three vectors are concatenated along the channel dimension to obtain an intermediate feature map Subject F c to dimensionality reduction and dimensionality increase in sequence through 1×1 convolution, ReLU activation function, 1×1 convolution, ReLU activation function, and 1×1 convolution, and then obtain the weights on the channel dimension through the Sigmoid activation function Subject F W1Decompose it into three vectors F with a scale of 1×1×C along the channel dimension W10 , F W11 , F W12 , and multiply them with the input feature maps F low_m , F spa_m , F fre_m of the frequency domain feature fusion sub-module respectively to obtain the output feature maps of channel attention The specific formula is as follows:
[0052] F c = Concat(Avgpooling s (F low_m ), Avgpooling s (F spa_m ), Avgpooling s (F fre_m ))
[0053] FW1 = Sigmoid(Conv1(ReLU(Conv1(ReLU(Conv1(F c ))))))
[0054] F low_c = F W10 × F low_m
[0055] F spa_c = F W11 × F spa_m
[0056] F fre_c = F W12 × F fre_m
[0057] Among them, Concat is the concatenation operation along the channel dimension, Avgpooling s is the global average pooling in the spatial dimension, ReLU is the activation function, Conv1 is the 1×1 convolution, and Sigmoid is the Sigmoid activation function;
[0058] Step B222: Design spatial attention. The input is the three feature maps F low_c , F spa_c , F fre_c obtained in step B221. After the three feature maps are respectively subjected to average pooling in the channel dimension, three feature maps with a scale of H×W×1 are obtained, and then the three feature maps are concatenated along the channel dimension to obtain the intermediate feature map Take F sAfter passing through an average pooling layer with a kernel size of 2×2 and a stride of 2, a ReLU activation function, and an upsampling layer in sequence, the weights in the spatial dimension are obtained through a Sigmoid activation function. Decompose F W2 into three feature maps F W20 、F W21 、F W22 with a scale of H×W×1, and multiply them with the input feature maps F low_c 、F spa_c 、F fre_c of the spatial attention respectively to obtain the output feature maps of the spatial attention. The specific formula is as follows:
[0059] F s =Concat(Avgpooling c (F low_c ),Avgpooling c (F spa_c ),Avgpooling c (F fre_c ))
[0060] F W2 =Sigmoid(Upsampling(ReLU(Avgpooling(F s ))))
[0061] F low_s =F W20 ×F low_c
[0062] F spa_s =F W21 ×F spa_c
[0063] F fre_s =F W22 ×F fre_c
[0064] Among them, Concat is the concatenation operation along the channel dimension, Avgpooling c is the average pooling in the channel dimension, ReLU is the activation function, Sigmoid is the Sigmoid activation function, Avgpooling is the average pooling layer with a kernel size of 2×2 and a stride of 2, and Upsampling is the nearest neighbor upsampling layer with a factor of 2;
[0065] Step B223: Design a frequency domain feature fusion sub-module, and the input feature maps are the feature maps F low_m 、F spa_m 、F fre_m obtained in step B23., the three feature maps first pass through the channel attention in step B221 to obtain the feature map F low_c , F spa_c , F fre_c , and then pass through the spatial attention in step B222 to obtain the feature map F low_s , F spa_s , F fre_s . After adding the three feature maps, the final output F f is obtained; the specific formula is as follows:
[0066] F f = F low_s + F spa_s + F fre_s .
[0067] Further, the specific implementation steps of step B3 are as follows:
[0068] Step B31: Design a feature aggregation convolution block to realize the fusion of detailed information and context information; the input feature map is the feature map F fre obtained in step B1 and the feature map F f obtained in step B2. After splicing them along the channel dimension and passing through a 3×3 convolution, the output feature map F conv is obtained; the specific formula is as follows:
[0069] F conv = Conv3(Concat(F fre , F f ))
[0070] where Conv3 is a 3×3 convolution and Concat is an operation of splicing along the channel dimension;
[0071] Step B32: Design a collaborative enhancer module to collaboratively enhance the fusion information of detailed information and context information; the input feature map is F conv obtained in step B31. After passing F conv sequentially through a 1×1 convolution, a ReLU6 activation function, a Dropout random inactivation layer, a 1×1 convolution, and a Dropout random inactivation layer, and adding it to F conv , the intermediate feature map F mid is obtained. Then, through a LeakyReLU activation function, after splicing with F conv along the channel dimension and passing through a 3×3 convolution, the output feature map F co is obtained; the specific formula is as follows:
[0072] F mid = Dropout(Conv1(Dropout(ReLU6(Conv1(F conv ))))) + Fconv
[0073] F co = Conv3(Concat(LeakyReLU(F mid ), F conv ))
[0074] Among them, Conv1 is a 1×1 convolution, Conv3 is a 3×3 convolution, Concat is an operation of concatenating along the channel dimension, Dropout is a random inactivation layer, ReLU6 is a ReLU6 activation function, and LeakyReLU is a LeakyReLU activation function;
[0075] Step B33: Design a feature aggregation and enhancement module, which consists of a feature aggregation convolutional block and a collaborative enhancement sub-module. The input feature map is the feature map F obtained in step B1 fre and the feature map F obtained in step B2 f . After passing through the feature aggregation convolutional block, the feature map F conv is obtained, and then after passing through the collaborative enhancement sub-module, the feature map F co is obtained.
[0076] Furthermore, the specific implementation method of step B4 is as follows:
[0077] Step B4: Design a full-resolution low-light image enhancement network for aggregating context and enhancing details, which is composed of integrating a full-resolution detail extraction module, a frequency-domain context information attention module, and a feature aggregation and enhancement module; Input the low-light image I, and after passing through the full-resolution detail extraction module in step B1, three feature maps F low 、F spa 、F fre are obtained. After passing through the frequency-domain context information attention module, the feature map F f is obtained, and then after passing through the feature aggregation and enhancement module, the feature map F co is obtained. Then F co is concatenated with the feature map F low in step B1 along the channel dimension, and then through a 3×3 convolution, the final enhanced image I out is obtained; The specific formula is expressed as follows:
[0078] I out = Conv3(Concat(F co , F low ))
[0079] Among them, Conv3 is a 3×3 convolution, and Concat is an operation of concatenating along the channel dimension.
[0080] Furthermore, the specific implementation method of step C is as follows:
[0081] Step C: Design a loss function, which consists of L2 loss and VGG perceptual loss. The total objective loss function of the network is as follows:
[0082] l = ω1||I out - G|| 2 + ω2||Φ(I out ) - Φ(G)||1
[0083] where Φ(·) represents the operation of extracting the features of the Conv4-1 layer using the pre-trained VGG-16 classification model on the ImageNet dataset; I out represents the enhanced image of the low-light image I, G represents the label image corresponding to the low-light image I, ||.||1 represents the L1 loss, and ||.|| 2 represents the L2 loss, and ω1 and ω2 are weights.
[0084] Furthermore, the specific implementation steps of Step D are as follows:
[0085] Step D1: Randomly divide the training dataset obtained in Step A into several batches, and each batch contains N pairs of images;
[0086] Step D2: Input the low-light image I, and after passing through the full-resolution low-light image enhancement network for aggregating context and enhancing details in Step B, obtain the enhanced image I out , and calculate the loss l using the formula in Step C;
[0087] Step D3: Calculate the gradients of the parameters in the network using the backpropagation method according to the loss, and update the network parameters using the Adam optimization method;
[0088] Step D4: Repeat Steps D1 to D3 in batches until the value of the objective loss function of the network converges to the Nash equilibrium, save the network parameters, and obtain the full-resolution low-light image enhancement model for aggregating context and enhancing details.
[0089] Compared with the prior art, the present invention and its preferred solutions extract detail features at full resolution, extract context features in the frequency domain and spatial domain, and aggregate the two types of features together for joint enhancement, which can better extract the two types of information and learn the relationship between the two types of features during the enhancement process. A full-resolution low-light image enhancement network for aggregating context and enhancing details is designed, and a full-resolution detail extraction module, a frequency-spatial domain context information attention module, and a feature aggregation and enhancement module are respectively set to extract detail features, extract frequency-spatial domain context features, aggregate the two features, and perform collaborative enhancement. Different from other methods that independently solve the problems existing in low-light images, the present invention can better extract detail features, context features, and achieve collaborative enhancement. Description of the Drawings
[0090] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments:
[0091] Figure 1 It is a flowchart for implementing the method of the embodiment of the present invention.
[0092] Figure 2 It is a structural diagram of a full-resolution low-light image enhancement network for aggregating context and enhancing details in the embodiment of the present invention.
[0093] Figure 3 It is a structural diagram of a multi-scale feature extraction sub-module in the embodiment of the present invention.
[0094] Figure 4 It is a structural diagram of a frequency-domain and spatial-domain feature fusion sub-module in the embodiment of the present invention.
[0095] Figure 5 It is a structural diagram of a collaborative enhancement sub-module in the embodiment of the present invention. Specific embodiments
[0096] To make the features and advantages of this patent more obvious and understandable, specific embodiments are given below for detailed description as follows:
[0097] It should be noted that the following detailed description is illustrative and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used in this specification have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs.
[0098] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0099] The following further specifically introduces the solution of this embodiment in conjunction with the accompanying drawings:
[0100] The present invention provides a full-resolution low-light image enhancement method for aggregating context and enhancing details, as Figures 1 - 5 shown, including the following steps:
[0101] Step A: Perform data preprocessing, including data pairing, data random cropping, and data enhancement processing, to obtain a training data set;
[0102] Step B: Design a full-resolution low-light image enhancement network that aggregates context and enhances details. This network consists of a full-resolution detail extraction module, a frequency-domain context information attention module, and a feature aggregation and enhancement module;
[0103] Step C: Design a loss function to guide the parameter optimization of the network designed in Step B;
[0104] Step D: Use the training dataset obtained in Step A to train the full-resolution low-light image enhancement network that aggregates context and enhances details in Step B, converge to the Nash equilibrium, and obtain a trained full-resolution low-light image enhancement model that aggregates context and enhances details;
[0105] Step E: Input the low-light image to be tested into the trained full-resolution low-light image enhancement model that aggregates context and enhances details, and output the enhanced normal-light image.
[0106] Furthermore, Step A includes the following steps:
[0107] Step A1: Pair the low-light images with their corresponding label images;
[0108] Step A2: Randomly crop each low-light image with a size of h×w×3 into an image with a size of p×p×3, and use the same random cropping method for its corresponding label image, where h and w are the height and width of the low-light image and the label image, and p is the height and width of the cropped image;
[0109] Step A3: Randomly use one of the following 8 enhancement methods to perform data enhancement on the paired images to be trained: keep the original image, vertical flip, rotate 90 degrees, vertical flip after rotating 90 degrees, rotate 180 degrees, vertical flip after rotating 180 degrees, rotate 270 degrees, vertical flip after rotating 270 degrees.
[0110] Furthermore, Step B includes the following steps:
[0111] Step B1: Construct a full-resolution detail extraction module, which consists of a shallow feature extraction sub-module, an attention sub-module based on CBAM, and a frequency-domain transformation sub-module, and use the designed network to extract detail features;
[0112] Step B2: Design a frequency-domain context information attention module, which consists of a multi-scale feature extraction sub-module and a frequency-domain feature fusion sub-module, and use the designed network to extract context features;
[0113] Step B3: Design a feature aggregation and enhancement module, which consists of a feature aggregation convolution block and a collaborative enhancement sub-module, aggregate the detail features extracted in Step B1 and the context features extracted in Step B2, and enhance the two types of features together;
[0114] Step B4: Design a full-resolution low-light image enhancement network that aggregates context and enhances details, including a full-resolution detail extraction module, a frequency domain context information attention module, and a feature aggregation and enhancement module.
[0115] Furthermore, step B1 includes the following steps:
[0116] Step B11: Design a shallow feature extraction sub-module. The input is a low-light image I, which undergoes a 3×3 convolution to obtain an initial feature map F ori and then enters three branches. The first branch contains one 3×3 convolution, the second branch contains two serial 3×3 convolutions, and the third branch contains three serial 3×3 convolutions. The processing results F B1 、F B2 、F B3 are concatenated along the channel dimension and then passed through a 3×3 convolution to obtain the feature map F low output by the shallow feature extraction sub-module. The specific formula is as follows:
[0117] F ori =Conv3(I)
[0118] F B1 =Conv3(F ori )
[0119] F B2 =Conv3(Conv3(F ori ))
[0120] F B3 =Conv3(Conv3(Conv3(F ori )))
[0121] F low =Conv3(Concat(F B1 ,F B2 ,F B3 ))
[0122] Among them, Conv3 is a 3×3 convolution, and Concat is an operation of concatenating along the channel dimension;
[0123] Step B12: Construct an attention sub-module based on CBAM. This module consists of a serial channel dimension attention Att c and a spatial dimension attention Att s . The input feature map is the feature map F low obtained in step B11, and the output feature map of the attention sub-module based on CBAM is F spa . The specific formula is as follows:
[0124] F spa = Att s (Att c (F low ))
[0125] where Att c is the attention of the channel dimension, and Att s is the attention of the spatial dimension.
[0126] Step B13: Design a frequency-domain transformation sub-module. The input feature map is the feature map F spa obtained in Step B12. After converting the spatial domain to the frequency domain using the Fourier transform function, passing through a 3×3 convolution, a normalization layer, and a ReLU activation function in sequence, and then using the inverse Fourier transform function to convert the frequency domain back to the spatial domain, the output feature map F fre is obtained. The specific formula is as follows:
[0127] F fre = idft(ReLU(BN(Conv3(dft(F spa )))))
[0128] where dft is the Fourier transform, idft is the inverse Fourier transform, ReLU is the ReLU activation function, BN is the batch normalization layer, and Conv3 is the 3×3 convolution.
[0129] Step B14: Construct a full-resolution detail extraction module, which consists of a shallow feature extraction sub-module, an attention sub-module based on CBAM, and a frequency-domain transformation sub-module. Let the input be the low-light image I processed in Step A. After passing through the shallow feature extraction sub-module, the attention sub-module, and the frequency-domain transformation sub-module in sequence, the feature maps F low , F spa , F fre are obtained.
[0130] Furthermore, Step B2 includes the following steps:
[0131] Step B21: Design a multi-scale feature extraction sub-module. The input feature map of this module is denoted as F, where H, W, and C are the height, width, and number of channels of the feature F respectively. After passing through an average pooling layer with a kernel size of 2×2 and a stride of 2, it is successively reduced in dimension through a 1×1 convolution, a ReLU activation function, a 1×1 convolution, and a ReLU activation function to obtain an intermediate feature map Then it is divided into two branches. After the upper branch is further reduced in dimension through a 1×1 convolution, the output of the upper branch is obtained through an upsampling layer a is the number of channels after dimensionality reduction. After another branch passes through an average pooling layer with a kernel size of 2×2 and a stride of 2, it successively passes through 1×1 convolution, ReLU activation function, 1×1 convolution, and ReLU activation function for dimensionality reduction to obtain an intermediate feature map Intermediate feature map F 121 Then, it successively passes through an upsampling layer, 1×1 convolution, ReLU activation function, and upsampling layer to obtain the output of the lower branch Let F 11 and F 12 After adding them and concatenating with F in the channel dimension, it passes through the SE module and then adjusts the channels through 1×1 convolution to obtain the feature map output by the multi-scale feature extraction sub-module The specific formula is as follows:
[0132] F1 = ReLU(Conv1(ReLU(Conv1(Avgpooling(F)))))
[0133] F 11 = Upsampling(Conv1(F1))
[0134] F 121 = ReLU(Conv1(ReLU(Conv1(Avgpooling(F1)))))
[0135] F 12 = Upsampling(EeLU(Conv1(Upsampling(F 121 ))))
[0136] F m = Conv1(SE(Concat(F 11 + F 12 , F)))
[0137] Among them, ReLU is the activation function, Conv1 is 1×1 convolution, SE(·) is the SE module, Avgpooling is the average pooling layer with a kernel size of 2×2 and a stride of 2, Upsampling is the two-fold nearest neighbor upsampling layer, and Concat is the concatenation operation along the channel dimension;
[0138] Step B22: Design a frequency-domain and spatial-domain feature fusion sub-module, which is composed of serial connection of channel attention and spatial attention;
[0139] Step B23: Design a frequency-domain and spatial-domain context information attention module, which is composed of three multi-scale feature extraction sub-modules and a frequency-domain and spatial-domain feature fusion sub-module. The inputs of the three multi-scale feature extraction sub-modules are the three feature maps F low 、F spa 、Ffre , after being processed by the multi-scale feature extraction sub-module designed in step B21 respectively, feature maps F with context information are obtained low_m 、F spa_m 、F fre_m , and then through the frequency-spatial domain feature fusion sub-module designed in step B22, the feature map F output by the frequency-spatial domain context information attention module is obtained f .
[0140] Furthermore, step B22 includes the following steps:
[0141] Step B221, design channel attention, and the input is the feature map obtained in step B23 The three feature maps are respectively subjected to global average pooling in the spatial dimension to obtain three vectors with a scale of 1×1×C, and then the three vectors are concatenated along the channel dimension to obtain an intermediate feature map Subject F c to dimensionality reduction and dimensionality increase in sequence through 1×1 convolution, ReLU activation function, 1×1 convolution, ReLU activation function, and 1×1 convolution, and then obtain the weights on the channel dimension through the Sigmoid activation function Decompose F W1 along the channel dimension into three vectors with a scale of 1×1×C, F W10 、F W11 、F W12 , and multiply them with the input feature maps F low_m 、F spa_m 、F fre_m of the frequency-spatial domain feature fusion sub-module respectively to obtain the output feature map of channel attention The specific formula is expressed as follows:
[0142] F c =Concat(Avgpooling s (F low_m ),Avgpooling s (F spa_m ),Avgpooling s (F fre_m ))
[0143] F W1 =Sigmoid(Conv1(ReLU(Conv1(ReLU(Conv1(F c ))))))
[0144] F low_c =F W10 ×F low_m
[0145] F spa_c = F W11 × F spa_m
[0146] F fre_c = F W12 × F fre_m
[0147] where Concat is the concatenation operation along the channel dimension, Avgpooling s is the global average pooling in the spatial dimension, ReLU is the activation function, Conv1 is the 1×1 convolution, and Sigmoid is the Sigmoid activation function;
[0148] Step B222: Design spatial attention. The input is the three feature maps F low_c , F spa_c , F fre_c obtained in Step B221. After performing average pooling on each of the three feature maps along the channel dimension, three feature maps with a scale of H×W×1 are obtained. Then, the three feature maps are concatenated along the channel dimension to obtain an intermediate feature map F s is successively passed through an average pooling layer with a kernel size of 2×2 and a stride of 2, a ReLU activation function, an upsampling layer, and then through a Sigmoid activation function to obtain the weight in the spatial dimension F W2 is decomposed into three feature maps with a scale of H×W×1, F W20 , F W21 , F W22 , which are respectively multiplied by the input feature maps F low_c , F spa_c , F fre_c of the spatial attention to obtain the output feature map of the spatial attention The specific formula is as follows:
[0149] F s = Concat(Avgpooling c (F low_ c), Avgpooling c (F spa_c ), Avgpooling c (F fre_c ))
[0150] F W2 = Sigmoid(Upsampling(ReLU(Avgpooling(F s ))))
[0151] F low_s = FW20 ×F low_c
[0152] F spa_s = F W21 ×F spa_c
[0153] F fre_s = F W22 ×F fre_c
[0154] Among them, Concat is the concatenation operation along the channel dimension, Avgpooling c is the average pooling in the channel dimension, ReLU is the activation function, Sigmoid is the Sigmoid activation function, Avgpooling is the average pooling layer with a kernel size of 2×2 and a stride of 2, and Upsampling is the nearest neighbor upsampling layer with a factor of 2;
[0155] Step B223: Design the frequency domain feature fusion sub-module. The input feature maps are the feature maps F low_m , F spa_m , F fre_m . The three feature maps first pass through the channel attention in Step B221 to obtain the feature maps F low_c , F spa_c , F fre_c , and then pass through the spatial attention in Step B222 to obtain the feature maps F low_s , F spa_s , F fre_s . The three feature maps are added together to obtain the final output F f . The specific formula is as follows:
[0156] F f = F low_s + F spa_s + F fre_s
[0157] Furthermore, Step B3 is implemented as follows:
[0158] Step B31: Design the feature aggregation convolution block to achieve the fusion of detailed information and context information. The input feature maps are the feature map F fre obtained in Step B1 and the feature map F f obtained in Step B2. After concatenating the two along the channel dimension, they pass through a 3×3 convolution to obtain the output feature map F conv . The specific formula is as follows:
[0159] F conv = Conv3(Concat(F fre , F f ))
[0160] Among them, Conv3 is a 3×3 convolution, and Concat is an operation of concatenating along the channel dimension;
[0161] Step B32: Design a collaborative enhancer module to collaboratively enhance the fused information of detailed information and context information. The input feature map is F obtained in step B31 conv , and F conv is sequentially passed through a 1×1 convolution, a ReLU6 activation function, a Dropout random inactivation layer, a 1×1 convolution, and a Dropout random inactivation layer, and then added to F conv to obtain an intermediate feature map F mid . Then, through a LeakyReLU activation function, it is concatenated with F conv along the channel dimension and then passed through a 3×3 convolution to obtain an output feature map F co . The specific formula is as follows:
[0162] F mid = Dropout(Conv1(Dropout(ReLU6(Conv1(F conv ))))) + F conv
[0163] F co = Conv3(Concat(LeakyReLU(F mid ), F conv ))
[0164] Among them, Conv1 is a 1×1 convolution, Conv3 is a 3×3 convolution, Concat is an operation of concatenating along the channel dimension, Dropout is a random inactivation layer, ReLU6 is a ReLU6 activation function, and LeakyReLU is a LeakyReLU activation function;
[0165] Step B33: Design a feature aggregation and enhancement module, which consists of a feature aggregation convolution block and a collaborative enhancer module. The input feature maps are the feature map F fre obtained in step B1 and the feature map F f obtained in step B2. After passing through the feature aggregation convolution block, the feature map F conv is obtained, and then after passing through the collaborative enhancer module, the feature map F co is obtained.
[0166] Furthermore, step B4 is implemented as follows:
[0167] Step B4: Design a full-resolution low-light image enhancement network that aggregates context and enhances details, which is composed of integrating a full-resolution detail extraction module, a frequency-domain context information attention module, and a feature aggregation and enhancement module. Input the low-light image I, and after passing through the full-resolution detail extraction module in Step B1, three feature maps F low 、F spa 、F fre are obtained. After passing through the frequency-domain context information attention module, the feature map F f is obtained. Finally, after passing through the feature aggregation and enhancement module, the feature map F co is obtained. Then, F co and the feature map F low in Step B1 are concatenated along the channel dimension, and then through a 3×3 convolution, the final enhanced image I out is obtained. The specific formula is as follows:
[0168] I out =Conv3(Concat(F co ,F low ))
[0169] where Conv3 is a 3×3 convolution, and Concat is an operation of concatenating along the channel dimension.
[0170] Further, Step C is implemented as follows:
[0171] Step C: Design a loss function, which consists of L2 loss and VGG perceptual loss. The total objective loss function of the network is as follows:
[0172] l=ω1||I out -G|| 2 +ω2||Φ(I out )-Φ(G)||1
[0173] where Φ(·) represents the operation of extracting the features of the Conv4-1 layer using the pre-trained VGG-16 classification model on the ImageNet dataset. I out represents the enhanced image of the low-light image I, G represents the labeled image corresponding to the low-light image I, ||.||1 represents the L1 loss, ||.|| 2 represents the L2 loss, and ω1, ω2 are weights.
[0174] Further, Step D is implemented as follows:
[0175] Step D1: Randomly divide the training dataset obtained in Step A into several batches, and each batch contains N pairs of images;
[0176] Step D2: Input the low-light image I, and obtain the enhanced image I after passing through the full-resolution low-light image enhancement network for aggregating context and enhancing details in Step B. out , and calculate the loss l using the formula in Step C;
[0177] Step D3: Calculate the gradients of the parameters in the network using the backpropagation method according to the loss, and update the network parameters using the Adam optimization method.
[0178] Step D4: Repeat Steps D1 to D3 in batches until the value of the objective loss function of the network converges to the Nash equilibrium, save the network parameters, and obtain the full-resolution low-light image enhancement model for aggregating context and enhancing details.
[0179] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0180] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0181] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0182] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are performed on the computer or other programmable apparatus to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one process or a plurality of processes and / or blocks. Figure 1 One process or a plurality of processes and / or blocks Figure 1 Steps for implementing the functions specified in one block or a plurality of blocks.
[0183] As described above, it is only the preferred embodiment of the present invention, and it is not intended to limit the present invention in any other form. Any person skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes. However, any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the technical solution content of the present invention still belong to the protection scope of the technical solution of the present invention.
[0184] This patent is not limited to the above best mode. Anyone inspired by this patent can obtain other various forms of full-resolution low-light image enhancement methods with aggregated context and enhanced details. All equivalent changes and modifications made according to the scope of the patent application of the present invention shall fall within the scope covered by this patent.
Claims
1. A full-resolution low-light image enhancement method for aggregating context and enhancing details, characterized in that: Step A: Perform data preprocessing, including data pairing, data random cropping, and data augmentation processing to obtain a training data set; Step B: Design a full-resolution low-light image enhancement network for aggregating context and enhancing details, including: a full-resolution detail extraction module, a frequency-domain context information attention module, and a feature aggregation and enhancement module; Step C: Design a loss function for guiding the parameter optimization of the network designed in Step B; Step D: Use the training data set obtained in Step A to train the full-resolution low-light image enhancement network for aggregating context and enhancing details in Step B, converge to the Nash equilibrium, and obtain a trained full-resolution low-light image enhancement model for aggregating context and enhancing details; Step E: Input the low-light image to be measured into the trained full-resolution low-light image enhancement model for aggregating context and enhancing details, and output the enhanced normal-light image; The specific implementation steps of Step B are as follows: Step B1: Construct a full-resolution detail extraction module, which consists of a shallow feature extraction sub-module, an attention sub-module based on CBAM, and a frequency-domain transformation sub-module, and use the designed network to extract detail features; Step B2: Design a frequency-domain context information attention module, which consists of a multi-scale feature extraction sub-module and a frequency-domain feature fusion sub-module, and use the designed network to extract context features; Step B3: Design a feature aggregation and enhancement module, which consists of a feature aggregation convolutional block and a collaborative enhancement sub-module, aggregate the detail features extracted in Step B1 and the context features extracted in Step B2, and enhance the two types of features together; Step B4: Design a full-resolution low-light image enhancement network for aggregating context and enhancing details, including a full-resolution detail extraction module, a frequency-domain context information attention module, and a feature aggregation and enhancement module; The specific implementation method of Step C is: Step C: Design a loss function, which consists of an L2 loss and a VGG perceptual loss. The total objective loss function of the network is as follows: l = ω1||I out -G|| 2 +ω2||Φ(I out ) - Φ(G)||1 Among them, Φ(·) represents the operation of extracting the features of the Conv4-1 layer using the VGG-16 classification model pre-trained on the ImageNet dataset; I out represents the enhanced image of the low-light image I, G represents the label image corresponding to the low-light image I, ||.||1 represents the L1 loss, ||.|| 2 represents the L2 loss, and ω1 and ω2 are weights.
2. The full-resolution low-light image enhancement method for aggregating context and enhancing details according to claim 1, wherein The specific implementation steps of Step A are as follows: Step A1: Pair the low-light image and the corresponding label image; Step A2: Randomly crop each low-light image with a size of h×w×3 into an image with a size of p×p×3, and perform the same random cropping method on its corresponding label image, where h and w are the heights and widths of the low-light image and the label image, and p is the height and width of the cropped image; Step A3: Randomly use one of the following 8 enhancement methods to perform data augmentation on the training paired images: keep the original image, vertical flip, rotate 90 degrees, vertical flip after rotating 90 degrees, rotate 180 degrees, vertical flip after rotating 180 degrees, rotate 270 degrees, vertical flip after rotating 270 degrees.
3. The full-resolution low-light image enhancement method for aggregating context and enhancing details according to claim 1, characterized in that: The specific implementation steps of Step B1 are as follows: Step B11: Design a shallow feature extraction sub-module. The input is a low-illumination image I, which is convolved with a 3×3 filter to obtain an initial feature map F ori After that, it enters three branches. The first branch contains one 3×3 convolution, the second branch contains two serial 3×3 convolutions, and the third branch contains three serial 3×3 convolutions. The processing results F B1 、F B2 、F B3 are concatenated along the channel dimension and then convolved with a 3×3 filter to obtain the feature map F output by the shallow feature extraction sub-module low ; The specific formula is as follows: F ori = Conv3(I) F B1 = Conv3(F ori ) F B2 = Conv3(Conv3(F ori )) F B3 = Conv3(Conv3(Conv3(F ori ))) F low = Conv3(Concat(F B1 , F B2 , F B3 )) Wherein, Conv3 is a 3×3 convolution, and Concat is an operation of splicing along the channel dimension; Step B12: Construct an attention sub-module based on CBAM, which consists of sequential channel-dimensional attention Att c and spatial-dimensional attention Att s . The input feature map is the feature map F obtained in Step B11 low , and the feature map output by the attention sub-module based on CBAM is F spa ; The specific formula is as follows: F spa = Att s (Att c (F low )) Among them, Att c is the attention in the channel dimension, and Att s is the attention in the spatial dimension; Step B13: Design a frequency-domain transformation sub-module. The input feature map is the feature map F obtained in step B12 spa , use the Fourier transform function to convert the spatial domain to the frequency domain, then successively pass through a 3×3 convolution, a normalization layer, and a ReLU activation function, and finally use the inverse Fourier transform function to convert the frequency domain back to the spatial domain to obtain the output feature map F fre ; The specific formula is as follows: F fre = idft(ReLU(BN(Conv3(dft(F spa )))) Among them, DFT is the Fourier transform, IDFT is the inverse Fourier transform, ReLU is the ReLU activation function, BN is the batch normalization layer, and Conv3 is a 3×3 convolution; Step B14: Construct a full-resolution detail extraction module, which consists of a shallow feature extraction sub-module, an attention sub-module based on CBAM, and a frequency domain transformation sub-module. Let the input be the low-illuminance image I processed in Step A. After passing through the shallow feature extraction sub-module, the attention sub-module, and the frequency domain transformation sub-module in sequence, the feature map F is obtained. low , F spa , F fre .
4. The full-resolution low-light image enhancement method for aggregating context and enhancing details according to claim 1, characterized in that: The specific implementation steps of step B2 are as follows: Step B21: Design a multi-scale feature extraction sub-module. Denote the input feature map as F. Let H, W, and C be the height, width, and number of channels of the feature F respectively. After passing through an average pooling layer with a kernel size of 2×2 and a stride of 2, it successively passes through a 1×1 convolution, a ReLU activation function, a 1×1 convolution, and a ReLU activation function for dimensionality reduction to obtain an intermediate feature map. Then it is divided into two branches. After the upper branch passes through a 1×1 convolution for further dimensionality reduction, the output of the upper branch is obtained through an upsampling layer. a is the number of channels after dimensionality reduction; the other branch passes through an average pooling layer with a kernel size of 2×2 and a stride of 2, and then successively passes through a 1×1 convolution, a ReLU activation function, a 1×1 convolution, and a ReLU activation function for dimensionality reduction to obtain an intermediate feature map. The intermediate feature map F 121 Then it successively passes through an upsampling layer, a 1×1 convolution, a ReLU activation function, and an upsampling layer to obtain the output of the lower branch. Add F 11 and F 12 After addition, concatenate it with F in the channel dimension, and then adjust the channels through a 1×1 convolution after passing through the SE module to obtain the feature map output by the multi-scale feature extraction sub-module. The specific formula is as follows: F1 = ReLU(Conv1(ReLU(Conv1(Avgpooling(F))))) F 11 = Upsampling(Conv1(F1)) F 121 = ReLU(Conv1(ReLU(Conv1(Avgpooling(F1))))) F 12 = Upsampling(ReLU(Conv1(Upsampling(F 121 )))) F m = Conv1(SE(Concat(F 11 + F 12 , F))) Among them, ReLU is the activation function, Conv1 is a 1×1 convolution, SE(·) is the SE module, Avgpooling is an average pooling layer with a kernel size of 2×2 and a stride of 2, Upsampling is a two-fold nearest neighbor upsampling layer, and Concat is a concatenation operation along the channel dimension; Step B22: Design a frequency-domain feature fusion sub-module, which consists of a serial connection of channel attention and spatial attention; Step B23: Design a frequency-domain context information attention module, which consists of three multi-scale feature extraction sub-modules and a frequency-domain feature fusion sub-module; the inputs of the three multi-scale feature extraction sub-modules are the three feature maps F low , F spa , F fre respectively. After being processed by the multi-scale feature extraction sub-modules designed in Step B21, the feature maps F low_m , F spa_m , F fre_m with context information are obtained. Then, through the frequency-domain feature fusion sub-module designed in Step B22, the feature map F f output by the frequency-domain context information attention module is obtained.
5. The full-resolution low-light image enhancement method for aggregating context and enhancing details according to claim 4, characterized in that: The specific implementation steps of step B22 are as follows: Step B221: Design channel attention, with the input being the feature map obtained in Step B23 The three feature maps respectively undergo global average pooling in the spatial dimension to obtain three vectors of size 1×1×C, and then the three vectors are concatenated along the channel dimension to obtain an intermediate feature map Let F c successively go through 1×1 convolution, ReLU activation function, 1×1 convolution, ReLU activation function, and 1×1 convolution for dimensionality reduction and dimensionality increase, and then obtain the weights on the channel dimension through the Sigmoid activation function Let F W1 be decomposed along the channel dimension into three vectors of size 1×1×C, namely F W10 , F W11 , F W12 , which are respectively multiplied by the input feature maps F low_m , F spa_m , F fre_m of the frequency domain feature fusion sub-module to obtain the output feature map of channel attention The specific formula is as follows: F c = Concat(Avgpooling s (F low_m ), Avgpooling s (F spa_m ), Avgpooling s (F fre_m )) F W1 = Sigmoid(Conv1(ReLU(Conv1(ReLU(Conv1(F c )))))) F low_c = F W10 × F low_m F spa_c = F W11 × F spa_m F fre_c = F W12 × F fre_m Among them, Concat is the concatenation operation along the channel dimension, Avgpooling s is the global average pooling in the spatial dimension, ReLU is the activation function, Conv1 is the 1×1 convolution, and Sigmoid is the Sigmoid activation function; Step B222: Design spatial attention. The input is the three feature maps F low_c 、F spa_c 、F fre_c . After performing average pooling on the three feature maps along the channel dimension respectively, three feature maps with a scale of H×W×1 are obtained. Then, the three feature maps are concatenated along the channel dimension to obtain an intermediate feature map Pass F s sequentially through an average pooling layer with a kernel size of 2×2 and a stride of 2, a ReLU activation function, an upsampling layer, and then obtain the weights in the spatial dimension through a Sigmoid activation function Decompose F W2 into three feature maps with a scale of H×W×1, namely F W20 、F W21 、F W22 . Multiply them with the input feature maps F low_c 、F spa_c 、F fre_c of the spatial attention respectively to obtain the output feature maps of the spatial attention The specific formula is as follows: F s = Concat(Avgpooling c (F low_c ), Avgpooling c (F spa_c ), Avgpooling c (F fre_c )) F W2 = Sigmoid(Upsampling(ReLU(Avgpooling(F s )))) F low_s = F W20 × F low_c F spa_s = F W21 × F spa_c F fre_s = F W22 × F fre_c Among them, Concat is the concatenation operation along the channel dimension, Avgpooling c is the average pooling in the channel dimension, ReLU is the activation function, Sigmoid is the Sigmoid activation function, Avgpooling is the average pooling layer with a kernel size of 2×2 and a stride of 2, and Upsampling is the nearest neighbor upsampling layer with a factor of 2; Step B223: Design a frequency domain feature fusion sub-module. The input feature maps are the feature maps F obtained in Step B23 low_m 、F spa_m 、F fre_m . First, the three feature maps pass through the channel attention in Step B221 to obtain the feature maps F low_c 、F spa_c 、F fre_c . Then, they pass through the spatial attention in Step B222 to obtain the feature maps F low_s 、F spa_s 、F fre_s . After adding the three feature maps, the final output F f is obtained. The specific formula is as follows: F f = F low_s + F spa_s + F fre_s .
6. The full-resolution low-light image enhancement method for aggregating context and enhancing details according to claim 1, characterized in that: The specific implementation steps of step B3 are as follows: Step B31: Design a feature aggregation convolutional block to achieve the fusion of detailed information and context information; the input feature map is the feature map F obtained in step B1 fre and the feature map F obtained in step B2 f . After concatenating the two along the channel dimension and passing through a 3×3 convolution, the output feature map F conv is obtained; the specific formula is as follows: F conv = Conv3(Concat(F fre , F f )) Among them, Conv3 is a 3×3 convolution, and Concat is a concatenation operation along the channel dimension; Step B32: Design a collaborative enhancement sub-module to collaboratively enhance the fused information of detailed information and context information; the input feature map is F obtained in Step B31 conv , pass F conv through a 1×1 convolution, ReLU6 activation function, Dropout random inactivation layer, 1×1 convolution, and Dropout random inactivation layer in sequence, and then add it to F conv to obtain an intermediate feature map F mid . Then, through the LeakyReLU activation function, concatenate it with F conv along the channel dimension and pass through a 3×3 convolution to obtain an output feature map F co ; the specific formula is as follows: F mid = Dropout(Conv1(Dropout(ReLU6(Conv1(F conv )))) + F conv F co = Conv3(Concat(LeakyReLU(F mid ), F conv )) Among them, Conv1 is a 1×1 convolution, Conv3 is a 3×3 convolution, Concat is a concatenation operation along the channel dimension, Dropout is a random inactivation layer, ReLU6 is the ReLU6 activation function, and LeakyReLU is the LeakyReLU activation function; Step B33: The design feature aggregation and enhancement module, which consists of a feature aggregation convolutional block and a collaborative enhancement sub-module. The input feature maps are the feature map F obtained in Step B1 fre and the feature map F obtained in Step B2 f . After passing through the feature aggregation convolutional block, the feature map F conv is obtained. After passing through the collaborative enhancement sub-module, the feature map F co is obtained.
7. The full-resolution low-light image enhancement method for aggregating context and enhancing details according to claim 1, characterized in that: The specific implementation manner of step B4 is: Step B4: Design a full-resolution low-light image enhancement network that aggregates context and enhances details, which is composed of integrating a full-resolution detail extraction module, a frequency-domain context information attention module, and a feature aggregation and enhancement module; input the low-light image I, and after passing through the full-resolution detail extraction module in Step B1, three feature maps F low , F spa , F fre are obtained. After passing through the frequency-domain context information attention module, the feature map F f is obtained, and then after passing through the feature aggregation and enhancement module, the feature map F co is obtained. Then, F co is concatenated with the feature map F low in Step B1 along the channel dimension, and then through a 3×3 convolution to obtain the final enhanced image I out ; the specific formula is as follows: I out = Conv3(Concat(F co , F low )) Among them, Conv3 is a 3×3 convolution, and Concat is a concatenation operation along the channel dimension.
8. The full-resolution low-light image enhancement method for aggregating context and enhancing details according to claim 1, wherein: The specific implementation steps of step D are as follows: Step D1: Randomly divide the training data set obtained in step A into several batches, and each batch contains N pairs of images; Step D2: Input the low-light image I, and obtain the enhanced image I after passing through the full-resolution low-light image enhancement network for aggregating context and enhancing details in Step B out , and calculate the loss l using the formula in Step C; Step D3: Calculate the gradients of the parameters in the network using the backpropagation method according to the loss, and update the network parameters using the Adam optimization method; Step D4: Repeat steps D1 to D3 in batches until the value of the target loss function of the network converges to the Nash equilibrium, save the network parameters, and obtain a full-resolution low-light image enhancement model for aggregating context and enhancing details.