Low-light aerial image enhancement method, terminal and storage medium
By extracting the context features of low-light aerial images and combining the attention discrete wavelet transformation module and frequency feature modulation module, the problems of noise increase, overexposure and color distortion in the existing low-light aerial image enhancement methods in the prior art are solved, and more effective image enhancement effects are achieved and the accuracy of target detection is improved.
Patent Information
- Application Number
- CN202510120031.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-25
- Publication Date
- 2025-06-27
AI Technical Summary
Existing low-light aerial image enhancement methods have problems with noise increase, overexposure and color distortion, and have failed to effectively consider the contribution of different frequency features to illumination estimation and image enhancement.
By extracting contextual features in low-light aerial images, multi-scale image features are learned using an encoder based on the attention discrete wavelet transform module, multi-scale illumination features are estimated in combination with a decoder based on the attention inverse discrete wavelet transform module, and multi-scale features are integrated through the frequency feature modulation module to perform convolution operations to obtain enhanced aerial images.
It effectively suppresses noise, avoids overexposure and color distortion, improves image discrimination and brightness, enhances the accuracy of object detection, and meets the image enhancement needs under low-light conditions.
Smart Images

Figure CN120219191A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and particularly relates to a low-light aerial image enhancement method, a terminal, and a storage medium. Background Art
[0002] To meet the requirements of round-the-clock combat, unmanned aerial vehicles (UAVs) usually conduct reconnaissance across day and night. The obtained images may be low-light or even dark images due to limited illumination, making the distinction between the target and the surrounding environment weaker and showing the characteristic of insufficient brightness, which seriously affects the subsequent target detection task.
[0003] In related technologies, low-light image enhancement methods are mainly divided into three categories: methods based on histogram equalization, methods based on the Retinex model, and methods based on deep learning. Although these methods have achieved good performance, there are still some problems: methods based on histogram equalization cannot well adapt to various scenarios and may cause an undesired increase in image noise; methods based on the Retinex model often have problems such as overexposure and color distortion in the enhanced images due to the ill-posedness of the decomposition problem and insufficient constraints on the reflection component, which is inconsistent with human visual perception; methods based on deep learning usually use attention modules to emphasize meaningful features along the channel and spatial dimensions, without considering that features of different frequencies have different contributions to illumination estimation and image enhancement. In view of the shortcomings of existing methods, there is a current need to design a more effective image enhancement method. Summary of the Invention
[0004] In view of this, embodiments of the present invention provide a low-light aerial image enhancement method, a terminal, and a storage medium to more effectively enhance low-light aerial images.
[0005] The first aspect of the embodiments of the present invention provides a low-light aerial image enhancement method, including:
[0006] Extracting context features from the low-light aerial image;
[0007] For the context features, learning multi-scale image features through an encoder based on an attention discrete wavelet transform module, estimating multi-scale illumination features through a decoder based on an attention inverse discrete wavelet transform module, and there is a skip connection between the encoder and the decoder;
[0008] Integrating the multi-scale image features and the multi-scale illumination features through a frequency feature modulation module to obtain integrated features;
[0009] Performing a convolution operation on the integrated features to obtain an enhanced aerial image.
[0010] As a possible implementation, each scale corresponds to a frequency feature modulation module; integrating the multi-scale image features and the multi-scale illumination features through the frequency feature modulation module to obtain integrated features, including:
[0011] Through each of the frequency feature modulation modules, performing an affine transformation on the illumination feature of the corresponding scale to obtain a scaling parameter and a shifting parameter;
[0012] Modulating the feature with the lowest frequency in the image feature of the corresponding scale according to the scaling parameter, and modulating the features of all frequencies in the image feature of the corresponding scale according to the shifting parameter to obtain the integrated features; wherein, the image features include features of multiple different frequencies.
[0013] As a possible implementation, the encoder includes: an attention discrete wavelet transform module corresponding to each scale and a first residual block; learning multi-scale image features through the encoder based on the attention discrete wavelet transform module, including:
[0014] Through each of the attention discrete wavelet transform modules, performing a separation operation on the input feature of the corresponding scale to obtain features of multiple different frequencies;
[0015] Inputting the features of multiple different frequencies into the first residual block to obtain the image features of the corresponding scale.
[0016] As a possible implementation, the attention discrete wavelet transform module includes: a discrete wavelet transform unit, a group of parallel first spatial attention units, a first channel attention unit, and a first convolution unit; performing a separation operation on the input feature of the corresponding scale through each of the attention discrete wavelet transform modules to obtain features of multiple different frequencies, including:
[0017] Through the discrete wavelet transform unit, performing a discrete wavelet transform on the input feature of the corresponding scale to obtain multiple different frequency sub-bands;
[0018] Performing spatial-level modulation on each frequency sub-band through the first spatial attention unit;
[0019] Performing channel-level modulation on each frequency sub-band through the first channel attention unit;
[0020] Performing a convolution operation through the first convolution unit to obtain features of multiple different frequencies.
[0021] As a possible implementation, the decoder includes: an attention inverse discrete wavelet transform module corresponding to each scale and a second residual block; estimating multi-scale illumination features through the decoder based on the attention inverse discrete wavelet transform module, including:
[0022] Through each of the attention inverse discrete wavelet transform modules, a reconstruction operation is performed on the input features of the corresponding scale to obtain reconstructed features;
[0023] The reconstructed features are input into the second residual block to obtain illumination features of the corresponding scale.
[0024] As a possible implementation manner, the attention inverse discrete wavelet transform module includes: an inverse discrete wavelet transform unit, a group of parallel second spatial attention units, a second channel attention unit, and a second convolution unit; the step of performing a reconstruction operation on the input features of the corresponding scale through each of the attention inverse discrete wavelet transform modules to obtain reconstructed features includes:
[0025] Performing spatial-level modulation on the input features of the corresponding scale through the second spatial attention unit;
[0026] Performing channel-level modulation on the input features of the corresponding scale through the second channel attention unit;
[0027] Performing an inverse discrete wavelet transform on the input features of the corresponding scale through the inverse discrete wavelet transform unit, and performing a convolution operation through the second convolution unit to obtain reconstructed features.
[0028] As a possible implementation manner, the extraction of context features in the low-light aerial image includes: extracting context features in the low-light aerial image through a 3×3 dilation convolution with a dilation rate of 2.
[0029] As a possible implementation manner, the method further includes:
[0030] Constructing a joint loss function of the image enhancement network according to the weighted sum of pixel loss, content loss, smooth loss, and color loss; wherein, the image enhancement network is the entire network for processing the low-light aerial image to obtain the enhanced aerial image;
[0031] Optimizing the parameters of the image enhancement network based on the joint loss function.
[0032] A second aspect of the embodiments of the present invention provides a terminal, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the method in the above first aspect or any one of the implementation manners of the first aspect are implemented.
[0033] A third aspect of the embodiments of the present invention provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps of the method in the above first aspect or any one of the implementation manners of the first aspect are implemented.
[0034] The beneficial effects of the embodiments of the present invention compared with the prior art are as follows:
[0035] In the embodiments of the present invention, the context features in the low-light aerial images are first extracted, and then the multi-scale image features are learned by an encoder based on the attention discrete wavelet transform module, and the multi-scale illumination features are estimated by a decoder based on the attention inverse discrete wavelet transform module. In the attention discrete wavelet transform module and the attention inverse discrete wavelet transform module, by combining the attention mechanism with the wavelet transform, important features can be captured and redundant noise can be suppressed. Further, through the frequency feature modulation module, the illumination information is combined with the image features to perform frequency feature modulation to avoid noise amplification. Therefore, through the image enhancement network designed in the embodiments of the present invention, the low-light aerial images can be more effectively enhanced. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the embodiments or the prior art descriptions. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0037] Figure 1 is a schematic structural diagram of the image enhancement network provided by the embodiments of the present invention;
[0038] Figure 2 is a schematic structural diagram of the attention discrete wavelet transform module provided by the embodiments of the present invention;
[0039] Figure 3 is a schematic structural diagram of the attention inverse discrete wavelet transform module provided by the embodiments of the present invention;
[0040] Figure 4 is a schematic structural diagram of the frequency feature modulation module provided by the embodiments of the present invention;
[0041] Figure 5 is a schematic structural diagram of the terminal provided by the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0042] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present invention. However, those skilled in the art should clearly understand that the present invention can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present invention.
[0043] To illustrate the technical solution of the present invention, the following will be described through specific embodiments.
[0044] The low-light aerial image enhancement method provided in this embodiment is implemented based on Figure 1 the constructed image enhancement network. The following will describe the method of this embodiment in conjunction with Figure 1 .
[0045] Step 1: Extract the context features in the low-light aerial image.
[0046] As shown in Figure 1 , in this embodiment, a 3×3 extended convolution with an expansion rate of 2 can be used to extract the context features of the target in the low-light aerial image.
[0047] Step 2: For the context features, use the encoder based on the attention discrete wavelet transform module to learn multi-scale image features, and use the decoder based on the attention inverse discrete wavelet transform module to estimate multi-scale illumination features. There is a skip connection between the encoder and the decoder.
[0048] In this embodiment, the encoder is used to learn multi-scale image features I L k , where I L k represents the features generated at the k-th scale of the encoder. As shown in Figure 3 , the encoder can be composed of an attention discrete wavelet transform module and a residual block. Each group of attention discrete wavelet transform module and residual block corresponds to processing image features at one scale. The attention discrete wavelet transform module is used to separate high-frequency noise and low-frequency content in the frequency domain, adaptively modulate features of different frequencies, similar to a downsampling operation, reducing the resolution of the input feature F in by two times and outputting the feature F out . The residual block can be composed of two 1×1 convolutions and a ReLU activation function, which is used to avoid the problem of gradient disappearance or gradient explosion during the network training process.
[0049] Combining Step 1 and Step 2, Figure 1 the image features at three scales in
[0050]
[0051] I L 2 are calculated as follows: L 1 ))
[0052] I L 3 =RB(ADWT(I L 2))
[0053] Among them, I L k represents the feature generated at the k-th scale of the encoder, where k = 1, 2, 3. RB() represents the operation of the residual block, and ADWT() represents the operation of the attention discrete wavelet transform module. represents a 3×3 dilation convolution with a dilation rate of 2.
[0054] Similarly, the decoder is used to estimate the multi-scale illumination features Among them represents the feature generated at the k-th scale of the decoder. The decoder can be composed of an attention inverse discrete wavelet transform module and a residual block. The attention inverse discrete wavelet transform module can also separate the high-frequency noise and low-frequency content in the frequency domain, adaptively modulate the features of different frequencies, similar to the upsampling operation, doubling the resolution of the input feature F in ' and outputting the feature F out '. The residual block can be composed of two 1×1 convolutions and a ReLU activation function to avoid the problem of gradient vanishing or gradient explosion during the network training process.
[0055] Combining Step 1 and Step 2, Figure 1 the illumination features at the three scales in
[0056] F I 3 = AIDWT(RB(I L 3 ))
[0057] F I 2 = AIDWT(RB(F I 3 + I L 2 ))
[0058] F I 1 = AIDWT(RB(F I 2 + I L 1 ))
[0059] Among them, I L k represents the feature generated at the k-th scale of the encoder, represents the feature generated at the k-th scale of the decoder, where k = 1, 2, 3. RB() represents the operation of the residual block, and AIDWT() represents the operation of the attention inverse discrete wavelet transform module.
[0060] Here, the image enhancement network can be implemented using a U-Net architecture, so skip connections can be designed between the encoder and the decoder, as Figure 1 shown by the yellow arrows in
[0061] Step S103: Integrate the multi-scale image features and multi-scale illumination features through a frequency feature modulation module to obtain integrated features.
[0062] Refer to Figure 1 shown. Each scale corresponds to a frequency feature modulation module. The frequency feature modulation module is used to integrate the image features and illumination features in a coarse-to-fine manner, and the illumination information is used to guide the low-light image enhancement. Combining the illumination information with the image features to perform frequency feature modulation can effectively avoid noise amplification.
[0063] Step S104: Perform a convolution operation on the integrated features to obtain an enhanced aerial image.
[0064] Refer to Figure 1 shown. The integrated features output by the frequency feature modulation module are input into a 3×3 convolution for convolution operation to generate an enhanced image I E .
[0065] In the embodiments of the present invention, the context features in the low-light aerial image are first extracted, then the multi-scale image features are learned through an encoder based on the attention discrete wavelet transform module, and the multi-scale illumination features are estimated through a decoder based on the attention inverse discrete wavelet transform module. The attention discrete wavelet transform module and the attention inverse discrete wavelet transform module can capture important features and suppress redundant noise by combining the attention mechanism with the wavelet transform. Further, through the frequency feature modulation module, the illumination information is combined with the image features to perform frequency feature modulation to avoid noise amplification. Therefore, the image enhancement network designed by the embodiments of the present invention can more effectively enhance the low-light aerial image.
[0066] In some embodiments, the attention discrete wavelet transform module and the attention inverse discrete wavelet transform module can be respectively implemented through the Figure 2 and Figure 3 structures.
[0067] To suppress the noise and artifacts in the enhanced image, this embodiment incorporates the wavelet transform into the deep convolutional neural network to separate the high-frequency noise and low-frequency content in the frequency domain. Considering the contributions of different frequency features, the attention mechanism is combined with the wavelet transform to propose an attention wavelet transform mechanism to capture the important content in different frequency features and suppress redundant noise. Similar to the downsampling and upsampling operations in the encoder-decoder structure, the attention wavelet transform mechanism can be further divided into an attention discrete wavelet transform module and an attention inverse discrete wavelet transform module.
[0068] Note that the discrete wavelet transform module is as Figure 2 shown. First, use the discrete wavelet transform unit to convert the input feature F in into four frequency sub-bands {I LL , I LH , I HL , I HH}:
[0069] I LL , I LH , I HL , I HH = DWT(F in )
[0070] where DWT() represents the operation of the discrete wavelet transform unit, and I xy represents different frequency sub-bands, L represents low frequency, and H represents high frequency.
[0071] Then, apply the spatial attention unit and the channel attention unit to modulate these frequency sub-bands. Features of different frequencies contain different information. A group of parallel spatial attention units and a channel attention unit are used to adaptively modulate features of different frequencies.
[0072] Finally, use 1×1 convolution to adjust the number of feature channels.
[0073] That is:
[0074] F out = Conv1(CA[SA(DWT(F in ))])
[0075] where Conv1 represents 1×1 convolution, CA() represents the operation of the channel attention unit, SA() represents the operation of the spatial attention unit, and [] represents the concatenation operation.
[0076] Similarly, note that the inverse discrete wavelet transform module is as Figure 3 shown. First, perform spatial-level and channel-level feature modulation on the input feature F in ', and then use the inverse discrete wavelet transform operation and 1×1 convolution to reconstruct the output feature F out '.
[0077] That is:
[0078] F out ' = Conv1(IDWT(CA[SA(F in ')]))
[0079] Among them, Conv1 represents a 1×1 convolution, IDWT() represents the operation of the inverse discrete wavelet transform unit, CA() represents the operation of the channel attention unit, SA() represents the operation of the spatial attention unit, and [] represents the concatenation operation.
[0080] In some embodiments, the frequency feature modulation module can be implemented by Figure 4 the following structure.
[0081] The frequency feature modulation module of this embodiment combines illumination information with image features in a coarse-to-fine manner. To avoid amplifying the noise in the high-frequency features while enhancing the contrast, the frequency feature modulation module performs shift modulation on the overall image features, but only performs scaling modulation on the low-frequency image features.
[0082] As Figure 4 shown, given the illumination feature F I ∈R C×H×W , first use two 1×1 convolutions to generate the affine transformation α and β∈R C×H×W . Then, the scaling parameter α is applied to the low-frequency image feature while the shift parameter β is applied to the overall image feature I L ∈R C×H×W .
[0083] That is:
[0084]
[0085] where S E represents the output feature, and FFM() represents the operation of the frequency feature modulation module.
[0086] The above-constructed image enhancement network needs to be trained before use to optimize the network parameters. In some embodiments, considering that the goal of low-light aerial image enhancement is to solve complex image degradation problems, including brightness enhancement, detail restoration, noise removal, and color correction, a joint loss function is designed in this embodiment to more effectively optimize the image enhancement network.
[0087] (1) Pixel loss: To obtain accurate reconstruction, pixel-level mean square error (MSE) is used to measure the similarity between the enhanced image I E and the normal-light ground truth image I N . The formula for pixel loss is:
[0088] L Pixel = L MSE (I E , I N )
[0089] (2) Content loss: To better measure the difference in brightness and contrast between the enhanced image and the normal-light ground truth image, negative multi-scale structural similarity (MS-SSIM) is used to minimize the distance between the enhanced image and the normal-light ground truth image. The formula for content loss is:
[0090] L Content =-L MS-SSIM (I E ,I N )
[0091] (3) Smoothness loss: To further refine texture details and suppress noise in the enhanced image, the L1 loss of the gradient between the enhanced image and the normal-light ground truth image is minimized. The formula for smoothness loss is:
[0092]
[0093] where is the gradient operator.
[0094] (4) Color loss: To prevent obvious color distortion in the enhanced image, the color loss between the enhanced image and the normal-light ground truth image is introduced. The formula for color loss is:
[0095] L Color =L CS (I E ,I N )
[0096] where L CS represents the cosine similarity between the enhanced image and the normal-light ground truth image.
[0097] Combining the above four loss functions, the joint loss function is expressed as:
[0098] L = λ1L Pixel +λ2L Content +λ3L Smooth +λ4L Color
[0099] where λ1, λ2, λ3, and λ4 are the trade-off parameters for each term in the joint loss function. For example, they are set to λ1 = 1.5, λ2 = 1, λ3 = 1, and λ4 = 0.5 respectively. The above loss functions fully consider different degradation factors and are convenient for improving the quality of low-light aerial images under uncertain lighting conditions.
[0100] To verify the enhancement effect and generalization performance of the image enhancement network for low-light aerial images, in this embodiment, a large number of comparative experiments were conducted not only on the LOL-v2 dataset but also on the DroneVehicle dataset. To quantitatively estimate the low-light image enhancement effect of different network models, peak signal-to-noise ratio (PSNR) and structural similarity (SSIM) were used to evaluate the enhanced images, which are metrics widely used for image quality assessment.
[0101] (1) Dataset
[0102] The LOL-v2 dataset is divided into two parts: real-world scenes and synthetic scenes to further enrich the training and evaluation processes. In the real-world scene dataset, there are 689 pairs of training images and 100 pairs of test images. These images cover various environmental conditions, including changes in natural and artificial lighting, scene complexity, and atmospheric effects. In the synthetic scene dataset, there are 900 pairs of training images and 100 pairs of test images. Using advanced rendering techniques, the synthetic scene dataset simulates a wide range of lighting conditions and introduces a controllable level of noise to replicate real-world imaging challenges. It should be noted that each pair of images consists of a low-light image and a corresponding normal-light ground truth image.
[0103] The DroneVehicle dataset carefully selects real low-light images of night scenes, including 4279 training images and 438 test images. These images contain five categories, namely Bus, Fright, Truck, Car, and Van.
[0104] It should be noted that not only the low-light and normal-light image pairs of the LOL-v2 dataset were used for low-light image enhancement experiments, but also the real low-light images of the DroneVehicle dataset were used for low-light object detection experiments.
[0105] (2) Run for 200 epochs using the Adam optimizer with default parameters β1 = 0.9, β2 = 0.999, and ε = 10 -8 The initial learning rate was set to 10 -4 in the first 100 epochs, and linearly decreased to 0 in the other 100 epochs. Except for the attention discrete wavelet transform module, attention inverse discrete wavelet transform module, and frequency feature modulation module, the kernel size of other convolutions was set to 3. The batch size was set to 32, and the input low-light images were randomly cropped to 256×256 pixels.
[0106] (3) A large number of comparative experiments were conducted between the image enhancement network and current popular low-light image enhancement methods. To ensure fairness, some methods were directly trained using the provided model parameters, while some other methods were retrained on the same dataset using the given code.
[0107] I. Quantitative Analysis
[0108] The image enhancement network of this embodiment was quantitatively compared with current popular low-light image enhancement methods. Tables 1 and 2 show the quantitative comparison of different methods on real captured images and synthetic images in the LOL-v2 dataset. Among these methods, the image enhancement network of this embodiment achieved the highest PSNR and SSIM on both real captured images and synthetic images in the LOL-v2 dataset, indicating that the network has the best low-light image enhancement effect.
[0109] Table 1 Quantitative Comparison Results on Real Captured Images in the LOL-v2 Dataset
[0110]
[0111] Table 2 Quantitative Comparison Results on Synthetic Images in the LOL-v2 Dataset
[0112]
[0113] II. Qualitative Analysis
[0114] On the LOL-v2 dataset, the image enhancement network of this embodiment was qualitatively compared with current popular representative low-light image enhancement methods for the enhanced images. The results show that the image enhancement network of this embodiment has the best effect in restoring the brightness of the image and suppressing redundant noise, fully demonstrating the superiority of the image enhancement network on real low-light images. In addition, compared with other methods, the image enhancement network of this embodiment not only has a visual advantage in enhancing brightness and contrast, but also has bright colors and clear edges in the enhanced image.
[0115] III. Efficiency Analysis
[0116] In Table 3, the running time, number of parameters, and number of floating-point operations of different methods on an image with a size of 600×400 pixels were evaluated. As shown in Table 3, the image enhancement network of this embodiment is faster and lighter than most methods and performs better in terms of efficiency.
[0117] Table 3 Efficiency Analysis of Running Time, Parameters, and Floating-Point Operations
[0118]
[0119] (4) Low-Light Object Detection Experiment
[0120] To evaluate the impact of low-light image enhancement on computer vision tasks, taking object detection as an example, a comparative study was conducted on real low-light images from the DroneVehicle dataset. The results show that the image enhancement network in this embodiment achieved a maximum object detection accuracy of 86.10% for low-light images. Moreover, the image enhancement network can not only avoid the problem of missed detection of vehicle targets, perfectly match the contour shape of vehicle targets, but also has the highest detection accuracy for all vehicle targets, fully demonstrating the superiority of the network in real drone low-light scenarios.
[0121] Table 4 mAP of different methods for real low-light images from the DroneVehicle dataset
[0122]
[0123]
[0124] (5) Ablation study
[0125] A series of ablation studies were conducted on real captured images of the LOL-v2 dataset to verify the effectiveness of the image enhancement network in this embodiment.
[0126] Note that the attention wavelet transform mechanism consists of an attention mechanism and a wavelet transform. To verify the effectiveness of the attention mechanism and the wavelet transform, they were removed respectively for ablation experiments. The results of the ablation study are shown in Table 5. The results show that without the attention mechanism or the wavelet transform, its performance will drop significantly. Moreover, the attention wavelet transform mechanism obtained the highest PSNR, proving that the attention mechanism and the wavelet transform are both important for low-light image enhancement and work interdependently.
[0127] Table 5 Ablation study of the attention wavelet transform mechanism
[0128]
[0129] The frequency feature modulation module was developed to integrate illumination information and image features. To verify the effectiveness of the illumination information, first, all frequency feature modulation modules in the image enhancement network were deleted. Then, the summation operation, concatenation operation, and spatial feature transformation were used to replace them respectively. From the ablation study results in Table 6, the performance of fusing features using the summation operation, concatenation operation, and spatial feature transformation is even lower than that of the network without the frequency feature modulation module, indicating that an inappropriate feature fusion method will lead to poor low-light image enhancement performance. However, the frequency feature modulation module that integrates illumination information and image features obtained the highest PSNR, proving the effectiveness of the frequency feature modulation module.
[0130] Table 6 Ablation study of the frequency feature modulation module
[0131]
[0132] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0133] Figure 5 It is a schematic diagram of a terminal 50 provided by an embodiment of the present invention. As Figure 5 shown, the terminal 50 of this embodiment includes: a processor 51, a memory 52, and a computer program 53 stored in the memory 52 and executable on the processor 51, such as a low-light aerial image enhancement program. When the processor 51 executes the computer program 53, the steps in the above-mentioned various embodiments of the low-light aerial image enhancement method are implemented.
[0134] Exemplarily, the computer program 53 can be divided into one or more modules / units. The one or more modules / units are stored in the memory 52 and executed by the processor 51 to complete the present invention. The one or more modules / units can be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program 53 in the terminal 50.
[0135] The terminal 50 can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The terminal 50 may include, but is not limited to, a processor 51 and a memory 52. Those skilled in the art can understand that Figure 5 this is only an example of the terminal 50 and does not constitute a limitation to the terminal 50. It may include more or fewer components than shown in the figure, or combine some components, or different components. For example, the terminal 50 may further include input / output devices, network access devices, a bus, etc.
[0136] The so-called processor 51 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0137] The memory 52 may be an internal storage unit of the terminal 50, such as a hard disk or memory of the terminal 50. The memory 52 may also be an external storage device of the terminal 50, such as a plug-in hard disk equipped on the terminal 50, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 52 may also include both an internal storage unit of the terminal 50 and an external storage device. The memory 52 is used to store the computer program and other programs and data required by the terminal 50. The memory 52 may also be used to temporarily store data that has been output or will be output.
[0138] Those skilled in the art can clearly understand that for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be described herein again.
[0139] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0140] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0141] In the embodiments provided by the present invention, it should be understood that the disclosed device / terminal and method can be implemented in other ways. For example, the device / terminal embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in electrical, mechanical or other forms.
[0142] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0143] In addition, in each embodiment of the present invention, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0144] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above method embodiments of the present invention, it can also be completed by a computer program instructing the relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device that can carry the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc.
[0145] The above-described embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A low-light aerial image enhancement method, characterized in that: include: Extracting contextual features from low-light aerial images; For the context features, multi-scale image features are learned by an encoder based on an attention discrete wavelet transform module, and multi-scale illumination features are estimated by a decoder based on an attention inverse discrete wavelet transform module, wherein the encoder and the decoder are connected via a jump connection; Integrating the multi-scale image features and the multi-scale illumination features through a frequency feature modulation module to obtain integrated features; A convolution operation is performed on the integrated features to obtain an enhanced aerial image.
2. The low-light aerial image enhancement method according to claim 1, characterized in that: Each scale corresponds to a frequency feature modulation module; The multi-scale image features and the multi-scale illumination features are integrated through a frequency feature modulation module to obtain integrated features, including: Through each of the frequency characteristic modulation modules, an affine transformation is performed on the illumination characteristics of the corresponding scale to obtain a scaling parameter and a shift parameter; The lowest frequency feature among the image features of the corresponding scale is modulated according to the scaling parameter, and the features of all frequencies among the image features of the corresponding scale are modulated according to the shift parameter to obtain the integrated feature; wherein the image feature includes features of multiple different frequencies.
3. The low-light aerial image enhancement method according to claim 1, characterized in that: The encoder comprises: an attention discrete wavelet transform module and a first residual block corresponding to each scale; The method learns multi-scale image features by an encoder based on an attention discrete wavelet transform module, comprising: By using each of the attention discrete wavelet transform modules, a separation operation is performed on the input features of the corresponding scale to obtain multiple features of different frequencies; The features of the multiple different frequencies are input into the first residual block to obtain image features of corresponding scales.
4. The low-light aerial image enhancement method according to claim 3, characterized in that: The attention discrete wavelet transform module includes: a discrete wavelet transform unit, a set of parallel first spatial attention units, a first channel attention unit and a first convolution unit; The separation operation is performed on the input features of the corresponding scale by each of the attention discrete wavelet transform modules to obtain multiple features of different frequencies, including: By means of the discrete wavelet transform unit, a discrete wavelet transform is performed on the input features of the corresponding scale to obtain a plurality of frequency sub-bands of different frequencies; Performing spatial level modulation on each frequency subband by the first spatial attention unit; Performing channel-level modulation on each frequency subband by the first channel attention unit; A convolution operation is performed by the first convolution unit to obtain features of multiple different frequencies.
5. The low-light aerial image enhancement method according to claim 1, characterized in that: The decoder includes: an attention inverse discrete wavelet transform module and a second residual block corresponding to each scale; The method estimates multi-scale illumination features through a decoder based on an attention inverse discrete wavelet transform module, comprising: Reconstructing the input features of the corresponding scale through each of the attention inverse discrete wavelet transform modules to obtain reconstructed features; The reconstructed features are input into the second residual block to obtain illumination features of corresponding scales.
6. The low-light aerial image enhancement method according to claim 5, characterized in that: The attention inverse discrete wavelet transform module includes: an inverse discrete wavelet transform unit, a set of parallel second spatial attention units, a second channel attention unit and a second convolution unit; The step of reconstructing the input features of the corresponding scale by each of the attention inverse discrete wavelet transform modules to obtain the reconstructed features includes: Performing spatial level modulation on the input features of the corresponding scale through the second spatial attention unit; Performing channel-level modulation on the input features of the corresponding scale through the second channel attention unit; The inverse discrete wavelet transform unit performs an inverse discrete wavelet transform on the input features of the corresponding scale, and the second convolution unit performs a convolution operation to obtain a reconstructed feature.
7. The low-light aerial image enhancement method according to claim 1, characterized in that: The extracting context features in the low-light aerial image includes: extracting the context features in the low-light aerial image by performing a 3×3 dilated convolution with a dilation rate of 2.
8. The low-light aerial image enhancement method according to any one of claims 1 to 7, characterized in that: The method further comprises: Constructing a joint loss function of an image enhancement network according to a weighted sum of pixel loss, content loss, smoothness loss and color loss; wherein the image enhancement network is the entire network that processes the low-light aerial image to obtain the enhanced aerial image; Based on the joint loss function, parameters of the image enhancement network are optimized.
9. A terminal comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.