Remote sensing and unmanned aerial vehicle image defogging method

Through the DWTMA-Net dehazing network, wavelet transform and multi-dimensional attention mechanism are used to solve the problem of feature information loss in remote sensing and UAV image dehazing, achieve better feature expression and dehazing effects, and improve image quality and the performance of downstream tasks.

CN120655535APending Publication Date: 2025-09-16ANHUI UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510822114.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Existing technologies in the remote sensing and UAV image dehazing process suffer from the loss of frequency domain feature information and the lack of frequency attention to enhance feature capture, resulting in insufficient feature expression and poor dehazing effect.

Method used

The dehazing network DWTMA-Net is adopted, which is based on a U-shaped architecture and includes discrete wavelet blocks DWB, multidimensional attention modules MAB and wavelet downsampling modules WDM. It enhances feature extraction and expression through Haar discrete wavelet transform and multidimensional attention mechanism, and captures global and local features by combining channel, pixel and Fourier frequency attention.

Benefits of technology

Effectively restore the frequency domain and spatial domain information of remote sensing images, improve feature representation, enhance dehazing performance, and improve image clarity and the reliability of downstream tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120655535A_ABST
    Figure CN120655535A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of computer vision and image processing, and particularly relates to a remote sensing and unmanned aerial vehicle image defogging method, a defogging network DWTMA-Net is adopted, and the DWTMA-Net is constructed based on a U-shaped architecture and comprises a discrete wavelet block DWB, a multi-dimensional attention module MAB and a wavelet down-sampling module WDM; downsampling of the encoder part is carried out by using Haar discrete wavelet transform through a WDM expansion downsampling method, and frequency information of wavelet transform is combined with spatial information of convolution downsampling; the DWB and the MAB are sequentially arranged between the encoder part and the decoder part, the DWB decomposes features into four frequency components by using Haar discrete wavelet transform (DWT), the low-frequency features are processed by a small AOD network to extract the features, and the high-frequency features are refined by using an expansion residual block; and then spatial information is reconstructed by applying inverse wavelet transform. According to the method, feature representation is improved, and the defogging performance of the network is effectively enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision and image processing, and in particular relates to a remote sensing and unmanned aerial vehicle (UAV) image defogging method. Background Art

[0002] The quality of remote sensing and drone imagery is often affected by complex atmospheric interference, including haze and translucent clouds. These adverse factors not only weaken signal quality but also distort visual information and obscure important details, thus limiting large-scale data collection and real-time monitoring. As a result, recovering information from remote sensing imagery becomes extremely challenging. High-quality, clear imagery can provide richer and more accurate visual information, which is crucial for improving the performance and reliability of various downstream tasks, such as temporal change analysis, disaster assessment, ecological evaluation, and defense-related observations.

[0003] Deep learning-based dehazing techniques have emerged as an effective alternative, demonstrating superior generalization and performance compared to methods guided by prior information. By using comprehensive datasets, such techniques can autonomously map hazy images to clear images without relying on predefined physical models. This allows them to excel even under variable atmospheric conditions, as they are able to effectively capture complex features and trends from the data. Although early deep learning-based dehazing methods were designed based on atmospheric scattering models, their practical application in real-world haze scenes remains challenging. This is primarily because physical scattering models cannot fully capture the complexity and diversity of real atmospheric conditions, limiting the effectiveness of these methods in complex environments.

[0004] To overcome the limitations of traditional learning-based dehazing techniques, recent methods have adopted end-to-end learning architectures, aiming to completely eliminate the reliance on physical modeling. Among these methods, multi-scale convolutional neural networks (CNNs) have attracted much attention due to their ability to directly learn the mapping from blurry images to clear images. The effectiveness of these methods is largely due to their ability to automatically extract rich and discriminative features through stacked convolutional layers. However, a fundamental limitation of the convolution operation is its inherent locality, which limits the model's ability to capture long-range dependencies and global contextual information. In terms of frequency domain feature extraction, the fast Fourier transform (FFT) and discrete cosine transform (DCT) techniques commonly used in existing technologies are different and may result in information loss. Existing technologies also lack the ability to effectively utilize frequency attention to enhance the capture of features in the image in terms of modeling, resulting in insufficient feature expression. These defects all affect the dehazing effect. Summary of the Invention

[0005] The purpose of the present invention is to provide a remote sensing and UAV image defogging method to solve the technical problems in the prior art of insufficient feature expression and poor defogging effect due to the loss of information on frequency domain features and the lack of enhancement effect of frequency attention on feature capture.

[0006] The remote sensing and UAV image defogging method adopts a defogging network DWTMA-Net, which is built based on a U-shaped architecture and includes a discrete wavelet block DWB, a multidimensional attention module MAB and a wavelet downsampling module WDM. The downsampling of the encoder part is performed by using the WDM extended downsampling method, and the frequency information of the wavelet transform is downsampled with the spatial information of the convolution downsampling. The DWB and MAB are sequentially arranged between the encoder part and the decoder part. The DWB uses the Haar discrete wavelet transform DWT to decompose the features into four frequency components, where the low-frequency features are processed by a small AOD network to extract features, and the high-frequency features are refined using a dilated residual block. Then, an inverse wavelet transform is applied to reconstruct the spatial information. The MAB uses depthwise separable convolution to extract deep features from the output of the DWB, and then uses convolutions of different kernel sizes to enhance feature diversity. Channel attention, pixel attention and Fourier frequency attention are integrated into the multidimensional attention mechanism to capture global and local features. Finally, the decoder part processes and outputs the result.

[0007] Preferably, the wavelet downsampling module WDM captures frequency information by combining discrete wavelet transform DWT; the input is first converted by Haar wavelet transform, HH and LL are a set of features, and HL and LH are a set of features; the two sets of features are connected and then 3×3 convolution is performed, and the result is added element by element with the result of 2×2 convolution of the input to obtain the output of WDM.

[0008] Preferably, the discrete wavelet block DWB first applies Haar discrete wavelet transform to obtain the frequency features of four sub-bands, namely HH, HL, LL, and LH; a small AOD network is used to extract features from the LL sub-band, while the other three sub-bands are refined using dilated residual blocks to enhance high-frequency features; the results obtained from the two processing methods are connected, and finally an inverse wavelet transform is applied to restore the features to the spatial domain to generate an output feature map.

[0009] Preferably, the physical model used in the small AOD network uses a learning-based method to estimate the deviation, uses global average pooling to compress the feature dimension, and calculates the average value of the feature map in its spatial dimension to obtain a one-dimensional feature vector consistent with the deviation value feature; then, the vector undergoes a 1×1 convolution for feature transformation, and then an S-type activation function is performed to obtain the deviation; a stacked convolution layer with a kernel size of 3×3 is used to promote feature learning.

[0010] Preferably, the multi-dimensional attention module MAB first applies deep convolution DWConv to the input features, then performs normalization, and then uses two convolutions with kernels of 1 and 3 to increase the number of channels; the convolution with kernel 1 mainly handles channel transformation and dimensionality adjustment, while the convolution with kernel 3 is used to capture local patterns; after extracting and normalizing the features, the convolved features are fused to provide a more comprehensive input representation; finally, the fused result is convolved with a kernel of 3 to further refine the features.

[0011] Preferably, the multi-dimensional attention module MAB applies a dual attention module and a frequency attention module. The dual attention module includes a pixel-level attention module and a channel attention module running in parallel. The outputs of the pixel-level attention module and the channel attention module are fused by element-by-element summation; frequency-aware modulation is performed through the frequency attention module, and then scaled by the Sigmoid activation function. The scaled result is fused with the result after fusion of the dual attention module by element-by-element multiplication, and then 1×1 convolution is used, followed by 3×3 convolution, and residual connection is combined to retain the original signal and generate the final output.

[0012] Preferably, the defogging method comprises the following steps:

[0013] Step 1. Establish a dehazing model;

[0014] Step 2. Establish a dehazing network DWTMA-Net that designs wavelet transform and multi-dimensional attention;

[0015] Step 3: Collect training data set;

[0016] Step 4: Train the dehazing network DWTMA-Net;

[0017] Step 5: Remote sensing image dehazing: The image to be dehazed is input into the dehazing network DWTMA-Net, and through forward propagation, a restored clear image is obtained at the output.

[0018] Preferably, the atmospheric scattering model is described as: I(x) = J(x)t(x) + A[1-t(x)], where I(x) is the image interfered with by haze; J(x) is the haze-free image to be restored; t(x) is the transmittance of light through the atmospheric medium; x represents the position of the pixel; and A is the global atmospheric light constant.

[0019] Preferably, the training adopts the L1 loss function:

[0020] L1=‖‖J-GT||1+α[||J real -GT real ||1+||J imag -GT imag ||1]

[0021] Among them, J represents the output remote sensing image after DWTMA-Net dehazing processing, GT represents its corresponding real image, and J real represents the real part of the generated image, GT real represents the real part of the real image, J imag Represents the imaginary part of the generated image, GT imag represents the imaginary part of the real image, and α represents the weight.

[0022] The advantages of the present invention are as follows: the dehazing network adopted by the present method can restore the frequency domain and spatial domain information in the remote sensing image; DWB obtains most of the signal energy from the LL subband to realize feature extraction, and the other three subbands are refined using dilated residual blocks to enhance and refine high-frequency features, and finally the inverse wavelet transform is applied to restore the features to the spatial domain; WDM uses Haar discrete wavelet transform for downsampling, and combines the frequency information of the wavelet transform with the spatial information of the convolution downsampling to improve feature representation; MAB can use convolutions of different kernel sizes to enhance feature diversity, and further apply channel attention, pixel attention and Fourier frequency attention, integrating them into a multi-dimensional attention mechanism to capture global and local features. Through the above improvements, the present invention overcomes the problems of information loss and insufficient feature details in frequency domain features, improves feature representation, and effectively enhances the dehazing performance of the network. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 This is a structural diagram of the defogging network DWTMA-Net in a remote sensing and UAV image defogging method in the present invention.

[0024] Figure 2 for Figure 1 Schematic diagram of the wavelet downsampling module WDM in the shown structure.

[0025] Figure 3 for Figure 1 Schematic diagram of the discrete wavelet block DWB in the shown structure.

[0026] Figure 4 for Figure 3 Schematic diagram of the small AOD network and dilated residual module in the shown architecture.

[0027] Figure 5 for Figure 1 Schematic diagram of the multi-dimensional attention module MAB in the shown structure.

[0028] Figure 6 It is a flow chart of the basic steps of the present invention.

[0029] Figure 7The visual comparison results of two images in the Haze1k-thin dataset before and after dehazing using the present invention and other existing technologies.

[0030] Figure 8 The visual comparison results of two images in the Haze1k-moderate dataset before and after applying the present invention and other existing technologies for dehazing.

[0031] Figure 9 The visual comparison results of two images in the Haze1k-thick dataset before and after dehazing using the present invention and other existing technologies.

[0032] Figure 10 The following are the visual comparison results of three images in the LHID dataset before and after applying the present invention and other existing technologies for dehazing.

[0033] Figure 11 The visual comparison results of three images in the DHID dataset before and after dehazing using the present invention and other existing technologies.

[0034] Figure 12 The following are the visual comparison results of four images in the HazyDet dataset before and after dehazing using the present invention and other existing technologies. DETAILED DESCRIPTION

[0035] The specific implementation methods of the present invention will be further explained in detail below through the description of embodiments with reference to the accompanying drawings, so as to help those skilled in the art to have a more complete, accurate and in-depth understanding of the inventive concept and technical solution of the present invention.

[0036] like Figure 1-Figure 5 As shown, the present invention provides a remote sensing and UAV image defogging method, which adopts a defogging network DWTMA-Net. The DWTMA-Net is constructed based on a U-shaped architecture and includes a discrete wavelet block DWB, a multidimensional attention module MAB and a wavelet downsampling module WDM; the downsampling of the encoder part is performed by using the WDM extended downsampling method, and the frequency information of the wavelet transform is downsampled using the Haar discrete wavelet transform, and the frequency information of the wavelet transform is combined with the spatial information of the convolution downsampling; the DWB and the MAB are sequentially arranged between the encoder part and the decoder part, and the DWB uses the Haar discrete wavelet transform DWT to decompose the features into four frequency components, where the low-frequency features are processed by a small AOD network to extract features, and the high-frequency features are refined using the dilated residual block; then the inverse wavelet transform is applied to reconstruct the spatial information; the MAB uses depthwise separable convolution to perform deep feature extraction on the output of the DWB, and then uses convolution of different kernel sizes to enhance feature diversity; the channel attention, pixel attention and Fourier frequency attention are integrated into the multidimensional attention mechanism to capture global and local features; finally, the decoder part processes and outputs the result.

[0037] 1. Wavelet Downsampling Module (WDM): The wavelet downsampling module extends traditional downsampling methods by incorporating discrete wavelet transforms (DWTs) to capture frequency information. This differs significantly from traditional methods that rely solely on convolution to reduce feature size. By integrating both spatial and frequency information, wavelet downsampling achieves an additive fusion of convolutional downsampling and wavelet transforms.

[0038] In WDM, the input is transformed through Haar wavelet transform, HH and LL form a set of features, and HL and LH form a set of features; the two sets of features are connected and then 3×3 convolved, and the result is added element-by-element with the result of 2×2 convolution of the input to obtain the output of WDM.

[0039] 2. Discrete Wavelet Block (DWB): Wavelet transform can reduce the spatial dimension by half with each transformation without sacrificing information, unlike other techniques such as Fast Fourier Transform (FFT) and Discrete Cosine Transform (DCT). Therefore, in this scheme, the frequency feature extraction block first extracts frequency domain features from the wavelet downsampling features, then performs spatial domain feature extraction, and finally generates the output feature map.

[0040] Specifically, the Haar discrete wavelet transform is first applied to obtain frequency features of four subbands: HH, HL, LL, and LH. The LL subband typically contains the majority of signal energy, while the other three subbands capture edge and detail information. A small AOD network is used to extract features from the LL subband, while the other three subbands are refined using dilated residual blocks to enhance high-frequency features. The results of these two processing methods are concatenated, and finally, an inverse wavelet transform is applied to restore the features to the spatial domain, generating the output feature map.

[0041] The physical model used by the small AOD network consists of a clear image, a foggy image, a bias, and a parameter K, which is estimated using a learning-based approach. In the small AOD network, global average pooling is used to compress feature dimensions and filter out duplicate or non-informative content from the representation space. Global average pooling calculates the average of the feature map across its spatial dimensions, resulting in a one-dimensional feature vector that is consistent with the bias value. This vector is then transformed by a 1×1 convolution, followed by a sigmoid activation function to obtain the bias. Because the parameter K is non-uniform, applying global average pooling results in information loss. Therefore, stacked convolutional layers with a kernel size of 3×3 are used to facilitate feature learning. The output of the stacked convolutional layer is multiplied element-wise with the input after a sigmoid activation function. The result is then element-wise subtracted from the output of the stacked convolutional layer. Finally, the bias obtained is added element-wise to produce the final result.

[0042] 3. Multi-dimensional Attention Module (MAB): In this module, deep convolution (DWConv) is first applied to the input features to ensure the preservation of feature details while improving computational efficiency. Normalization is then performed to ensure consistent feature scales, thereby accelerating network training. Next, to enhance feature diversity, two types of convolution are used, one with a kernel of 1 and one with a kernel of 3, to increase the number of channels. The kernel of 1 primarily handles channel transformation and dimensionality adjustment, while the kernel of 3 is used to capture local patterns. After extracting and normalizing the features, the convolved features are fused to provide a more comprehensive representation of the input. Finally, the fused result is subjected to a convolution with a kernel of 3 to further refine the features.

[0043] According to FFA-Net, the pixel-level attention (PA) mechanism is crucial for separating scale-related structures by enhancing important pixel information. The pixel-level attention mechanism aims to accurately locate and enhance key spatial regions in the image - this ability is particularly important in the dehazing task where maintaining local visibility is crucial. In addition to the pixel-level attention (PA) module, this scheme introduces a channel attention (CA) module running in parallel in MAB. The pixel-level attention (PA) module emphasizes local clarity, while the channel branch selectively enhances haze-related responses based on global contextual cues. This combined attention strategy enables the model to effectively combine detailed textures with high-level semantic features.

[0044] The outputs of the dual attention modules (PA and CA) are fused through element-wise summation, resulting in a richer and more informative feature representation. Inspired by strategies from previous research, this module also uses a frequency attention module (FAM) for frequency-aware modulation and scales with a sigmoid activation function to adaptively adjust feature strength and improve representational expressiveness. The scaled result is fused with the result of the dual attention module through element-wise multiplication. To further refine the features, the module uses 1×1 convolution to reduce dimensionality, compress key information, and reduce the risk of overfitting. 3×3 convolutions are then used to expand contextual understanding, combined with residual connections to preserve the original signal and produce the final output. The final output effectively combines improved enhancements with preserved input, ensuring robustness to downstream tasks.

[0045] Next, we will verify the solution through specific experiments:

[0046] In the experiment, the computer configuration used is: Intel (R) Xeon (R) CPU E5-2620 v4 processor, Nvidia GeForce GTX 1080Ti graphics processor, main frequency 2.10GHz, memory 12GB, operating system Ubuntu 18.04. The dehazing method is implemented based on the Pytorch framework. The specific steps are as follows Figure 6 As shown, details are as follows.

[0047] Step 1. Establish a dehazing model.

[0048] The atmospheric scattering model used in this method is described as: I(x) = J(x)t(x) + A[1-t(x)], where I(x) is the image interfered with by haze; J(x) is the haze-free image to be restored; t(x) is the transmittance of light through the atmospheric medium; x represents the position of the pixel; and A is the global atmospheric light constant.

[0049] After the above model conversion, let the parameter K(x) be:

[0050]

[0051] Then we have: J(x)=K(x)I(x)-K+b, which is the physical model used by the small AOD network, where b is the bias.

[0052] Step 2. Establish a dehazing network DWTMA-Net that designs wavelet transform and multi-dimensional attention.

[0053] This method uses a convolutional network based on wavelet transform and multidimensional attention to fit the mapping relationship F(x) between foggy images and clear images, that is, y=F(x); the mapping relationship between foggy images and clear images fitted by the convolutional network based on wavelet transform and multidimensional attention is expressed as F(x), where y represents the restored clear image, the function F represents the mapping relationship between the foggy image and the corresponding clear image, and x represents the foggy image.

[0054] The network input is a foggy image, and the output is the corresponding clear image. The entire network has an end-to-end structure. Each layer extracts multi-scale and multi-dimensional features, which are then processed by the atmospheric scattering model unit. For each layer, the number of blocks [N1, N2, N3, N4, N5] is set to [2, 2, 4, 2, 2], and the corresponding embedding channels are [24, 48, 96, 48, 24]. Multiple sub-maps in the dehazing model are learned separately. The final layer of the network maps the feature map back to the original image.

[0055] Step 3: Collect training dataset.

[0056] The performance of the proposed DWTMA-Net was evaluated on two synthetic remote sensing (RS) haze datasets (SateHaze1k and HRSD) and a real drone-based haze dataset (HazyDet). SateHaze1k is divided into three subsets corresponding to different haze densities: light haze, moderate haze, and heavy haze. Each subset contains 320 training samples and 45 test images. Light haze scenes are generated using haze masks extracted from real clouds, while moderate haze images combine features of light and moderate haze. Heavy haze uses a transmittance map to simulate dense atmospheric conditions. The HRSD dataset consists of two subsets: LHID and DHID. LHID contains 30,517 training images and 500 test images. These images are generated using an atmospheric scattering model to simulate different levels of haze, thereby improving the model's robustness to varying haze intensities. In contrast, DHID contains 14,990 images synthesized using real haze images, which can more realistically characterize haze characteristics. Of these, 14,490 images were used for training and 500 images for testing. These subsets contain both synthetic images and real haze features, providing a comprehensive platform for evaluating DWTMA-Net's dehazing capabilities. To demonstrate the generalization ability of our model in real-world environments, we evaluate its performance using the newly released HazyDet dataset, which contains haze-affected images captured by drones. This dataset consists of a training set of 8,000 images, a validation set of 1,000 images, and a test set of 2,000 images. It contains real hazy images captured under natural fog conditions, as well as artificial hazy images generated using an atmospheric scattering model (ASM). In addition, HazyDet comes with a dedicated Real Foggy Drone Detection Test Set (RDDTS) designed to evaluate the model's robustness in real-world scenarios. An example training sample from this dataset is provided.

[0057] Step 4: Train a multi-scale and multi-dimensional physical dehazing network.

[0058] Learning-based dehazing methods require labeled fog samples for training.

[0059] Learning the mapping relationship between images. In the field of dehazing, studies have shown that L1 loss is more conducive to dehazing. The L1 loss function is:

[0060] L1=||J-GT||1+α[||J real -GT real ||1+||J imag -GT imag ||1]

[0061] Among them, J represents the output remote sensing image after DWTMA-Net dehazing processing, GT represents its corresponding real image, and J real represents the real part of the generated image, GT reaI represents the real part of the real image, J imag Represents the imaginary part of the generated image, GT imag Represents the imaginary part of the real image, α represents the weight, which is 0.1. The present invention uses the PyTorch framework for training on a system equipped with four NVIDIA GeForce GTX 1080 TiGPUs. In order to enhance the training data, we apply random rotations of 90°, 180°, and 270° and horizontal flipping. The input image is RGB remote sensing data, resized to 240 pixels. We use the Adam optimizer (β1=0.9, β2=0.999) and a batch size of 12 to train each sub-dataset. The initial learning rate is set to 0.0002, and we use the cosine annealing strategy to gradually reduce it to 0. The present invention selects the stochastic gradient descent method to optimize the loss function, uses foggy images to iteratively learn the network, updates the network parameters, and ends the training when the network loss value tends to be stable. The network parameters saved at this time are the trained defogging network model.

[0062] Step 5: Dehazing remote sensing images.

[0063] The image to be dehazed is input into the network, and through the forward propagation of the network, a restored clear image can be obtained at the output.

[0064] We used the test set from the dataset to test the effectiveness. To demonstrate the competitiveness of our dehazing method, we compared it with advanced dehazing methods: DCP, AOD-Net, FCTF-Net, GridDehaze-Net, FFA-Net, MixDehaze-Net, OK-Net, and MMPD-Net.

[0065] Tables 1-3 show the performance metrics of our method and other existing methods for the SateHaze1k, HRSD, and HazyDet datasets, with the best and second-best methods marked in bold and underlined. Comparison shows that our method achieves the best dehazing effect and performance.

[0066] Table 1: Performance table of SateHaze1k data

[0067]

[0068] Table 2: Performance table of HRSD data

[0069]

[0070] Table 3: Performance table of HazyDet data

[0071]

[0072] After the experiment, the visual comparison of two images in different datasets is as follows: Figure 7-12 .

[0073] The present invention is described above by way of example in conjunction with the accompanying drawings. It is obvious that the specific implementation of the present invention is not limited to the above-mentioned method. As long as various non-substantial improvements are made using the inventive concept and technical solution of the present invention, or the inventive concept and technical solution are directly applied to other occasions without improvement, they are all within the scope of protection of the present invention.

Claims

1. A remote sensing and drone image defogging method, characterized by: A dehazing network DWTMA-Net is used. DWTMA-Net is built based on a U-shaped architecture and includes a discrete wavelet block DWB, a multidimensional attention module MAB, and a wavelet downsampling module WDM. The downsampling of the encoder part is performed through the WDM extended downsampling method, using Haar discrete wavelet transform for downsampling, combining the frequency information of the wavelet transform with the spatial information of the convolution downsampling. DWB and MAB are sequentially set between the encoder and decoder parts. DWB uses Haar discrete wavelet transform DWT to decompose the features into four frequency components, where low-frequency features are processed by a small AOD network to extract features, and high-frequency features are refined using dilated residual blocks. Then, the inverse wavelet transform is applied to reconstruct the spatial information. MAB uses depthwise separable convolution to perform deep feature extraction on the output of DWB, and then uses convolution of different kernel sizes to enhance feature diversity. Channel attention, pixel attention, and Fourier frequency attention are integrated into the multidimensional attention mechanism to capture global and local features. Finally, the decoder part processes and outputs the result.

2. The remote sensing and UAV image defogging method according to claim 1, characterized in that: The wavelet downsampling module (WDM) captures frequency information by combining discrete wavelet transform (DWT). The input is first converted using Haar wavelet transform, with HH and LL forming one set of features and HL and LH forming another set of features. The two sets of features are concatenated and then subjected to a 3×3 convolution. The result is then added element-wise to the result of the 2×2 convolution of the input to obtain the output of the WDM.

3. The remote sensing and UAV image defogging method according to claim 1, characterized in that: The discrete wavelet block DWB first applies Haar discrete wavelet transform to obtain the frequency features of four sub-bands, namely HH, HL, LL, and LH; a small AOD network is used to extract features from the LL sub-band, while the other three sub-bands are refined using dilated residual blocks to enhance high-frequency features; the results of the two processing methods are connected, and finally the inverse wavelet transform is applied to restore the features to the spatial domain to generate an output feature map.

4. The remote sensing and UAV image defogging method according to claim 3, characterized in that: The physical model used in the small AOD network uses a learning-based method to estimate the deviation and uses global average pooling to compress the feature dimension. Global average pooling calculates the average value of the feature map in its spatial dimension to obtain a one-dimensional feature vector consistent with the deviation value characteristics; then, the vector undergoes feature transformation through 1×1 convolution and then performs an S-type activation function to obtain the deviation; stacked convolution layers with a kernel size of 3×3 are used to promote feature learning.

5. The remote sensing and UAV image defogging method according to claim 1, characterized in that: The multi-dimensional attention module MAB first applies deep convolution DWConv to the input features, then performs normalization, and then uses two convolutions with kernels of 1 and 3 to increase the number of channels; the convolution with kernel 1 mainly handles channel transformation and dimensionality adjustment, while the convolution with kernel 3 is used to capture local patterns; after extracting and normalizing the features, the convolved features are fused to provide a more comprehensive input representation; finally, the fused result is subjected to a convolution with a kernel of 3 to further refine the features.

6. The remote sensing and UAV image defogging method according to claim 5, characterized in that: The multi-dimensional attention module MAB applies a dual attention module and a frequency attention module. The dual attention module includes a pixel-level attention module and a channel attention module running in parallel. The outputs of the pixel-level attention module and the channel attention module are fused by element-wise summation; frequency-aware modulation is performed through the frequency attention module, and then scaled by the Sigmoid activation function. The scaled result is fused with the result of the dual attention module fusion by element-wise multiplication, and then 1×1 convolution is used, followed by 3×3 convolution, and combined with residual connection to retain the original signal and produce the final output.

7. A remote sensing and UAV image defogging method according to any one of claims 1 to 6, characterized in that: The following steps are involved: Step 1. Establish a dehazing model; Step 2. Establish a dehazing network DWTMA-Net that designs wavelet transform and multi-dimensional attention; Step 3: Collect training data set; Step 4: Train the dehazing network DWTMA-Net; Step 5: Remote sensing image dehazing: The image to be dehazed is input into the dehazing network DWTMA-Net, and through forward propagation, a restored clear image is obtained at the output.

8. The remote sensing and UAV image defogging method according to claim 7, characterized in that: The atmospheric scattering model is described as: I(x) = J(x)t(x) + A[1-t(x)], where I(x) is the image affected by haze; J(x) is the haze-free image to be restored; t(x) is the transmittance of light through the atmospheric medium; x represents the position of the pixel; and A is the global atmospheric light constant.

9. The remote sensing and UAV image defogging method according to claim 7, characterized in that: The training uses the L1 loss function: L1=||J-GT||1+α[||J real -GT real ||1+||J imag -GT imag ||1] Among them, J represents the output remote sensing image after DWTMA-Net dehazing processing, GT represents its corresponding real image, and J real represents the real part of the generated image, GT real represents the real part of the real image, J imag Represents the imaginary part of the generated image, GT imag represents the imaginary part of the real image, and α represents the weight.

Citation Information

Cited By

  • Traffic image defogging method based on wavelet convolution and semantic-content guide fusion

    CN121353135A

  • Slender crack semantic segmentation method and system based on hybrid architecture support

    CN121788844A