Method for constructing remote sensing image defogging network based on wavelet frequency domain heterogeneous enhancement

By constructing a remote sensing image defog network with heterogeneity enhanced wavelet frequency domain, separating and enhancing high and low frequency characteristics, the problem of insufficient processing of high-frequency features in the existing methods is solved, and the high-quality remote sensing image defog effect is achieved.

CN120355601APending Publication Date: 2025-07-22CHINA THREE GORGES UNIV
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510471601.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing remote sensing image defog removal methods fail to fully distinguish the physical characteristics of high and low frequency characteristics when processing high-frequency characteristics, resulting in blurred image details or incomplete defog removal after defog removal.

Method used

A remote sensing image defogging network based on wavelet frequency domain heterogeneous enhancement is constructed. Through U-shaped architecture, multi-scale deformable convolution and fast Fourier transform, high and low frequency features are separated and enhanced, combining jump connections and multi-loss constraint training to achieve efficient defogging.

Benefits of technology

Improve the effects of high-frequency detail recovery and low-frequency haze removal, achieve high-quality remote sensing image defog removal results, significantly improving image clarity and texture retention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355601A_ABST
    Figure CN120355601A_ABST
Patent Text Reader

Abstract

The invention discloses a method for constructing a remote sensing image defogging network based on wavelet frequency domain heterogeneous enhancement, the network adopts a U-shaped architecture as a basic framework, the network input is a foggy image, and firstly, shallow layer features are extracted through a convolution block; then, a symmetric codec structure is adopted to learn layered representation, a codec comprises up and down sampling and a wavelet frequency domain heterogeneous enhancement module, and the resolution of up and down sampling is controlled through step convolution and deconvolution; the wavelet frequency domain heterogeneous enhancement module separates the high and low frequency features of the image through discrete wavelet transform, and performs heterogeneous enhancement on the separated high and low frequency features by combining the dynamic receptive field advantage of deformable convolution and the global perception capability of Fourier transform; therefore, the recovery of high-frequency local texture details and the removal of low-frequency global haze are effectively promoted. And finally, reconstructing a clear fogless image through the convolution block. According to the research algorithm, the texture features and natural colors of the scene can be precisely reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image processing, and particularly relates to a method for constructing a remote sensing image dehazing network based on wavelet frequency domain heterogeneous enhancement. Background Art

[0002] With the continuous progress of remote sensing technology, remote sensing images are increasingly deeply applied in the fields of ecological monitoring, climate analysis, topographic mapping, marine environment monitoring, disaster assessment, and deep space exploration. However, during the process of image acquisition by satellites and unmanned aerial vehicles, suspended particles such as aerosols, clouds, and dust in the atmosphere will have a significant impact on the imaging effect, resulting in problems such as image blurring, uneven brightness, color deviation, and texture information loss. These interferences not only weaken the intuitive expressiveness of the image, but also seriously restrict the subsequent image analysis and application effects, hindering the effective utilization of remote sensing data in actual scenarios. Therefore, how to effectively eliminate the atmospheric interference in remote sensing images, improve the clarity, brightness uniformity, and color restoration degree of the images, while maintaining rich texture features, has become a key research topic in the field of remote sensing image processing.

[0003] Currently, remote sensing image dehazing techniques are mainly divided into two categories: dehazing methods based on traditional prior knowledge and dehazing methods based on deep learning. Early studies focused more on using artificially designed prior conditions, such as Non-Local Prior and Dark Channel Prior. These methods usually combine with the Atmospheric Scattering Model (ASM) to achieve the dehazing effect. With the rise of deep learning technology, data-driven dehazing methods have gradually shown their unique advantages in remote sensing image analysis. These methods can be further divided into two types: one is the physically driven method based on parameter estimation, and the other is the end-to-end learning method. The former mainly uses deep learning models to infer the atmospheric light and transmission rate map, and then combines these inferred results with the atmospheric scattering model to restore a clear image, such as models like AOD-Net and DenseNet. However, although these prior-based methods perform well under specific conditions, their dehazing effects are often not ideal when facing complex scenarios such as uneven haze distribution or rapidly changing lighting conditions. To overcome these limitations, researchers have begun to focus on developing end-to-end dehazing algorithms that do not rely on the atmospheric scattering model, and among them, spatial domain feature extraction methods have received extensive attention and research. The paper "Dense haze removal based on dynamic collaborative inference learning for remote sensing images" published by Zhang et al. proposed a new dynamic collaborative inference learning framework and a twin network structure with shared weights to effectively restore real surface information from dense foggy remote sensing images. The paper "Depth Information Assisted Collaborative Mutual Promotion Network for Single Image Dehazing" published by Zhang et al. proposed a dual-task collaborative mutual promotion framework that combines a depth estimation network and a dehazing network to achieve mutual enhancement of their performances. However, due to the fact that spatial domain extraction algorithms ignore the global structure of the image, in recent years, methods based on frequency domain extraction have been widely proposed. The paper "Frequency and spatial dual guidance for image dehazing" published by Yu et al. proposed a spatial-frequency domain dual guidance network, and through the Fast Fourier Transform, it was explored that the degradation characteristics of haze are mainly contained in the amplitude spectrum. And slight phase changes are compensated through haze residuals."WSAMF-Net: Wavelet Spatial Attention-Based MultiStream Feedback Network for SingleImage Dehazing" published by Song et al. proposed wavelet spatial attention, which utilizes information in the frequency domain and spatial domain to enhance the extracted features, thereby obtaining better structures and edges. However, existing methods do not fully consider the different characteristics of high-frequency and low-frequency features when processing high-frequency features. Only a unified operation strategy or simple design is adopted to process high-frequency and low-frequency information, which may lead to insufficient frequency-domain feature learning ability, thus causing problems such as blurred image details or incomplete haze removal.

[0004] Therefore, a method for constructing a remote sensing image dehazing network based on wavelet frequency-domain heterogeneous enhancement is needed to solve the above problems. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a method for constructing a remote sensing image dehazing network based on wavelet frequency-domain heterogeneous enhancement, aiming to solve the key problem that existing frequency-domain methods fail to fully distinguish the physical characteristics of high-frequency and low-frequency features when processing high-frequency features. In the prior art, since high-frequency features usually exhibit complex deformations and significant edge changes, while low-frequency features mainly reflect the overall haze distribution and large-scale structure of the image, adopting a unified operation strategy or simple design often leads to insufficient feature extraction ability, thereby causing problems such as blurred image details or incomplete haze removal after dehazing. This method effectively removes haze interference in remote sensing images by fully considering the differences between high-frequency and low-frequency features, while effectively retaining the detail information of the image.

[0006] To solve the above technical problems, the technical solution adopted by the present invention is:

[0007] A method for constructing a remote sensing image dehazing network based on wavelet frequency-domain heterogeneous enhancement, comprising the following steps:

[0008] S1, constructing a U-shaped remote sensing image dehazing network, including an encoding layer, a decoding layer, skip connections, and a wavelet frequency-domain heterogeneous enhancement module;

[0009] S2, constructing multiple upsampling and downsampling modules and a skip connection mechanism, using 2D convolution for downsampling, transposed convolution for upsampling, and pixel-by-pixel addition for skip connections;

[0010] S3, constructing multiple wavelet frequency-domain heterogeneous enhancement modules, including wavelet transform, multi-scale deformable convolution, and fast Fourier transform, for heterogeneous enhancement of high-frequency and low-frequency features;

[0011] S4. Feed the hazy remote sensing image into the U-shaped image dehazing network. Through multiple serial wavelet frequency domain heterogeneous enhancement modules and upsampling and downsampling modules, finally output a clear haze-free image;

[0012] S5. Calculate the loss through the output clear image to constrain the training of the network.

[0013] Preferably, the U-shaped remote sensing image dehazing network constructed in step S1 includes:

[0014] Input the hazy image Hazy, go through shallow layer extraction, to the first layer of the encoding layer, to the second layer of the encoding layer, to the third layer of the encoding layer, to the first layer of the decoding layer, to the second layer of the decoding layer, to the third layer of the decoding layer, and output the dehazing image Dehazing at the reconstructed image.

[0015] Preferably, in step S2, construct multiple upsampling and downsampling modules and skip connection mechanisms as follows:

[0016] The specific operation of downsampling (taking the first layer of the encoding layer as an example): For the encoding layer features (H, W, C), go through 3*3 Convolution, to the features after downsampling (H / 2, W / 2, 2C), where H, W, and C are the height, width, and number of channels of the image respectively;

[0017] The specific operation of upsampling (taking the first layer of the decoding layer as an example): For the decoding layer features (H / 4, W / 4, 4C), go through 3*3 DeConvolution, to the features after upsampling (H / 2, W / 2, 2C). Where H, W, and C are the height, width, and number of channels of the image respectively;

[0018] The steps of the skip connection mechanism are as follows:

[0019] Perform pixel-by-pixel addition on the output features of the third layer of the encoding layer ((H / 4, W / 4, 4C)) and the first layer of the decoding layer ((H / 4, W / 4, 4C));

[0020] Perform pixel-by-pixel addition on the output features of the second layer of the encoding layer ((H / 2, W / 2, 2C)) and the second layer of the decoding layer (H / 2, W / 2, 2C);

[0021] Perform pixel-by-pixel addition on the output features of the first layer of the encoding layer (H, W, C) and the third layer of the decoding layer (H, W, C).

[0022] Preferably, in step S3, construct the serial wavelet frequency domain heterogeneous enhancement module WFHE as follows:

[0023] S301. Separate the frequency information of the image by performing discrete wavelet transform on the features L of the encoding and decoding to obtain three high-frequency sub-bands L HL , L LH , LHH and a low-frequency sub-band L LL ;

[0024] S302, splice and fuse the three high-frequency sub-bands, and then extract multi-scale irregular textures and high-frequency features with directional changes through deformable convolution;

[0025] Use spatial attention and channel attention mechanisms to select the most representative high-frequency features for enhancement, and weight the high-frequency features at different scales by generating a pixel-level weight matrix, thereby strengthening the expression of key-scale features;

[0026] S303, perform a fast Fourier transform on the low-frequency sub-band to separate the amplitude and phase; the haze information of the haze image is mainly contained in the Fourier amplitude spectrum, while the Fourier phase spectrum carries more structural information;

[0027] Restore the amplitude information, and use the amplitude residual to generate a haze distribution degradation map, thereby guiding the reconstruction of the structural information;

[0028] S304, obtain the final output feature L through the inverse wavelet transform of the heterogeneous enhanced high- and low-frequency features out .

[0029] Preferably, in step S4, the U-shaped network is constructed as follows:

[0030] Input the hazy image Hazy, to shallow feature extraction, to the encoder, to the decoder, to the reconstructed haze-removed image, to the output haze-removed image, to the loss constraint.

[0031] Preferably, each encoding and decoding layer is embedded with upsampling and downsampling and wavelet frequency domain heterogeneous enhancement modules, and information flow is ensured through skip connections; trained using four loss constraint networks, including L1 loss, adversarial loss, perceptual loss, and multi-scale structural similarity loss.

[0032] Preferably, the specific formula of the L1 loss is:

[0033]

[0034] where F represents the haze-removed image output by the network, x i and y i represent the values of the haze-removed image and the clear image at pixel p, respectively; K represents the number of pixels in the image.

[0035] Preferably, the adversarial loss constructs a loss function using the adversarial training mechanism of the generative adversarial network GAN; the generator G is responsible for converting the hazy image into a clear image, and the discriminator D distinguishes the distribution difference between the generated image and the real haze-free image through adversarial learning, and its adversarial loss function can be expressed as:

[0036]

[0037] Among them, X represents the defogged image, and N represents the number of image pixel points.

[0038] Preferably, the perceptual loss measures the difference in the deep feature space, effectively preserving the texture details and structural information of the image; using the pre-trained VGG-16 network to extract high-level semantic features, the perceptual loss is defined as follows:

[0039]

[0040] Among them, O i represents the size of the feature map of the i-th layer of the VGG16 pre-trained model; K represents the number of feature layers of the VGG16 pre-trained model used in the perceptual loss; where x and y represent the foggy image and the clear image respectively.

[0041] Preferably, the multi-scale structural similarity loss is used to constrain the network to make the structural similarity between the defogged image and the clear image closer; the specific expression is as follows: the specific formula is:

[0042]

[0043] Among them, μ i , μ j represent the means of the image after defogging and the clear image respectively, σ i , σ j represent the standard deviations of the image after defogging and the clear image respectively, σ ij is used to represent the covariance of the image after defogging and the clear image, k m , l m are two important terms in the equation; C1 and C2 are constant terms;

[0044] The overall network loss function is expressed as:

[0045] L total = λ1L adv + λ2L content + λ3L clear + λ4L BGCC + λ5L BLCC ;

[0046] Among them, λ1, λ2, λ3, λ4, and λ5 are hyperparameters of each function.

[0047] The beneficial effects of the present invention are as follows:

[0048] 1. The present invention proposes a remote sensing image dehazing network based on wavelet frequency domain heterogeneous enhancement. By deeply exploring the physical characteristics of high and low frequency features and performing heterogeneous enhancement on the high and low frequency information of the image, the restoration of high frequency details and the restoration of low frequency global haze are improved. High-quality remote sensing image dehazing results are achieved, and it is proven that the algorithm proposed by us achieves the best performance on two public remote sensing image datasets, StateHaze1k and RSID.

[0049] 2. The wavelet frequency domain heterogeneous enhancement module proposed by the present invention performs multi-scale extraction of irregular high frequency features through deformable convolution, and generates a pixel-level weight matrix according to the contribution degree of each layer of the network to weight different scales, thereby highlighting the expression of important scales. In terms of low frequency information processing, the fast Fourier transform is used to separate the amplitude and phase. Among them, the amplitude represents the haze information, and the phase represents the structural information. By restoring the amplitude information and using the amplitude residual to guide the reconstruction of the phase information, the restoration effect of the image structural information is effectively improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 is the overall network structure diagram of the embodiment of the present invention;

[0051] Figure 2 is Figure 1 the structural schematic diagram of the wavelet frequency domain heterogeneous enhancement module (WFHE) in DETAILED DESCRIPTION OF THE EMBODIMENTS

[0052] Example 1:

[0053] As Figure 1 shown, a construction method of a remote sensing image dehazing network based on wavelet frequency domain heterogeneous enhancement includes the following steps:

[0054] S1. Construct a U-shaped remote sensing image dehazing network, including an encoding layer, a decoding layer, skip connections, and a wavelet frequency domain heterogeneous enhancement module;

[0055] S2. Construct multiple upsampling and downsampling modules and a skip connection mechanism. Downsampling uses 2D convolution, upsampling uses transposed convolution, and skip connections use pixel-by-pixel addition;

[0056] S3. Construct multiple wavelet frequency domain heterogeneous enhancement modules, including wavelet transform, multi-scale deformable convolution, and fast Fourier transform, for heterogeneous enhancement of high and low frequency features;

[0057] S4. Send the hazy remote sensing image into the U-shaped image dehazing network, and finally output a clear haze-free image through multiple serial wavelet frequency domain heterogeneous enhancement modules and upsampling and downsampling modules;

[0058] S5. Calculate the loss through the output clear image to constrain the training of the network.

[0059] Preferably, the U-shaped remote sensing image dehazing network constructed in step S1 includes:

[0060] Input the hazy image Hazy, go through shallow extraction, reach the first layer of the encoding layer, the second layer of the encoding layer, the third layer of the encoding layer, the first layer of the decoding layer, the second layer of the decoding layer, the third layer of the decoding layer, and output the dehazed image Dehazing from the reconstructed image.

[0061] Preferably, in step S2, construct multiple upsampling and downsampling modules and skip connection mechanisms as follows:

[0062] The specific operation of downsampling (taking the first layer of the encoding layer as an example): For the features of the encoding layer (H, W, C), go through 3*3 Convolution, and reach the features after downsampling (H / 2, W / 2, 2C), where H, W, and C are the height, width, and number of channels of the image respectively;

[0063] The specific operation of upsampling (taking the first layer of the decoding layer as an example): For the features of the decoding layer (H / 4, W / 4, 4C), go through 3*3 DeConvolution, and reach the features after upsampling (H / 2, W / 2, 2C). Where H, W, and C are the height, width, and number of channels of the image respectively;

[0064] The steps of the skip connection mechanism are as follows:

[0065] The output features of the third layer of the encoding layer ((H / 4, W / 4, 4C)) and the first layer of the decoding layer ((H / 4, W / 4, 4C)) are added pixel by pixel;

[0066] The output features of the second layer of the encoding layer ((H / 2, W / 2, 2C)) and the second layer of the decoding layer (H / 2, W / 2, 2C) are added pixel by pixel;

[0067] The output features of the first layer of the encoding layer (H, W, C) and the third layer of the decoding layer (H, W, C) are added pixel by pixel.

[0068] Preferably, in step S3, construct a serial wavelet frequency domain heterogeneous enhancement module WFHE as follows:

[0069] S301, Separate the frequency information of the image by performing discrete wavelet transform on the features L of the encoding and decoding, and obtain three high-frequency subbands L HL , L LH , L HH and a low-frequency subband L LL ;

[0070] S302, splice and fuse the three high-frequency subbands, and then extract high-frequency features with multi-scale irregular textures and directional changes through deformable convolution;

[0071] The most representative high-frequency features are selected for enhancement by using spatial attention and channel attention mechanisms, and different-scale high-frequency features are weighted by generating a pixel-level weight matrix, thereby strengthening the expression of key-scale features;

[0072] S303. Perform a fast Fourier transform on the low-frequency subband to separate the amplitude and phase; the haze information of the haze image is mainly contained in the Fourier amplitude spectrum, while the Fourier phase spectrum carries more structural information;

[0073] By restoring the amplitude information and using the amplitude residual to generate a haze distribution degradation map, thereby guiding the reconstruction of the structural information;

[0074] S304. Obtain the final output feature L through the inverse wavelet transform of the heterogeneous-enhanced high- and low-frequency features out 。

[0075] Preferably, in step S4, the U-shaped network is constructed as follows:

[0076] Input the hazy image Hazy, go to shallow feature extraction, go to the encoder, go to the decoder, go to the reconstructed haze-removed image, go to the output haze-removed image, go to loss constraint.

[0077] Preferably, each encoding and decoding layer is embedded with upsampling and downsampling and wavelet frequency domain heterogeneous enhancement modules, and the information flow is ensured through skip connections; four loss constraint networks are used for training, including L1 loss, adversarial loss, perceptual loss, and multi-scale structural similarity loss.

[0078] Preferably, the specific formula of the L1 loss is:

[0079]

[0080] where F represents the haze-removed image output by the network, x i and y i represent the values of the haze-removed image and the clear image at pixel p, respectively; K represents the number of pixels in the image.

[0081] Preferably, the adversarial loss constructs a loss function by adopting the adversarial training mechanism of the generative adversarial network GAN; the generator G is responsible for converting the hazy image into a clear image, and the discriminator D distinguishes the distribution difference between the generated image and the real haze-free image through adversarial learning, and its adversarial loss function can be expressed as:

[0082]

[0083] where X represents the haze-removed image and N represents the number of image pixel points.

[0084] Preferably, the perceptual loss effectively preserves the texture details and structural information of the image by measuring the difference in the deep feature space; the high-level semantic features are extracted using the pre-trained VGG-16 network, and the perceptual loss is defined as follows:

[0085]

[0086] where O i represents the size of the feature map of the i-th layer of the VGG16 pre-trained model; K represents the number of feature layers of the VGG16 pre-trained model used in the perceptual loss; where x and y represent the hazy image and the clear image respectively.

[0087] Preferably, the multi-scale structural similarity loss is used to constrain the network to make the structural similarity between the dehazed image and the clear image closer; the specific expression is as follows: the specific formula is:

[0088]

[0089] where μ i , μ j represent the means of the dehazed image and the clear image respectively, σ i , σ j represent the standard deviations of the dehazed image and the clear image respectively, σ ij is used to represent the covariance between the dehazed image and the clear image, k m , l m are two important terms in the equation; C1 and C2 are constant terms;

[0090] The overall network loss function is expressed as:

[0091] L total = λ1L adv + λ2L content + λ3L clear + λ4L BGCC + λ5L BLCC ;

[0092] where λ1, λ2, λ3, λ4, and λ5 are the hyperparameters of each function.

[0093] Embodiment 2:

[0094] The specific method provided in this embodiment includes the following steps:

[0095] Step S1 specifically includes:

[0096] This network uses a three-level processing flow to achieve end-to-end remote sensing image dehazing (as Figure 1As shown in the figure. First, the network extracts shallow features from the input hazy image through the convolutional layer. Subsequently, the feature stream enters a deep processing architecture composed of six groups of symmetric encoder-decoder units: each encoder-decoder layer integrates an innovative wavelet frequency-domain heterogeneous enhancement module (WFHE) to implement a heterogeneous enhancement strategy for high- and low-frequency features. Among them, the skip connection mechanism runs through the encoder-decoder process to ensure the efficient transmission of feature information. Finally, the network reconstructs a clear haze-free image through the convolutional layer.

[0097] Step S2 specifically includes:

[0098] Downsampling: Taking the first layer of the encoding layer as an example, downsampling can be expressed by the following formula:

[0099] E2 = Convolution 3×3 (E1);

[0100] where E1 ∈ (H, W, C), H, W, and C represent the height, width, and number of channels of the feature. After the feature of the encoding layer passes through the 2D convolution, the resolution is reduced to half of the original, and at the same time, the number of channels is expanded to twice the original.

[0101] Upsampling: Taking the first layer D1 of the decoding layer as an example, upsampling can be expressed by the following formula:

[0102] D2 = DeConvolution 3×3 (D1);

[0103] where H, W, and C represent the height, width, and number of channels of the feature. After the feature of the decoding layer passes through the transposed convolution, the resolution is expanded to twice the original, and at the same time, the number of channels is reduced to half of the original.

[0104] Skip connection mechanism: Taking the third layer of the encoding layer and the first layer of the decoding layer as an example, the skip connection can be expressed by the following formula:

[0105] skip1 = E3 + D1;

[0106] where H, W, and C represent the height, width, and number of channels of the feature. By constructing a dense cross-layer connection architecture, hierarchical aggregation and gradient propagation optimization of multi-level features in the encoding stage are achieved. This design not only alleviates the feature degradation problem in deep networks but also ensures the bidirectional transmission of high-dimensional semantic information and low-level texture features.

[0107] Step S3 specifically includes:

[0108] Step 1) Construct a wavelet frequency-domain heterogeneous enhancement module. First, for the input feature L, the low-frequency and high-frequency sub-bands of the image are separated by the Haar discrete wavelet. The Haar wavelet consists of a low-pass filter L and a high-pass filter H, as shown below:

[0109]

[0110] In this embodiment, four sub-bands can be obtained, which can be expressed as:

[0111] L LL ,{L LH ,L HL ,L HH}=DWT(L i );

[0112] Where respectively represent the input low-frequency component and the high-frequency components in the vertical, horizontal, and diagonal directions.

[0113] L H =Concat(L LL ,L HL ,L HH );

[0114] Step 2) Concatenate and fuse the three high-frequency sub-bands, and use deformable convolution to extract high-frequency features at different scales. The formula is expressed as follows:

[0115]

[0116] Among them, the sizes of Kernel1, Kernel2, and Kernel3 are 3, 5, and 7 respectively. Then, in this embodiment, a concatenation operation is used to fuse these features, and an attention mechanism is used to adaptively enhance important features in both the spatial dimension and the channel dimension, realizing cross-scale interaction and generating three corresponding pixel-level weight matrices:

[0117] E=expand(PA(Concat(L1,L2,L3)));

[0118] F=expand(CA(Concat(L1,L2,L3)));

[0119] W1,W2,W3=spilt(E+F);

[0120] Among them, PA represents the pixel attention mechanism, CA represents the channel attention mechanism, E represents the channel attention weight, F represents the spatial attention weight, spilt represents the channel splitting operation, and expand represents the broadcasting operation. W1, W2, and W3 respectively represent the pixel-level weight matrices of different scales, and the network adaptively generates corresponding weight values according to the contribution degrees of features of different scales. The features of different scales are weighted by these weight matrices to highlight the detailed information of the key scale, and the formula is expressed as follows:

[0121]

[0122] Finally, these features are fused to obtain the enhanced high-frequency information:

[0123]

[0124] Step 3) For the processing of low-frequency information, in this embodiment, the fast Fourier transform is performed on the low-frequency features:

[0125]

[0126] where h and w are the coordinates in the space, and u and v are the coordinates in the Fourier space. The complex components in the Fourier space can be represented by the amplitude spectrum and the amplitude spectrum:

[0127]

[0128] where R(x) and I(x) respectively represent the real part and the imaginary part of LLL, and A(X(u, v)) and P(X(u, v)) respectively correspond to the amplitude spectrum and the phase spectrum of the frequency-domain representation. Subsequently, in this embodiment, two layers of 1×1 convolution and the ReLU activation function are used to recover the haze information. Then, the recovered haze information is subtracted from the original haze information to obtain the haze residual, and the formula is expressed as follows:

[0129] A(X(u, v))' = Conv 1×1 (ReLU(Conv 1×1 (A(X(u, v)))));

[0130] A res (u, v) = A(X(u, v))' - A(X(u, v));

[0131] where A res (u, v) represents the haze residual, and then two haze degradation distribution maps are generated using the haze residual:

[0132] W = (Sigmoid(MLP(GAP(A(X(u, v))'))));

[0133] M = (Sigmoid(MLP(GAP(A(X(u, v))'))));

[0134] W and M are the haze weight distribution maps. W reflects the degradation degree of haze in different channels, and M represents the degradation degree of haze in different regions. Finally, these two haze degradation maps are used to reconstruct the phase information:

[0135]

[0136] Then, A(X(u, v))′ and P(u, v)′ are mapped back to the image space through the inverse Fourier transform to obtain the final output result Finally, the differentiated high-frequency and low-frequency information is obtained through the inverse wavelet transform to get the result enhanced in the frequency domain:

[0137]

[0138] Through the decoupled learning of the frequency domain, the network adopts a differentiated design strategy for the different physical characteristics of high-frequency and low-frequency information, thereby effectively improving the restoration effect of high-frequency texture details and the ability to remove low-frequency global haze.

[0139] Example Three:

[0140] This example discloses the test process as follows:

[0141] The experimental evaluation of this study was carried out on two publicly available remote sensing dehazing datasets, SateHaze1k and RSID. Among them, the SateHaze1k dataset contains three subsets: SateHaze1k Thin, SateHaze1k Moderate, and SateHaze1k Thick, and each subset contains 400 pairs of synthetic haze images (320 pairs in the training set, 35 pairs in the validation set, and 45 pairs in the test set). The RSID dataset contains 1000 pairs of synthetic remote sensing image pairs with a resolution of 256×256 pixels. In this study, it was divided into a training set (900 pairs) and a test set (100 pairs) at a ratio of 9:1 for validation. In addition, the present invention uses the Adam optimizer to optimize the proposed network, with the momentum decay exponents β1 = 0.9, β2 = 0.999, and the learning rate and batch size are set to 0.0001 and 4 respectively. The initial learning rate is set to 0.001, and the MultiStepLR is used to dynamically adjust the learning rate between them. During the training process, the parameters of the loss function in the network model of this example are set to ω1 = 1, ω2 = 0.0005, ω3 = 0.01, and ω4 = 0.5. This example compares the dehazing network of this example with the other 7 recent excellent algorithms. The main algorithms for comparison are DCP, FFA, FSDGN, Restormer, DEA-Net, MIMO, and SFAN.

[0142] 2. Experimental Results

[0143] The invention of this embodiment was quantitatively evaluated with other seven algorithms on the SateHaze1k and RISID datasets. Table 1 shows the average PSNR and SSIM values of the test methods. As shown in Table 1, the PSNR of DCP on the SateHaze1k dataset is only 11.37dB, while the PSNR of the remaining algorithms is higher than 20dB, proving that the end-to-end defogging algorithm is superior to the traditional parameter estimation algorithm. In this embodiment, it is observed that on the SateHaze1k dataset, whether in the case of light fog, medium fog or thick fog, the method of this embodiment is significantly superior to other existing algorithms. Specifically, on the SateHaze 1k Thin, SateHaze 1k Moderate and SateHaze 1k Thick datasets, compared with the second-ranked method, the algorithm of this embodiment improves by 1.11dB, 1.89dB and 0.69dB respectively in PSNR, and improves by 0.003, 0.008, 0.010 respectively in SSIM. In addition, in the RSID dataset, the method proposed in this paper has better performance compared with other algorithms.

[0144] Table 1: Comparison results of this application with each method on the StateHaze1k and RICE1 datasets;

[0145]

[0146] The comprehensive experimental results show that the defogging network proposed in this paper shows significant advantages in different fog concentration scenarios (light fog, medium fog, thick fog) and real haze conditions. Through the wavelet frequency domain heterogeneous enhancement module, the network can effectively separate the high and low frequency information of the image, and perform heterogeneous processing on the physical characteristics of the high and low frequency features, so as to promote the restoration of high frequency local details and the defogging of global haze. Through quantitative analysis and visual comparison verification, it is proved that the method is superior in key indicators such as removing haze degradation features, improving image clarity and restoring scene colors, providing a highly robust solution for the remote sensing image defogging task.

[0147] 3. Ablation Experiments

[0148] In order to evaluate the effectiveness of each component, the present invention mainly conducts ablation experiments on the wavelet frequency domain heterogeneous enhancement module. It includes three experiments:

[0149] (1) Base: The basic U-shaped framework is mainly composed of six symmetric encoders and decoders. Each encoding and decoding layer contains upsampling, downsampling, residual blocks, channel attention and pixel attention. Among them, skip connections are used in the encoding and decoding layers to ensure information flow and transmission.

[0150] (2) Base + High - freq Branch: On the basis of the basic U - shaped framework, the upper branch (high - freq branch) of the wave frequency - domain heterogeneous enhancement module is embedded in each encoding - decoding layer, without including the lower branch (low - freq branch).

[0151] (3) Base + High - freq Branch + Low - freq Branch: On the basis of (2), the lower branch (low - freq branch) is added, that is, the complete wavelet frequency - domain heterogeneous enhancement module is added to the Base network.

[0152] For the fairness of the experiment, in this embodiment, the above three experiments are trained in the same way under the RSID dataset, and the PSNR and SSIM results are shown in Table 2.

[0153] Table 2: Quantitative results of ablation experiments at each stage on the StateHaze1kThin dataset;

[0154]

[0155] First, the basic framework achieved 27.65 db and 0.946 db in PSNR and SSIM. Adding the upper branch to the basic framework, the PSNR and SSIM increased by 0.75 db and 0.005 db respectively; the experimental results show the effectiveness of the upper branch proposed in this embodiment. Then, in this embodiment, the lower branch is added to the network, and it can be seen that both PSNR and SSIM have been greatly improved. Compared with only the upper branch, the PSNR and SSIM are increased by 0.8 db and 0.005 respectively. The experimental results show that the present invention has a certain effect on remote - sensing image defogging and verifies the effectiveness of the wavelet heterogeneous enhancement module (WFHE).

Claims

1. A construction method of a remote sensing image dehazing network based on wavelet frequency domain heterogeneous enhancement, characterized in that It includes the following steps: S1. Construct a U-shaped remote sensing image dehazing network, including an encoding layer, a decoding layer, skip connections, and a wavelet frequency domain heterogeneous enhancement module; S2. Construct multiple upsampling and downsampling modules and a skip connection mechanism. Use 2D convolution for downsampling, transposed convolution for upsampling, and pixel-wise addition for skip connections; S3. Construct multiple wavelet frequency domain heterogeneous enhancement modules, including wavelet transform, multi-scale deformable convolution, and fast Fourier transform, for heterogeneous enhancement of high- and low-frequency features; S4. Feed the hazy remote sensing image into the U-shaped image dehazing network, and finally output a clear haze-free image through multiple serial wavelet frequency domain heterogeneous enhancement modules and upsampling and downsampling modules; S5. Calculate the loss through the output clear image to constrain the training of the network.

2. The construction method of a remote sensing image defogging network based on wavelet frequency domain heterogeneous enhancement according to claim 1, characterized in that, The U-shaped remote sensing image dehazing network constructed in step S1 includes: Input the hazy image Hazy, perform shallow extraction, go to the first layer of the encoding layer, the second layer of the encoding layer, the third layer of the encoding layer, the first layer of the decoding layer, the second layer of the decoding layer, the third layer of the decoding layer, and reconstruct the image to output the dehazed image.

3. The construction method of a remote sensing image dehazing network based on wavelet frequency domain heterogeneous enhancement according to claim 1, characterized in that, In step S2, in constructing multiple upsampling and downsampling modules and a skip connection mechanism, the specific operations for the first layer of the decoding layer are as follows: The specific operation of downsampling is: for the features of the encoding layer (H, W, C), go to 3*3 Convolution, and then to the features after downsampling (H / 2, W / 2, 2C), where H, W, and C are the height, width, and number of channels of the image respectively; The specific operation of upsampling is: for the features of the decoding layer (H / 4, W / 4, 4C), go to 3*3 DeConvolution, and then to the features after upsampling (H / 2, W / 2, 2C); where H, W, and C are the height, width, and number of channels of the image respectively; The steps of the skip connection mechanism are as follows: Perform pixel-wise addition on the output features of the third layer of the encoding layer ((H / 4, W / 4, 4C)) and the first layer of the decoding layer ((H / 4, W / 4, 4C)); Perform pixel-wise addition on the output features of the second layer of the encoding layer ((H / 2, W / 2, 2C)) and the second layer of the decoding layer (H / 2, W / 2, 2C); Perform pixel-wise addition on the output features of the first layer of the encoding layer (H, W, C) and the third layer of the decoding layer (H, W, C).

4. The construction method of a remote sensing image dehazing network based on wavelet frequency domain heterogeneous enhancement according to claim 1, characterized in that, In step S3, construct the serial wavelet frequency domain heterogeneous enhancement module WFHE as follows: S301, separate the frequency information of the image by performing discrete wavelet transform on the encoded and decoded feature L to obtain three high-frequency subbands L HL , L LH , L HH and a low-frequency subband L LL ; S302. Concatenate and fuse the three high-frequency subbands, and then extract multi-scale irregular textures and high-frequency features with directional changes through deformable convolution; Use spatial attention and channel attention mechanisms to select the most representative high-frequency features for enhancement, and generate a pixel-level weight matrix to weight the high-frequency features at different scales, so as to strengthen the expression of key scale features;; S303. Perform fast Fourier transform on the low-frequency subband to separate the amplitude and phase; the haze information of the haze image is mainly contained in the Fourier amplitude spectrum, while the Fourier phase spectrum carries more structural information; Restore the amplitude information, and use the amplitude residual to generate a haze distribution degradation map, so as to guide the reconstruction of the structural information; S304, the high and low frequency features after heterogeneous enhancement are subjected to inverse wavelet transform to obtain the final output feature L out 。 5. The construction method of a remote sensing image dehazing network based on wavelet frequency domain heterogeneous enhancement according to claim 1, characterized in that, In step S4, the U-shaped network is constructed as follows: Input the hazy image Hazy, to shallow feature extraction, to the encoder, to the decoder, to reconstruct the dehazed image, to output the dehazed image, to loss constraint.

6. The construction method of a remote sensing image defogging network based on wavelet frequency domain heterogeneous enhancement according to claim 5, characterized in that, Each encoding and decoding layer is embedded with upsampling and downsampling and wavelet frequency domain heterogeneous enhancement modules, and information flow is ensured through skip connections; four loss constraint networks are used for training, including L1 loss, adversarial loss, perceptual loss, and multi-scale structural similarity loss.

7. The construction method of a remote sensing image dehazing network based on wavelet frequency domain heterogeneous enhancement according to claim 6, characterized in that The specific formula of L1 loss is: where F represents the dehazed image output by the network, and x i and y i represent the values of the dehazed image and the clear image at pixel p, respectively; K represents the number of pixels in the image.

8. A method for constructing a remote sensing image dehazing network based on wavelet frequency domain heterogeneous enhancement according to claim 7, characterized in that The adversarial loss constructs a loss function using the adversarial training mechanism of the generative adversarial network GAN; the generator G is responsible for converting the hazy image into a clear image, and the discriminator D distinguishes the distribution difference between the generated image and the real haze-free image through adversarial learning. Its adversarial loss function can be expressed as: Among them, X represents the input image, and N represents the number of image pixel points.

9. A method for constructing a remote sensing image dehazing network based on wavelet frequency domain heterogeneous enhancement, characterized in that, The perceptual loss effectively preserves the texture details and structural information of the image by measuring the difference in the deep feature space; the pre-trained VGG-16 network is used to extract high-level semantic features, and the perceptual loss is defined as follows: Among them, O i represents obtaining the feature map size of layer i of the VGG16 pre-trained model; K represents the number of feature layers of the VGG16 pre-trained model used in the perceptual loss; where x and y represent the foggy image and the clear image respectively.

10. The construction method of a remote sensing image dehazing network based on wavelet frequency domain heterogeneous enhancement according to claim 9, characterized in that, The multi-scale structural similarity loss is used to constrain the network to make the structural similarity between the dehazed image and the clear image closer; the specific expression is as follows. The specific formula is: where μ i and μ j represent the means of the dehazed image and the clear image respectively, σ i and σ j represent the standard deviations of the dehazed image and the clear image respectively, σ ij is used to represent the covariance between the dehazed image and the clear image, k m and l m are two important terms in the equation; C1 and C2 are constant terms; The overall network loss function is expressed as: L total = λ1L adv + λ2L content + λ3L clear + λ4L BGCC + λ5L BLCC ; Among them, λ1, λ2, λ3, λ4, and λ5 are the hyperparameters of each function.

Citation Information

Cited By

  • Image defogging system and method for low-altitude scene

    CN120598820A

  • An image dehazing system and method for low-altitude scenes

    CN120598820B

  • Remote sensing image defogging method and device based on wavelet multi-scale decomposition

    CN120876239A

  • Low-light remote sensing image restoration method and system based on double-frequency-domain processing

    CN121563821A

  • A low-light remote sensing image restoration method and system based on dual-frequency domain processing

    CN121563821B