Construction method of remote sensing image defogging network based on CNN and Transform bidirectional modulation

By constructing a U-shaped remote sensing image defog network, combining parallel convolutional blocks and multi-head self-attention blocks, the modulation feature distribution differences is solved by using the bidirectional modulation module guided by differential experts, the problem of span modulation feature mismatch in the CNN and Transformer fusion architecture is solved, and the high-quality remote sensing image defog effect is achieved.

CN120355600APending Publication Date: 2025-07-22CHINA THREE GORGES UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510471599.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

There is a cross-modal feature space-semantic mismatch problem in the existing CNN and Transformer fusion architecture, resulting in poor effect of remote sensing image defogging method in color and texture detail reconstruction.

Method used

A U-shaped remote sensing image defogging network is constructed, combining parallel convolutional blocks and multi-head self-attention blocks, modulate the feature distribution differences through the two-way modulation module guided by differential experts, and improve feature robustness and generalization capabilities using the atmospheric scattering model.

Benefits of technology

It significantly improves the effect of remote sensing images to remove fog, restores high-quality clarity and detailed information, and shows excellent robustness and generalization capabilities, especially under different haze conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355600A_ABST
    Figure CN120355600A_ABST
Patent Text Reader

Abstract

The invention discloses a construction method of a remote sensing image defogging network based on CNN and Transform bidirectional modulation. The network mainly comprises a feature extraction mechanism and a differential expert guided bidirectional modulation module (DGBM). The feature extraction mechanism extracts local features and global features through a CNN block and a Transform block which are parallel to each other. The DGBM is mainly composed of a difference expert, a physical inversion model and bidirectional affine transformation. The differential experts are used for modeling unique representations of CNN and Transform features, and suppressing expressions of redundant information at the same time. A physical inversion model is used for exploring the degradation characteristics of haze, so that the robustness and generalization ability of the characteristics are improved. And finally, affine transformation is used for modulating the difference of the feature distribution of the two, so that the semantic and spatial distribution inconsistency of the features of the two are effectively aligned. According to the method, the effectiveness of the proposed framework and module is verified, and the clear remote sensing image with higher quality can be recovered while the network performance is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image processing, and particularly relates to a method for constructing a remote sensing image dehazing network based on bidirectional modulation of CNN and Transformer. Background Technique

[0002] With the rapid development of remote sensing technology, remote sensing images are increasingly widely used in the fields of environmental monitoring, meteorological forecasting, ground observation, ocean resource management, disaster warning, and space exploration. However, during the process of image acquisition by remote sensing satellites and unmanned aerial vehicles, minute particles such as haze, smoke, and dust in the atmosphere will cause significant interference to the imaging quality, resulting in problems such as image blurring, decreased contrast, color distortion, and detail loss. These problems not only reduce the visual effect of the image but also seriously affect the subsequent image interpretation and analysis accuracy, limiting the potential of remote sensing data in practical applications; therefore, how to efficiently remove the haze interference in remote sensing images, restore the clarity, contrast, and color authenticity of the images, and at the same time retain rich detail information has become an important research direction in the field of remote sensing image processing.

[0003] Currently, the methods for remote sensing image dehazing can be mainly divided into two categories: dehazing methods based on handcrafted priors and dehazing methods based on deep learning. Most early methods mainly relied on hand-designed prior assumptions, such as the Dark Channel Prior (DCP) and the Color Attenuation Prior (CAP), and achieved dehazing by combining the Atmospheric Scattering Model (ASM). In recent years, with the development of deep learning technology, data-driven dehazing methods have shown significant advantages in remote sensing image interpretation. Learning-based remote sensing image dehazing methods are divided into two categories: physically-driven methods based on parameter estimation and end-to-end methods. The former usually relies on deep learning models to predict the atmospheric light and transmission rate maps, and then combines these prediction results with the Atmospheric Scattering Model (ASM) to generate clear images, such as AOD-Net and Dehaze-Net. However, although these prior-based methods can achieve good results in some scenarios, when the assumed conditions are significantly different from the real situation, such as uneven haze distribution or drastic changes in illumination, the dehazing performance is often greatly affected. To solve this problem, researchers have begun to explore end-to-end image dehazing algorithms without relying on the atmospheric scattering model. Among them, Convolutional Neural Networks (CNNs) and Transformer architectures have been widely explored in end-to-end dehazing algorithms. In terms of CNN methods, "Partial siamese with multiscale bi-codec networks for remote sensing image haze removal" published by Sun et al. proposed a siamese network to enhance the constraint ability for haze regions and utilized multi-scale information to improve the network's reconstruction ability for color and texture information. In terms of Transformer methods, "Dehaze-tggan: Transformer-guide generative adversarial networks with spatial-spectrum attention for unpaired remote sensing dehazing" published by Zheng et al. further expanded the applicability of the Transformer architecture by introducing an additional attention mechanism in the frequency domain and the total variation loss. However, CNN-based feature extraction methods have limitations in capturing long-range dependencies, while Transformer methods have weak local perception ability. In recent years, some dehazing methods have effectively improved the dehazing performance by combining the complementary advantages of CNNs and Transformers, thus overcoming the limitations of single-method dependence."Vision transformers for single image dehazing" published by Song et al. combines Swin-Transformer with Unet and significantly improves the defogging performance by optimizing the normalization layer, activation function, and other components. "Ucpsmbformer: Unified transformer with semantically contrastive learning for image dehazing" published by Wang et al. designs a joint Transformer module that cleverly integrates the advantages of CNN and Transformer by adding the CNN and Transformer features pixel by pixel. However, these methods usually adopt an implicit compensation mechanism of simple feature superposition and fail to effectively solve the spatial misalignment problem of cross-modal features, which has become a key bottleneck restricting the improvement of defogging performance.

[0004] Therefore, a method for constructing a remote sensing image defogging network based on bidirectional modulation of CNN and Transformer is needed to solve the above problems. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a method for constructing a remote sensing image defogging network based on bidirectional modulation of CNN and Transformer, aiming to solve the key problem of cross-modal feature space-semantic mismatch existing in the current CNN and Transformer fusion architectures; due to the inherent difference in feature extraction paradigms between the local inductive bias of CNN and the global attention mechanism of Transformer in existing methods, the feature expression ability after fusion is insufficient, resulting in deviations in the color and texture details of the generated images.

[0006] To solve the above technical problems, the technical solution adopted by the present invention is:

[0007] A method for constructing a remote sensing image defogging network based on bidirectional modulation of CNN and Transformer, comprising the following steps:

[0008] S1, construct a U-shaped remote sensing image defogging network, including an encoding layer, a decoding layer, a feature extraction mechanism, a differential expert-guided bidirectional modulation module DGBM, and skip connections;

[0009] S2, construct multiple feature extraction mechanisms, including parallel convolutional blocks CB and multi-head self-attention blocks TB, which are respectively used to extract local features and long-range dependencies;

[0010] S3. Construct multiple parallel differential expert-guided bidirectional modulation modules (DGBMs). Each DGBM includes multiple differential expert modules, a physical model, and a bidirectional affine transformation, which are used to modulate the difference in the feature distributions of the two;

[0011] S4. Feed the hazy remote sensing image into the U-shaped image dehazing network. Through multiple parallel feature extraction and differential expert-guided bidirectional modulation modules, finally output a clear haze-free image;

[0012] S5. Calculate the loss based on the output clear image to constrain the training of the network.

[0013] Preferably, the U-shaped remote sensing image dehazing network constructed in step S1 includes:

[0014] Input the hazy image Hazy, go through shallow extraction, to the first layer of the encoding layer, to the second layer of the encoding layer, to the bottleneck layer, to the first layer of the decoding layer, to the second layer of the decoding layer, and output the dehazed image Dehazing through the reconstructed image.

[0015] Preferably, in step S2, construct multiple feature extraction mechanisms as follows:

[0016] The specific operation of CB is: for the features of the encoding layer Perform 3*3 Convolution, then go through BatchNorm, ReLU, 3*3 Convolution, BatchNorm, ReLU, 3*3 Convolution, BatchNorm, ReLU, 3*3 Convolution, BatchNorm, ReLU in sequence to obtain local features

[0017] The specific operation of TB is: for the features of the encoding layer Perform BatchNorm to obtain projection features Then go through 1*1 Convolution and 3*3 Convolution in sequence to obtain the query vector Q, key vector K, and value vector V;

[0018] For the query vector Q and key vector K, perform matrix multiplication (Q·K T ), then go through the activation function Sigmoid to obtain the attention matrix M;

[0019] For the value vector V and attention matrix M, perform matrix multiplication (V·M) to obtain global features

[0020] Preferably, in step S3, construct multiple parallel differential expert-guided bidirectional modulation modules (DGBMs) as follows:

[0021] S301, Subtract the extracted CNN features C in and the Transformer features T in pixel by pixel to obtain the differential features C d and T d ;

[0022] S302, Input C d and T d into the differential expert to model the unique representations of CNN and Transformer and suppress the shared information to obtain C0 and T0;

[0023] S303, Use C0 to generate the local atmospheric light A1 and the local transmission map T1, T0 to generate the global atmospheric light A2 and T2, and use the atmospheric scattering model to obtain the enhanced features C * and T * , thereby improving the robustness and generalization ability of the features;

[0024] S304, Predict the affine parameters (X, Y) and (M, N) through mutual conditional information, and use affine transformation to modulate the difference in the feature distributions of the two to obtain the modulated features C res and T res ;

[0025] S305, Add the two modulated features C res and T res element by element to obtain the output feature L of DGBM.

[0026] Preferably, in step S4, the U-shaped network is constructed as follows:

[0027] Input the hazy image Hazy, go to shallow feature extraction, go to the encoder, go to the bottleneck layer, go to the decoder, go to reconstruct the dehazed image, go to output the dehazed image, go to loss constraint.

[0028] Preferably, each encoding and decoding layer and the bottleneck layer are embedded with a feature extraction mechanism and a differential expert-guided bidirectional modulation module; trained using four loss constraint networks, including L1 loss, adversarial loss, perceptual loss, and multi-scale structural similarity loss.

[0029] Preferably, the specific formula of the L1 loss is:

[0030]

[0031] where F represents the dehazed image output by the network, x i and y i represent the values of the dehazed image and the clear image at pixel i, respectively. K represents the number of pixels in the image.

[0032] Preferably, the anti-loss constructs a loss function by adopting the adversarial training mechanism of the generative adversarial network GAN; the generator G is defined to realize the domain conversion from the foggy image to the clear image, and the discriminator D distinguishes the distribution difference between the generated image and the real fog-free image through adversarial learning. Its adversarial loss function can be expressed as:

[0033]

[0034] where X represents the defogged image and N represents the number of image pixels.

[0035] Preferably, the perceptual loss effectively retains the texture details and structural information of the image through the metric in the deep feature space; the pre-trained VGG-16 network is used to extract high-level semantic features, and the perceptual loss is defined as follows:

[0036]

[0037] where M i represents the size of the feature map of the i-th layer of the VGG16 pre-trained model; K represents the number of feature layers of the VGG16 pre-trained model used in the perceptual loss; where x and y represent the foggy image and the clear image respectively.

[0038] Preferably, the multi-scale structural similarity loss is used to constrain the network. In order to make the structural similarity between the defogged image and the clear image closer, the multi-scale structural similarity loss is introduced, and the specific expression is as follows: The specific formula is:

[0039]

[0040] where μ i , μ j represent the means of the image after defogging and the clear image respectively, σ i , σ j represent the standard deviations of the image after defogging and the clear image respectively, σ ij is used to represent the covariance of the image after defogging and the clear image, β m , a m are two important terms in the equation; C1 and C2 are constant terms;

[0041] The overall network loss function is expressed as:

[0042] L total = ω1L adv + ω2L content + ω3L clear + ω4L BGCC + ω5L BLCC

[0043] where ω1, ω2, ω3, ω4 and ω5 are the hyperparameters of each function.

[0044] The beneficial effects of the present invention are as follows:

[0045] 1. The present invention proposes a bidirectional modulation network based on CNN and Transformer. By effectively modulating the differences in the feature distributions of CNN and Transformer, the expression ability after fusion and feature recalibration are improved. The network's ability to reconstruct clear images is strengthened from both the global and local aspects. High-quality remote sensing image dehazing results are achieved, and it is proven that the algorithm we proposed achieves the best performance on two public remote sensing image datasets, StateHaze1k and RICE.

[0046] 2. The differential expert-guided bidirectional modulation module proposed in the present invention uses its differential features to effectively model the unique representations of CNN and Transformer features and suppress the expression of shared information. In addition, this module deeply explores the degradation characteristics of haze in the feature space, thus effectively blocking the transmission of haze information to the features of both, and significantly improving the robustness and generalization ability of the features. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 is the overall network structure diagram of the embodiment of the present invention;

[0048] Figure 2 is Figure 1 the structural schematic diagram of the CNN extraction module (CB) of the feature extraction mechanism in

[0049] Figure 3 is Figure 1 the structural schematic diagram of the Transformer extraction module (TB) of the feature extraction mechanism in

[0050] Figure 4 is Figure 1 the structural schematic diagram of the differential expert-guided bidirectional modulation module (DGBM) in DETAILED DESCRIPTION OF THE EMBODIMENTS

[0051] Embodiment 1:

[0052] As Figure 1 shown, a method for constructing a remote sensing image dehazing network based on bidirectional modulation of CNN and Transformer includes the following steps:

[0053] S1. Construct a U-shaped remote sensing image dehazing network, including an encoding layer, a decoding layer, a feature extraction mechanism, a differential expert-guided bidirectional modulation module DGBM, and skip connections;

[0054] S2. Construct multiple feature extraction mechanisms, including parallel convolutional blocks CB and multi-head self-attention blocks TB, which are used to extract local features and long-range dependencies respectively;

[0055] S3. Construct multiple parallel differential expert-guided bidirectional modulation modules DGBM. DGBM includes multiple differential expert modules, physical models, and bidirectional affine transformations, which are used to modulate the differences in the feature distributions of the two;

[0056] S4. Feed the hazy remote sensing image into the U-shaped image dehazing network. Through multiple parallel feature extraction and differential expert-guided bidirectional modulation modules, finally output a clear haze-free image;

[0057] S5. Calculate the loss through the output clear image to constrain the training of the network.

[0058] Preferably, the U-shaped remote sensing image dehazing network constructed in step S1 includes:

[0059] Input the hazy image Hazy, go to shallow extraction, go to the first layer of the encoding layer, go to the second layer of the encoding layer, go to the bottleneck layer, go to the first layer of the decoding layer, go to the second layer of the decoding layer, and output the reconstructed image to obtain the dehazing image Dehazing.

[0060] Preferably, in step S2, construct multiple feature extraction mechanisms as follows:

[0061] The specific operation of CB is: for the features of the encoding layer Perform 3*3 Convolution, then go to BatchNorm, then to ReLU, then to 3*3 Convolution, then to BatchNorm, then to ReLU, then to 3*3 Convolution, then to BatchNorm, then to ReLU, then to 3*3 Convolution, then to BatchNorm, then to ReLU to obtain local features

[0062] The specific operation of TB is: for the features of the encoding layer Perform BatchNorm to obtain the projected features Then go to 1*1 Convolution, then to 3*3 Convolution to obtain the query vector Q, key vector K, and value vector V;

[0063] The query vector Q, the key vector K, go to matrix multiplication (Q·K T )), then to the activation function Sigmoid to obtain the attention matrix M;

[0064] The value vector V, the attention matrix M, go to matrix multiplication (V·M) to obtain the global features

[0065] Preferably, in step S3, multiple parallel differential expert-guided bidirectional modulation modules DGBM are constructed as follows:

[0066] S301, subtract the extracted CNN feature C in and the Transformer feature T in pixel by pixel to obtain the differential features C d and T d ;

[0067] S302, input C d and T d into the differential expert to model the unique representations of CNN and Transformer and suppress the shared information to obtain C0 and T0;

[0068] S303, use C0 to generate the local atmospheric light A1 and the local transmission map T1, use T0 to generate the global atmospheric light A2 and T2, and use the atmospheric scattering model to obtain the enhanced features C * and T * , thereby enhancing the robustness and generalization ability of the features;

[0069] S304, predict the affine parameters (X, Y) and (M, N) through mutual conditional information, and use affine transformation to modulate the difference in the feature distributions of the two to obtain the modulated features C res and T res ;

[0070] S305, add the two modulated features C res and T res element by element to obtain the output feature L of DGBM.

[0071] Preferably, in step S4, the U-shaped network is constructed as follows:

[0072] Input the hazy image Hazy, to shallow feature extraction, to the encoder, to the bottleneck layer, to the decoder, to reconstruct the dehazed image, to output the dehazed image, to loss constraint.

[0073] Preferably, each encoding and decoding layer and the bottleneck layer are embedded with a feature extraction mechanism and a differential expert-guided bidirectional modulation module; trained using four loss constraint networks, including L1 loss, adversarial loss, perceptual loss, and multi-scale structural similarity loss.

[0074] Preferably, the specific formula for L1 loss is:

[0075]

[0076] where F represents the dehazed image output by the network, x i and yi They respectively represent the values of the defogged image and the clear image at pixel i. K represents the number of pixels in the image.

[0077] Preferably, the loss against loss uses the adversarial training mechanism of the generative adversarial network GAN to construct the loss function; it is defined that the generator G realizes the domain conversion from the foggy image to the clear image, and the discriminator D distinguishes the distribution difference between the generated image and the real fog-free image through adversarial learning. Its adversarial loss function can be expressed as:

[0078]

[0079] Among them, X represents the defogged image, and N represents the number of image pixel points.

[0080] Preferably, the perceptual loss effectively retains the texture details and structural information of the image through the measurement in the deep feature space; the pre-trained VGG-16 network is used to extract high-level semantic features, and the perceptual loss is defined as follows:

[0081]

[0082] Among them, M i represents the size of the feature map obtained from the i-th layer of the VGG16 pre-trained model; K represents the number of feature layers of the VGG16 pre-trained model used in the perceptual loss; where x and y respectively represent the foggy image and the clear image.

[0083] Preferably, the multi-scale structural similarity loss is used to constrain the network. In order to make the structural similarity between the defogged image and the clear image closer, the multi-scale structural similarity loss is introduced, and the specific expression is as follows: the specific formula is:

[0084]

[0085] Among them, μ i , μ j respectively represent the means of the image after defogging and the clear image, σ i , σ j respectively represent the standard deviations of the image after defogging and the clear image, σ ij is used to represent the covariance of the image after defogging and the clear image, β m , a m are two important terms in the equation; C1 and C2 are constant terms;

[0086] The overall network loss function is expressed as:

[0087] L total = ω1L adv + ω2L content + ω3L clear + ω4L BGCC + ω5LBLCC

[0088] Among them, ω1, ω2, ω3, ω4, and ω5 are hyperparameters of each function.

[0089] Example 2:

[0090] The specific method provided in this example includes the following steps:

[0091] Step S1 specifically includes:

[0092] As Figure 1 shown, the network first performs shallow feature extraction on the input foggy image through a convolutional layer; subsequently, the features pass through five symmetric encoder-decoder layers to gradually learn hierarchical deep feature representations; among them, the skip connection mechanism runs through the encoder-decoder process to ensure the efficient transmission of feature information; each encoder-decoder layer contains a multi-branch feature extraction mechanism and a differential expert-guided bidirectional modulation module (DGBM) to achieve the collaborative optimization of local features and global dependencies. Finally, the network reconstructs a clear fog-free image through a convolutional layer.

[0093] Step S2 specifically includes:

[0094] Construct a multi-branch feature extraction mechanism, including two branches. The upper branch uses a convolutional block (CB) to extract features of the CNN, and the lower branch uses a multi-head self-attention block (TB) to capture long-range dependencies.

[0095] As Figure 2 shown, the input feature passes through 4 layers of 3*3 Convolution, BatchNorm (BN) batch normalization, and ReLU activation function, and finally obtains local features which can be expressed by the following formula:

[0096]

[0097] At the same time, as Figure 3 shown, the input feature passes through a normalization layer to obtain and then obtains query vector Q, key vector K, and value vector V through three projection layers. Then, the transpose of Q and K is multiplied matrix-wise, and then passed through Softmax to obtain attention matrix M. Finally, the attention matrix M and V are weighted to obtain global features The formula is expressed as follows:

[0098]

[0099] M = Softmax(Q × K T ));

[0100] Tin =V×M.

[0101] Step S3 specifically includes:

[0102] Construct a differential expert-guided two-way modulation module, and subtract the output results T in and C in from each other pixel by pixel to obtain two differential features T d and C d , and then model the unique representations of the two features through two differential expert mechanisms and suppress the transmission of shared information; the formula is expressed as follows:

[0103] C d =C in -T in ;

[0104] T d =T in -C in ;

[0105]

[0106] θ i and η i represent the i-th differential expert, and represent randomly generated weights. Subsequently, the modeled features are used to mine the degradation characteristics of haze in the feature space through a physical model to improve the robustness and generalization ability of the features. The formula is expressed as follows:

[0107] A1,A2 = δ1(MLP(GAP(C0,T0);

[0108] T1,T2 = δ2(Conv 1×1 (Unet(Conv 1×1 (C0,T0)))));

[0109] C * =C0×T1+(1 - T1)A1;

[0110] T * =C0×T2+(1 - T2)A2;

[0111] A1,A2 represent local atmospheric light and global atmospheric light, and T1,T2 represent local transmission map and global transmission map. Finally, the precise affine parameters are predicted through mutual conditional information, so as to effectively modulate the distribution difference of the two features. The formula is expressed as follows:

[0112] X,Y = P(C * ));

[0113] M,N = P(T* );

[0114] C res = C * × M + N;

[0115] T res = T * × X + Y;

[0116] Where P represents the prediction module, which consists of two layers of convolution. Each layer of convolution generates an affine parameter.

[0117] Finally, the two modulated features are added pixel by pixel to obtain the fused feature; the formula is expressed as follows:

[0118] L = C res + T res .

[0119] Example 3:

[0120] This example discloses the test process as follows:

[0121] 1. Parameter setting:

[0122] The experimental evaluation of this example is carried out on two public remote sensing dehazing datasets, SateHaze1k and RICE. Among them, the SateHaze1k dataset contains three subsets: SateHaze1k Thin, SateHaze1k Moderate, and SateHaze1k Thick. Each subset contains 400 pairs of synthetic haze images (320 pairs in the training set, 35 pairs in the validation set, and 45 pairs in the test set). The RICE1 dataset is constructed based on the Google Earth platform and contains 500 pairs of 512×512 resolution paired images (clear images and corresponding fogged images). In this study, it is divided into a training set (450 pairs) and a test set (50 pairs) at a ratio of 9:1 for verification. In addition, the present invention uses the Adam optimizer to optimize the proposed network, with the momentum decay exponents β1 = 0.9, β2 = 0.999, and the learning rate and batch size are set to 0.0001 and 4 respectively. The initial learning rate is set to 0.001, and the MultiStepLR is used to dynamically adjust the learning rate between them. During the training process, the parameters of the loss function in the network model of this example are set to ω1 = 1, ω2 = 0.0005, ω3 = 0.01, ω4 = 0.5. This example compares the dehazing network of this example with the other 7 recent excellent algorithms. The comparison algorithms mainly include DCP, FFA, FSDGN, DeHamer, DEA-Net, FSNet, and DCMP.

[0123] 2. Experimental results:

[0124] The method provided in this embodiment is quantitatively evaluated with seven other algorithms on the SateHaze1k and RICE1 datasets. Table 1 shows the average PSNR and SSIM values of the tested methods.

[0125]

[0126] Table 1: Comparison results of this application and each method on StateHaze1k and RICE1 datasets;

[0127] As shown in the fifth column of Table 1, DCP has a PSNR of only 11.37dB on the SateHaze1k dataset, while the rest of the algorithms have PSNRs higher than 20dB, proving that the end-to-end dehazing algorithm is superior to the traditional parameter estimation algorithm. This embodiment notes that on the SateHaze1k dataset, the network of this embodiment is much higher than other methods in terms of thin fog, medium fog, and thick fog. Among them, the invention of this embodiment has a significant improvement over the second place in the SateHaze 1k Thin, SateHaze 1kMederate, and SateHaze 1kThick datasets, with PSNR improvements of 0.49db, 1.14db, and 0.46db, respectively, and SSIM improvements of 0.005, 0.001, and 0.016, respectively. In addition, in the RICE1 dataset, the method proposed in this question achieved better performance than other algorithms.

[0128] Comprehensive experimental results show that the proposed defogging network has significant advantages in different fog density scenes (light fog, medium fog, dense fog) and real haze conditions. Through the differential expert-guided bidirectional modulation mechanism and multi-scale feature fusion strategy, the network can effectively suppress fog concentration interference, enhance texture detail fidelity, and accurately restore scene color information. Quantitative analysis and visual comparison verify the advancedness of this method in core indicators such as elimination of haze degradation features, image clarity improvement, and scene color restoration, providing a highly robust solution for remote sensing image defogging tasks.

[0129] 3. Ablation experiment

[0130] In order to evaluate the effectiveness of each component, this embodiment mainly designs an ablation experiment for the innovation of the feature extraction mechanism and the differential expert-guided bidirectional modulation module. It includes three experiments:

[0131] (1) Base: The basic U-shaped framework is mainly composed of five symmetrical codecs. Each codec layer contains up and down sampling, residual blocks, channel attention, and pixel attention. The codec layer uses skip connections to ensure information flow and transmission.

[0132] (2) Base + Feature Extraction Mechanism + Add: On the basis of the basic U-shaped framework, a feature extraction mechanism is embedded in each encoding and decoding layer, and the fusion method is pixel-by-pixel addition.

[0133] (3) Base + Feature Extraction Mechanism + Differential Expert-guided Bidirectional Modulation Module (DGBM): On the basis of (2), the innovative module DGBM of this embodiment is added.

[0134] For the fairness of the experiment, the above three experiments were trained in the same way on the SateHaze 1k Thin dataset, and the PSNR and SSIM results are shown in Table 2.

[0135]

[0136] Table 2: Quantitative results of ablation experiments at each stage on the StateHaze1kThin dataset;

[0137] First, the basic framework achieved 26.36 and 0.914 in PSNR and SSIM; adding the feature extraction mechanism to the basic framework, the PSNR and SSIM increased by 0.98 dB and 0.009 respectively; the experimental results show the effectiveness of the feature extraction mechanism proposed in this embodiment. Then, adding DGBM to the network in this embodiment, it can be seen that both PSNR and SSIM have been greatly improved. Compared with the Add method, the PSNR and SSIM have increased by 0.76 and 0.014 respectively. The experimental results show that the present invention has a certain effect on remote sensing image defogging and verifies the effectiveness of the DGBM module.

Claims

1. A construction method of a remote sensing image dehazing network based on bidirectional modulation of CNN and Transformer, characterized in that, It includes the following steps: S1. Construct a U-shaped remote sensing image dehazing network, including an encoding layer, a decoding layer, a feature extraction mechanism, a differential expert-guided bidirectional modulation module DGBM, and skip connections; S2. Construct multiple feature extraction mechanisms, including parallel convolutional blocks CB and multi-head self-attention blocks TB, which are used to extract local features and long-range dependencies respectively; S3. Construct multiple parallel differential expert-guided bidirectional modulation modules DGBM. DGBM includes multiple differential expert modules, physical models, and bidirectional affine transformations, which are used to modulate the difference in the feature distributions of the two; S4. Send the hazy remote sensing image into the U-shaped image dehazing network, and finally output a clear haze-free image through multiple parallel feature extraction and differential expert-guided bidirectional modulation modules; S5. Calculate the loss through the output clear image to constrain the training of the network.

2. The construction method of a remote sensing image defogging network based on bidirectional modulation of CNN and Transformer according to claim 1, characterized in that, The U-shaped remote sensing image dehazing network constructed in step S1 includes: Input the hazy image Hazy, go through shallow extraction, to the first layer of the encoding layer, to the second layer of the encoding layer, to the bottleneck layer, to the first layer of the decoding layer, to the second layer of the decoding layer, and output the dehazed image as the reconstructed image.

3. The construction method of a remote sensing image dehazing network based on bidirectional modulation of CNN and Transformer according to claim 1, characterized in that, In step S2, construct multiple feature extraction mechanisms as follows: The specific operation of CB is as follows: for the encoded layer features Perform 3*3 Convolution, followed by BatchNorm, ReLU, 3*3 Convolution, BatchNorm, ReLU, 3*3 Convolution, BatchNorm, ReLU, 3*3 Convolution, BatchNorm, ReLU to obtain local features The specific operation of TB is as follows: for the encoded layer features perform BatchNorm until the projected features are obtained successively go through 1*1 Convolution and then 3*3 Convolution to obtain the query vector Q, key vector K, and value vector V; Query vector Q, key vector K, to matrix multiplication (Q·K T ), to activation function Sigmoid, to attention matrix M; Value vector V, attention matrix M, matrix multiplication (V·M), to obtain global features 4. The construction method of a remote sensing image dehazing network based on bidirectional modulation of CNN and Transformer according to claim 1, characterized in that, In step S3, construct multiple parallel differential expert-guided bidirectional modulation modules DGBM as follows: S301, subtract the extracted CNN feature C in and the Transformer feature T in pixel by pixel to obtain the differential features C d and T d ; S302, input C d and T d into the differential expert to model the unique representations of CNN and Transformer and suppress the shared information to obtain C0 and T0; S303, generate local atmospheric light A1 and local transmission map T1 using C0, generate global atmospheric light A2 and T2 using T0, and obtain the enhanced feature C using the atmospheric scattering model * and T * , thereby enhancing the robustness and generalization ability of the feature; S304, predict the affine parameters (X, Y) and (M, N) by using each other as conditional information, and use affine transformation to modulate the difference in the feature distributions of the two to obtain the modulated feature C res and T res ; S305, perform element-wise addition on the two modulated features C res and T res to obtain the output feature L of DGBM.

5. A method for constructing a remote sensing image dehazing network based on bidirectional modulation of CNN and Transformer, characterized in that, In step S4, construct the U-shaped network as follows: Input the hazy image Hazy, go through shallow feature extraction, to the encoder, to the bottleneck layer, to the decoder, to the reconstructed dehazed image, to the output dehazed image, and to the loss constraint.

6. The construction method of a remote sensing image dehazing network based on bidirectional modulation of CNN and Transformer according to claim 5, characterized in that, Each encoding and decoding layer and the bottleneck layer are embedded with a feature extraction mechanism and a differential expert-guided bidirectional modulation module; four losses are used to constrain the network training, including L1 loss, adversarial loss, perceptual loss, and multi-scale structural similarity loss.

7. The construction method of a remote sensing image defogging network based on bidirectional modulation of CNN and Transformer according to claim 6, characterized in that, The specific formula of the L1 loss is: where F represents the dehazed image output by the network, and x i and y i represent the values of the dehazed image and the clear image at pixel i, respectively; K represents the number of pixels in the image.

8. A method for constructing a remote sensing image dehazing network based on bidirectional modulation of CNN and Transformer, characterized in that, The adversarial loss constructs a loss function using the adversarial training mechanism of the generative adversarial network GAN; define the generator G to realize the domain conversion from the hazy image to the clear image, and the discriminator D distinguishes the distribution difference between the generated image and the real haze-free image through adversarial learning. Its adversarial loss function can be expressed as: Among them, X represents the dehazed image, and N represents the number of pixel points.

9. The construction method of a remote sensing image dehazing network based on bidirectional modulation of CNN and Transformer according to claim 8, characterized in that, The perceptual loss measures through the deep feature space and effectively retains the texture details and structural information of the image; use the pre-trained VGG-16 network to extract high-level semantic features, and the perceptual loss is defined as follows: Among them, M i represents obtaining the feature map size of layer i of the VGG16 pre-trained model; K represents the number of feature layers of the VGG16 pre-trained model used in the perceptual loss; where x and y represent the foggy image and the clear image respectively.

10. A method for constructing a remote sensing image dehazing network based on bidirectional modulation of CNN and Transformer, characterized in that, The multi-scale structural similarity loss is used to constrain the network. In order to make the structural similarity between the dehazed image and the clear image closer, the multi-scale structural similarity loss is introduced, and the specific expression is as follows: The specific formula is: where μ i and μ j represent the means of the dehazed image and the clear image respectively, σ i and σ j represent the standard deviations of the dehazed image and the clear image respectively, σ ij is used to represent the covariance of the dehazed image and the clear image, β m , a m are two important terms in the equation; C1 and C2 are constant terms; The overall network loss function is expressed as: L total = ω1L adv + ω2L content + ω3L clear + ω4L BGCC + ω5L BLCC Among them, ω1, ω2, ω3, ω4, and ω5 are the hyperparameters of each function.