Remote sensing image defogging method based on multi-expert cooperation model
By adopting a multi-expert collaborative model method in the field of remote sensing image defog removal, combined with the U-net architecture and attention mechanism, the problem of poor defog removal in complex environments is solved, and a higher quality and more adaptable image defog removal effect is achieved.
Patent Information
- Application Number
- CN202510003158.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-02
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-02
AI Technical Summary
The prior art has shortcomings in handling remote sensing images defogging in complex and dynamic environments, especially in dealing with non-uniform fog and maintaining image details.
Using a remote sensing image defogging method based on a multi-expert collaborative model, a multi-expert gated network and attention unit are designed through U-net encoding and decoding architecture, combining variable core convolution and feature fusion convolution to enhance feature extraction and image recovery capabilities.
It significantly improves the accuracy and image quality of remote sensing images, enhances the adaptability and generalization capabilities of the model, can perform well under a variety of fog conditions, and improves the defog effect and overall performance of the network.
Smart Images

Figure CN119941566A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to digital image processing and computer vision, and in particular to the field of remote sensing image processing and enhancement. Background Art
[0002] With the exponential growth of global aerospace and aviation, remote sensing technology has become the main method for obtaining surface data. Remote sensing images play a key role in many fields such as environmental monitoring, urban planning, and agricultural management. However, the quality of remote sensing images is often affected by adverse weather conditions, such as fog and haze, which reduce the contrast and color saturation of images, resulting in loss of details, and bring major challenges to the analysis and applications that rely on high-quality images. In order to improve image contrast and sharpen details, dehazing algorithms have emerged and become an important tool for enhancing image clarity. Therefore, developing innovative dehazing algorithms to restore remote sensing images is a research area of great significance and practical value in the field of computer vision.
[0003] Image dehazing algorithms have made significant progress in the past few decades. Early methods were mainly based on physical models and assumptions. Although effective in specific scenarios, these methods usually rely on overly simplified assumptions, such as constant global atmospheric light intensity or uniform fog distribution, which limits their applicability in complex and dynamic environments. Therefore, researchers have been looking for more adaptable dehazing techniques that can provide high-quality image restoration under a variety of environmental conditions.
[0004] With the rise of artificial intelligence, deep learning algorithms have shown great potential in the field of image dehazing. These techniques excel in feature extraction and adaptability without the need for complex physical models or prior knowledge. They directly map the transformation from foggy images to clear images, showing remarkable flexibility. In particular, the integration of convolutional neural networks (CNNs) and visual transformers (ViTs) enables the model to have enhanced performance in complex foggy scenes.
[0005] Although existing deep learning-based dehazing methods have made progress in some aspects, they are still insufficient in dealing with the complexity of image features. For example, existing U-net-based dehazing algorithms, which focus on a single network structure, may not be able to adequately cope with the complexity of image features. In addition, existing methods still have challenges in dealing with non-uniform fog and maintaining image details. Summary of the invention
[0006] The purpose of the present invention is to solve the problems in the prior art and to propose a remote sensing image defogging method based on a multi-expert collaboration model.
[0007] A remote sensing image defogging method based on a multi-expert collaborative model is based on the U-net encoding and decoding architecture, specifically including the following steps:
[0008] Step S1: Establish a defogging model;
[0009] Step S2: Design a multi-expert collaborative dehazing network;
[0010] Step S3: Collect training data set;
[0011] Step S4: training a multi-expert collaborative dehazing network;
[0012] Step S5: remote sensing image defogging;
[0013] The multi-expert collaborative dehazing network is divided into three stages. The first stage is four downsampling stages. By reducing the resolution of the feature map, the computational burden of the subsequent layers is reduced, the processing speed of the network is improved, and the perception field of the network is increased. The second stage includes a multi-expert gating network and an attention unit. The function is to assign the most appropriate weights to the high-level features extracted by downsampling in different channels and different positions. The attention unit includes a multi-channel attention unit and a pixel attention unit. The third stage is the upsampling stage. In this stage, the features of different stages of the image are fused and processed, and the size and detail information of the image are restored by using information of different scales.
[0014] In addition, the multi-expert collaborative dehazing network also includes variable kernel convolution and feature fusion convolution, wherein the variable kernel convolution replaces the initial convolution block of the downsampling part of the U-net, and the feature fusion convolution replaces the upsampling part of the convolution block of the U-net, wherein part of the input in the feature fusion convolution is obtained by jump connection from the downsampling output result.
[0015] In the above-mentioned remote sensing image defogging method based on the multi-expert collaboration model, the multi-expert gating network specifically includes an edge expert module, a color expert module and a gating mechanism module. The edge experts and color experts process the edge and color information of the image respectively, and then integrate this information through the gating mechanism module to obtain more comprehensive image features, thereby improving the accuracy of defogging and the quality of the image. The edge expert module adopts a 3×3 convolution kernel and batch normalization processing, and forms a feature map through ReLU activation. The color expert module adopts a 1×1 convolution kernel and batch normalization processing, and then performs upsampling processing after ReLU activation to form a feature map. The gating mechanism module adopts a 3×3 convolution kernel and batch normalization processing, and then after ReLU activation to form a feature map, the feature map is compressed through a fully connected layer to generate gating weights equal to the number of experts, and then normalized through the Softmax function to ensure that the sum of the weights is 1.
[0016] In the above-mentioned remote sensing image dehazing method based on the multi-expert collaboration model, the multi-channel attention unit compresses the spatial dimension of each channel of the input feature map to 1×1 through an adaptive average pooling layer to capture global spatial information, and after forming the feature map through ReLU activation, a 1×1 convolution layer is used to reduce and restore the feature dimension, introduce nonlinearity, and finally use the CSigmoid function to assign channel weights to achieve feature enhancement.
[0017] In the above-mentioned remote sensing image dehazing method based on the multi-expert collaboration model, the pixel attention unit reduces the feature dimension and maintains the spatial size through a 3×3 convolution layer, forms a feature map through ReLU activation, and then reduces the feature dimension to a single channel through another 3×3 convolution layer, including the global spatial attention weight, and finally multiplies the weight processed by the CSigmoid activation function with the input feature map to achieve feature enhancement.
[0018] In the above remote sensing image defogging method based on the multi-expert collaboration model, the CSigmoid activation function is an improvement on the Sigmoid function, and its expression is:
[0019]
[0020] This function expands the weight range to (-2, 2) through a linear transformation, providing a wider dynamic range for the attention weights of different channels. This extension allows the model to show greater flexibility and adaptability in adjusting the strength of feature map channels.
[0021] In the above-mentioned remote sensing image dehazing method based on the multi-expert collaboration model, the initial stage of the feature fusion convolution uses a 1×1 convolution kernel to interleave the channels of the input feature map to integrate inter-channel information and enhance the expressiveness of the features. Subsequently, through the processing of a 3×3 convolution kernel and a batch normalization layer, local spatial details are further integrated, the distribution of the feature map is stabilized, and spatial information is mined. Through multiple rounds of convolution processing, the network's ability to recognize local structures is enhanced, so that the spatial information of the feature map is more fully utilized.
[0022] In the above-mentioned remote sensing image dehazing method based on the multi-expert collaboration model, in step S3, two synthetic remote sensing haze datasets SateHaze1k and HRSD are used. SateHaze1k includes three subsets: light haze, moderate haze and thick haze. Each subset contains multiple training sets and test sets of images. The light haze images use the haze mask from the real cloud layer, while the moderate haze samples combine the characteristics of fog and moderate haze. The thick fog images are generated using the transmission map of dense fog. The HRSD dataset is divided into two subsets: LHID and DHID. Each subset contains multiple training images and test images. The foggy images in LHID are generated using the atmospheric scattering model, and the foggy images in DHID are created using the real haze map.
[0023] Compared with the prior art, the present invention has the following advantages:
[0024] 1. The multi-expert collaborative defogging network in the present invention is improved on the basis of the U-net encoding and decoding architecture. The kernel-variable (LD) convolution adopted allows the convolution kernel to have any number of parameters and any sampling shape, breaking the limitation of traditional convolution being limited to a fixed local window and a fixed sampling shape. This adaptability allows more precise adjustment for different data sets and target locations, thereby improving the accuracy of feature extraction.
[0025] 2. The edge and color information of the image are processed separately by edge experts and color experts, and then integrated through the gating mechanism to obtain more comprehensive image features, thereby improving the accuracy of dehazing and the quality of the image.
[0026] 3. Part of the input in the feature fusion convolution is obtained by jump connection from the down-sampled output results, which enables the network to take into account the global context information, enhances the generalization ability of the model, enables it to perform well under a variety of different fog conditions, and significantly improves the dehazing effect and the overall performance of the network. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 This is an overall architecture diagram of a multi-expert collaborative defogging network in a remote sensing image defogging method based on a multi-expert collaborative model proposed in the present invention.
[0028] Figure 2 This is a structural diagram of the expert gating network in the remote sensing image dehazing method based on a multi-expert collaboration model proposed in the present invention.
[0029] Figure 3 This is a structural diagram of a multi-channel attention unit in a remote sensing image dehazing method based on a multi-expert collaboration model proposed in the present invention.
[0030] Figure 4This is a structural diagram of the pixel attention unit in the remote sensing image dehazing method based on a multi-expert collaboration model proposed in the present invention.
[0031] Figure 5 This is a structural diagram of feature fusion convolution in a remote sensing image dehazing method based on a multi-expert collaboration model proposed in the present invention.
[0032] Figure 6 The present invention provides a flowchart of a remote sensing image defogging method based on a multi-expert collaboration model.
[0033] Figure 7 A comparison chart of the visualization results of different methods on one of the images in the SateHaze1k-Thin dataset.
[0034] Figure 8 A comparison chart of the visualization results of different methods on one of the images in the SateHaze1k-Moderate dataset.
[0035] Fig. 9 A comparison chart of the visualization results of different methods on one of the images in the SateHaze1k-Thick dataset.
[0036] Fig.10 A comparison chart of the visualization results of different methods on one of the images in the HRSD-LHID dataset.
[0037] Fig.11 A comparison chart of the visualization results of different methods on one of the images in the HRSD-DHID dataset. DETAILED DESCRIPTION
[0038] Reference Figure 1-11 A remote sensing image defogging method based on a multi-expert collaborative model uses a multi-expert collaborative defogging network for defogging. The overall network architecture is as follows Figure 1As shown, it is improved on the basis of the U-net encoding architecture. The invention adopts a multi-expert collaborative defogging network (MCDNet) divided into three stages. The first stage is four downsampling stages. By reducing the resolution of the feature map, the computational burden of the subsequent layers is reduced, the processing speed of the network is improved, and the network's perception field of view is increased. The second stage includes a multi-expert gated network (ECEGN) and an attention unit. The multi-expert gated network captures more complex features and patterns to improve the performance of the defogging task and assigns the most appropriate weights. The attention unit refocuses on key information and strengthens the generalization of the model on various tasks and data sets. The attention unit includes a multi-channel attention unit (MCAU) and a pixel attention unit (PAU). The third stage is the upsampling stage. This stage restores the size and detail information of the image by fusing the features of different stages of the image and using information of different scales. In addition, the multi-expert collaborative defogging network The network also includes variable kernel convolution (LDConv) and feature fusion convolution (FFC), where variable kernel convolution replaces the initial convolution block of the downsampling part of U-net, and variable kernel convolution allows the convolution kernel to have any number of parameters and any sampling shape, breaking the limitation of traditional convolution being limited to fixed local windows and fixed sampling shapes. This adaptability allows more precise adjustment of different data sets and target positions, thereby improving the accuracy of feature extraction, and feature fusion convolution replaces the upsampling part of the convolution block of U-net, where part of the input in feature fusion convolution is obtained by jump connection of the downsampling output result, and feature fusion convolution can not only effectively combine feature maps from different layers or sources to obtain richer and higher-level feature representations, but also this fusion produces richer and more complex feature representations, enhancing the model's understanding of the input data, and enhancing its expressiveness and generalization capabilities.
[0039] The multi-expert gating network processes the edge and color information of the image respectively through edge experts and color experts, and then integrates this information through the gating mechanism to obtain more comprehensive image features, thereby improving the accuracy of dehazing and the quality of the image. Figure 2 As shown in the figure, a 3×3 convolution kernel and batch normalization are used, and a feature map is formed through ReLU activation. The edge expert can capture the edge features in the image by using a 3×3 convolution kernel, which is crucial for restoring the details and structure of the image. Edge information is an important indication of the contour and shape of objects in the image. By accurately extracting this information, the clarity and visual perception of the dehazed image can be significantly improved.
[0040] At the same time, color experts focus on processing color channels, using a 1×1 convolution kernel and batch normalization, and then upsampling after ReLU activation to form a feature map. The use of a 1×1 convolution kernel can adjust and extract color features to correct color deviations caused by atmospheric scattering and maintain color authenticity. This step is crucial for restoring the true color of the image and improving the authenticity of the image after dehazing.
[0041] The gating mechanism is the key to achieve dynamic integration of expert outputs. The mechanism first compresses the feature map through a fully connected layer to generate gating weights equal to the number of experts, and then normalizes it through the Softmax function to ensure that the sum of the weights is 1. This design allows the network to dynamically adjust the importance of edge and color information according to the image content, so that the model can flexibly adjust its focus according to different fog scenes, thereby improving the adaptability of the model to different fog scenes. Through the above-mentioned specialized processing and dynamic integration methods, the details and colors of the image can be more effectively restored, and the defogging performance can be significantly improved. In addition, this design also enhances the robustness of the model, enabling it to better handle various complex fog conditions, including non-uniform fog and dynamic fog environments. In practical applications, this means that the multi-expert gating network can provide clearer and more accurate defogging results for remote sensing images, which is of great practical significance for subsequent image analysis and applications, such as environmental monitoring, urban planning, and agricultural management.
[0042] The attention unit includes a multi-channel attention unit and a pixel attention unit, which enhance the feature representation at the channel and pixel levels respectively to improve the dehazing performance. The unit structure is as follows: Figure 3 and Figure 4 As shown in the figure, the multi-channel attention unit compresses the spatial dimension of each channel of the input feature map to 1×1 through an adaptive average pooling layer to capture global spatial information, and reduces and restores the feature dimension through a 1×1 convolution layer, introducing nonlinearity, and finally uses the CSigmoid function to assign channel weights to achieve feature enhancement. The CSigmoid function is an improvement on the Sigmoid function. The function expands the weight range to (-2,2) through linear transformation, providing a wider dynamic range for the attention weights of different channels. This extension allows the model to show greater flexibility and adaptability in adjusting the strength of the feature map channel, and its expression is:
[0043]
[0044] The pixel attention unit focuses on pixel-level features, reduces the feature dimension through a 3×3 convolution layer while maintaining the spatial size, and then reduces the feature dimension to a single channel through another convolution layer, including the global spatial attention weight, and finally multiplies the weight processed by the CSigmoid activation function with the input feature map to achieve feature enhancement. These two attention units work together, the multi-channel attention unit strengthens the global features, and the pixel attention unit strengthens the local features, so that the multi-expert collaborative dehazing network can more accurately restore the details and colors of the image, improve the quality and realism of the dehazed image, and also enhance the adaptability and robustness of the model to complex foggy environments. Through this dual-viewpoint attention mechanism, the multi-expert collaborative dehazing network can effectively improve the dehazing effect of remote sensing images and provide clearer and more accurate image data for subsequent image analysis and applications.
[0045] Feature fusion convolution achieves deep fusion of feature maps from different network layers and different sources through a series of convolution and batch normalization operations. Its structure is as follows: Figure 5 As shown in the figure, the initial stage of the feature fusion convolution uses a 1×1 convolution kernel to interleave the channels of the input feature map to integrate inter-channel information and enhance the expressiveness of the features. Subsequently, through the processing of the 3×3 convolution kernel and the batch normalization layer, the local spatial details are further integrated, the distribution of the feature map is stabilized, and the spatial information is mined. Multiple rounds of convolution processing enhance the network's ability to recognize local structures, making the spatial information of the feature map more fully utilized. The deep fusion operation of the feature fusion convolution not only enhances the model's understanding and expression of the input data, but also improves the performance of the defogging network, especially in restoring image details and colors. In addition, since part of the input of the feature fusion convolution is derived from the jump connection of the downsampled output result, this enables the network to take into account the global context information, enhances the generalization ability of the model, and enables it to perform well under a variety of different fog conditions, significantly improving the defogging effect and the overall performance of the network.
[0046] The flow chart of the present invention is as follows Figure 6 As shown, the computer configuration uses:
[0047] Intel(R)Xeon(R)CPU E5-2620 v4 processor;
[0048] Nvidia GeForce GTX 1080Ti graphics processor;
[0049] Main frequency: 2.10GHz;
[0050] Memory: 12GB;
[0051] The operating system is Ubuntu 18.04.
[0052] The implementation of the defogging method is based on the Pytorch framework. The present invention is a defogging method based on a convolutional neural network, which specifically includes the following steps:
[0053] Step 1: Build a dehazing model
[0054] Let h represent the foggy image, r represent the restored clear image, and function F represent the mapping relationship between the foggy image and the corresponding clear image. Then the defogging problem is modeled as follows (i.e., the defogging model):
[0055] r=F(h)
[0056] According to the above formula, once the mapping relationship F is obtained, given a foggy image h, a clear image can be obtained through functional relationship mapping, thereby achieving image defogging.
[0057] Step 2: Design a multi-expert collaborative dehazing network
[0058] According to the dehazing model established in step 1, the design is as follows Figure 1 The multi-expert collaborative dehazing network shown.
[0059] Step 3: Collect training dataset
[0060] The present invention uses two synthetic remote sensing (RS) haze datasets, SateHaze1k and HRSD. SateHaze1k includes three subsets: thin haze, moderate haze, and thick haze. Each subset contains a training set of 320 images and a test set of 45 images. The thin haze images use haze masks from real clouds, while the moderate haze samples combine the features of fog and moderate haze. The thick haze images are generated using the transmission map of dense fog. The HRSD dataset is divided into two subsets: LHID and DHID. LHID contains a total of 30,517 training images and 500 test images, all of which are generated using an atmospheric scattering model. This subset simulates various haze conditions to train the network to handle different degrees of haze in remote sensing images. On the other hand, DHID consists of a total of 14,990 images, which are created using real haze maps and can present haze more realistically. Among them, 14,490 images are used for training and 500 are used for testing. The combination of these two subsets ensures a comprehensive evaluation of dehazing methods in both synthetic haze scenes and more realistic haze scenes.
[0061] Step 4: Train a multi-expert collaborative dehazing network
[0062] Learning-based dehazing methods require labeled samples for training. Currently, the most commonly used remote sensing dehazing image datasets are SateHaze1k and HRSD.
[0063] In this step, the network is trained with samples from the dataset, and the model parameters in the network are updated through the optimizer so that it can better learn the mapping relationship between foggy images and clear images. In the field of defogging, studies have shown that L1 loss is more conducive to defogging. The L1 loss function is:
[0064] L1=‖J-GT‖
[0065] Wherein, J is the actual output result of the network, GT is the true value image, and the present invention uses the PyTorch framework for training on a system equipped with four NVIDIAGeForce GTX 1080Ti GPUs. In order to enhance the training data set, we randomly rotated the input images by 90, 180 and 270 degrees and flipped them horizontally during training. The network input is an RGB remote sensing image with a size of 256×256 pixels that is randomly cropped from the original image. We use the Adam optimizer to iteratively update the parameters in the network and train each sub-dataset with a batch size of 4. The initial learning rate is set to 1.0×10-4, and the cosine annealing strategy is used to gradually reduce the learning rate to 0. The present invention selects the stochastic gradient descent method to optimize the loss function, iteratively learns the network with paired foggy and fog-free images, updates the network parameters, and ends the training when the loss value of the network tends to be stable. At this time, the saved network parameters are the trained defogging network model.
[0066] Step 5: Remote sensing image dehazing
[0067] The remote sensing image dehazing method designed by the present invention is end-to-end. Once the network model is trained, it is only necessary to input the foggy image into the network. Through the forward propagation of the network, the restored dehazed image can be obtained at the output. The present invention uses the test set in the data set to conduct effect testing experiments. In order to demonstrate the competitiveness of the dehazing method of the present invention, the present invention compares the network with seven advanced dehazing methods: DCP, AOD-Net, GridDehaze-Net, FFA-Net, FCTF-Net, SCA-Net and PhD-Net. To ensure a fair comparison, we used the official implementations of these deep learning models during the training process. These models are provided by their respective authors. For intuitive visual effects, please refer to the attached. Figure 7-11 .
[0068] In the following table, PSNR is peak signal-to-noise ratio and SSIM is result similarity.
[0069] Table 1 shows the quantitative comparison of the dehazing effects of different methods on the SateHaze1k dataset. To distinguish the performance, we use bold to indicate the best data and underline to indicate the suboptimal data.
[0070]
[0071] Table 1
[0072] Table 2 shows the quantitative comparison of the dehazing effects of different methods on the HRSD dataset. To distinguish the performance, we use bold to indicate the best data and underline to indicate the suboptimal data.
[0073]
[0074] Table 2
[0075] Figure 7 This is a visualization comparison of different methods on one of the images in the SateHaze1k-Thin dataset. In the figure, (a) (j) represent the foggy image and the real fog-free image respectively, (b) (c) (d) (e) (f) (g) (h) (i) represent the dehazed images of DCP, AOD-Net, GridDehaze-Net, FFA-Net, FCTF-Net, SCA-Net, PhD-Net and MCD-Net respectively. Similarly, Figure 8-Figure 11 These are comparison diagrams of defogging images using the above methods for different data sets. It can be seen that the defogging method of the present invention has great advantages in processing details, defogging accuracy, image quality and enhanced clarity.
[0076] It is known from common technical knowledge that the present invention can be implemented by other embodiments that do not deviate from its spirit or essential features. Therefore, the above disclosed embodiments are only illustrative in all respects and are not exclusive. All changes within the scope of the present invention or within the scope equivalent to the present invention are included in the present invention.
Claims
1. A remote sensing image defogging method based on a multi-expert collaboration model, characterized in that: The architecture of U-net encoding and decoding includes the following steps: Step S1: Establish a defogging model and use a multi-expert collaborative convolutional network to fit the mapping relationship F(h) between the foggy image and the clear image; Step S2: Design a multi-expert collaborative dehazing network; Step S3: Collect training data set; Step S4: training a multi-expert collaborative dehazing network; Step S5: remote sensing image defogging; The multi-expert collaborative dehazing network is divided into three stages. The first stage is four downsampling stages, which reduces the resolution of the feature map, reduces the computational burden of the subsequent layers, improves the processing speed of the network, and increases the network's perception field of view. The second stage includes a multi-expert gating network and an attention unit, both of which are used to assign the most appropriate weights to the high-level features extracted by downsampling in different channels and different positions. The attention unit includes a multi-channel attention unit and a pixel attention unit. The third stage is an upsampling stage, which restores the size and detail information of the image by fusing the features of different stages of the image and using information of different scales. In addition, the multi-expert collaborative dehazing network also includes variable kernel convolution and feature fusion convolution, wherein the variable kernel convolution replaces the initial convolution block of the downsampling part of the U-net, and the feature fusion convolution replaces the upsampling part of the convolution block of the U-net, wherein part of the input in the feature fusion convolution is obtained by jump connection from the downsampling output result.
2. The remote sensing image defogging method based on a multi-expert collaboration model according to claim 1, characterized in that: The multi-expert gating network specifically includes an edge expert module, a color expert module and a gating mechanism module. The edge experts and color experts process the edge and color information of the image respectively, and then integrate the information through the gating mechanism module to obtain more comprehensive image features, thereby improving the accuracy of dehazing and the quality of the image. The edge expert module adopts a 3×3 convolution kernel and batch normalization processing, and forms a feature map through ReLU activation. The color expert module adopts a 1×1 convolution kernel and batch normalization processing, and then performs upsampling processing after ReLU activation to form a feature map. The gating mechanism module adopts a 3×3 convolution kernel and batch normalization processing, and then after ReLU activation to form a feature map, compresses the feature map through a fully connected layer to generate gating weights equal to the number of experts, and then performs normalization processing through the Softmax function to ensure that the sum of the weights is 1.
3. The remote sensing image defogging method based on a multi-expert collaboration model according to claim 1, characterized in that: The multi-channel attention unit compresses the spatial dimension of each channel of the input feature map to 1×1 through an adaptive average pooling layer to capture global spatial information, and uses a 1×1 convolution layer to reduce and restore the feature dimension after forming the feature map through ReLU activation, introducing nonlinearity, and finally uses the CSigmoid function to assign channel weights to achieve feature enhancement.
4. The remote sensing image defogging method based on a multi-expert collaboration model according to claim 1, characterized in that: The pixel attention unit reduces the feature dimension and maintains the spatial size through a 3×3 convolution layer, forms a feature map through ReLU activation, and then reduces the feature dimension to a single channel through another 3×3 convolution layer, including the global spatial attention weight, and finally multiplies the weight processed by the CSigmoid activation function with the input feature map to achieve feature enhancement.
5. A remote sensing image defogging method based on a multi-expert collaboration model according to claim 3 or 4, characterized in that: The CSigmoid activation function is an improvement on the Sigmoid function, and its expression is: This function expands the weight range to (-2, 2) through a linear transformation, providing a wider dynamic range for the attention weights of different channels. This extension allows the model to show greater flexibility and adaptability in adjusting the strength of feature map channels.
6. The remote sensing image defogging method based on a multi-expert collaboration model according to claim 1, characterized in that: In the initial stage of the feature fusion convolution, a 1×1 convolution kernel is used to interleave the channels of the input feature map to integrate the information between channels and enhance the expressiveness of the features. Subsequently, through the processing of a 3×3 convolution kernel and a batch normalization layer, local spatial details are further integrated, the distribution of the feature map is stabilized, and spatial information is mined. Through multiple rounds of convolution processing, the network's ability to recognize local structures is enhanced, so that the spatial information of the feature map is more fully utilized.
7. The remote sensing image defogging method based on a multi-expert collaboration model according to claim 1, characterized in that: In step S3, two synthetic remote sensing haze datasets SateHaze1k and HRSD are generated. SateHaze1k includes three subsets: light haze, moderate haze and thick haze. Each subset contains multiple training sets and test sets of images. The light haze images use the haze mask from the real cloud layer, while the moderate haze samples combine the features of light haze and moderate haze. The thick haze images are generated using the transmission map of dense fog. The HRSD dataset is divided into two subsets: LHID and DHID. Each subset contains multiple training images and test images. The foggy images in LHID are generated using the atmospheric scattering model, and the foggy images in DHID are created using the real haze map.
Citation Information
Patent Citations
Single image defogging method and system
CN117036182A
Image defogging method based on self-attention coding and decoding
CN117151990A
Single image dehazing method based on detail recovery
US20240289928A1
Cited By
Optical remote sensing image defogging method based on lightweight parallel attention network
CN121746248A