A method for constructing a remote sensing image defogging network based on heterogeneous packet expert prompts
By constructing the Heterogeneous Grouping Expert Prompt Module (HGEP) and the Adaptive Amplitude-Guided Loss (AASG-Loss), the problem of insufficient modeling of haze distribution characteristics in remote sensing images under complex atmospheric conditions was solved, thus improving the defogging effect and image quality.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA THREE GORGES UNIV
- Filing Date
- 2026-02-28
- Publication Date
- 2026-06-09
Smart Images

Figure CN122175820A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of remote sensing image technology, and specifically to a method for constructing a remote sensing image dehazing network based on heterogeneous grouping expert prompts. Background Technology
[0002] Early remote sensing image dehazing methods primarily relied on hand-designed prior knowledge or deep learning methods to estimate parameters in atmospheric scattering models. For example, Wang et al.'s "Selfpromer: Self-prompt dehazing transformers with depth-consistency" used self-prompts to dehaze images using depth information, and Zhang et al.'s "Depth information assisted collaborative mutual promotion network for single image dehazing" used a pre-trained depth cueing network to optimize depth information and guide the network for dehazing. However, these methods often failed to adapt to diverse and complex scenarios, resulting in unsatisfactory performance. In recent years, a series of end-to-end methods have been proposed, significantly improving dehazing performance by directly learning the mapping between hazy and clear images. However, in the absence of task-specific cues, these networks mainly rely on implicit feature extraction to approximate the degradation process, which may lead to misunderstandings of semantic content under complex atmospheric conditions. For example, brightness attenuation or color deviation caused by haze may be incorrectly interpreted as illumination changes, resulting in color shifts, structural distortions, and excessive smoothing of image details. These misunderstandings impair the model's generalization ability and stability. To address these issues, researchers have introduced the concept of cue learning from Natural Language Processing (NLP) into image restoration tasks, attempting to improve the dehazing process through guidance information. Furthermore, some studies have used depth maps as degradation cues to capture spatial variations in haze distribution and guide the network to focus on areas heavily affected by haze. However, under complex atmospheric conditions, depth information is often affected by haze obstruction and scale changes, making depth cues unreliable in guiding dehazing models. Inspired by expert mechanisms, which can model the key information required for different tasks, designing high-quality cues that can effectively represent haze features and guide the dehazing process remains a pressing problem. Summary of the Invention
[0003] The technical problem to be solved by the present invention is to provide a method for constructing a remote sensing image dehazing network based on heterogeneous grouping expert prompts. This method solves the problem that existing dehazing techniques fail to effectively model the distribution characteristics of haze, resulting in poor dehazing effects under complex atmospheric conditions, especially in terms of detail texture restoration and color consistency.
[0004] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows: A method for constructing a remote sensing image dehazing network based on heterogeneous grouping expert prompts includes the following steps: Step 1: Construct a heterogeneous grouped expert prompting (HGEP) module. Through multi-dimensional distributed perception and heterogeneous expert aggregation mechanism, it clarifies the statistical changes related to haze and enhances the robustness and dehazing ability of the network. The HGEP module guides the network to capture key features of haze areas and improves the accuracy of dehazing by modeling multiple expert groups in different statistical spaces. Step 2: Introduce Adaptive Amplitude Guided Loss (AASG-Loss). The loss function of AASG-Loss optimizes the network's learning in the frequency-amplitude domain, enabling the network to effectively learn cue representations related to haze. The cue representations maintain consistency across multiple frequency sub-bands and statistical spaces, ensuring that the network's perception of haze areas is more accurate, thereby improving the image dehazing effect.
[0005] The specific process of Step 1 above is as follows: Haze degradation in remote sensing images exhibits significant nonlinearity, spatial heterogeneity, and complex distribution variations. These characteristics hinder deep networks from accurately capturing the spatial distribution and dynamic statistical features of haze. To address these challenges, a Heterogeneous Grouped Expert Hints (HGEP) module is proposed. This module explicitly models the statistical variations associated with haze through multidimensional distribution perception and heterogeneous expert aggregation mechanisms. This design enables HGEP to generate physically consistent and statistically robust cues, thereby providing enhanced guidance for effective haze removal. Assumptions As the input features, the input is first processed through a Random Fourier Feature (RFF) layer to capture the nonlinear variations in the haze distribution, mapping it to a low-dimensional explicit space, as shown in the following formula: ; Where the bias parameter From uniform distribution Mid-sampling, weight vector From Gaussian distribution Mid-sampling; This represents the feature dimension of the projection; this transformation explicitly transforms the input features. Mapping to a low-dimensional explicit embedding space allows us to capture the nonlinear changes in haze distribution, providing a solid foundation for subsequent spatial adaptive modeling. To fully model the spatial misalignment and non-uniformity of haze distribution and enhance the model's ability to perceive degradation patterns at multiple scales, a multi-scale deformable convolutional structure is introduced. This structure adaptively adjusts the sampling position, flexibly capturing spatial variations and the refinement of haze boundary regions. Three deformable convolutional kernels are applied to extract hierarchical feature representations, thereby enhancing the Random Fourier Feature (RFF). ; in This represents deformable convolutions with kernel sizes of 3, 5, and 7; the responses of these kernels are fused using learnable soft weights. ; in, These are learnable parameters that control the relative importance of each convolutional kernel scale; To better capture the complex regional variations in haze distribution across different spatial regions, a heterogeneous grouping expert mechanism is proposed. This mechanism constructs perception experts, integration experts, and identification experts through multiple dimensions and operates within different statistical domain feature spaces, thereby achieving refined feature representation.
[0006] In the aforementioned heterogeneous grouping expert mechanism, specifically, the expert set is divided into three statistical groups: Each group consists of multiple parallel sub-experts; the mechanism includes two key processes: (1) intra-group adaptive modeling and (2) inter-group dynamic selection and fusion; (1) Intra-group adaptive modeling aims to achieve adaptive fusion and modulation of features within each statistical group; to this end, the entire expert set is divided into three heterogeneous statistical expert groups, each group is responsible for learning the statistical features of different haze distributions. (2) Intergroup dynamic selection and fusion: Since the three heterogeneous statistical expert groups have captured different aspects of feature perception, direct splicing or averaging may lead to feature bias. Therefore, an intergroup dynamic selection mechanism is introduced to achieve adaptive fusion of features between statistical fields.
[0007] The above-described within-group adaptive modeling process is as follows: exist middle, These correspond to the perception, integration, and discrimination statistical spaces, respectively; each group is responsible for modeling haze characteristics at different levels: Focus on capturing low-frequency trends. Emphasizing the integration of globally consistent features, while The modeling of local biases and structural details has been enhanced; to clearly represent these differences from a statistical perspective, for each group Associated pair of statistical computations This is used to extract different statistical features, and the process can be defined as follows: ; Here, avg, std, and max represent average pooling, standard deviation pooling, and max pooling, respectively, used to capture haze features of different statistical dimensions.
[0008] Each of the above statistical expert groups Internally, construct Parallel statistical sub-experts Each sub-expert shares features. And apply a pair of statistical calculators To extract feature representations in different statistical spaces: ; in This represents the features obtained after applying statistical calculations.
[0009] The aforementioned inter-group dynamic selection mechanism specifically includes: Aggregation characteristics of each statistical expert group First, the corresponding group attention weights are obtained through global encoding: ; in, This represents the normalized inter-group dynamic weights. The outputs of the perception expert, ensemble expert, and discriminator expert are represented by GAP, which stands for Global Average Pooling. Then, the group outputs are weighted and fused according to these weights to obtain the final feature representation. ; in This represents the weight of the k-th sub-expert in expert group g. This represents the output of the k-th sub-expert in expert group g. This represents the weight of expert g in the group, and at the same time .
[0010] The aforementioned perception experts aim to capture the global distribution and low-frequency trends of haze, emphasizing large-scale uniformity and smooth spatial transitions. To this end, each sub-expert learns a structured representation within an independent statistical domain by jointly modeling two complementary cues: boundary response features of the haze-clearness transition and global trend features describing the large-scale smooth haze distribution; integrating these cues yields a stable haze distribution representation for each sub-expert.
[0011] After extracting the statistical features, they are passed as input to the cross-attention mechanism to help the model more effectively capture the global trend of haze distribution and its boundary transitions, thereby enhancing the model's ability to perceive haze information in different regions; this process is expressed as: ; Here, Projector represents depthwise convolution and pointwise convolution operations; after obtaining and After feature identification, the next step is to apply a cross-attention mechanism, which enables the model to more effectively capture the global trend of haze distribution and its boundary transitions, enhancing the model's ability to perceive haze information in different regions.
[0012] in This represents the output of the k-th sub-expert within group P. Since each sub-expert corresponds to a different statistical domain, the cross-attention mechanism establishes a global dependency between boundary transition cues and global trend cues, fusing them in a complementary manner to generate a more stable and accurate representation of haze distribution.
[0013] The aforementioned ensemble experts focus on achieving globally consistent feature fusion. To achieve this, they combine multiple statistical perspectives, enabling the model to capture both local and global features of haze. Specifically, each sub-expert handles local variations and global trends in haze, ensuring the model effectively combines these two features to provide a comprehensive representation of haze distribution. Once the statistical features are extracted, they serve as input for subsequent processing steps. This allows the model to dynamically capture and combine haze distribution features from both local and global perspectives. ; ; in, This represents the fusion feature map generated by the k-th sub-expert. Through The model employs dynamically calculated adaptive weights that balance the contributions of local and global features, allowing it to adjust its focus on haze features based on the input context. To obtain a comprehensive haze distribution, the model integrates local and global features through a feature fusion step based on the weighted contributions of local and global features, thereby generating the final haze distribution representation. ; in, This represents the output of the k-th sub-expert within group I.
[0014] The aforementioned identification experts focus on capturing local differences and structural details in the distribution of haze, particularly the boundaries and transitions between haze and clear areas; these experts aim to capture the fine-grained spatial characteristics of haze; subsequently, the extracted statistical features are used as conditional information to predict fine-grained modulation parameters, which can be expressed as: ; ; in, This represents the result of the k-th expert in group D after the downsampling operation; This indicates a downsampling operation. This represents the k-th expert projection module, used to generate the modulation weight matrix. and bias terms for each sub-expert ; To mitigate perception bias caused by the heterogeneity of haze distribution and enhance the consistency of haze information across different statistical perspectives, a cross-statistical modulation mechanism is introduced, particularly within each sub-expert. Specifically, two types of statistical features within the sub-expert serve as mutual conditional inputs, generating modulation parameters through cross-interaction; the form is as follows: ; ; ; in, and These are the modulated features, and Up represents the upsampling operation. To further enable the model to adaptively select the sub-expert that best matches the current haze feature distribution, a gating network is introduced to assign different weights to each sub-expert based on the input representation, thereby dynamically selecting the appropriate expert. The process is defined as follows: ; ; ; GAP and GMP represent global average pooling and global max pooling, respectively, used to extract global feature descriptors. , and This represents a fully connected layer and the Sigmoid activation function. and This is a learnable weight matrix; next, a Top-K operation is applied to select the top m experts with the largest weights, and zero weights are assigned to the remaining experts; finally, the Sigmoid function is applied to obtain the expert weights: .
[0015] The specific process of Step 2 above is as follows: To more accurately depict the distribution of smog in the generated alert information, and inspired by the fact that smog degradation is mainly manifested in the amplitude spectrum, AASG-Loss is introduced to optimize the distribution of the alert information in the amplitude domain: ; Where h and w represent coordinates in the spatial domain, and u and v correspond to coordinates in the Fourier domain; symbols Represents the inverse Fourier transform; complex components in Fourier space. It can be represented as amplitude components. and phase components : ; ; Where R(x) and I(x) represent the real and imaginary parts of X(u,v), respectively. Based on the magnitude representation, the AASG loss is defined as: ; in, This represents the i-th layer feature extracted from the pre-trained VGG-19. This represents the hierarchical weight coefficient.
[0016] By reconstructing the perceptual loss in the amplitude spectrum, the optimization process is guided in the Fourier amplitude domain, enabling the model to better learn cue representations consistent with haze distribution. The proposed AASG-Loss demonstrates good performance in both physical interpretability and spectral guidance. It optimizes the distribution of the generated cue information to better align with the energy characteristics of haze in the frequency-amplitude domain, thereby guiding the model to learn cue representations that are physically consistent and perceptually reflect haze.
[0017] The present invention discloses a method for constructing a remote sensing image dehazing network based on heterogeneous grouping expert prompts (HGEP) module and adaptive amplitude-guided AASG loss. By proposing a heterogeneous grouping expert prompt HGEP module and an adaptive amplitude-guided AASG loss, the present invention can integrate multi-dimensional haze degradation features and accurately capture key information of haze areas, thereby significantly improving the quality, detail recovery and color fidelity of dehazed images, and overcoming the problems of feature confusion and inaccurate prompts in existing methods. Attached Figure Description
[0018] The present invention will be further described below with reference to the accompanying drawings and embodiments: Figure 1 This is an overall network structure diagram of an embodiment of the present invention; Figure 2 for Figure 1 The structural diagram of HGEP is shown in the expert's suggestion for heterogeneous grouping. Detailed Implementation
[0019] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0020] like Figure 1 and 2 As shown, a method for constructing a remote sensing image dehazing network based on heterogeneous grouping expert prompts includes the following steps: Step S1: Construct a heterogeneous grouped expert prompting (HGEP) module. Through multi-dimensional distributed perception and heterogeneous expert aggregation mechanism, this module clarifies and models statistical changes related to haze, enhancing the network's robustness and dehazing ability. This module guides the network to capture key features of haze areas and improves dehazing accuracy through modeling by multiple expert groups in different statistical spaces.
[0021] Step S2: Introduce the adaptive amplitude-guided loss AASG-Loss. The loss function of the guided loss AASG-Loss optimizes the network's learning in the frequency-amplitude domain, enabling the network to effectively learn cue representations related to haze. The cue representations maintain consistency across multiple frequency sub-bands and statistical spaces, ensuring that the network's perception of haze areas is more accurate, thereby improving the image dehazing effect.
[0022] Step S1 specifically includes: Haze degradation in remote sensing images exhibits significant nonlinearity, spatial heterogeneity, and complex distribution variations. These characteristics hinder deep networks from accurately capturing the spatial distribution and dynamic statistical features of haze. To address these challenges, a Heterogeneous Grouping Expert Hint Module (HGEP) is proposed. This module explicitly models the statistical variations associated with haze through multidimensional distribution perception and heterogeneous expert aggregation mechanisms. This design enables HGEP to generate physically consistent and statistically robust cues, thereby providing enhanced guidance for effective haze removal. Assumptions As the input features, the input is first processed through a Random Fourier Feature (RFF) layer to capture the nonlinear variations in the haze distribution, mapping it to a low-dimensional explicit space, as shown in the following formula: ; Where the bias parameter From uniform distribution Mid-sampling, weight vector From Gaussian distribution Mid-sampling. Here, This represents the feature dimension of the projection. This transformation explicitly transforms the input features... By mapping to a low-dimensional explicit embedding space, the nonlinear changes in haze distribution can be captured, providing a solid foundation for subsequent spatial adaptive modeling.
[0023] To fully model the spatial misalignment and non-uniformity of haze distribution and enhance the model's ability to perceive degradation patterns at multiple scales, a multi-scale deformable convolutional structure is introduced. This structure adaptively adjusts the sampling position, flexibly capturing spatial variations and the refinement of haze boundary regions. Three deformable convolutional kernels are applied to extract hierarchical feature representations, thereby enhancing the Random Fourier Feature Free (RFF). ; in This represents deformable convolutions with kernel sizes of 3, 5, and 7; the responses of these kernels are fused using learnable soft weights. ; in, These are learnable parameters that control the relative importance of each convolutional kernel scale.
[0024] To better capture the complex regional variations in haze distribution across different spatial regions, a heterogeneous grouping expert mechanism is proposed. This mechanism constructs perceptual experts across multiple dimensions and operates within different statistical domain feature spaces, thereby achieving refined feature representation. Specifically, the expert set is divided into three statistical groups: Each group consists of multiple parallel sub-experts. The mechanism comprises two key processes, as described below: (1) Intra-group adaptive modeling: This process aims to achieve adaptive fusion and modulation of features within each statistical group. To this end, the entire expert set is divided into three heterogeneous statistical expert groups, each responsible for learning the statistical features of different haze distributions.
[0025] exist middle, These correspond to the perception, integration, and discrimination statistical spaces, respectively. Each group is responsible for modeling haze characteristics at different levels: Focus on capturing low-frequency trends. Emphasizing the integration of globally consistent features, while The modeling of local biases and structural details has been enhanced. To clearly represent these differences from a statistical perspective, for each group... Associated pair of statistical computations This is used to extract different statistical features. The process can be defined as follows: ; Here, avg, std, and max represent average pooling, standard deviation pooling, and max pooling, respectively, used to capture haze features of different statistical dimensions.
[0026] In each group Inside, it was built Parallel statistical sub-experts Each sub-expert shares features. And apply a pair of statistical calculators To extract feature representations in different statistical spaces: ; in This represents the features obtained after applying statistical calculations.
[0027] (a) Perception Experts: Perception experts aim to capture the global distribution and low-frequency trends of haze, emphasizing large-scale uniformity and smooth spatial transitions. To this end, each sub-expert learns a structured representation within an independent statistical domain by jointly modeling two complementary cues: boundary response features of the haze-clearness transition and global trend features describing the large-scale smooth haze distribution. Integrating these cues yields a stable haze distribution representation for each sub-expert.
[0028] After extracting the statistical features, they are passed as input to the cross-attention mechanism, helping the model more effectively capture the global trend of haze distribution and its boundary transitions, thereby enhancing the model's ability to perceive haze information in different regions. This process can be expressed as: ; Here, Projector represents depthwise convolution and pointwise convolution operations; after obtaining and After feature extraction, the next step is to apply a cross-attention mechanism, which enables the model to more effectively capture the global trend of haze distribution and its boundary transitions, enhancing the model's ability to perceive haze information in different regions.
[0029] ; in This represents the output of the k-th sub-expert within group P. Since each sub-expert corresponds to a different statistical domain, the cross-attention mechanism establishes a global dependency between boundary transition cues and global trend cues, fusing them in a complementary manner to generate a more stable and accurate representation of haze distribution.
[0030] (b) Ensemble Experts: Ensemble experts focus on achieving globally consistent feature fusion. To achieve this, ensemble experts combine multiple statistical perspectives, enabling the model to capture both local and global features of haze. Specifically, each sub-expert handles local variations and global trends in haze, ensuring the model effectively combines these two features to provide a comprehensive representation of haze distribution. Once the statistical features are extracted, they serve as input for subsequent processing steps. This allows the model to dynamically capture and combine haze distribution features from both local and global perspectives: ; ; in, This represents the fusion feature map generated by the k-th sub-expert. Through Dynamically calculated adaptive weights are applied. These weights balance the contributions of local and global features, allowing the model to adjust its focus on haze features based on the input context. To obtain a comprehensive haze distribution, the model integrates local and global features through a feature fusion step based on the weighted contributions of local and global features, thereby generating the final haze distribution representation. ; in, This represents the output of the k-th sub-expert within group I.
[0031] (c) Discriminant Experts: Discriminant experts focus on capturing local differences and structural details in the distribution of haze, particularly the boundaries and transitions between haze and clear areas. These experts aim to capture the fine-grained spatial characteristics of haze. Subsequently, the extracted statistical features are used as conditional information to predict fine-grained modulation parameters, which can be expressed as: ; ; in, This represents the result of the k-th expert in group D after the downsampling operation. This indicates a downsampling operation. This represents the k-th expert projection module, used to generate the modulation weight matrix. and bias terms (For each sub-expert).
[0032] To mitigate perception bias caused by the heterogeneity of haze distribution and enhance the consistency of haze information across different statistical perspectives, a cross-statistical modulation mechanism is introduced, particularly within each sub-expert. Specifically, two types of statistical features within the sub-expert serve as mutual conditional inputs, generating modulation parameters through cross-interaction. Its form is as follows: ; ; ; in, and These are the modulated features, and Up represents the upsampling operation. To further enable the model to adaptively select the sub-expert that best matches the current haze feature distribution, a gating network is introduced to assign different weights to each sub-expert based on the input representation, thereby dynamically selecting the appropriate expert. This process is defined as follows: ; ; ; GAP and GMP represent global average pooling and global max pooling, respectively, used to extract global feature descriptors. , and This represents a fully connected layer and the Sigmoid activation function. and This is a learnable weight matrix. Next, a Top-K operation is applied to select the top m experts with the largest weights, and zero weights are assigned to the remaining experts. Finally, the Sigmoid function is applied to obtain the expert weights: ; (2) Dynamic Selection and Fusion Between Groups: Since the three statistical groups P, I, and D capture different aspects of feature perception, direct splicing or averaging may lead to feature bias. Therefore, a dynamic selection mechanism between groups is introduced to achieve adaptive fusion of features across statistical domains. Specifically, the aggregated features of each group... First, the corresponding group attention weights are obtained through global encoding: ; in, This represents the normalized inter-group dynamic weights. The outputs of the perception expert, ensemble expert, and discriminator expert are represented by GAP, which stands for Global Average Pooling. Then, the group outputs are weighted and fused according to these weights to obtain the final feature representation. ; in This represents the weight of the k-th sub-expert in expert group g. This represents the output of the k-th sub-expert in expert group g. This represents the weight of expert g in the group, and at the same time .
[0033] Step S2 specifically includes: To enable the generated prompts to more accurately depict the distribution of smog, and inspired by the fact that smog degradation is mainly manifested in the amplitude spectrum, AASG-Loss is introduced to optimize the distribution of prompts in the amplitude domain.
[0034]
[0035] Here, h and w represent coordinates in the spatial domain, while u and v correspond to coordinates in the Fourier domain. (Symbols) This represents the inverse Fourier transform. The complex components in Fourier space. It can be represented as amplitude components. and phase components :
[0036]
[0037] Where R(x) and I(x) represent the real and imaginary parts of X(u,v), respectively. Based on the magnitude representation, the AASG loss is defined as:
[0038] in, This represents the i-th layer feature extracted from the pre-trained VGG-19. The hierarchical weight coefficients are represented. By reconstructing the perceptual loss in the amplitude spectrum, the optimization process is guided in the Fourier amplitude domain, enabling the model to better learn cue representations consistent with haze distribution. The proposed AASG-Loss exhibits good performance in both physical interpretability and spectral guidance. It optimizes the distribution of the generated cue information to be more consistent with the energy characteristics of haze in the frequency-amplitude domain, thereby guiding the model to learn cue representations that are physically consistent and perceptually reflect haze.
[0039] Example 1: Parameter settings To verify the effectiveness of the proposed dehazing method, this study used the RICE and RSID datasets, with all images having a resolution of 512×512. The RICE dataset consists of two subsets, RICE1 and RICE2, where RICE1 contains 500 image pairs and RICE2 contains 450 image pairs, both sourced from Landsat 8 OLI / TIRS, and randomly divided into training and testing sets in a 9:1 ratio. The RSID dataset contains 1000 image pairs, with 900 pairs used for training and 100 pairs for testing, also with a resolution of 512×512.
[0040] This experiment was conducted on an NVIDIA RTX 3090 GPU and implemented using the PyTorch deep learning framework. During training, the AdamW optimizer was used with a learning rate of 0.0001 and a batch size of 2. Momentum decay exponents were set to β1=0.9 and β2=0.999. The initial learning rate was set to 0.001, and CosineAnnealingLR was used to dynamically adjust the learning rate. During training, the input image was randomly cropped to a size of 256×256. To better evaluate the proposed method, Peak Signal-to-Noise Ratio (PSNR) and Structure Similarity Index (SSIM) were used to quantitatively evaluate the dehazing results of all algorithms. Higher PSNR and SSIM values indicate that the dehazed image is closer to the true ground truth (GT).
[0041] Experimental results Table 1. Comparison results of the method of the present invention with each other on various dehazed image datasets.
[0042] Table 1 shows the performance comparison results of the proposed remote sensing image dehazing method based on heterogeneous grouping expert prompts on three dehazed image datasets: RICE1, RICE2, and RSID. In the experiments, three mainstream dehazing algorithms—DCP, FFA, and DEA—were selected as comparison methods. Experimental results show that the proposed method significantly outperforms the comparison algorithms in all evaluation metrics.
Claims
1. A method for constructing a remote sensing image dehazing network based on heterogeneous grouping expert prompts, characterized in that, Includes the following steps: Step 1: Construct a heterogeneous grouped expert prompt HGEP module. The HGEP module guides the network to capture the characteristics of hazy areas and improves the accuracy of defogging by modeling multiple expert groups in different statistical spaces. Step 2: Introduce Adaptive Amplitude Guided Loss (AASG-Loss). The loss function of AASG-Loss optimizes the network's learning in the frequency-amplitude domain, enabling the network to learn cue representations related to haze. The cue representations maintain consistency across multiple frequency sub-bands and statistical spaces, ensuring that the network's perception of haze areas is more accurate, thereby improving the image dehazing effect.
2. The method for constructing a remote sensing image dehazing network based on heterogeneous grouping expert prompts according to claim 1, characterized in that, The specific process of Step 1 is as follows: Assumption As the input features, the input is first processed through a Random Fourier Feature (RFF) layer to capture the nonlinear variations in the haze distribution, mapping it to a low-dimensional explicit space, as shown in the following formula: ; Where the bias parameter From uniform distribution Mid-sampling, weight vector From Gaussian distribution Mid-sampling; Indicates the feature dimension of the projection; This transformation explicitly transforms the input features Mapping to a low-dimensional explicit embedding space allows us to capture the nonlinear changes in haze distribution, providing a foundation for subsequent spatial adaptive modeling. A multi-scale deformable convolutional structure is introduced to adaptively adjust the sampling position, flexibly capturing spatial variations and refining the boundary regions of haze. Three deformable convolutional kernels are applied to extract hierarchical feature representations, thereby enhancing the random Fourier transform feature (RFF). ; in This represents deformable convolutions with kernel sizes of 3, 5, and 7; the responses of these kernels are fused using learnable soft weights. ; in, These are learnable parameters that control the relative importance of each convolutional kernel scale; To better capture the complex regional variations in haze distribution across different spatial regions, a heterogeneous grouping expert mechanism is proposed. This mechanism constructs perception experts, integration experts, and identification experts across multiple dimensions, and allows them to operate within different statistical domain feature spaces, thereby achieving refined feature representation.
3. The method for constructing a remote sensing image dehazing network based on heterogeneous grouping expert prompts according to claim 2, characterized in that, In the aforementioned heterogeneous grouping expert mechanism, the expert set is divided into three statistical groups: Each group consists of multiple parallel sub-experts; It includes two processes: (1) intra-group adaptive modeling and (2) inter-group dynamic selection and fusion; (1) Intra-group adaptive modeling aims to achieve adaptive fusion and modulation of features within each statistical group; To this end, the entire expert set was divided into three heterogeneous statistical expert groups, each responsible for learning the statistical characteristics of different haze distributions; (2) Intergroup dynamic selection and fusion: Since the three heterogeneous statistical expert groups have captured different aspects of feature perception, direct splicing or averaging may lead to feature bias. Therefore, an intergroup dynamic selection mechanism is introduced to achieve adaptive fusion of features between statistical fields.
4. The method for constructing a remote sensing image dehazing network based on heterogeneous grouping expert prompts according to claim 3, characterized in that, The intra-group adaptive modeling process is as follows: exist middle, These correspond to the perception, integration, and discrimination statistical spaces, respectively; each group is responsible for modeling haze characteristics at different levels: Focus on capturing low-frequency trends. Emphasizing the integration of globally consistent features, while The modeling of local biases and structural details has been enhanced; to clearly represent these differences from a statistical perspective, for each group Associated pair of unified computational sub-units This is used to extract different statistical features, and the process can be defined as follows: ; Here, avg, std, and max represent average pooling, standard deviation pooling, and max pooling, respectively, used to capture haze features of different statistical dimensions.
5. The method for constructing a remote sensing image dehazing network based on heterogeneous grouping expert prompts according to claim 4, characterized in that, Each of the aforementioned statistical expert groups Internally, construct Parallel statistical sub-experts Each sub-expert shares features. And apply a pair of statistical calculators To extract feature representations in different statistical spaces: ; in This represents the features obtained after applying statistical calculations.
6. The method for constructing a remote sensing image dehazing network based on heterogeneous grouping expert prompts according to claim 3, characterized in that, The aforementioned inter-group dynamic selection mechanism specifically includes: Aggregation characteristics of each statistical expert group First, the corresponding group attention weights are obtained through global encoding: ; in, This represents the normalized inter-group dynamic weights. The outputs of the perception expert, ensemble expert, and discriminator expert are represented by GAP, which stands for Global Average Pooling. Then, the group outputs are weighted and fused according to these weights to obtain the final feature representation. ; in This represents the weight of the k-th sub-expert in expert group g. This represents the output of the k-th sub-expert in expert group g. This represents the weight of expert g in the group, and at the same time .
7. The method for constructing a remote sensing image dehazing network based on heterogeneous grouping expert prompts according to claim 2, characterized in that, The perception experts are designed to capture the global distribution and low-frequency trends of haze. To this end, each sub-expert learns a structured representation in an independent statistical domain by jointly modeling two complementary cues: boundary response features of haze-clear transition and global trend features describing large-scale smooth haze distribution. A stable haze distribution representation is obtained for each sub-expert; After extracting the statistical features, they are passed as input to the cross-attention mechanism to enhance the model's ability to perceive haze information in different regions; this process is expressed as: ; Here, Projector represents depthwise convolution and pointwise convolution operations; after obtaining and After feature identification, the next step is to apply a cross-attention mechanism, enabling the model to more effectively capture the global trend of haze distribution and its boundary transitions, thereby enhancing the model's ability to perceive haze information in different regions. ; in This represents the output of the k-th sub-expert within group P. Since each sub-expert corresponds to a different statistical domain, the cross-attention mechanism establishes a global dependency between boundary transition cues and global trend cues, fusing them in a complementary manner to generate a more stable and accurate representation of haze distribution.
8. The method for constructing a remote sensing image dehazing network based on heterogeneous grouping expert prompts according to claim 2, characterized in that, The ensemble experts focus on achieving globally consistent feature fusion. To achieve this, they combine multiple statistical perspectives, enabling the model to capture both local and global features of haze. Specifically, each sub-expert handles local variations and global trends in haze, ensuring the model effectively combines these two features to provide a comprehensive representation of haze distribution. Once the statistical features are extracted, they serve as input for subsequent processing steps. This allows the model to dynamically capture and combine haze distribution features from both local and global perspectives. ; ; in, This represents the fusion feature map generated by the k-th sub-expert. Through The model dynamically calculates the corresponding adaptive weights. To obtain a comprehensive haze distribution, it integrates local and global features through a feature fusion step based on the weighted contributions of local and global features, thereby generating the final haze distribution representation. ; in, This represents the output of the k-th sub-expert within group I.
9. The method for constructing a remote sensing image dehazing network based on heterogeneous grouping expert prompts according to claim 2, characterized in that, The aforementioned identification experts focus on capturing local differences and structural details in the distribution of haze, particularly the boundaries and transitions between haze and clear areas; these experts aim to capture the fine-grained spatial characteristics of haze; subsequently, the extracted statistical features are used as conditional information to predict fine-grained modulation parameters, which can be expressed as: ; ; in, This represents the result of the k-th expert in group D after the downsampling operation; This indicates a downsampling operation. This represents the k-th expert projection module, used to generate the modulation weight matrix. and bias terms for each sub-expert ; To mitigate perception bias caused by the heterogeneity of haze distribution and enhance the consistency of haze information across different statistical perspectives, a cross-statistical modulation mechanism is introduced. Specifically, two types of statistical features within a sub-expert serve as mutual conditional inputs, generating modulation parameters through cross-interaction; the form is as follows: ; ; ; in, and These are the modulated features, and Up represents the upsampling operation. To further enable the model to adaptively select the sub-expert that best matches the current haze feature distribution, a gating network is introduced to assign different weights to each sub-expert based on the input representation, thereby dynamically selecting the appropriate expert. The process is defined as follows: ; ; ; GAP and GMP represent global average pooling and global max pooling, respectively, used to extract global feature descriptors. , and This represents a fully connected layer and the Sigmoid activation function. and This is a learnable weight matrix; next, a Top-K operation is applied to select the top m experts with the largest weights, and zero weights are assigned to the remaining experts; finally, the Sigmoid function is applied to obtain the expert weights: 。 10. The method for constructing a remote sensing image dehazing network based on heterogeneous grouping expert prompts according to claim 2, characterized in that, The specific process of Step 2 is as follows: To enable the generated alerts to more accurately depict the distribution of smog, and inspired by the fact that smog degradation is primarily characterized by its amplitude spectrum, AASG-Loss is introduced to optimize the distribution of alerts in the amplitude domain: ; Where h and w represent coordinates in the spatial domain, and u and v correspond to coordinates in the Fourier domain; symbols Represents the inverse Fourier transform; complex components in Fourier space. It can be represented as amplitude components. and phase components : ; ; Where R(x) and I(x) represent the real and imaginary parts of X(u,v), respectively; based on the magnitude representation, the AASG loss is defined as: ; in, This represents the i-th layer feature extracted from the pre-trained VGG-19. This represents the hierarchical weight coefficient.