Synthetic aperture radar guided optical remote sensing image defogging method
The dual-branch network with HPDM and PFM modules addresses the challenges of optical-SAR fusion by adaptively weighting and progressively fusing features to enhance dehazing performance.
Patent Information
- Application Number
- CN202510469363.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-07-15
AI Technical Summary
The existing optical-SAR fusion defogging method introduces SAR information in the fogg-free area, resulting in a decline in quality, and the mode difference between optical and SAR data leads to a domain shift that easily leads to direct fogging, affecting the defogging effect.
The optical and SAR features are extracted separately by a dual-branch network, and the haze area is dynamically identified and decoupled through the HPDM module. The PFM module is combined with the progressive feature fusion, and the adaptive weighting mechanism is used to adjust the fusion ratio, retain high confidence information, and realize the precise fusion of cross-modal information.
It effectively avoids the quality of the foggy-free area, improves the fogging effect and model stability, enhances the fogging detail recovery ability in complex scenes, reduces domain offset problems, and improves the visual quality of the fogging image and the applicability of the downstream tasks.
Smart Images

Figure CN120318115A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of remote sensing image processing and computer vision, and particularly relates to a method for dehazing SAR (Synthetic Aperture Radar) - guided optical remote sensing images based on an adaptive state - space model. Background Art
[0002] Optical remote sensing image dehazing is an important task in computer vision, aiming to restore clear images from haze - affected images. Due to the non - uniform distribution of haze in large - scale space, remote sensing image dehazing faces great challenges. With the wide application of optical remote sensing technology in fields such as military, forestry, and agriculture, the degradation of image quality caused by haze seriously affects the performance of various remote sensing analysis tasks. Therefore, obtaining high - resolution, haze - free remote sensing images is crucial for subsequent tasks.
[0003] Early dehazing methods based on the atmospheric scattering model (ASM) relied on artificially set prior information, such as the dark channel prior (DCP) proposed by Han et al. and the saturation line prior proposed by Ling et al. However, the applicability of these methods in complex scenarios is limited, which may lead to incorrect estimation and affect the dehazing effect. In recent years, with the rapid development of deep learning technology, dehazing methods based on convolutional neural networks (CNNs) and Transformers have made significant progress. For example, Li et al. and Chen et al. respectively developed an end - to - end image dehazing network AODNet and GCA - Net to improve the dehazing effect through a self - supervised learning framework. Song et al. proposed the first dehazing network based on Transformer, which achieved good dehazing effects on both natural image and remote sensing image processing. However, CNNs are limited by local receptive fields and are difficult to effectively model global information, while although Transformers have the ability to model globally, their quadratic computational complexity limits the processing efficiency of high - resolution images.
[0004] Single - image defogging methods lack reference information, often prone to artifacts, and difficult to restore clear details when dealing with complex haze. Synthetic aperture radar (SAR) images are not affected by haze and provide important reference information for optical image defogging. However, optical - SAR fusion defogging methods still face two major challenges: (1) Due to the non - uniform characteristics of remote - sensing haze, there may be haze - free areas in the optical image that do not require SAR information. Introducing too much heterogeneous SAR information in these haze - free areas may affect the defogging effect; (2) There are significant modal differences between optical and SAR images. Direct fusion is difficult to fully utilize SAR features and may even lead to a decrease in defogging quality due to noise interference. To address these problems, researchers have tried to integrate SAR information through deep learning to improve defogging results. For example, Li et al. proposed a feature fusion method based on bilinear pooling to enhance the complementarity of optical and SAR features. In addition, Fu et al. used the GAN framework for modal conversion and fusion to fill in the missing information in optical images. However, existing methods are still relatively simple in the fusion strategy and fail to fully exploit the defogging potential of SAR images.
[0005] In summary, optical - SAR fusion defogging methods are of great value in improving the quality of remote - sensing images. The existing SAR - guided defogging methods mainly have the following two key problems: (1) Directly fusing SAR information may lead to a decrease in the image quality of haze - free areas and cannot accurately distinguish haze - affected areas; (2) There are significant modal differences between optical and SAR data. Direct fusion is prone to domain shift, resulting in unstable features and affecting the defogging effect. Summary of the Invention
[0006] The technical problem to be solved by the present invention is how to efficiently fuse optical and SAR information and avoid the degradation of the quality of haze - free areas after introducing SAR information.
[0007] The present invention solves the above - mentioned technical problems by the following technical means: A synthetic aperture radar - guided optical remote - sensing image defogging method, comprising:
[0008] First, a dual - branch network is used to obtain the optical image and the SAR image respectively, and preliminary feature extraction and channel adjustment are performed through convolutional layers, and then deep feature extraction and downsampling are carried out to obtain high - order optical features and SAR features;
[0009] Perform at least one round of feature fusion on the high-order optical features and SAR features. Each round of feature fusion includes separately performing deep feature extraction on the high-order optical features and SAR features. After extraction, the features are each divided into two branches. One branch performs cross-modal feature fusion. During the cross-modal feature fusion process, first calculate the high-order feature difference to remove redundant information, and then adjust the fusion ratio of the high-order features through an adaptive weighting mechanism to retain high-confidence information and achieve fusion. The fused features, the other optical features after extraction, and the other SAR features after extraction are sent to the next round of feature fusion;
[0010] Upsample and perform defogging calculation on the final fused features to restore image details and output the final result.
[0011] As a further optimized technical solution, use a dual-branch network to separately obtain an optical image and a SAR remote sensing image, and perform preliminary feature extraction and channel adjustment through a convolutional layer, and then perform deep feature extraction and downsampling to obtain high-order optical features and SAR features, including:
[0012] The SAR image features are first subjected to deep feature extraction. After extraction, the deep feature data is downsampled in scale to obtain high-order SAR feature Si;
[0013] The optical image is first subjected to deep feature extraction. After extraction, it is divided into two branches: one branch is downsampled in scale to obtain high-order optical feature Ri, and the other branch performs adaptive feature fusion.
[0014] As a further optimized technical solution, perform 2 rounds of feature fusion on the high-order optical features and SAR features, including:
[0015] The high-order SAR feature Si and the high-order optical feature Ri are each subjected to further deep feature extraction and then each generate two feature branches. One of the branches is sequentially input into the first HPDM and the first PFM for cross-modal feature fusion. The first HPDM calculates the difference between the high-order SAR feature Si and the high-order optical feature Ri to remove redundant information, and the first PFM adjusts the fusion ratio of the high-order SAR feature Si and the high-order optical feature Ri through an adaptive weighting mechanism to retain high-confidence information;
[0016] The fused features obtained after the first HPDM and the first PFM processing are added to another branch after further deep feature extraction from the high-order optical feature Ri, and then sent to the optical image channel for further refined deep feature extraction. At the same time, another branch of the SAR image undergoes independent deep feature extraction. One of the two feature branches generated after the refined deep feature extraction in the optical image channel and the output after the independent deep feature extraction enter the second HPDM and the second PFM simultaneously to further extract the high-order difference features between modalities and enhance the complementarity. The features processed by the secondary HPDM and PFM are finally fused with the optical features of the other branch after the refined deep feature extraction and sent to the subsequent steps for upsampling and defogging calculation.
[0017] As a further optimized technical solution, in the HPDM, it includes a fog perception stage and a decoupling stage in sequence. In the fog perception stage, the difference information between the optical features and the SAR features is extracted through feature extraction and the CSS2D module, and the fog and haze information is characterized by the difference information. In the decoupling stage, the difference information is converted into a fog distribution weight map W1 through normalization, a joint gating mechanism, a linear layer, and an activation layer, and W1 is used to decouple the features of the fog area and the non-fog area in the optical and SAR features.
[0018] As a further optimized technical solution, the feature extraction in the fog perception stage specifically includes linear projection, depth convolution, and the SiLU activation function. The linear projection maps the input features to a high-dimensional space, the depth convolution further extracts local features, and the SiLU activation function introduces a non-linear transformation to enhance the expression ability of the features, and finally the output features are obtained.
[0019] The output features are subsequently input into the CSS2D module for modal feature extraction. In the CSS2D module, the linear layer processes the features to generate the parameter matrices A, B, C, and D required for the state space equations of their respective modalities. Subsequently, the matrices A and B are discretized to obtain the discretized matrices of their respective modalities and The discretized matrices are used to iterate the hidden state h k , the C matrix is adjusted through modal difference calculation, and the adjusted C matrix decodes the hidden state h k to obtain the optical modal decoded features and the AR modal decoded features Finally, the absolute difference between the optical modal and SAR modal decoded features is used as the output feature of the CSS2D.
[0020] As a further optimized technical solution, the specific processing process in the CSS2D module includes:
[0021] Generate the parameter matrices A, B, C, and D required for the state space equations of their respective modalities:
[0022] For the RGB modality, generate the matrix Δ rgb , A rgb , B rgb , C rgb , D rgb
[0023] For the SAR modality, generate the matrix Δ sar , A sar , B sar , C sar , D sar ;
[0024] Discretize the matrix A and the matrix B to obtain the discretized matrices of their respective modalities and The process is expressed as:
[0025]
[0026] The discretized matrices are used to iterate the hidden state h k of each modality, and the iteration formula for the hidden state h k is:
[0027]
[0028] Where and respectively represent the input features of the optical modality and the SAR modality of CSS2D at the time step t; the C matrix is adjusted through modal difference calculation to obtain the difference-aware C matrix, and the calculation method is as follows:
[0029] C = |C rgb - C sar |
[0030] The adjusted C matrix decodes the hidden state h k to obtain the output features, and the decoding formula is:
[0031]
[0032] Finally, the absolute difference between the decoded features of the optical modality and the SAR modality is used as the output feature of CSS2D:
[0033]
[0034] As a further optimized technical solution, in the decoupling stage, first normalize the difference information extracted in the perception stage, adopt a joint gating mechanism, and generate joint gating information by using the information between optical features and SAR features at the same time to control the flow of the difference information. After the difference information adjusted by the joint gating is projected through a linear layer, it is then activated using a Sigmoid activation function to obtain a weight map W1 representing the haze distribution. Finally, W1 and 1 - W1 are respectively assigned to the SAR features and the optical features, and the information in the non-haze area of the SAR features and the information in the haze area of the optical features are respectively suppressed through pixel-wise multiplication to obtain the decoupled optical features and SAR features.
[0035] As a further optimized technical solution, the decoupled optical features and SAR features are gradually fused in two stages according to the information quality. Let the decoupled optical features and SAR features be F0 and F s :
[0036] In the first stage, F0 and F s are concatenated along the channel dimension, and then the number of channels is adjusted through a convolutional layer to generate a rough fusion feature F cf ;
[0037] In the second stage, first use a Sigmoid activation function to convert F cf into a weight map W2, and assign W2 and 1 - W2 to the optical feature F o and the SAR feature F s , respectively, and obtain F ao and F as through pixel-wise multiplication, respectively;
[0038] Concatenate F ao , F as and F cf along the channel dimension, and use VSS and a convolutional layer for fusion to output a fusion feature F out .
[0039] As a further optimized technical solution, the upsampling and defogging calculation of the final fusion feature to restore image details and output the final result includes:
[0040] The fusion feature F out is upsampled to restore the resolution of the feature map;
[0041] The upsampled features are adaptively feature-fused;
[0042] The features after adaptive feature fusion are then subjected to final depth feature extraction;
[0043] Finally, the features after the final depth feature extraction are adjusted in channels through a convolutional layer and converted into the final dehazed image.
[0044] As a further optimized technical solution, the depth feature extraction adopts a DMBlock module, and the DMBlock module includes a first LN unit, a VSS unit, a first Scale unit, a second LN unit, an MLP unit, and a second Scale unit;
[0045] The features after the preliminary feature extraction are simultaneously input into the first LN unit and the first Scale unit. The first LN unit is used to normalize the features, and the first Scale unit transforms the feature scale. The normalized features are input into the VSS unit. The VSS unit extracts the global context information of the features with linear computational overhead. The output of the VSS unit and the output of the first Scale unit are added to obtain the fused features;
[0046] The fused features simultaneously enter the second LN unit and the second Scale unit. The second LN unit continues to normalize the data and then inputs it into the MLP unit. The MLP unit is responsible for enhancing the local detail features and improving the expression ability of the features through non-linear transformation. The second Scale unit further adjusts the feature scale. Finally, the features output by the second Scale unit are added to the output of the MLP.
[0047] The advantages of the present invention are as follows:
[0048] 1. The present invention proposes a novel SAR-guided remote sensing image dehazing system. This system improves the clarity of optical remote sensing images through a progressive fog perception and decoupling fusion strategy, and effectively avoids the quality degradation problem caused by existing methods in fog-free areas. Specifically, the present invention designs two core modules: a fog perception and decoupling module (HPDM) and a progressive fusion module (PFM). The HPDM dynamically identifies the haze area by analyzing the feature differences between optical and SAR images, and finely decouples the fog information, making the dehazing process more targeted. The PFM introduces an adaptive adjustment mechanism based on feature quality during the feature fusion process, and fuses cross-modal information in stages, thereby effectively alleviating the domain shift problem and improving the dehazing effect.
[0049] 2. Compared with the prior art, the present invention has significant advantages in terms of defogging effect, feature fusion strategy, and model stability. Traditional single-image defogging methods are difficult to handle large-scale and highly non-uniform haze distributions, resulting in limited defogging effects. In contrast, the present invention can effectively enhance the fog-obscured details in optical images under complex scenarios by combining SAR fog-free information for guidance. In addition, existing SAR-guided defogging methods often lead to a decrease in the quality of the fog-free region after introducing SAR information. However, the present invention realizes refined fog information decoupling through the pixel-level dynamic weighting mechanism of the HPDM module, ensuring that the fog-free region maintains its original quality. At the same time, the PFM module adopts a progressive fusion strategy, which can more effectively integrate cross-modal information and reduce the domain shift problem compared with traditional single-step fusion methods, greatly improving the visual quality of the final defogged image and its applicability to downstream tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 is the overall framework diagram of the synthetic aperture radar-guided optical remote sensing image defogging system in an embodiment of the present invention;
[0051] Figure 2 is the architecture diagram of the fog perception and decoupling module (HPDM) in an embodiment of the present invention;
[0052] Figure 3 is the architecture diagram of the progressive fusion module (PFM) in an embodiment of the present invention;
[0053] Figure 4 is the architecture diagram of the DMBlock in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0054] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0055] The state space model (SSM) demonstrates superior performance in processing sequence data, and its low linear complexity can effectively reduce the computational cost. The latest research shows that the SSM structure can break through the computational bottleneck of the Transformer while maintaining the global modeling ability. Gu et al. developed a defogging network combining SSM, achieving a good balance between computational complexity and defogging quality. With the rise of Mamba, some Mamba-based defogging methods have achieved remarkable results in research. Therefore, by combining the advantages of SSM and introducing Mamba into the defogging task, the computational amount can be reduced while the defogging quality is improved.
[0056] The present invention proposes a synthetic aperture radar (SAR)-guided optical remote sensing image dehazing system (DehazeMamba) based on Mamba, which adopts a progressive haze decoupling and fusion strategy to effectively improve the dehazing quality and fusion stability. The SAR-guided optical remote sensing image dehazing system mainly includes two key modules:
[0057] (1) Haze Perception and Decoupling Module (HPDM), by calculating the differences between optical and SAR features, dynamically identifies the haze-affected areas, realizes accurate separation of the haze areas, and avoids interference of SAR information on the haze-free areas. HPDM utilizes the global receptive field and multi-directional sequence modeling ability of Mamba to more accurately extract the optical-SAR feature differences, and realizes efficient haze perception and separation.
[0058] (2) Progressive Fusion Module (PFM), adaptively adjusts the fusion method according to the feature quality, and adopts a two-stage fusion strategy to maximize the dehazing effect. In the first stage, PFM first uses the low-resolution global features for preliminary optical-SAR feature matching to reduce the influence of the modality differences brought by direct fusion, ensure the preliminary alignment of the features at a large scale, and reduce the domain shift risk. In the second stage, PFM adaptively adjusts the fusion ratio of the optical-SAR information according to the feature quality to ensure that the final fusion result reaches the optimal.
[0059] Specifically, the present invention proposes a synthetic aperture radar-guided optical remote sensing image dehazing method, including the following steps:
[0060] Step 1: Multi-modal feature extraction and downsampling (D-Layer processing)
[0061] In the dehazing processing of the optical image and the SAR image, first, preprocess and extract features from the two input images. Use a dual-branch network to obtain the optical image and the SAR image respectively, and perform preliminary feature extraction and channel adjustment through the convolutional layer (Conv). Subsequently, these features are input into their respective D-Layers for deep feature extraction and downsampling.
[0062] In the SAR image channel, the D-Layer mainly consists of the first DMBlock and the first D-sample. First, the SAR image features are subjected to deep feature extraction through the first DMBlock. After being extracted by the first DMBlock, the deep feature data directly enters the first D-sample for scale downsampling to obtain the high-order SAR feature Si to ensure the consistency of subsequent processing.
[0063] In the optical image channel, the D-Layer structure is basically the same as that in the SAR channel, mainly composed of the second DMBlock and the second D-sample. However, after the second DMBlock, the features are divided into two branches: one branch enters the second D-sample for scale downsampling to obtain the high-order optical feature Ri, and the other branch enters the first SKFusion (Selective Attention Fusion Module) for adaptive feature fusion. The first SKFusion enhances the attention to key regions through an adaptive feature selection mechanism, suppresses irrelevant information, and ensures that effective optical features are maximally retained during the defogging process.
[0064] The finally extracted high-order optical feature Ri and high-order SAR feature Si are input into the subsequent Fusion Layer for fog perception and defogging calculation.
[0065] As Figure 4 shown, the DMBlock module consists of multiple computing units, including the first LN (LayerNorm) unit, VSS (Visual State Space Module) unit, first Scale unit, second LN unit, MLP (Multi-Layer Perceptron) unit, and second Scale unit.
[0066] The features after preliminary feature extraction are simultaneously input into the first LN unit and the first Scale unit. Among them, the first LN unit is used to normalize the features, eliminate the differences in feature distribution, and improve the model stability; the first Scale unit transforms the feature scale to dynamically adjust the intensity of the original features in the residual connection. The normalized features are input into the VSS unit, which extracts the global context information of the features with linear computational overhead and improves the model's ability to globally perceive the haze area. The output of the VSS unit and the output of the first Scale unit are added to obtain the fused feature, which allows the VSS unit to focus on learning haze-related information while adaptively retaining the original information.
[0067] The fused features simultaneously enter the second LN unit and the second Scale unit. The second LN unit continues to normalize the data and then inputs it into the MLP unit. The MLP unit is responsible for enhancing local detail features and improving the feature expression ability through non-linear transformation. The second Scale unit further adjusts the feature scale to ensure feature consistency. Finally, the features output by the second Scale unit are added to the output of the MLP to ensure stable gradient propagation while retaining the original information, making the defogging calculation more robust.
[0068] Step 2: Feature Interaction and Fusion Enhancement (Processed by Fusion Layer)
[0069] After the D-Layer processing is completed, the high-order SAR feature Si and the high-order optical feature Ri are respectively input into the Fusion Layer for feature fusion and enhancement. The Fusion Layer mainly consists of DMBlock, HPDM, and PFM, and is responsible for cross-modal feature alignment and fusion, enhancing the image defogging ability. The specific process is as follows:
[0070] The high-order SAR feature Si and the high-order optical feature Ri obtained after D-Layer processing are respectively input into the third DMBlock and the fourth DMBlock for further feature extraction. The structure of the DMBlock is the same as that of the DMBlock introduced in step 1. After passing through the DMBlock, the high-order SAR feature Si and the high-order optical feature Ri each generate two feature branches, and one of the branches is sequentially input into the first HPDM and the first PFM for cross-modal feature fusion. The first HPDM calculates the difference between the high-order SAR feature Si and the high-order optical feature Ri, removes redundant information, and enhances the ability to distinguish haze regions. The first PFM (Progressive Fusion Module) adjusts the fusion ratio of the high-order SAR feature Si and the high-order optical feature Ri through an adaptive weighting mechanism, retains high-confidence information, and achieves optimal fusion.
[0071] The fused feature obtained after being processed by the first HPDM and the first PFM is added to the optical feature output by the fourth DMBlock and then sent to the fifth DMBlock of the optical image for refinement. At the same time, the other branch of the SAR image enters the sixth DMBlock for independent feature extraction to enhance the independence of the modal features. One of the two feature branches generated by the fifth DMBlock processing the optical feature, together with the output generated by the sixth DMBlock processing the SAR feature, then enter the second HPDM and the second PFM at the same time to further extract the high-order difference features between modalities and enhance complementarity. The feature processed by the secondary HPDM and PFM is finally fused with the optical feature processed by the fifth DMBlock and sent to the subsequent U-Layer for upsampling and defogging calculation.
[0072] As Figure 2 shown, in the HPDM module, it includes a fog perception stage and a decoupling stage that are carried out in sequence. In the fog perception stage, the difference information between the optical feature and the SAR feature is extracted through feature extraction and the CSS2D module, and the haze information is characterized by the difference information. In the decoupling stage, the difference information is converted into a fog distribution weight map W1 through normalization, a joint gating mechanism, a linear layer, and an activation layer, and W1 is used to decouple the features of the fog region and the non-fog region in the optical and SAR features.
[0073] First, the fog perception stage is carried out, and in this stage, the semantic difference information between optical features and SAR features is extracted. Since the same target in the aligned optical image and SAR image has similar semantics, and the haze in the optical image will destroy this similarity, the haze information can be characterized by the semantic difference information between optical features and SAR features. In this stage, the optical features and SAR features respectively go through feature extraction operations (Extract), which specifically include linear projection (Linear), depth convolution (DConv), and the SiLU activation function: linear projection maps the input features into a high-dimensional space, depth convolution further extracts local features, and the SiLU activation function introduces a non-linear transformation to enhance the expressive ability of the features, and finally the output features are obtained This process can be expressed as (index i represents the "RGB" or "SAR" modality):
[0074]
[0075] The features processed by the feature extraction operation are then input into the CSS2D module for modal feature extraction. In the CSS2D module, the linear layer processes the input features to generate the parameter matrices required for the state space equation:
[0076] For the RGB modality, matrices Δ rgb , A rgb , B rgb , C rgb , D rgb
[0077] For the SAR modality, matrices Δ sar , A sar , B sar , C sar , D sar
[0078] Subsequently, the matrices A and B are discretized to obtain the discretized matrices of their respective modalities and This process can be expressed as:
[0079]
[0080] The discretized matrix is used to iterate the hidden state h k of its respective modality, and the iteration formula for the hidden state h k is:
[0081]
[0082] where and They respectively represent the input features of the optical modality and the SAR modality of CSS2D at time step t. The C matrix is adjusted through modal difference calculation to obtain the difference-aware C matrix, and the calculation method is as follows:
[0083] C = |C rgb - C sar |
[0084] The adjusted C matrix decodes the hidden state h k to obtain the output features. The decoding formula is:
[0085]
[0086] Finally, the absolute difference between the decoded features of the optical modality and the SAR modality is used as the output feature of CSS2D to highlight the difference between modalities:
[0087]
[0088] In the decoupling stage, the difference information extracted in the perception stage is first normalized to ensure numerical stability. To optimize the information flow control, a joint gating mechanism (Gate operation) is adopted, and the information between the optical features and the SAR features is used to generate joint gating information to more finely control the flow of difference information. After the difference information adjusted by the joint gating is projected through a linear layer, it is then activated using the Sigmoid activation function to obtain the weight map W1 representing the haze distribution. Finally, W1 and 1 - W1 are respectively assigned to the SAR features and the optical features, and the information in the non-haze region of the SAR features and the haze region of the optical features is suppressed through per-pixel multiplication to obtain the decoupled optical features and SAR features. W1 has two functions: multiplying W1 by the SAR features can suppress the information in the non-haze region of the SAR features to ensure that as little SAR information that is not helpful for defogging is introduced when fusing the optical-SAR information subsequently; multiplying 1 - W1 by the optical features can known the features in the haze region of the optical features, that is, weaken the haze information to prepare for subsequent defogging.
[0089] As Figure 3 shown, in the PFM (Progressive Fusion Module), the decoupled optical features and SAR features are progressively fused in two stages according to the information quality to improve the defogging effect.
[0090] Let the decoupled optical features and SAR features be F0 and F s . In the first stage, F0 and F s are concatenated along the channel dimension, and then the number of channels is adjusted through a convolutional layer (Conv) to generate the coarsely fused feature F cf .
[0091] In the second stage, to adaptively adjust the feature weights to match the quality differences of different modalities, first, the Sigmoid activation function is used to convert F cf into the weight map W2. Since F cf contains information from both optical and SAR features, the weight W2 can measure the relative information of the quality strengths of optical and SAR features. The larger W2 is, the higher the quality of the optical features compared to the SAR features. W2 and 1 - W2 are respectively assigned to the optical feature F o and the SAR feature F s , and by pixel-wise multiplication, F ao and F as are obtained respectively, which can dynamically adjust the weights of the two during subsequent fusion according to the relative information of the quality of optical and SAR features.
[0092] After obtaining F ao and F as , F ao , F as and F cf are concatenated along the channel dimension, and are fused using VSS (Visual State Space Module) and convolutional layers to output the fused feature F out . In VSS, the input features first pass through two branches respectively. In the first branch, the input features pass through a linear layer and a SiLU activation function in sequence to generate gating information. In the second branch, the input features pass through a linear layer, a depth convolution, a SiLU activation function, SS2D (2D Selective Scanning Operation), and a normalization layer in sequence to extract global context information. After the outputs of the two branches are multiplied, a linear layer refines the multiplied result to obtain the output feature of VSS. After the concatenated features are processed by VSS, a convolutional layer adjusts the number of channels to obtain the output fused feature F out .
[0093] Step 3. Upsampling and Final Image Reconstruction (U-Layer Processing)
[0094] After the Fusion Layer processing is completed, the fused feature F out is input into the U-Layer for upsampling and defogging calculation to restore image details and output the final result. First, the fused feature F out output by the Fusion Layer enters the U-Layer and is upsampled by U-sample to restore the resolution of the feature map, gradually approaching the size of the original image. Upsampling can expand the spatial information of the feature map, enabling the defogged image to retain more detailed information and improve the overall clarity.
[0095] The features after upsampling enter the second SKFusion. The second SKFusion assigns weights according to the importance of feature channels, enhances key features and suppresses useless information. Through the second SKFusion, the most representative features can be adaptively selected, enabling subsequent processing to more accurately restore image details and further strengthening the useful information in the defogging process.
[0096] The features processed by the second SKFusion are then sent to the seventh DMBlock for final feature extraction and detail enhancement. The structure of the DMBlock is the same as the DMBlock structure introduced in step 1. The seventh DMBlock extracts global information through the Visual State Space module (VSS), enhances local details using the Multi-Layer Perceptron (MLP), and combines residual connections to ensure stable gradient propagation. During this process, the texture details and structural information of the feature map are further optimized, making the defogged image more natural and clear.
[0097] Finally, the features after U-Layer processing are adjusted in channels through a convolutional layer (Conv) and converted into the final defogged image. The Conv is responsible for optimizing the channel distribution of the features, adapting the pixel value range of the image to the final output format, and ensuring a high-quality fog-free optical remote sensing image.
[0098] Based on the complementary characteristics of optical remote sensing and Synthetic Aperture Radar (SAR) data, the present invention constructs a collaborative defogging framework based on the Mamba architecture. Its working principle follows the technical route of "fog information perception-modal feature decoupling-cross-domain progressive fusion", and realizes image enhancement through the collaborative action of the fog perception and decoupling module (HPDM) and the progressive feature fusion module (PFM).
[0099] In the data preprocessing stage, the high-order optical image feature R i and the SAR image feature S i are subjected to feature extraction through the DMBlock module in the dual-branch Mamba encoder. This module innovatively combines the Visual State Space module (VSS) and the Multi-Layer Perceptron (MLP). While capturing global context information through the VSS, it uses the MLP to enhance local detail features. A learnable parameter controller is introduced between the VSS and the MLP, combined with the residual connection mechanism, which not only ensures the stability of the gradient flow but also realizes the dynamic adjustment of feature channels, laying a foundation for subsequent processing.
[0100] Based on the high-level features extracted by DMBlock, the HPDM module generates a fog distribution weight map W1 through Sigmoid transformation, which quantitatively characterizes the spatial distribution characteristics of fog concentration. By decoupling the strategy of assigning the W1 weight to the SAR feature and the complementary weight (1-W1) to the optical feature, the pixel-level separation of fog pollution information and clean features is achieved. In this process, the CSS2D component constructs the parameter matrix (Δ rgb , A rgb , B rgb , C rgb , D rgb , Δ sar , A sar , B sar , C sar , D sar ), after discretization, the hidden state h is generated k Finally, the C matrix is adjusted through differential calculation, and the absolute difference operation is combined to highlight the difference characteristics between modes, thus completing the accurate analysis of fog information.
[0101] To solve the domain offset problem of optical and SAR features, the PFM module adopts a two-stage progressive fusion strategy. First, the decoupled optical feature F0 and SAR feature F s Perform channel splicing and generate coarse fusion feature F through convolution dimensionality reduction cf ; Then, the adaptive weight map W2 is generated through Sigmoid transformation, and the contribution ratio of the dual-mode features is dynamically adjusted to achieve fine alignment of the feature space. In this process, the microwave penetration characteristics provided by the SAR data complement the multi-spectral information of the optical data, and the detailed features obscured by the haze are gradually restored through iterative optimization at the feature level.
[0102] In order to improve the performance of the algorithm, the AdamW optimizer is used in conjunction with the cosine annealing strategy during the optimization process. The model convergence stability is enhanced by dynamically adjusting the learning rate. The parameter update formula of the AdamW optimizer is as follows:
[0103]
[0104] θ t is the model parameter, η t is the learning rate, and are the first-order moment and second-order moment estimates of the gradient, ∈ is the smoothing term, and λ is the weight decay coefficient. The learning rate η t Dynamic adjustment through cosine annealing strategy:
[0105]
[0106] Among them, η max and ηmin are the maximum and minimum values of the learning rate, respectively, where T cur is the current training step, and T max is the total number of training steps.
[0107] The loss function design adopts a spatial-frequency dual-domain constraint mechanism: in the spatial domain the loss ensures the integrity of the overall structure of the image, and in the frequency domain the loss (weight coefficient 0.1) enhances the restoration of high-frequency details. The total loss function is defined as:
[0108]
[0109] In the spatial domain the loss and in the frequency domain the loss are calculated respectively as:
[0110]
[0111] I pred and I gt are the predicted image and the ground truth image respectively, denotes the Fourier transform, and Ν is the total number of pixels. Their combined effect ensures the balanced performance of the dehazed image in terms of texture sharpness and radiometric fidelity.
[0112] The whole process forms a complete technical closed-loop from fog area detection, feature decoupling to cross-modal fusion, and finally outputs high-quality remote sensing images that not only maintain a natural visual effect but also have rich ground object details.
[0113] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A method for defogging synthetic aperture radar-guided optical remote sensing images, characterized in that: Including: First, use a dual-branch network to obtain optical images and SAR images respectively, perform preliminary feature extraction and channel adjustment through convolutional layers, then perform deep feature extraction and downsampling to obtain high-order optical features and SAR features; Perform at least one round of feature fusion on the high-order optical features and SAR features. Each round of feature fusion includes separately performing deep feature extraction on the high-order optical features and SAR features. After extraction, the features are each divided into two branches. One branch performs cross-modal feature fusion. During the cross-modal feature fusion process, first calculate the high-order feature difference to remove redundant information, and then adjust the fusion ratio of the high-order features through an adaptive weighting mechanism to retain high-confidence information and achieve fusion. The fused features, the fused features obtained after fusing with the other optical feature after extraction, and the other SAR feature after extraction are sent to the next round of feature fusion; Upsample the final fused features and perform defogging calculation to restore image details and output the final result.
2. The method for dehazing synthetic aperture radar-guided optical remote sensing images according to claim 1, wherein: Using a dual-branch network to obtain optical images and SAR remote sensing images respectively, perform preliminary feature extraction and channel adjustment through convolutional layers, then perform deep feature extraction and downsampling, and obtaining high-order optical features and SAR features includes: The SAR image features first perform deep feature extraction. After extraction, the deep feature data is downsampled in scale to obtain high-order SAR feature Si; The optical image first performs deep feature extraction. After extraction, it is divided into two branches: one branch performs downsampling in scale to obtain high-order optical feature Ri, and the other branch performs adaptive feature fusion.
3. The synthetic aperture radar-guided optical remote sensing image defogging method according to claim 1, wherein: Perform 2 rounds of feature fusion on the high-order optical features and SAR features, including: The high-order SAR feature Si and the high-order optical feature Ri are respectively further deep feature extracted and each generate two feature branches. One of the branches is sequentially input into the first HPDM and the first PFM for cross-modal feature fusion. The first HPDM calculates the difference between the high-order SAR feature Si and the high-order optical feature Ri to remove redundant information, and the first PFM adjusts the fusion ratio of the high-order SAR feature Si and the high-order optical feature Ri through an adaptive weighting mechanism to retain high-confidence information; The fused features obtained after being processed by the first HPDM and the first PFM are added to the other branch after the high-order optical feature Ri is further deep feature extracted and then sent to the optical image channel for further refined deep feature extraction. At the same time, the other branch of the SAR image performs independent deep feature extraction. One of the two feature branches generated after the refined deep feature extraction in the optical image channel and the output after the independent deep feature extraction enter the second HPDM and the second PFM at the same time to further extract the high-order difference features between modalities and enhance complementarity. The features processed by the secondary HPDM and PFM are finally fused with the other optical feature after the refined deep feature extraction and sent to the subsequent steps for upsampling and defogging calculation.
4. The method for removing haze from synthetic aperture radar-guided optical remote sensing images according to claim 3, wherein: In HPDM, it includes a fog perception stage and a decoupling stage carried out in sequence. In the fog perception stage, the difference information between the optical features and SAR features is extracted through feature extraction and the CSS2D module, and the fog and haze information is characterized by the difference information. In the decoupling stage, the difference information is converted into a fog distribution weight map W1 through normalization, a joint gating mechanism, a linear layer, and an activation layer, and W1 is used to decouple the features of the fog and non-fog regions in the optical and SAR features.
5. The method for dehazing optical remote sensing images guided by synthetic aperture radar according to claim 4, characterized in that: The feature extraction in the fog perception stage specifically includes linear projection, depth convolution, and the SiLU activation function. The linear projection maps the input features to a high-dimensional space. The depth convolution further extracts local features. The SiLU activation function introduces a non-linear transformation to enhance the expressive ability of the features, and finally the output features are obtained Output feature Subsequently, it is input into the CSS2D module for modal feature extraction. In the CSS2D module, the linear layer processes the features to generate the parameter matrices A, B, C, and D required for the state space equations of their respective modes. Subsequently, the matrices A and B are discretized to obtain the discretized matrices of their respective modes and The discretized matrices are used to iterate the hidden state h of their respective modes k , and the C matrix is adjusted through modal difference calculation. The adjusted C matrix decodes the hidden state h k to obtain the optical modal decoding feature and the AR modal decoding feature Finally, the absolute difference between the optical modal and SAR modal decoding features is used as the output feature of CSS2D.
6. The method for removing haze from synthetic aperture radar-guided optical remote sensing images according to claim 5, characterized in that: The specific processing process in the CSS2D module includes: Generating the parameter matrices A, B, C, and D required for the state space equations of their respective modalities: For the RGB modality, generate the matrix Δ rgb 、A rgb 、B rgb 、C rgb 、D rgb For the SAR mode, generate matrices Δ sar , A sar , B sar , C sar , D sar ; Discretize matrix A and matrix B to obtain the discretized matrices of their respective modes and The process is expressed as: The discretized matrix is used to iterate the hidden state h of each mode k , and the iterative formula for the hidden state h k is as follows: Among them, and respectively represent the input features of the optical mode and the SAR mode of CSS2D at the time step t; Adjusting the C matrix through modal difference calculation to obtain a difference-aware C matrix, and the calculation method is as follows: C = |C rgb -C sar | The adjusted C matrix decodes the hidden state h k to obtain the output features, and the decoding formula is: Finally, the absolute difference between the decoded features of the optical modality and the SAR modality is used as the output feature of CSS2D:
7. The synthetic aperture radar-guided optical remote sensing image defogging method according to any one of claims 4 to 6, characterized in that: In the decoupling stage, first normalize the difference information extracted in the perception stage, adopt a joint gating mechanism, and generate joint gating information by using the information between the optical features and SAR features to control the flow of the difference information. After the difference information adjusted by the joint gating is projected through a linear layer, it is then activated using the Sigmoid activation function to obtain a weight map W1 representing the fog and haze distribution. Finally, W1 and 1 - W1 are respectively assigned to the SAR features and optical features, and the information in the non-fog and haze regions of the SAR features and the fog and haze regions of the optical features is suppressed respectively through pixel-wise multiplication to obtain the decoupled optical features and SAR features.
8. The method for dehazing synthetic aperture radar-guided optical remote sensing images according to claim 3, wherein: In PFM, the decoupled optical features and SAR features are progressively fused in two stages according to the information quality. Let the decoupled optical features and SAR features be F0 and F respectively s : In the first stage, F0 and F s are concatenated along the channel dimension, and then the number of channels is adjusted through a convolutional layer to generate the rough fusion feature F cf ; In the second stage, first use the Sigmoid activation function to convert F cf into the weight map W2, and assign W2 and 1-W2 to the optical feature F o and the SAR feature F s , respectively, and obtain F ao and F as ; Concatenate F ao , F as and F cf along the channel dimension, and fuse them using VSS and convolutional layers to output the fused feature F out .
9. The method for removing haze from synthetic aperture radar-guided optical remote sensing images according to claim 1, characterized in that: The upsampling and dehazing calculation of the final fused features to restore image details and output the final result includes: Fused feature F out Perform upsampling to restore the resolution of the feature map; Performing adaptive feature fusion on the upsampled features; The features after adaptive feature fusion are then subjected to final depth feature extraction; Finally, the features after the final depth feature extraction are adjusted in channels through a convolutional layer and converted into the final dehazed image.
10. The method for removing haze from synthetic aperture radar-guided optical remote sensing images according to any one of claims 1, 2, 3, and 9, characterized in that: The depth feature extraction uses the DMBlock module, and the DMBlock module includes a first LN unit, a VSS unit, a first Scale unit, a second LN unit, an MLP unit, and a second Scale unit; The features after preliminary feature extraction are simultaneously input into the first LN unit and the first Scale unit. The first LN unit is used to normalize the features, and the first Scale unit transforms the feature scale. The normalized features are input into the VSS unit, and the VSS unit extracts the global context information of the features with linear computational overhead. The output of the VSS unit and the output of the first Scale unit are added to obtain a fused feature; The fused features enter the second LN unit and the second Scale unit simultaneously. The second LN unit continues to normalize the data and then inputs it into the MLP unit. The MLP unit is responsible for enhancing the local detail features and improving the expression ability of the features through non-linear transformation. The second Scale unit further adjusts the feature scale. Finally, the features output by the second Scale unit are added to the output of the MLP.
Citation Information
Cited By
Image defogging method, device, equipment, medium and program product
CN120807334A