A SAR image segmentation method and system based on frequency-space prior migration

By using the frequency-space prior transfer method, optical priors are introduced into SAR images through a dual-path coding structure and an optically guided cross-domain fusion module. Combined with multi-resolution anisotropic coherent units and coherent sensing gated skip connections, the boundary ambiguity and scale inconsistency problems in SAR image segmentation are solved, achieving high-precision and robust target segmentation results.

CN122199953APending Publication Date: 2026-06-12HANGZHOU DIANZI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HANGZHOU DIANZI UNIV
Filing Date
2026-03-03
Publication Date
2026-06-12

AI Technical Summary

Technical Problem

Existing SAR image segmentation techniques suffer from problems such as boundary ambiguity caused by speckle noise, propagation of jump-connection pseudo-features, scale inconsistency, and lack of semantic priors in the optical domain. These issues can lead to jagged boundaries, structural breaks, and category confusion in complex scenes.

Method used

A frequency-space prior transfer-based approach is adopted to obtain multi-scale SAR structural features and pseudo-optical prior features through a dual-path joint coding structure. Geometric semantic constraints are introduced under SAR single-mode input through an optically guided cross-domain fusion module. Combined with multi-resolution anisotropic coherent units and coherent sensing gated jump connections, the diffusion of coherent pseudo-features is suppressed, and the directional continuity and multi-scale sensing capabilities are enhanced.

Benefits of technology

Achieving high-precision and robust target segmentation under SAR single-mode input improves boundary accuracy and structural integrity, reduces the risk of false detection and missed detection in complex scenarios, and enhances the reliability and deployability of segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122199953A_ABST
    Figure CN122199953A_ABST
Patent Text Reader

Abstract

The application discloses a SAR image segmentation method and system based on frequency-space prior migration. A dual-driven spectrum is used to construct a cross-modal GAN confrontation segmentation network, multi-scale SAR structural features and multi-scale pseudo-optical prior features are obtained through a double-path joint coding structure, and feature alignment and fusion are performed on a multi-scale level through an optical guiding cross-domain fusion module. The input features of the decoder upsampling module are direction consistency enhanced and multi-scale completed through a multi-resolution anisotropic coherent unit. The SAR image segmentation result is obtained through the decoder upsampling module, and a coherent perception gating jump connection is introduced in the decoding process. The application learns the optical domain structural semantic distribution through the confrontation generation branch, injects the SAR feature space through the optical guiding cross-domain fusion module, alleviates the SAR geometric semantic prior missing problem, and does not depend on strict registration optical-SAR paired data, thereby improving the deployability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent interpretation of remote sensing images and computer vision. Specifically, it relates to a coherent structure SAR (Synthetic Aperture Radar) image segmentation method, system and device based on adversarial optical prior transfer and multi-resolution anisotropic decoding. It belongs to the cross-technical direction of deep learning, cross-modal prior transfer, GAN adversarial learning, SAR coherent noise suppression and multi-scale semantic segmentation. Background Technology

[0002] SAR possesses the advantages of all-day, all-weather imaging, enabling stable acquisition of surface observation information under complex weather conditions such as nighttime, clouds, rain, fog, and haze. Therefore, it has irreplaceable engineering value in fields such as remote sensing mapping, resource surveys, marine monitoring, and emergency management. Current SAR image interpretation tasks, including building extraction, road and river network segmentation, ship and port target detection, land-sea segmentation, disaster range assessment (such as floods / landslides), and vegetation and bare land monitoring, all place higher demands on pixel-level segmentation accuracy. Unlike optical remote sensing images, which primarily rely on visible light reflection to form surface textures, SAR images mainly characterize the scattering characteristics of electromagnetic waves and ground objects. They are influenced by factors such as imaging geometry, incident angle, polarization, and surface roughness. Furthermore, SAR imaging is a coherent imaging mechanism, and its grayscale and phase information often do not have a one-to-one correspondence with the true semantic boundaries. Specifically, the same type of ground features may exhibit drastically different brightness and darkness distributions under different scattering mechanisms (specular scattering, volume scattering, dual reflection, etc.), while different types of ground features may produce similar texture responses under specific incident conditions, resulting in the phenomenon of "different spectra for the same object, and similar spectra for different objects." At the same time, SAR commonly suffers from geometric distortions such as layover, shadowing, and foreshortening, causing misalignment or loss of target outlines from the true geometric boundaries. Therefore, SAR semantic segmentation models often face inherent difficulties such as "lack of prior geometric semantics" and "weak boundary discernibility," leading to problems like jagged boundaries, structural breaks, and category confusion in complex scenes.

[0003] In summary, existing SAR segmentation techniques still suffer from the following core bottlenecks: speckle noise and false edge propagation, scale inconsistency and scale drift of similar targets, and lack of optical domain geometric / semantic priors.

[0004] To address the aforementioned issues, a method is needed that can: introduce structural priors from the optical domain even when only SAR is input, thereby alleviating the problems of weak SAR semantic boundaries and lack of geometric priors; suppress the spread of coherent pseudo-features during the encoding-decoding jump link transfer process, achieving "selective detail recovery" rather than "indiscriminate high-frequency amplification"; simultaneously enhance directional continuity and multi-scale perception capabilities during the decoding stage, improving connectivity for directional structural targets such as roads, shorelines, and ships, and robustness to scale drift; thereby achieving stable and reliable SAR semantic segmentation in complex scenarios and meeting the comprehensive requirements of engineering applications for accuracy, robustness, and deployability. Summary of the Invention

[0005] This invention aims to solve the problems of boundary ambiguity caused by speckle noise, propagation of jump-connection pseudo-features, scale inconsistency, and lack of semantic prior in the optical domain in existing SAR image segmentation. It provides a SAR image segmentation method and system based on frequency-space prior transfer, which achieves high-precision and robust target segmentation under SAR single-mode input conditions.

[0006] The technical solution adopted by this invention to solve its technical problem is as follows:

[0007] In a first aspect, embodiments of this application provide a SAR image segmentation method based on frequency-space prior migration, comprising the following steps:

[0008] Step 1: Pre-training dataset construction:

[0009] Several SAR images and their corresponding optical images were acquired as a pre-training dataset.

[0010] Step 2: Preprocess the training dataset.

[0011] Step 3: Construct a dual-driven spectral-oriented cross-modal GAN ​​adversarial segmentation network:

[0012] The network includes: a dual-path joint coding structure, an optically guided cross-domain fusion module, a coherent sensing gated skip connection, a multi-resolution anisotropic coherent unit, and several decoder upsampling modules.

[0013] Step 4: Obtain multi-scale SAR structural features and multi-scale pseudo-optical prior features through a dual-path joint coding structure:

[0014] The dual-path joint coding structure includes a segmentation backbone encoder branch and an optical prior adversarial generation branch. The segmentation backbone encoder branch is used to extract multi-scale SAR structural features from the SAR image. The optical prior adversarial generation branch generates multi-scale pseudo-optical prior features based on the input SAR image.

[0015] Step 5: The SAR structural features and pseudo-optical prior features are aligned and fused at multiple scale levels through the optically guided cross-domain fusion module to obtain the corresponding fused features, realizing the "transfer of optical priors to SAR latent space" and still obtaining geometric semantic constraints under SAR single-mode input.

[0016] Step 6: Enhance the directional consistency and multi-scale completion of the input features of the decoder upsampling module by using a multi-resolution anisotropic coherent unit, thereby highlighting the real structure and suppressing speckle noise interference.

[0017] Step 7: Obtain the SAR image segmentation result through the decoder upsampling module.

[0018] Step 8: In the decoding process, a coherent sensing gated skip connection is introduced. The shallow features of the same level are combined with the high-level semantic features obtained by the decoder upsampling module after wavelet decomposition and anti-spoofing processing. The resulting fused features are then used to perform gated filtering on the original shallow features. The filtered features are added to the output of the decoder upsampling module and processed by a multi-resolution anisotropic coherent unit before being used as the input of the next decoder upsampling module.

[0019] Step 9, Model Training and Inference:

[0020] Based on the preprocessed training dataset, an end-to-end training approach is adopted, using the AdamW optimizer and a preset learning rate strategy to complete the training; during the inference stage, only SAR images are input, and corresponding pixel-level segmentation results can be output.

[0021] In one possible implementation, an existing SAR image segmentation dataset is used as the training dataset, and it is divided into a training set, a validation set, and a test set according to a predetermined ratio to ensure the fairness and stability of model training and evaluation. Each SAR image sample in the SAR image segmentation dataset is accompanied by a corresponding pixel-level annotation file.

[0022] Normalization, size unification (cropping / scaling to a fixed resolution), random cropping, random intensity perturbation, and affine enhancement operations such as rotation / translation / scaling are sequentially performed on the SAR images in the training set to improve the model's generalization ability to scale drift and noise perturbation.

[0023] In one possible implementation, the generator's encoder is obtained by pre-training the generator and discriminator on a GAN network using SAR images and optical images of the same region from a pre-training dataset, which serves as the optical prior adversarial generation branch.

[0024] In one possible implementation, the optically guided cross-domain fusion module includes the following steps:

[0025] 1-1. Align the spatial resolution and channel dimension of the same-layer features of the segmentation backbone encoder branch and the optical prior adversarial generation branch.

[0026] 1-2. Channel compression is performed on the aligned features, and mutual exclusive attention weights that sum to one are generated for the segmentation backbone encoder branch and the optical prior adversarial generation branch through normalized weights. The attention weights are then weighted channel-by-channel with the compressed features of the corresponding branches, and the weighted features of the two branches are added together to obtain the global perception features.

[0027] 1-3. Spatial-level adaptive fusion: Based on global perception features, spatial context information is further extracted through cascaded convolution and activation functions to generate fine-grained pixel-level spatial attention weights. Subsequently, the pixel-level spatial attention weights are applied to SAR structural features and pseudo-optical prior features, and adaptive spatial fusion of cross-level features is achieved through weighted aggregation to obtain fused features.

[0028] In one possible implementation, the coherent sensing gated hop connection includes the following steps:

[0029] 2-1. Perform discrete wavelet transform on the shallow features output by the optically guided cross-domain fusion module, which has the same spatial resolution as the high-level semantic features output by the decoder upsampling module, to decompose the shallow features into low-frequency sub-bands and high-frequency sub-bands, so as to distinguish between real contours and coherent pseudo-textures.

[0030] 2-2. Pack the features of each subband (merge the subband dimension into the channel dimension) and refine the subband features through depthwise separable convolution and learnable scaling function to suppress high-frequency pseudo edges and pseudo textures.

[0031] 2-3. Perform inverse wavelet transform on the refined sub-band features to reconstruct shallow features. Then, perform residual fusion between the reconstructed shallow features and the basic features obtained by convolving the original shallow features to obtain enhanced shallow features.

[0032] 2-4. The high-level semantic features obtained by the decoder upsampling module with the same spatial resolution as the shallow features are convolved and added to the enhanced shallow features for fusion. After normalization and Sigmoid, a gating signal is generated. The gating signal is multiplied pixel by pixel with the original shallow features to achieve precise control of the proportion of skip connection information to obtain the gated features.

[0033] In one possible implementation, the multi-resolution anisotropic coherent unit includes anisotropic sensing branches and multi-resolution sensing branches, as detailed below:

[0034] 3-1. The anisotropic sensing branch extracts the directional consistency structure of the segmented target by directional convolution in the horizontal and vertical directions of the input features (the maximum scale fusion feature output by the optically guided cross-domain fusion module or the result of adding the output of the previous coherent sensing gated skip connection with the output of the previous decoder upsampling module), enhances the continuous contour and smooths random high-frequency speckle noise, and finally outputs an anisotropic coherent weight map.

[0035] 3-2. The multi-resolution perception branch aggregates the input features into different receptive fields through parallel multi-scale depth separable convolutions. The small kernel focuses on the details of small targets, while the large kernel focuses on macrostructures such as roads / building complexes, and finally outputs multi-resolution features.

[0036] 3-3. The anisotropic coherent weight map output by the anisotropic sensing branch is multiplied element-wise with the input features to achieve structured enhancement of the original features. Subsequently, the enhanced features are added element-wise with the multi-resolution features output by the multi-resolution sensing branch, and then residually connected with the input features to finally obtain output features that have both directional continuity and multi-scale robustness.

[0037] In one possible implementation, the first decoder upsampling module takes the maximum-scale fusion features output by the optically guided cross-domain fusion module after processing by the multi-resolution anisotropic coherence unit as input. The output of the first decoder upsampling module is then processed by the multi-resolution anisotropic coherence unit and input into the next decoder upsampling module, and so on. Through layer-by-layer upsampling, an accurate SAR image segmentation result is obtained through the last decoder upsampling module.

[0038] Secondly, embodiments of this application provide a SAR image segmentation system based on frequency-space prior transfer, comprising the following modules:

[0039] Preprocessing module: An existing SAR image segmentation dataset is used as the training dataset, and it is pre-divided into training, validation, and test sets to ensure the fairness and stability of model training and evaluation. Each SAR image sample in the SAR image segmentation dataset is accompanied by a corresponding pixel-level annotation file.

[0040] Normalization, size unification (cropping / scaling to a fixed resolution), random cropping, random intensity perturbation, and affine enhancement operations such as rotation / translation / scaling are sequentially performed on the SAR images in the training set to improve the model's generalization ability to scale drift and noise perturbation.

[0041] Feature extraction module: Obtains multi-scale SAR structural features and multi-scale pseudo-optical prior features through a dual-path joint coding structure.

[0042] The dual-path joint coding structure includes a segmentation backbone encoder branch and an optical prior adversarial generation branch. The segmentation backbone encoder branch is used to extract multi-scale SAR structural features from the SAR image. The optical prior adversarial generation branch generates multi-scale pseudo-optical prior features based on the input SAR image.

[0043] Feature fusion module: Through the optically guided cross-domain fusion module, SAR structural features and pseudo-optical prior features are aligned and fused at multiple scale levels to obtain corresponding fused features, realizing "optical prior to SAR latent space migration", and still obtaining geometric semantic constraints under SAR single-mode input.

[0044] Feature enhancement module: By using multi-resolution anisotropic coherent units to enhance the directional consistency and multi-scale completion of the input features of the decoder upsampling module, the module highlights the real structure and suppresses speckle noise interference.

[0045] Segmentation output module: Obtains SAR image segmentation results through the decoder upsampling module.

[0046] In the decoding process, coherent sensing gated skip connections are introduced. Shallow features at the same level are decomposed by wavelet to remove falsehoods and then combined with high-level semantic features obtained by the decoder upsampling module element by element. The resulting fused features are then used to screen the original shallow features. The screened features are added to the output of the decoder upsampling module and processed by a multi-resolution anisotropic coherent unit before being used as the input of the next decoder upsampling module.

[0047] Training and Reasoning Module:

[0048] Based on the preprocessed training dataset, an end-to-end training approach is adopted, using the AdamW optimizer and a preset learning rate strategy to complete the training; during the inference stage, only SAR images are input, and corresponding pixel-level segmentation results can be output.

[0049] Thirdly, embodiments of this application provide an electronic device, including a processor and a memory;

[0050] The memory is used to store computer programs.

[0051] When the processor executes the program stored in the memory, it implements any of the SAR image segmentation methods described in this application.

[0052] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements any of the SAR image segmentation methods described in this application.

[0053] Fifthly, embodiments of this application provide a computer program product containing instructions that, when run on a computer, cause the computer to execute any of the SAR image segmentation methods described in this application.

[0054] Compared with the prior art, the present invention has at least the following beneficial effects:

[0055] 1. Optical prior injection under SAR single-mode input: The optical domain structure semantic distribution is learned through adversarial generative branch and injected into the SAR feature space through optically guided cross-domain fusion module, which alleviates the problem of missing SAR geometric semantic prior and does not rely on strictly registered optical-SAR paired data, thus improving deployability.

[0056] 2. Suppressing coherent pseudo-feature jump propagation: Combining wavelet decomposition and semantic gating, shallow pseudo-edges / pseudo-textures are filtered out, and only semantically consistent true boundaries are transmitted, significantly improving boundary accuracy and structural integrity.

[0057] 3. Synergistic enhancement of directional continuity and multi-scale robustness: Multi-resolution anisotropic coherent units simultaneously model horizontal / vertical directional structures and multi-scale contexts, improving the ability to continuously segment directional texture targets such as ships, roads, and building complexes, and enhancing adaptability to scale drift.

[0058] 4. Stable generalization in complex scenes: In remote sensing tasks with strong speckle noise and large scale variations, such as densely built areas, port and ship scenes, and sea-land boundaries, it can effectively reduce the risk of false detection, missed detection and boundary breakage, and improve the overall segmentation reliability. Attached Figure Description

[0059] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other embodiments can be obtained based on these drawings.

[0060] Figure 1 This is a schematic diagram of the overall framework structure of the dual-driven spectral cross-modal GAN ​​adversarial segmentation network according to an embodiment of the present invention.

[0061] Figure 2 This is a schematic diagram of the coherent sensing gating jump connection in an embodiment of the present invention.

[0062] Figure 3 This is a schematic diagram of the structure of the multi-resolution anisotropic coherent unit in an embodiment of the present invention.

[0063] Figure 4This section describes the correspondence between SAR images and visible light images, as well as the generation results. (a) is a SAR image from the SARBuD dataset without a corresponding real visible light image; (b) is a visible light style image obtained through a GAN generator; (c)–(e) are results from datasets with corresponding real visible light images, where (c) is the original SAR image, (d) is the result after GAN generator conversion, and (e) is the corresponding real visible light image. Detailed Implementation

[0064] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art based on this application are within the scope of protection of this application.

[0065] This invention provides a SAR image segmentation method, system, and device based on adversarial optical prior transfer and multi-resolution anisotropic decoding. When acquiring SAR images to be segmented, the SAR coherent imaging mechanism introduces multiplicative speckle noise and geometric distortion, resulting in phenomena such as broken target boundaries, pseudo-textures, and pseudo-edge enhancement. Furthermore, similar targets exhibit significant scale drift under different imaging geometries and resolutions, making it difficult for segmentation networks relying solely on SAR echo intensity to obtain stable geometric semantic priors and consistent structural contours. This invention utilizes a dual-driven spectral cross-modal GAN ​​adversarial segmentation network. In the network encoding stage, composed of a dual-path joint encoding structure consisting of a segmentation backbone encoder branch and an optical prior adversarial generation branch, the SAR structural semantic features extracted from the ResNet backbone and the pseudo-optical prior representation generated by the optical prior adversarial generation branch are jointly modeled. An optically guided cross-domain fusion module establishes scale alignment and semantic correspondence at multiple scale levels, thus implicitly introducing optical domain geometric semantic priors to compensate for the structural deficiencies in the SAR representation when only SAR is input. Furthermore, in the decoding stage, which consists of alternating cascades of multiple decoder upsampling modules and multi-resolution anisotropic coherent units, coupled with coherent sensing gated skip connections, a two-stage feature refinement strategy is adopted: First, coherent sensing gated skip connections are introduced to perform wavelet decomposition on shallow features to separate low-frequency contours from high-frequency interference. A selective mask is then generated under the guidance of high-level semantic gate signals to suppress the spread of erroneous boundaries caused by coherent pseudo-features propagating through skip connections. Subsequently, the gated features are input into the multi-resolution anisotropic coherent units. The directional convolution of the anisotropic sensing branches enhances the continuity of the horizontal / vertical structure, and the multi-scale depth-separable convolution of the multi-resolution sensing branches aggregates contextual information from different receptive fields, achieving robust suppression of speckle noise perturbations and adaptive modeling of scale changes. Importantly, this invention does not rely on strictly registered optical-SAR paired inputs during the inference stage; it only requires an input SAR image to output the corresponding pixel-level segmentation results, thus balancing accuracy improvement and engineering deployment feasibility. The specific implementation steps of this method are as follows:

[0066] Step 1: Pre-training dataset construction: Paired SAR-optical image data covering various ground features were obtained and filtered from public remote sensing data sources for pre-training of the generative adversarial network (GAN).

[0067] Step 2: Training Dataset Preprocessing: Obtain SAR images and their corresponding pixel-level annotations for the semantic segmentation task, and construct a training sample set. In this embodiment, public data is used as an example, and SARBuD 1.0 (256×256 sample-level annotations) is used for building extraction. The obtained data is divided into training and test sets proportionally; when validation and parameter tuning are required, a validation set can be further divided. In this embodiment, SARBuD divides the training / test sets by region (e.g., 9:1) and performs normalization and size alignment processing on the obtained samples. Specifically, this includes: normalizing the input SAR image and adjusting it to a fixed resolution. The processed training samples are then enhanced to improve generalization ability. In this embodiment, random cropping, random intensity perturbation, and affine transformations such as rotation / translation / scaling can be used to increase the diversity of SAR texture, scale, and geometric deformation.

[0068] Step 3: Construct a dual-driven spectral-oriented cross-modal GAN ​​adversarial segmentation network:

[0069] The overall network architecture is designed based on a dual-path joint coding structure, an optically guided cross-domain fusion module, a coherent sensing gated skip connection, a multi-resolution anisotropic coherent unit, and several decoder upsampling modules.

[0070] Step 4: Obtain multi-scale SAR structural features and multi-scale pseudo-optical prior features through a dual-path joint coding structure:

[0071] The dual-path joint coding structure consists of a ResNet34-based segmentation backbone encoder branch and an optical prior adversarial generative branch based on a GAN generator encoder (U-Net structure). The segmentation backbone encoder branch is used to extract multi-scale SAR structural features from SAR images. The optical prior adversarial generative branch uses SAR images from a pre-training dataset and optical images of the same region to pre-train the generator and discriminator on a GAN network to obtain the generator's encoder. This allows the generator to learn the structural semantic distribution in the optical domain and generate multi-scale pseudo-optical prior features based on the input SAR image.

[0072] Step 5: The SAR structural features and pseudo-optical prior features are aligned and fused at multiple scale levels through the optically guided cross-domain fusion module to obtain the corresponding fused features, thereby achieving geometric semantic constraints under SAR single-mode input.

[0073] The optically guided cross-domain fusion module first aligns the spatial resolution and channel dimension of the same-layer features of the segmentation backbone encoder branch and the optical prior adversarial generation branch. Then, it performs channel compression on the aligned features and generates mutually exclusive attention weights that sum to one for both the segmentation backbone encoder branch and the optical prior adversarial generation branch using normalized weights. These weights are then weighted channel-by-channel with the compressed features of the corresponding branches. Finally, the weighted features from both branches are added together and fused to obtain the globally perceived features.

[0074] Finally, spatial-level adaptive fusion is performed based on global perception features. Spatial context information is further extracted through cascaded convolution and activation functions to generate fine-grained pixel-level spatial attention weights, which are then applied to SAR structural features and pseudo-optical prior features. The final fused features are obtained by weighted aggregation to achieve adaptive spatial fusion of cross-level features.

[0075] The optically guided cross-domain fusion module is implemented as follows:

[0076] For the feature pairs at each scale output by the dual-path joint coding structure, firstly, the SAR structural features at the same scale are... With pseudo-optical prior features Apply Convolution is used to achieve channel dimensionality reduction, resulting in compressed features. and The number of channels is uniformly set to Then, the two compressed features are integrated by summing element-wise to obtain the aggregated feature. Next, an activation operation is performed based on the aggregated feature. Generate global context description vector The activation operation involves performing a non-linear mapping on the aggregated feature encodings through two consecutive convolutions, as expressed by the following formula:

[0077]

[0078] In the formula: This is a global average pooling operation; For batch normalization; It is the ReLU activation function. This is an element-wise multiplication. The resulting vector... Global contextual information of the fused features was captured for subsequent attention weight calculation.

[0079] global context description vector Transformed into specific channel attention weights in two branches (split branch and GAN prior branch) and By The transformation is then converted into a tensor with branching dimensions, allowing the weights of each branch to be separated. The transformation process is expressed by the following formula:

[0080]

[0081] In the formula: This is a reshape operation; This represents the number of branches. It is a split and reshape operation.

[0082] The learned channel attention weights are multiplied by the compressed features and then summed to obtain features enhanced by the selective kernel mechanism. During fusion, competitive collaboration across branches is achieved, enabling the network to dynamically adjust the intensity of optical prior information introduction in different scenarios (such as water bodies, mountains, and cities), thus realizing collaborative feature modeling of "optical guidance and SAR dominance".

[0083]

[0084] To achieve adaptive spatial fusion, based on features enhanced by a selective kernel mechanism, weights are used... and A cascaded convolution and a softmax operation generate fine-grained pixel-level spatial attention weights. These pixel-level spatial attention weights are used in... and Weighted aggregation is performed on top. The aggregation result is passed through a convolutional layer (with weights of ). Further refinement yielded the fusion characteristics. The formula is expressed as follows:

[0085]

[0086] Step 6: Enhance the directional consistency and multi-scale completion of the input features of the decoder upsampling module by using a multi-resolution anisotropic coherent unit, thereby highlighting the real structure and suppressing speckle noise interference.

[0087] This unit comprises an anisotropic sensing branch and a multi-resolution sensing branch. First, the input features (the sum of the output of the previous coherent sensing gated skip connection and the output of the previous decoder upsampling module) are processed. Then, the anisotropic sensing branch extracts the directional consistency structure of the segmented target through horizontal and vertical directional convolutions, enhancing continuous contours and smoothing random high-frequency noise, outputting an anisotropic coherent weight map. Simultaneously, the multi-resolution sensing branch establishes feature aggregation across different receptive fields through parallel multi-scale depthwise separable convolutions. Small kernels focus on details of small targets, while large kernels focus on macrostructures such as roads and building complexes, outputting multi-resolution features. Finally, the anisotropic coherent weight map is element-wise multiplied with the input features to achieve structural enhancement, and then element-wise added to the multi-resolution features. Finally, a residual connection is formed with the input features to obtain output features that possess both directional continuity and multi-scale robustness.

[0088] In the upsampling decoding stage of the decoder upsampling module, a multi-resolution anisotropic coherent unit is set up. This unit consists of anisotropic sensing branch and multi-resolution sensing branch, which are modeled in parallel and jointly output decoding enhancement features.

[0089] (1) Anisotropic perception branch:

[0090] Specifically, the input feature X is first processed by average pooling (AvgPool) to smooth the noise, and then undergoes preliminary feature transformation through a module that includes convolution, batch normalization (BN), and SiLU activation to obtain intermediate features:

[0091]

[0092] in, These are the convolution weights.

[0093] Next, the intermediate features are fed into two parallel directional convolutional layers to extract structural information in the horizontal and vertical directions, respectively:

[0094]

[0095]

[0096] Among them, This indicates the size of the horizontal and vertical convolution kernels. and They represent the first The depth of each channel in the horizontal and vertical directions can be separable from the convolution kernel parameters. No. The weight of each position. : Represents the spatial coordinates of the output feature map.

[0097] Finally, the feature maps in the horizontal and vertical directions ( , They represent in , The specific numerical values ​​at the index spatial location are concatenated along the channel dimension and then convolved to generate the final anisotropic coherent weight map:

[0098]

[0099] in, For splicing operations, These are the convolution weights.

[0100] (2) Multi-resolution sensing branch:

[0101] Specifically, the input feature X is first channel-expanded through a pointwise convolutional layer. This pointwise convolutional layer contains a channel expansion layer using... Weighted convolution, batch normalization (BN), and ReLU6 activation function.

[0102]

[0103] The channel-expanded features are fed into a multi-scale convolutional module, which consists of n parallel deep convolutional branches. Each deep convolutional branch contains a deep convolution (DWConv), batch normalization (BN), and ReLU6 activation function.

[0104]

[0105]

[0106] in, Indicates output tensor The Each channel in spatial coordinates The value at that location. This represents a depthwise convolution kernel. Channels representing the input feature tensor. Indicates the first The kernel size of each depthwise convolutional branch. This indicates the stride of the convolution operation. This represents the amount of padding to maintain the feature map size. An index used for summing along the dimensions of the convolution kernel space.

[0107] This yields n scale-specific features output from the depthwise convolution branch. each Both methods capture spatial information of different objects at specific scales. The outputs of the n depthwise convolutional branches are first summed element-wise, and then channel projection is performed using a 1×1 convolution + BN to obtain multi-resolution features. The overall process is as follows:

[0108]

[0109] in This is the result of adding each element one by one.

[0110] Final feature fusion and output:

[0111] The anisotropic coherence weight map obtained from the anisotropic sensing branch is compared with the original input features. Element-wise multiplication is performed to enhance the structure of the original features. Then, the multiplication result is added element-wise to the multi-resolution features obtained from the multi-resolution perceptual branch. Finally, features from the original input are added. Residual connections are used to ensure the complete transmission of information and enhance features. :

[0112]

[0113] Step 7: Obtain the SAR image segmentation result through the decoder upsampling module.

[0114] The first decoder upsampling module receives the maximum-scale fused features output from the optically guided cross-domain fusion module after processing by the multi-resolution anisotropic coherence unit (MRCU) as input. It then performs a 2× upsampling operation using bilinear interpolation to improve the spatial resolution of the feature map. The features output from this module are then processed by the MRCU and input to the next decoder upsampling module for another 2× upsampling, and this process continues for subsequent layers. Through this layer-by-layer upsampling method, the spatial details of the features are gradually restored and structural continuity is enhanced. The final decoder upsampling module then yields an accurate SAR image segmentation result.

[0115] Step 8: In the decoding process, a coherent sensing gated skip connection is introduced to combine the shallow features of the same level with the high-level semantic features obtained by the decoder upsampling module after wavelet decomposition and anti-spoofing processing.

[0116] First, a discrete wavelet transform is performed on the shallow features output by the optically guided cross-domain fusion module, which has the same spatial resolution as the high-level semantic features output by the decoder upsampling module. This decomposes the features into low-frequency and high-frequency sub-bands to distinguish between real contours and coherent pseudo-textures. Next, the features of each sub-band are packaged and refined using depthwise separable convolution and a learnable scaling function to suppress high-frequency pseudo-edges and pseudo-textures. Then, an inverse wavelet transform is performed on the refined sub-band features to reconstruct the shallow features, and residual fusion is performed with the base features obtained by convolving the original shallow features to obtain enhanced shallow features. Finally, the high-level semantic features obtained by the decoder upsampling module are convolved and added to the enhanced shallow features. After normalization and a sigmoid function, a gating signal is generated. This signal is multiplied pixel-by-pixel with the shallow features to precisely control the proportion of skip connection information, resulting in gated features. The gated features are added to the output of the decoder upsampling module and processed by a multi-resolution anisotropic coherent unit before being used as the input to the next decoder upsampling module.

[0117] To suppress the unselective amplification of shallow speckle noise and pseudo-edges during hop connections, this invention proposes a coherent sensing gated hop connection, which includes the following steps:

[0118] Given the output of the optically guided cross-domain fusion module, i.e. shallow features x, and the output of the decoder upsampling module, i.e. high-level semantic features g, these two types of features are first processed by convolutional layers.

[0119] Convolution and normalization operations are performed on the high-level semantic features g to generate high-level global semantic features. :

[0120]

[0121] Among them, Indicates the convolution kernel weights. This indicates batch normalization operation.

[0122] For shallow features :

[0123] For a given four-dimensional shallow feature Zero-padding yields an augmented tensor of dimension B×C×(H+2P)×(W+2P). The padding width P is determined by the kernel size K. The padding operation formula is as follows:

[0124]

[0125] in for Features after zero padding Indicates in No. The first sample Coordinates of each channel The specific value at that location.

[0126] The zero-padded four-dimensional shallow features are processed using a predefined wavelet filter bank. The mapping is performed as a five-dimensional output tensor. The newly added dimension represents the multiple subbands generated by wavelet decomposition. The calculation formula for this mapping transformation is as follows:

[0127]

[0128] in, Indicates the weights of the wavelet filter bank; Indicates the offset index; is the spatial dimension of the convolution kernel; m and n represent the spatial coordinate indices of the feature map after wavelet transform. Represents the weight tensor No. The first channel The coordinates of the wavelet subband on its unique input channel 0 The specific value. For a five-dimensional output tensor, Indicates in No. The first sample Spatial coordinates of each channel The specific value at that location. Indicates in No. The first sample One channel Coordinates of each offspring The value.

[0129] Following the above process, the multiple sub-bands generated by wavelet decomposition are further processed to enhance the feature representation. The processing formula is as follows:

[0130]

[0131] in It is a reshape operator used to merge subband dimensions with channel dimensions. Representing the The weights of each convolutional layer. This is the activation function.

[0132] After performing the above enhancement operations on the subband features, a higher resolution spatial domain feature map is reconstructed by performing inverse wavelet transform (IWT).

[0133] The formula for calculating the inverse wavelet transform is as follows:

[0134]

[0135] in These represent the spatial height and width of the input feature map, respectively. This represents the wavelet subband index of the input. and Represents the input subband feature map Spatial coordinate index. Represents the filter (weight) tensor of the inverse wavelet transform. Indicates in No. The first sample One channel Coordinates of each offspring The value. express In the middle, used for the first The first input channel The spatial location of the filter corresponding to each subband on its unique input channel 0 within its group. The coefficient at that location. Where... The padding width is consistent with the forward transform. To reconstruct the feature map, where Indicates in No. The first sample Channel coordinates The value.

[0136] To effectively integrate multi-scale features from the wavelet domain with shallow information from the spatial domain, the model uses residual connections to integrate the shallow features. Features reconstructed from wavelet domain Combining these elements yields refined structural features in the wavelet domain. .

[0137]

[0138] Among them, This represents the weight kernel of the convolutional layer. This represents a non-linear processing function.

[0139] After processing the two types of features separately:

[0140] To suppress speckle artifacts in shallow features of SAR images, a gating mechanism is used to fuse high-level global semantic features. Structural features refined in the wavelet domain Semantic modulation is applied to shallow features x to enhance the real-world information. The formula is defined as follows:

[0141]

[0142] in, As a gating signal This represents the weights of the convolutional layer. Represents the composite function, that is .

[0143] Step 9: Model Training Configuration and Performance Evaluation

[0144] The network constructed in step 5 is trained using the training set obtained in step 2, and evaluated on the test set. In this embodiment, to ensure reproducibility, a unified data preprocessing and training configuration is adopted: the initial learning rate is set to 0.0001, the AdamW optimizer is used, the DiceCELoss loss function is employed, the number of training epochs is 100, and the batch size is 2. The hardware and software environment can be Python 3.10.0, PyTorch 2.0.0, and CUDA 11.8, and training is completed on a dual-GPU workstation. Evaluation metrics may include 95% Hausdorff distance (HD95), Dice, IoU, Kappa, and MCC, used to measure boundary consistency, region overlap, and classification consistency.

[0145] The results show that this invention has achieved good results on various neural network models.

[0146] Table 1 Comparison of segmentation results on the SARBUD dataset

[0147]

[0148] To verify the effectiveness of the proposed model, we evaluated it on the SARBUD dataset and compared it with 18 state-of-the-art baseline models. All experiments were conducted under identical conditions to ensure the fairness and comparability of the evaluation results. As shown in Table 1, our model achieves state-of-the-art performance across all metrics on this dataset.

[0149] To verify the effectiveness of each key module, we conducted ablation experiments on the SARBuD dataset, evaluating the contribution of the optical prior adversarial generative branch to segmentation performance for two sub-modules: coherent sensing gated skip connections and multi-resolution anisotropic coherent units. Table 2 shows the quantitative results for each module combination.

[0150] In this application, A–E represent, in order: A is the decoder, B is the anisotropic sensing branch, C is the multi-resolution sensing branch, D is the optical prior adversarial generation branch, and E is the coherent sensing gated skip connection.

[0151] Table 2 Ablation experiments on key modules of the SARBUD dataset.

[0152]

[0153] As shown in Table 2, the performance improvement of the dual-driven spectral cross-modal GAN ​​adversarial segmentation network stems from the synergistic effect of its three core modules. First, the multi-resolution anisotropic coherent unit, through joint modeling of anisotropic and multi-resolution sensing branches, achieves low-frequency structural enhancement and multi-scale feature fusion at the signal level, improving the model's Dice and IoU metrics by approximately 1.6% and 1.2%, respectively. This improvement arises from the fact that directional convolution strengthens the continuity of the main structure of artificial targets in both horizontal and vertical directions, while smoothing and suppressing random high-frequency speckle noise; multi-scale depthwise separable convolution balances local details and global structure across different receptive fields, enhancing the model's adaptability to targets of diverse scales. Second, the coherent sensing gated skip connection, through a dual-driven gating mechanism, fuses high-level semantics and wavelet despoofing features, selectively transmitting true boundary information consistent with semantics, effectively suppressing the propagation of pseudo-textures caused by coherent interference, resulting in a decrease in 95HD and an improvement in Kappa and MCC of approximately 1.9%, demonstrating its significant advantages in boundary discrimination and structural consistency maintenance. Finally, the optical prior adversarial generative branch introduces optical domain geometric priors through the SAR→optical generation task, enabling the model to retain natural semantic perception capabilities under pure SAR input conditions. This significantly enhances geometric stability and semantic discriminative ability in complex scenes, resulting in an additional improvement of approximately 1% in both Dice and IoU. Furthermore, when simple upsampling is used to replace the decoder, overall performance declines, indicating that the decoder plays a crucial role in maintaining structural continuity and suppressing false boundary propagation. In summary, the synergistic effect of these three components enables the network to exhibit superior comprehensive performance in structural reconstruction, noise suppression, and semantic consistency modeling.

[0154] Table 3 Ablation experiments on the optically guided cross-domain fusion module using the SARBUD dataset.

[0155]

[0156] Table 3 presents the ablation results for different feature fusion strategies. F represents the optically guided cross-domain fusion module. It can be seen that the optically guided cross-domain fusion module maintains advantages over recent methods such as TFAM / MAFM / CAFM (Dice at least +0.74, IoU at least +1.12), accompanied by lower variance, demonstrating more robust convergence and generalization. The core lies in "selective recalibration" for cross-domain features: first, aligning the two representations at multiple levels; then, through adaptive allocation at both the channel and spatial levels, establishing a "competition-cooperation" relationship between branches; dynamically assigning weights to each position and channel based on semantic consistency, thereby prioritizing the retention of clear boundaries and stable structural information related to the target category, and suppressing interference introduced by coherent noise or texture artifacts. Compared to indiscriminate fusion methods such as direct splicing or element-wise addition, the optically guided cross-domain fusion module can explicitly filter and inject cross-domain information, improving the consistency between structural perception and semantic parsing, thus obtaining better and more robust segmentation results in complex scenes.

[0157] This application also provides a SAR image segmentation system based on frequency-space prior transfer, including the following modules:

[0158] Preprocessing module: An existing SAR image segmentation dataset is used as the training dataset, and it is pre-divided into training, validation, and test sets to ensure the fairness and stability of model training and evaluation. Each SAR image sample in the SAR image segmentation dataset is accompanied by a corresponding pixel-level annotation file.

[0159] Normalization, size unification (cropping / scaling to a fixed resolution), random cropping, random intensity perturbation, and affine enhancement operations such as rotation / translation / scaling are sequentially performed on the SAR images in the training set to improve the model's generalization ability to scale drift and noise perturbation.

[0160] Feature extraction module: Obtains multi-scale SAR structural features and multi-scale pseudo-optical prior features through a dual-path joint coding structure.

[0161] The dual-path joint coding structure includes a segmentation backbone encoder branch and an optical prior adversarial generation branch. The segmentation backbone encoder branch is used to extract multi-scale SAR structural features from the SAR image. The optical prior adversarial generation branch generates multi-scale pseudo-optical prior features based on the input SAR image.

[0162] Feature fusion module: Through the optically guided cross-domain fusion module, SAR structural features and pseudo-optical prior features are aligned and fused at multiple scale levels to obtain corresponding fused features, realizing "optical prior to SAR latent space migration", and still obtaining geometric semantic constraints under SAR single-mode input.

[0163] Feature enhancement module: By using multi-resolution anisotropic coherent units to enhance the directional consistency and multi-scale completion of the input features of the decoder upsampling module, the module highlights the real structure and suppresses speckle noise interference.

[0164] Segmentation output module: Obtains SAR image segmentation results through the decoder upsampling module.

[0165] In the decoding process, coherent sensing gated skip connections are introduced. Shallow features at the same level are decomposed by wavelet to remove falsehoods and then combined with high-level semantic features obtained by the decoder upsampling module element by element. The resulting fused features are then used to screen the original shallow features. The screened features are added to the output of the decoder upsampling module and processed by a multi-resolution anisotropic coherent unit before being used as the input of the next decoder upsampling module.

[0166] Training and Reasoning Module:

[0167] Based on the preprocessed training dataset, an end-to-end training approach is adopted, using the AdamW optimizer and a preset learning rate strategy to complete the training; during the inference stage, only SAR images are input, and corresponding pixel-level segmentation results can be output.

[0168] In one possible implementation, the generator's encoder is obtained by pre-training the generator and discriminator on a GAN network using SAR images and optical images of the same region from a pre-training dataset, which serves as the optical prior adversarial generation branch.

[0169] In one possible implementation, the optically guided cross-domain fusion module includes the following steps:

[0170] 1-1. Align the spatial resolution and channel dimension of the same-layer features of the segmentation backbone encoder branch and the optical prior adversarial generation branch.

[0171] 1-2. Channel compression is performed on the aligned features, and mutual exclusive attention weights that sum to one are generated for the segmentation backbone encoder branch and the optical prior adversarial generation branch through normalized weights. The attention weights are then weighted channel-by-channel with the compressed features of the corresponding branches, and the weighted features of the two branches are added together to obtain the global perception features.

[0172] 1-3. Spatial-level adaptive fusion: Based on global perception features, spatial context information is further extracted through cascaded convolution and activation functions to generate fine-grained pixel-level spatial attention weights. Subsequently, the pixel-level spatial attention weights are applied to SAR structural features and pseudo-optical prior features, and adaptive spatial fusion of cross-level features is achieved through weighted aggregation to obtain fused features.

[0173] In one possible implementation, the coherent sensing gated hop connection includes the following steps:

[0174] 2-1. Perform discrete wavelet transform on the shallow features output by the optically guided cross-domain fusion module, which has the same spatial resolution as the high-level semantic features output by the decoder upsampling module, to decompose the shallow features into low-frequency sub-bands and high-frequency sub-bands, so as to distinguish between real contours and coherent pseudo-textures.

[0175] 2-2. Pack the features of each subband (merge the subband dimension into the channel dimension) and refine the subband features through depthwise separable convolution and learnable scaling function to suppress high-frequency pseudo edges and pseudo textures.

[0176] 2-3. Perform inverse wavelet transform on the refined sub-band features to reconstruct shallow features. Then, perform residual fusion between the reconstructed shallow features and the basic features obtained by convolving the original shallow features to obtain enhanced shallow features.

[0177] 2-4. The high-level semantic features obtained by the decoder upsampling module with the same spatial resolution as the shallow features are convolved and added to the enhanced shallow features for fusion. After normalization and Sigmoid, a gating signal is generated. The gating signal is multiplied pixel by pixel with the original shallow features to achieve precise control of the proportion of skip connection information to obtain the gated features.

[0178] In one possible implementation, the multi-resolution anisotropic coherent unit includes anisotropic sensing branches and multi-resolution sensing branches, as detailed below:

[0179] 3-1. The anisotropic sensing branch extracts the directional consistency structure of the segmented target by directional convolution in the horizontal and vertical directions of the input features (the maximum scale fusion feature output by the optically guided cross-domain fusion module or the result of adding the output of the previous coherent sensing gated skip connection with the output of the previous decoder upsampling module), enhances the continuous contour and smooths random high-frequency speckle noise, and finally outputs an anisotropic coherent weight map.

[0180] 3-2. The multi-resolution perception branch aggregates the input features into different receptive fields through parallel multi-scale depth separable convolutions. The small kernel focuses on the details of small targets, while the large kernel focuses on macrostructures such as roads / building complexes, and finally outputs multi-resolution features.

[0181] 3-3. The anisotropic coherent weight map output by the anisotropic sensing branch is multiplied element-wise with the input features to achieve structured enhancement of the original features. Subsequently, the enhanced features are added element-wise with the multi-resolution features output by the multi-resolution sensing branch, and then residually connected with the input features to finally obtain output features that have both directional continuity and multi-scale robustness.

[0182] In one possible implementation, the first decoder upsampling module takes the maximum-scale fusion features output by the optically guided cross-domain fusion module after processing by the multi-resolution anisotropic coherence unit as input. The output of the first decoder upsampling module is then processed by the multi-resolution anisotropic coherence unit and input into the next decoder upsampling module, and so on. Through layer-by-layer upsampling, an accurate SAR image segmentation result is obtained through the last decoder upsampling module.

[0183] This application also provides an electronic device, including a processor and a memory.

[0184] The memory is used to store computer programs.

[0185] When the processor executes a program stored in the memory, it implements any of the methods described in this application.

[0186] In one possible implementation, the electronic device of this application embodiment further includes a communication interface and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus.

[0187] The communication bus mentioned in the above electronic devices can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc.

[0188] The communication interface is used for communication between the aforementioned electronic devices and other devices.

[0189] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0190] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0191] In another embodiment provided in this application, a computer-readable storage medium is also provided, which stores a computer program that, when executed by a processor, implements any of the methods described in this application.

[0192] In another embodiment provided in this application, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to perform any of the methods described in this application.

[0193] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)).

[0194] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0195] The various embodiments in this specification are described in a related manner. Each embodiment focuses on the differences from other embodiments, and the same or similar parts between the various embodiments can be referred to each other.

[0196] The above description is merely a preferred embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application are included within the scope of protection of this application.

Claims

1. A SAR image segmentation method based on frequency-space prior transfer, characterized in that, Includes the following steps: Step 1: Pre-training dataset construction: Acquire several SAR images and their corresponding optical images as a pre-training dataset; Step 2: Preprocess the training dataset; Step 3: Construct a dual-driven spectral-oriented cross-modal GAN ​​adversarial segmentation network: The network includes: a dual-path joint coding structure, an optically guided cross-domain fusion module, a coherent sensing gated skip connection, a multi-resolution anisotropic coherent unit, and several decoder upsampling modules; Step 4: Obtain multi-scale SAR structural features and multi-scale pseudo-optical prior features through a dual-path joint coding structure: The dual-path joint coding structure includes a segmentation backbone encoder branch and an optical prior adversarial generation branch; the segmentation backbone encoder branch is used to extract multi-scale SAR structural features from SAR images; the optical prior adversarial generation branch generates multi-scale pseudo-optical prior features based on the input SAR image. Step 5: Align and fuse SAR structural features with pseudo-optical prior features at multiple scale levels using the optically guided cross-domain fusion module to obtain the corresponding fused features; Step 6: Enhance the orientation consistency and perform multi-scale completion on the input features of the decoder upsampling module by using a multi-resolution anisotropic coherent unit; Step 7: Obtain the SAR image segmentation result through the decoder upsampling module; Step 8: In the decoding process, a coherent sensing gated skip connection is introduced. The shallow features of the same level are combined with the high-level semantic features obtained by the decoder upsampling module after wavelet decomposition and anti-spoofing processing. The resulting fused features are then used to perform gated filtering on the original shallow features. The filtered features are added to the output of the decoder upsampling module and processed by a multi-resolution anisotropic coherent unit before being used as the input of the next decoder upsampling module. Step 9: Model training and inference.

2. The SAR image segmentation method based on frequency-space prior transfer according to claim 1, characterized in that, The existing SAR image segmentation dataset is used as the training dataset, and it is divided into training set, validation set and test set according to a preset ratio; Normalization, size unification, random cropping, random intensity perturbation, and affine enhancement operations are sequentially performed on the SAR images in the training set.

3. The SAR image segmentation method based on frequency-space prior transfer according to claim 1, characterized in that, The generator's encoder is obtained by pre-training the generator and discriminator on the GAN network using SAR images and optical images of the same region from the pre-training dataset. This encoder serves as the optical prior adversarial generative branch.

4. The SAR image segmentation method based on frequency-space prior transfer according to claim 1, characterized in that, The optically guided cross-domain fusion module includes the following steps: 1-1. Align the spatial resolution and channel dimension of the same-layer features of the segmentation backbone encoder branch and the optical prior adversarial generation branch. 1-2. Channel compression is performed on the aligned features, and mutual exclusive attention weights that sum to one are generated for the segmentation backbone encoder branch and the optical prior adversarial generation branch through normalized weights; the attention weights are then weighted channel by channel with the compressed features of the corresponding branches, and the weighted features of the two branches are added together and fused to obtain the global perception features. 1-3. Spatial-level adaptive fusion: Based on global perception features, spatial context information is further extracted through cascaded convolution and activation functions to generate fine-grained pixel-level spatial attention weights. Subsequently, the pixel-level spatial attention weights are applied to SAR structural features and pseudo-optical prior features, and the adaptive spatial fusion of cross-level features is achieved through weighted aggregation to obtain fused features.

5. The SAR image segmentation method based on frequency-space prior transfer according to claim 1, characterized in that, The coherent sensing gated skip connection includes the following steps: 2-1. Perform discrete wavelet transform on the shallow features output by the optically guided cross-domain fusion module, which has the same spatial resolution as the high-level semantic features output by the decoder upsampling module, and decompose the shallow features into low-frequency sub-bands and high-frequency sub-bands. 2-2. Package the features of each sub-band, that is, merge the sub-band dimensions into the channel dimensions and refine the sub-band features using depthwise separable convolution and learnable scaling functions; 2-3. Perform inverse wavelet transform on the refined sub-band features to reconstruct shallow features. Then, perform residual fusion between the reconstructed shallow features and the basic features obtained by convolving the original shallow features to obtain enhanced shallow features. 2-4. The high-level semantic features obtained by the decoder upsampling module with the same spatial resolution as the shallow features are convolved and added to the enhanced shallow features for fusion. After normalization and Sigmoid, a gating signal is generated. The gating signal is multiplied pixel by pixel with the original shallow features to achieve precise control of the proportion of skip connection information to obtain the gated features.

6. The SAR image segmentation method based on frequency-space prior transfer according to claim 1, characterized in that, The multi-resolution anisotropic coherent unit includes anisotropic sensing branches and multi-resolution sensing branches, as detailed below: 3-1. The anisotropic perception branch extracts the directional consistency structure of the segmentation target through directional convolution in the horizontal and vertical directions of the input features, and finally outputs an anisotropic coherence weight map. 3-2. The multi-resolution perceptual branch aggregates the input features by establishing different receptive fields through parallel multi-scale depthwise separable convolution, and finally outputs multi-resolution features; 3-3. Multiply the anisotropic coherence weight map output by the anisotropic perception branch with the input features element by element to achieve structured enhancement of the original features; Subsequently, the enhanced features are added element-wise to the multi-resolution features output by the multi-resolution perception branch, and then residually connected to the input features to obtain the final output features.

7. The SAR image segmentation method based on frequency-space prior transfer according to claim 1, characterized in that, The first decoder upsampling module takes the maximum scale fusion feature output by the optically guided cross-domain fusion module after processing by the multi-resolution anisotropic coherence unit as input. The output of the first decoder upsampling module is processed by the multi-resolution anisotropic coherence unit and then input into the next decoder upsampling module, and so on. Through layer-by-layer upsampling, the accurate SAR image segmentation result is obtained through the last decoder upsampling module.

8. The SAR image segmentation method based on frequency-space prior transfer according to claim 1, characterized in that, Based on the preprocessed training dataset, an end-to-end training approach is adopted, using the AdamW optimizer and a preset learning rate strategy to complete the training; during the inference stage, only SAR images are input, and corresponding pixel-level segmentation results can be output.

9. A SAR image segmentation system based on frequency-space prior transfer, characterized in that, Includes the following modules: Preprocessing module: The existing SAR image segmentation dataset is used as the training dataset, and it is divided into training set, validation set and test set according to a preset ratio; Normalization, size unification, random cropping, random intensity perturbation, and affine enhancement operations are sequentially performed on the SAR images in the training set. Feature extraction module: Obtains multi-scale SAR structural features and multi-scale pseudo-optical prior features through a dual-path joint coding structure; The dual-path joint coding structure includes a segmented backbone encoder branch and an optical prior adversarial generation branch. The segmentation backbone encoder branch is used to extract multi-scale SAR structural features from SAR images; the optical prior adversarial generation branch generates multi-scale pseudo-optical prior features based on the input SAR image. Feature fusion module: Through the optically guided cross-domain fusion module, SAR structural features and pseudo-optical prior features are aligned and fused at multiple scale levels to obtain the corresponding fused features; Feature enhancement module: Enhances the orientation consistency and completes the multi-scale features of the decoder upsampling module by using multi-resolution anisotropic coherent units; Segmentation output module: Obtains SAR image segmentation results through the decoder upsampling module; In the decoding process, a coherent sensing gated skip connection is introduced. Shallow features at the same level are decomposed by wavelet to remove falsehoods and then combined with high-level semantic features obtained by the decoder upsampling module element by element. The resulting fused features are then used to screen the original shallow features. The screened features are added to the output of the decoder upsampling module and processed by a multi-resolution anisotropic coherent unit before being used as the input of the next decoder upsampling module. Training and Reasoning Module: Based on the preprocessed training dataset, an end-to-end training approach is adopted, using the AdamW optimizer and a preset learning rate strategy to complete the training; during the inference stage, only SAR images are input, and corresponding pixel-level segmentation results can be output.

10. An electronic device, characterized in that, Including processor and memory; The memory is used to store computer programs; When the processor executes the program stored in the memory, it implements the SAR image segmentation method according to any one of claims 1-8.