A Remote Sensing Image Domain Generalization Semantic Segmentation Method Based on Multi-Scale Instance Decoupling

Through multi-scale instance coding and domain-invariant instance decoupling technology, the problem of insufficient generalization ability of the remote sensing image segmentation model under the differences between domains is solved, efficient semantic segmentation on the target domain has been achieved, and the cross-domain adaptability and generalization ability of the model are improved.

CN119992289BActive Publication Date: 2025-07-22GUIZHOU SURVEY & DESIGN RES INST FOR WATER RESOURCES & HYDROPOWER
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510484907.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-07-22
Estimated Expiration
2045-04-17

AI Technical Summary

Technical Problem

The inter-domain difference between the existing remote sensing image segmentation models between training data and inference data leads to insufficient generalization capabilities, making it difficult to effectively perform semantic segmentation on unseen target domains, and the universality and practical application of unsupervised domain adaptation methods are limited.

Method used

Multi-scale instance coding (MSIE) and domain-invariant instance decoupling (DII) technology are used to construct domain generalization semantic segmentation methods through multi-scale feature interaction, domain-invariant feature reconstruction and covariance optimization to achieve robustness and generalization capabilities of cross-domain remote sensing images.

Benefits of technology

Efficient semantic segmentation of remote sensing images is achieved on the unseen target domain, improving the cross-domain adaptability and generalization capabilities of the model, reducing the dependence on target domain data, and suitable for various remote sensing images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992289B_ABST
    Figure CN119992289B_ABST
Patent Text Reader

Abstract

The present invention relates to a remote sensing image domain generalization semantic segmentation method based on multi-scale instance decoupling, which relates to the technical field of remote sensing image interpretation, and includes multi-scale instance encoding, domain-invariant instance decoupling, and domain generalization semantic decoding. Taking the image features in the segmentation backbone network as input, the multi-scale instance encoding module effectively perceives the feature responses of cross-domain instances by using depth convolution, multi-scale collaboration, and multi-path perception, and is not affected by scale differences, spatial deformations, and noises. A domain-invariant instance decoupling module is designed to decouple the cross-domain style from the image representation. Finally, the decoupled features are input into the domain generalization semantic decoding module to predict the segmented image on unseen target domains. This segmentation method can generalize various unseen remote sensing images and obtain good segmentation capabilities.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of remote sensing image interpretation, and particularly relates to a remote sensing image domain generalization semantic segmentation method based on multi-scale instance decoupling. Background Art

[0002] Remote sensing image interpretation is a basic task of earth observation, and it plays an important role in many applications, such as land use classification, geospatial object extraction, and surface time series monitoring. In the task of remote sensing image interpretation, remote sensing semantic segmentation occupies a unique position, which can realize the recognition and classification of each ground object pixel. With the development of deep learning technology, remote sensing image interpretation methods based on deep learning have been widely studied, which has greatly promoted the development of remote sensing image semantic segmentation methods. However, most existing remote sensing image segmentation methods still assume that the training data and the inference data follow the same and independent statistical distributions. These segmentation methods are usually trained on a remote sensing training set of a specific benchmark and then verified on the test set of the same benchmark, where the remote sensing images within the same benchmark usually have the same feature distribution. In the actual application process, remote sensing image data often comes from different observation platforms and sensors, and factors such as surface landscapes, spatial resolutions, illumination conditions, and sensor parameters will cause significant differences in the feature distributions between the inference data (also called the target domain) and the training data (also called the source domain). At this time, the segmentation model trained on the training data set will not be able to effectively achieve the semantic segmentation of the target domain images. This deviation in the feature distribution between the training data and the inference data is called the inter-domain difference, and the existence of the inter-domain difference will greatly reduce the generalization ability of the remote sensing image semantic segmentation model.

[0003] In the prior art, various efforts have been made to improve the cross-domain generalization ability of remote sensing image segmentation models. Most of these methods are based on the idea of domain adaptation, and by aligning the feature distributions of the source domain images and the target domain images during training, the cross-domain generalization ability of the segmentation model is improved. Although this method can achieve an obvious improvement in the segmentation performance on the target domain images, in practical applications, it is difficult to obtain the labeled data of the target domain images at the same time. To solve this weakness, unsupervised domain adaptation remote sensing image semantic segmentation methods have been proposed, which no longer require the ground truth of the target domain images during training, but reduce the inter-domain deviation between the source domain and the target domain through methods such as data generation and low-cost functions. Since the target domain is involved in the training process, the segmentation models constructed by domain adaptation and unsupervised domain adaptation methods can only be applicable to the remote sensing data of a specific target domain. For other target domain data that have not participated in the training, the generalization segmentation ability of the model will be affected. This weakness limits the generality and practical application of domain adaptation methods and unsupervised domain adaptation methods. Ideally, the remote sensing image segmentation model should be able to perform inference on various remote sensing images that have not been seen during training.

[0004] To solve this problem, the invention team introduced a domain generalization method for remote sensing image segmentation, so that the target domain does not need to exist during training and can be generalized to other target domain images. Summary of the Invention

[0005] To solve the above problems, the object of the present invention is to provide a remote sensing image domain generalization semantic segmentation method based on multi-scale instance decoupling, including the following steps:

[0006] S10 Multi-scale Instance Encoding (MSIE): The multi-scale instance encoding module takes the image features in the segmentation backbone network as input, uses depth convolution, cross-scale interaction encoding, integrates multi-scale features, constructs a deformation-frequency domain unit, and perceives the feature responses of cross-domain instances;

[0007] S20 Domain-Invariant Instance Decoupling (DII): The domain-invariant instance decoupling module is based on the reconstruction of mixed distribution features and covariance optimization, decouples at the feature level, and realizes the learning of domain-invariant features;

[0008] S30 Domain Generalization Semantic Decoding: Input the features after S20 domain-invariant instance decoupling into the domain generalization semantic decoding module, and use the Hamburger module to model the global context;

[0009] S40 Generalized Semantic Segmentation: Input the remote sensing image of the unseen target domain into the trained framework to realize the prediction of the segmentation map on the unseen target domain.

[0010] Furthermore, the S10 multi-scale instance encoding (MSIE) specifically includes the following steps:

[0011] S11 Input Image Features: Input the original image from the source domain, divide the original image into squares (patches) of size 16*16*3 (length * width * channels), and map each small block into a one-dimensional vector to form an original image feature sequence and input it into the network. The operation of generating a one-dimensional image block sequence is simply referred to as Patch Embedding;

[0012] S12 Cross-Scale Interaction Encoding (CSE): Its core operation is defined mathematically as,

[0013] ;

[0014] where F represents the input image features, is a scaling operation to adapt to features with different receptive fields, represents element-wise multiplication, and DW-Conv represents depth convolution, It is a cross-scale interaction mechanism; this design ensures that information from different scales can be synergistically integrated, solving the problem of cross-domain scale variability. The output of the deep convolution undergoes cross-scale interaction operations, reducing the semantic fragmentation caused by independent processing of features at each scale. Scaling operations are used to adapt to multiple receptive fields, enabling the encoder to capture local and global context information.

[0015] The expression of the cross-scale interaction mechanism is as follows:

[0016] ;

[0017] ;

[0018] ;

[0019] Among them, is the high-level scale feature, is the low-level scale feature, is bilinear upsampling, represents the sigmoid function, (.) is the aggregation and concatenation operation; the high-level scale feature undergoes upsampling and deep convolution operations, and the low-level scale feature undergoes convolution operation to generate feature maps that can match the current scale feature and . After cross-scale interaction of the scale features at three levels, a feature map that can replace the current scale feature is generated. The cross-scale interaction mechanism can establish topological dependence relationships among multi-scale features. High-level semantics can be propagated from top to bottom to guide low-level features, and low-level details can correct high-level predictions from bottom to top, realizing bidirectional propagation of high- and low-level features.

[0020] S13 Multi-scale feature integration: The context attention mechanism is used to integrate feature information at different scales, dynamically assigning importance to features from different scales, and further integrating feature information at different scales. Its mathematical expression is:

[0021] ;

[0022] ;

[0023] Among them, is the result of cross-scale interaction encoding in step S12, is a learnable convolutional layer for generating attention weights, represents element-wise multiplication, represents the weight, Denote the attention-weighted feature map; by applying these attention weights, the encoder emphasizes the most informative features while suppressing irrelevant information, thus achieving more effective aggregation of multi-scale features. These re-weighted features are summed and passed through an additional convolutional layer to generate the final encoded representation.

[0024] S14 Deformation-Frequency Domain Unit (DFSU) Design: Design a deformation-frequency domain non-uniform convolution group to enhance the robustness of cross-domain image segmentation through differential processing of multi-branch features.

[0025] S15 Scale Normalization: In cross-domain scenarios, differences in spatial resolution often lead to domain shifts. To alleviate this, the encoder introduces a scale normalization step to ensure the consistency of feature representations at different scales, which is expressed as:

[0026] ;

[0027] ;

[0028] where μ and σ are the mean and standard deviation of the features respectively, and γ and β are learnable parameters. is the output feature map of the deformation-frequency domain unit in the S14 step; the features are normalized, reducing domain-specific biases and enhancing the encoder's generalization ability on aerial and satellite images.

[0029] Furthermore, in the S14 Deformation-Frequency Domain Unit (DFSU) design, a deformation-frequency domain non-uniform convolution group is designed to enhance the robustness of cross-domain image segmentation through differential processing of multi-branch features, and the combination of deformable convolution and adaptive dilation rate is used to enhance the perception ability of geometric deformations of cross-domain features.

[0030] Its expression is:

[0031] ;

[0032] where is the learnable offset, is the adaptive dilation rate, , is a multi-layer perceptron with 2 hidden layers, is the attention-weighted feature map is the result after the collaborative combination of deformable convolution and adaptive dilation rate; it solves the problem that the geometric shapes of ground objects in images from different sensor platforms and different regions are significantly different, and affected by the atmosphere and the stability of imaging hardware, in cross-domain remote sensing image segmentation tasks, in addition to the influence of multi-scale features, there are also problems such as sensor noise and feature deformation.

[0033] Using the two-dimensional discrete cosine transform (DCT2D) operation, the high-frequency texture and low-frequency semantics are separated, sensor noise is suppressed, and detailed information is retained. The core operation is as follows:

[0034] ;

[0035] Among them, DCT2D is the two-dimensional discrete cosine transform, IDCT2D is its inverse transform, Conv is the standard convolution, and the attention-weighted feature map After the two-dimensional discrete cosine transform (DCT2D), a transformed feature map is generated , is the transformed feature map The result after 5x5 standard convolution, and after inverse transformation, a two-dimensional discrete cosine transform weighted feature map is generated .

[0036] Furthermore, the domain-invariant instance decoupling (DII) in the S20 specifically includes the following steps:

[0037] S21 Instance Feature Trend Calculation: Given the output feature map of the encoder , calculate the central tendency and dispersion of the instance-level features,

[0038] ;

[0039] ;

[0040] reflects the overall intensity of the feature map, while quantifies the variability between spatial elements, , , are the height, width, and number of channels of the input feature map F in step S12 respectively; these metrics effectively capture the spatial complexity and domain-specific characteristics of remote sensing data.

[0041] S22 Constructing a Remote Sensing Instance Feature Reconstruction Model: Remote sensing data usually involves complex relationships between objects, textures, and environmental backgrounds. To eliminate domain-specific biases, and are rescaled to the interval [0,1] and randomly generated using statistical distribution methods; a Bayesian probability framework is introduced to construct a remote sensing instance feature reconstruction model based on the Beta-Student mixture distribution, and the original and are replaced. Through the mixture distribution operation, the domain adaptability of cross-domain features can be improved to a certain extent. Specifically,

[0042] ;

[0043] ;

[0044] wherein, is the mixing weight, is the degrees of freedom of the Student T distribution, is the shape parameter of the Beta distribution, is the truncation distribution, is the output feature map of the given encoder, Rectified Linear Unit; this feature reconstruction method eliminates the influence of domain-specific style information and forces the network to learn domain-independent features that focus on the underlying spatial structure and semantic content of remote sensing images. This is particularly important for remote sensing tasks because domain variations often obscure meaningful information.

[0045] The adjusted statistical data is integrated back into the feature map to generate augmented feature embeddings :

[0046] ;

[0047] S23 Feature Decoupling: To further enhance the feature generalization ability, the original feature embeddings and the augmented feature embeddings are decoupled through a shared encoder-decoder network. At each layer, the intermediate feature map and its augmented corresponding feature map are subjected to domain-invariant instance decoupling, where , , are the height, width, and number of channels of the intermediate feature map respectively.

[0048] Further, for each feature map, hierarchical dynamic covariance decomposition is performed. By semantic-guided channel grouping, such as semantic categories like edges, textures, shapes, etc., the channels of the feature map are divided into K groups, and each group of features is with weights dynamically assigned based on feature importance :

[0049] ;

[0050] where GAP(.) represents global average pooling, is a learnable parameter vector corresponding to the weights of the K groups of channels, exp(.) represents the exponential function with the natural constant e as the base, and the same operation is performed on the augmented feature map to generate the grouped feature map and the weights of the dynamically assigned feature importance;

[0051] Calculate the hierarchical covariance matrix for each group of features separately and perform weighted fusion, feature map 、 The covariance matrices are respectively:

[0052] ;

[0053] ;

[0054] where and are the means of each group of features 、 respectively, 、 、 are the row, column, and channel grouping numbers of the feature map respectively, 、 are the covariance matrices of the feature maps and respectively.

[0055] These covariance matrices contain the spatial relationships of the feature maps, which are crucial for remote sensing applications that require a detailed understanding of the spatial distribution. The traditional global calculation method of the covariance matrix will lead to the coupling of domain-related and domain-unrelated features. By applying semantic-guided channels for hierarchical covariance calculation, the robustness of domain-invariant feature decoupling is improved, and the computational complexity is reduced.

[0056] To improve variability, a difference metric is defined based on the covariance matrix:

[0057] ;

[0058] where represents the average covariance matrix. This metric ensures that the decoupled features are consistent with the domain-invariant characteristics while retaining spatial semantic consistency.

[0059] Furthermore, to improve the method's ability to select cross-domain features of remote sensing images, the Gumbel-Softmax differentiable mask is applied to replace the hard binary mask with learnable parameters to enhance the feature extraction ability in weakly textured regions:

[0060] ;

[0061] Gumbel(0,1).

[0062] where represents a random coefficient following the Gumbel(0,1) distribution, is a learnable control coefficient that controls the sparsity of the mask.

[0063] The co - design of dynamic group covariance decomposition and differentiable masks is optimized and improved in terms of computational efficiency, feature decoupling, and domain generalization stability, enabling it to better extract and retain domain - invariant feature information existing in different image domains.

[0064] Furthermore, a new mask joint constraint term is added:

[0065] ;

[0066] Applying the joint constraint term to prevent uncontrollable mask sparsity and combining it with the whitening loss function to construct a joint decoupling loss function:

[0067] .

[0068] Finally, the loss function for feature decoupling and the cross - entropy loss function together constitute the joint loss function of the domain - generalization semantic segmentation method for remote sensing images:

[0069] .

[0070] where is the output feature of the encoder in the fourth stage, is the label data of the source - domain image.

[0071] By using the domain - invariant instance decoupling (DII) module, this method effectively addresses the challenges brought by complex spatial relationships and domain variability, achieving powerful and domain - invariant feature extraction for remote sensing applications.

[0072] Furthermore, the learning of domain - invariant features from the input multi - scale instance encoding module to the domain - invariant instance decoupling module in remote sensing image data is regarded as one stage. Four stages of feature learning are repeated before domain - generalization semantic decoding, and the features extracted in the second, third, and fourth stages are jointly used as the feature input for domain - generalization semantic decoding.

[0073] Furthermore, the features of the last three stages of the encoder are aggregated and concatenated, and a lightweight head module is used to model the global context, and then a multi - layer perceptron (MLP) with two hidden layers is used for per - pixel classification. By focusing on the last three stages, the over - detailed low - level information in the first stage is avoided, which is less relevant to high - level semantic tasks.

[0074] Furthermore, S10 - S30 is the training stage, where only remote sensing images from the source domain are input into the framework, and S40 is the inference stage, where remote sensing images from unseen target domains are input into the trained framework.

[0075] The beneficial effects of the present invention are as follows:

[0076] 1. Based on the general idea of multi-scale extraction, a multi-scale feature interaction mechanism and a deformation-frequency domain unit are constructed to learn stable instance representations at multiple scales, suppress the influence of noise, and maintain semantic consistency.

[0077] 2. A domain-invariant instance decoupling (DII) method based on hybrid distribution feature reconstruction and covariance optimization is constructed to reduce the influence of domain-specific features in the image appearance on the generalization ability of semantic segmentation.

[0078] 3. Domain-invariant feature representation is achieved by integrating features from multiple stages, enabling the capture of local details and global context, which is crucial for understanding complex spatial relationships in remote sensing scenes. It does not require the presence of the target domain during training and can be generalized to other target domain images to predict segmentation images on unseen target domains. At the same time, the reduction in computational load makes it possible to process larger datasets. Description of the Drawings

[0079] Figure 1 is the network framework diagram of the present invention (Source Domain is the source domain, Unseen Domain is the unseen target domain, Patch Embedding is the operation of generating a one-dimensional image patch sequence, MSIE Block 1, MSIE Block 2, MSIE Block 3, and MSIE Block 4 are the multi-scale instance encoding modules 1, 2, 3, and 4 in four stages respectively, DIIDisentanglement is domain-invariant instance decoupling, Concatenate is the feature aggregation and splicing operation, Head is a lightweight hamburger module, MLP is a multi-layer perceptron with 2 hidden layers, Label is the label data of the source domain image, Loss is the joint loss function, Unseen Prediction is the predicted segmentation result of the unseen target domain image, and Stage1-4 are stages 1-4 respectively);

[0080] Figure 2 is the structural diagram of the multi-scale instance encoding module of the present invention (GAT is the cross-scale interaction gate, DCT and IDCT are the discrete cosine transform and its inverse transform);

[0081] Figure 3 is the structural diagram of the domain-invariant instance decoupling module of the present invention ( is the covariance )

[0082] Figure 4 is the image segmentation result of each ground object type of Embodiment 1 of the present invention on the ISPRS dataset;

[0083] Figure 5 This is the image segmentation result of building types in the WHU dataset for Embodiment 1 of the present invention. Detailed implementation manners

[0084] The present invention will be described in detail below in conjunction with the accompanying drawings and specific embodiments. Here, the illustrative embodiments of the present invention and the descriptions are used to explain the present invention, but do not limit the present invention.

[0085] Embodiment 1

[0086] A remote sensing image domain generalization semantic segmentation method based on multi-scale instance decoupling, comprising the following steps:

[0087] S10 Multi-scale Instance Encoding (MSIE): The multi-scale instance encoding module takes the image features in the segmentation backbone network as input, utilizes depth convolution, cross-scale interaction encoding, integrates multi-scale features, constructs a deformation-frequency domain unit, and perceives the feature responses of cross-domain instances;

[0088] S11 Input image features: Input the original image from the source domain, divide the original image into squares (patches) of size 16*16*3 (length * width * channels), and map each small block into a one-dimensional vector to form an original image feature sequence and input it into the network. The process of generating a one-dimensional image block sequence is simply referred to as Patch Embedding;

[0089] S12 Cross-scale Interaction Encoding (CSE): Its core operation is defined mathematically as

[0090] ;

[0091] where F represents the input image features, is a scaling operation to adapt to features with different receptive fields, represents element-wise multiplication, DW-Conv represents depth convolution, is a cross-scale interaction mechanism; this design ensures that information from different scales can be synergistically integrated, solving the problem of cross-domain scale variability. The output of the depth convolution undergoes cross-scale interaction operations, reducing the semantic fragmentation caused by independent processing of features at each scale. Scale scaling operations to adapt to multiple receptive fields, enabling the encoder to capture local and global context information.

[0092] The expression of the cross-scale interaction mechanism is as follows:

[0093] ;

[0094] ;

[0095] ;

[0096] is bilinear upsampling, represents the sigmoid function, (.) is the aggregation and splicing operation; the high-level scale features After upsampling and depth convolution operations, the low-level scale features are Convolution operations are respectively performed to generate feature maps that can match the current scale features of and . After cross-scale interaction of the scale features at three levels, a feature map that can replace the current scale features is generated . The cross-scale interaction mechanism can establish the topological dependence relationship between multi-scale features. High-level semantics can be propagated from top to bottom to guide low-level features, and low-level details can correct high-level predictions from bottom to top, realizing the two-way propagation of high and low-level features.

[0097] S13 Multi-scale feature integration: The context attention mechanism is used to integrate feature information at different scales, dynamically assign importance to features from different scales, and further integrate feature information at different scales. Its mathematical expression is:

[0098] ;

[0099] ;

[0100] where is the result of cross-scale interaction encoding in step S12, is a learnable convolutional layer for generating attention weights, represents element-wise multiplication, represents the weight, represents the attention-weighted feature map; by applying these attention weights, the encoder emphasizes the most informative features while suppressing irrelevant information, thus achieving more effective aggregation of multi-scale features. These re-weighted features are summed and passed through an additional convolutional layer to generate the final encoded representation.

[0101] S14 Deformation-Frequency Domain Unit (DFSU) design: Design a deformation-frequency domain heterogeneous convolution group to improve the robustness of cross-domain image segmentation by differentiating the processing of multi-branch features.

[0102] Utilize the adaptive collaborative combination of deformable convolution and dilation rate to enhance the perception ability of geometric deformation of cross-domain features. Its expression is:

[0103] ;

[0104] where is the learnable offset, is the hole rate adaption, , is the multi-layer perceptron with 2 hidden layers, is the attention weighted feature map which is the result of the collaborative combination of deformable convolution and hole rate adaption; it solves the problem that the geometric shapes of ground objects in images from different sensor platforms and different regions are significantly different, and affected by the atmosphere and the stability of imaging hardware. In the cross-domain remote sensing image segmentation task, in addition to the influence of multi-scale features, there are also the influences of sensor noise and feature deformation, etc.

[0105] Using the two-dimensional discrete cosine transform (DCT2D) operation to separate high-frequency textures from low-frequency semantics, suppress sensor noise and retain detailed information. Its core operation is

[0106] ;

[0107] where DCT2D is the two-dimensional discrete cosine transform, IDCT2D is its inverse transform, Conv is the standard convolution, and the attention weighted feature map generates a transformed feature map after the two-dimensional discrete cosine transform (DCT2D) , is the transformed feature map which is the result after a 5x5 standard convolution, and generates a two-dimensional discrete cosine transform weighted feature map after the inverse transform .

[0108] S15 Scale normalization: In the cross-domain scenario, the difference in spatial resolution often leads to domain shift. To alleviate this situation, the encoder introduces a scale normalization step to ensure the consistency of feature representations at different scales, which is expressed as:

[0109] ;

[0110] ;

[0111] where μ and σ are the mean and standard deviation of the features respectively, and γ and β are learnable parameters, is the output feature map of the deformation-frequency domain unit in the S14 step; the features are normalized, reducing the domain-specific bias and enhancing the encoder's generalization ability on aerial and satellite images.

[0112] S20 Domain-invariant instance decoupling (DII): The domain-invariant instance decoupling module is based on the reconstruction of the mixture distribution features and covariance optimization, decouples at the feature level, and realizes the learning of domain-invariant features;

[0113] S21 Instance feature trend calculation: Given the output feature map of the encoder Calculate the central tendency of instance-level features and dispersion ,

[0114] ;

[0115] ;

[0116] reflects the overall intensity of the feature map, while quantifies the variability between spatial elements, , , are the height, width, and number of channels of the input feature map F in step S12 respectively; these metrics effectively capture the spatial complexity and domain-specific characteristics of remote sensing data.

[0117] S22 Construct a remote sensing instance feature reconstruction model: Remote sensing data usually involves complex relationships between objects, textures, and environmental backgrounds. To eliminate domain-specific biases, and are rescaled to the interval [0,1] and randomly generated using statistical distribution methods; introduce a Bayesian probability framework, construct a remote sensing instance feature reconstruction model based on the Beta-Student mixture distribution, and replace the original and . Through the mixture distribution operation, the domain adaptability of cross-domain features can be improved to a certain extent. Specifically,

[0118] ;

[0119] ;

[0120] where, is the mixing weight, is the degree of freedom of the StudentT distribution, is the shape parameter of the Beta distribution, is the truncation distribution, is the output feature map of the given encoder, rectified linear unit; this feature reconstruction method eliminates the influence of domain-specific style information, forcing the network to learn domain-independent features that focus on the underlying spatial structure and semantic content of remote sensing images. This is particularly important for remote sensing tasks because domain variations often obscure meaningful information.

[0121] The adjusted statistical data is integrated back into the feature map to generate augmented feature embeddings :

[0122] ;

[0123] S23 Feature Decoupling: To further enhance the feature generalization ability, the original feature embedding and the augmented feature embedding are decoupled through a shared encoder-decoder network. At each layer, domain-invariant instance decoupling is performed on the intermediate feature map and its corresponding augmented feature map , where , , are the height, width, and number of channels of the intermediate feature map , respectively.

[0124] Furthermore, for each feature map, hierarchical dynamic covariance decomposition is performed. By semantically guiding channel grouping, such as semantic categories like edges, textures, shapes, etc., the channels of the feature map are divided into K groups, and each group of features is , and weights are dynamically assigned based on feature importance :

[0125] ;

[0126] where GAP(.) represents global average pooling, is a learnable parameter vector corresponding to the weights of the K groups of channels, exp(.) represents the exponential function with the natural constant e as the base, and the same operation is performed on the augmented feature map to generate the grouped feature map and the dynamically assigned weights of feature importance ;

[0127] For each group of features, hierarchical covariance matrix calculation and weighted fusion are performed. The covariance matrix contains the spatial relationship of the feature map. The covariance matrices of the feature maps , are respectively:

[0128] ;

[0129] ;

[0130] where and are the means of each group of features , , respectively, , , are the row, column, and channel grouping serial numbers of the feature map, , are the covariance matrices of the feature maps and , respectively.

[0131] These covariance matrices contain the spatial relationships of the feature maps, which are crucial for remote sensing applications that require a detailed understanding of the spatial distribution. The traditional global calculation method of covariance matrices leads to the coupling of domain-related and domain-unrelated features. By applying semantic-guided channels for hierarchical covariance calculation, the robustness of domain-invariant feature decoupling is improved, and the computational complexity is reduced.

[0132] To improve variability, a dissimilarity measure is defined based on the covariance matrix:

[0133] ;

[0134] where represents the average covariance matrix. This metric ensures that the decoupled features are consistent with domain-invariant characteristics while retaining spatial semantic consistency.

[0135] Furthermore, to improve the method's ability to select cross-domain features of remote sensing images, the Gumbel-Softmax differentiable mask is applied to replace the hard binary mask with learnable parameters to enhance the feature extraction ability in weakly textured regions:

[0136] ;

[0137] Gumbel(0,1).

[0138] where represents a random coefficient following the Gumbel(0,1) distribution, a learnable control coefficient that controls the sparsity of the mask.

[0139] The co-design of dynamic group covariance decomposition and differentiable mask is optimized and enhanced in terms of computational efficiency, feature decoupling, and domain generalization stability, enabling it to better extract and retain the domain-invariant feature information existing in different image domains.

[0140] Furthermore, a new mask joint constraint term is added:

[0141] ;

[0142] Applying the joint constraint term prevents the uncontrollable sparsity of the mask and combines it with the whitening loss function to construct a joint decoupling loss function:

[0143] ;

[0144] Finally, the loss function of feature decoupling and the cross-entropy loss function together constitute the joint loss function of the domain generalization semantic segmentation method for remote sensing images:

[0145] .

[0146] wherein is the output feature of the fourth-stage encoder, is the label data of the source domain image.

[0147] By utilizing the domain-invariant instance decoupling (DII) module, the method effectively addresses the challenges brought about by complex spatial relationships and domain variability, enabling powerful, domain-invariant feature extraction for remote sensing applications.

[0148] The learning of domain-invariant features from remote sensing image data from the input multi-scale instance encoding module to the domain-invariant instance decoupling module is considered one stage, and the feature learning of four stages is repeated before performing domain generalization semantic decoding.

[0149] S30 Domain Generalization Semantic Decoding: The features decoupled by the S20 domain-invariant instance are input into the domain generalization semantic decoding module. The features of the last three stages of the encoder are aggregated and concatenated, and a lightweight head module is used to model the global context. Then, a multi-layer perceptron (MLP) with two hidden layers is used for per-pixel classification. By focusing on the last three stages, the over-detailing of the low-level information in the first stage is avoided, which is less relevant to high-level semantic tasks.

[0150] S40 Generalized Semantic Segmentation: The remote sensing images of unseen target domains are input into the trained framework to predict segmentation maps on unseen target domains.

[0151] S10 - S30 are the training stages, where only remote sensing images from the source domain are input into the framework. S40 is the inference stage, where remote sensing images from unseen target domains are input into the trained framework.

[0152] All experiments in this embodiment are conducted using Pytorch, and its implementation relies on the timm and mMSIEgmentation libraries to complete the segmentation task. The encoder in the segmentation model is pre-trained on the ImageNet-1K dataset.

[0153] These models are trained on nodes equipped with 8 RTX 3090 graphics processing units (GPUs). In the segmentation experiments, common data augmentation techniques are applied, including random horizontal flipping, random scaling between 0.5 and 2, and random cropping. The batch size is set to 8. The AdamW optimizer with an initial learning rate of 0.00006 and a polynomial learning rate decay strategy are used.

[0154] For the ISPRS dataset, the training consists of 160K iterations; for the WHU dataset, the training consists of 80K iterations. For the WHU dataset, data augmentation is applied to each sub-dataset by randomly flipping and rotating the training samples; in contrast, for the ISPRS dataset, 10,000 images of size 512×512 pixels are generated using traditional data augmentation techniques and randomly cropped from the augmented samples for training.

[0155] For evaluation, precision, recall, F1-score, and mean intersection over union (mIoU) are used to evaluate the performance of the model, specifically as follows:

[0156] IoU and mIoU are used as the main accuracy metrics for evaluating performance on the ISPRS and WHU datasets. Specifically, IoU evaluates the cross-over union of a single class, while mIoU represents the average IoU of all classes. For the WHU dataset, additional metrics such as precision, recall, and the harmonic mean Fi-score are used to comprehensively evaluate the accuracy of building extraction.

[0157] Precision is defined as the ratio of true positives (TP) to the sum of true positives and false positives (FP):

[0158] ;

[0159] Recall measures the ratio of true positives to the sum of true positives and false negatives (FN):

[0160] ;

[0161] Fi-score provides a balanced measure by combining precision and recall, and is calculated as follows:

[0162] ;

[0163] IoU is a key metric in the segmentation task for evaluating the overlap between the predicted region and the ground truth region for a given class. It is calculated by dividing the number of true positive pixels by the total number of pixels in the union of the predicted region and the ground truth region. For a multi-class dataset, the mean IoU (mIoU) aggregates by averaging the IoU values of all classes, providing an overall measure of segmentation accuracy. Here, N represents the total number of classes, as follows:

[0164] ;

[0165] Among these metrics, TP and TN (true negative) represent the number of samples correctly predicted as positive and negative examples respectively, while FP and FN represent the number of times negative and positive examples are mispredicted. These metrics together provide an assessment of the model's performance under different datasets and evaluation criteria.

[0166] The proposed method was validated on four experimental sub-datasets of the ISPRS and WHU datasets respectively. The overall evaluation result metrics are shown in Tables 1, 2, and 3.

[0167] Table 1 IoU of each ground object classification under the ISPRS dataset

[0168] ;

[0169] As shown in Table 1, taking the images of Vaihingen and Postdam in the ISPRS dataset as the source domain and the target domain respectively, the segmentation performance of six types of elements such as impervious surfaces, buildings, low vegetation, trees, cars, and background was evaluated. The method has the best segmentation effect on three types of elements: impervious surfaces, buildings, and cars, and the classification accuracy on the tree category is relatively low. The overall evaluation MIoU reached 53.9 and 62.1 respectively, and cross-domain remote sensing image segmentation can be better achieved on the ISPRS dataset.

[0170] Table 2 IoU of each ground object classification under the WHU dataset

[0171] ;

[0172] As shown in Table 2, taking Aerial and Satellite II of the WHU dataset as the source domain and the target domain respectively, the accuracy, recall, F1, and IoU of the proposed method for cross-domain classification of each ground object were evaluated. The results show that even in the face of significant differences in spatial resolution, spectral characteristics, etc., the proposed method still has high cross-domain segmentation performance for remote sensing images.

[0173] Table 3 Precision evaluation of each module in the model ablation experiment

[0174] ;

[0175] As shown in Table 3, through the ablation experiment, the influence of each module on the segmentation performance in the proposed method was analyzed. The configuration with all components (MSIE, DII, and SD) has the best performance, and the IoU is 44.3%.

[0176] It is 4.0% higher than the baseline model. The addition of MSIE enhances the feature representation by capturing multi-scale information, while DII and SD further improve the model's ability to effectively separate and decode semantic features. When the semantic decoding (SD) component is excluded, the performance drops to an IoU of 43.8%, a decrease of 0.5%. This highlights the importance of SD in the final stage of the model, which helps to accurately interpret features for segmentation. Excluding the domain-invariant instance decoupling (DII) component results in a further performance drop, with an IoU of 41.7%, a significant decrease of 2.1%. This shows that DII plays a key role in decoupling domain-specific features, which is crucial for improving performance.

[0177] Although the present invention has been described in detail above with general descriptions and specific embodiments, some modifications or improvements can be made based on the present invention, which are obvious to those skilled in the art. Therefore, these modifications or improvements made without departing from the spirit of the present invention fall within the scope of the present invention claimed.

Claims

1. A remote sensing image domain generalization semantic segmentation method based on multi-scale instance decoupling, characterized in that It includes the following steps: S10 Multi-scale instance encoding: The multi-scale instance encoding module takes the image features in the segmentation backbone network as input, uses depth convolution, cross-scale interaction encoding, integrates multi-scale features, constructs a deformation-frequency domain unit, and perceives the feature responses of cross-domain instances; S20 Domain-invariant instance decoupling: The domain-invariant instance decoupling module is based on the reconstruction of the mixed distribution features and covariance optimization, decouples at the feature level, and realizes the learning of domain-invariant features; S30 Domain generalization semantic decoding: The features after S20 domain-invariant instance decoupling are input into the domain generalization semantic decoding module, and the hamburger module is used to model the global context; S40 Generalized semantic segmentation: The remote sensing image of the unseen target domain is input into the trained framework to realize the prediction of the segmentation map on the unseen target domain.

2. The remote sensing image domain generalization semantic segmentation method according to claim 1, characterized in that The S10 multi-scale instance encoding specifically includes the following steps: S11 Input image features: Input the original image from the source domain, divide the original image into squares with a length * width * channel size of 16 * 16 * 3, map each small block into a one-dimensional vector, form the original image feature sequence and input it into the network to generate a one-dimensional image block sequence; S12 Cross-scale interaction encoding: Its core operation is defined mathematically as, ; Among them, F represents the input feature, is a scaling operation to adapt features with different receptive fields, and DW-Conv represents depthwise convolution, is a cross-scale interaction mechanism; The expression of the cross-scale interaction mechanism is as follows: ; ; ; Among them, is the high-level scale feature, is the low-level scale feature, is the bilinear upsampling, represents the sigmoid function, (.) is the aggregation and splicing operation; the high-level scale feature is subjected to upsampling and depth convolution operations, and the low-level scale feature is subjected to convolution operation to generate feature maps that can match the current scale feature and respectively. After cross-scale interaction of the scale features at three levels, a feature map that can replace the current scale feature is generated; S13 Multi-scale feature integration: The context attention mechanism is used to integrate the feature information of different scales, and dynamically assigns importance to the features from different scales. Its mathematical expression is: ; ; Among them, is the result of cross-scale interaction coding in step S12, is a learnable convolutional layer for generating attention weights, represents element-wise multiplication, represents weights, represents the attention-weighted feature map. The encoder emphasizes the most informative features while suppressing irrelevant information. These re-weighted features are summed and passed through an additional convolutional layer to generate the final encoded representation; S14 Deformation-frequency domain unit design: Design a deformation-frequency domain special-shaped convolution group to differentially process multi-branch features; S15 Scale normalization: Normalize the features, which is expressed as: ; ; where μ and σ are the mean and standard deviation of the feature respectively, and γ and β are learnable parameters, is the output feature map of the deformation-frequency domain unit in the S14 step.

3. The remote sensing image domain generalization semantic segmentation method according to claim 2, wherein The S14 deformation-frequency domain unit design is specifically that the deformable convolution and the dilation rate are adaptively combined in a coordinated manner, and its expression is: ; Among them, is the learnable offset, is the dilation rate adaption, , is a multi-layer perceptron with 2 hidden layers, is the attention-weighted feature map which is the result after the collaborative combination of deformable convolution and dilation rate adaption; Using the two-dimensional discrete cosine transform operation, separate the high-frequency texture and the low-frequency semantics, suppress the sensor noise and retain the detail information. Its core operation is: ; Among them, DCT2D is the two-dimensional discrete cosine transform, IDCT2D is its inverse transform, Conv is the standard convolution, and the attention-weighted feature map generates a transformed feature map after two-dimensional discrete cosine transform , is the transformed feature map which is the result after 5x5 standard convolution, and generates a two-dimensional discrete cosine transform weighted feature map after inverse transform .

4. The remote sensing image domain generalization semantic segmentation method according to claim 1, characterized in that The S20 domain-invariant instance decoupling specifically includes the following steps: S21 Instance Feature Trend Calculation: Given the output feature map of the encoder , calculate the central tendency and dispersion of the instance-level features ; ; reflects the overall intensity of the feature map, while quantifies the variability between spatial elements, , , are the height, width, and number of channels of the input feature map F in step S12, respectively; S22 Construct a remote sensing instance feature reconstruction model: Construct a remote sensing instance feature reconstruction model based on the Beta-Student mixture distribution to replace the original and , specifically, ; ; Among them, is the mixing weight, is the degree of freedom of the Student T distribution, is the shape parameter of the Beta distribution, is the truncation distribution, is the output feature map of the given encoder, rectified linear unit; The adjusted statistical data is integrated back into the feature map to generate an augmented feature embedding : ; S23 Feature decoupling: The original feature embedding and the augmented feature embedding are decoupled through a shared encoder-decoder network, and the intermediate feature maps and their augmented corresponding feature maps are subjected to domain-invariant instance decoupling, where , , are the height, width, and number of channels of the intermediate feature map respectively.

5. The remote sensing image domain generalization semantic segmentation method according to claim 4, wherein The S23 feature decoupling is specifically, For each feature map perform hierarchical dynamic covariance decomposition, divide the channels of the feature map into K groups, and each group of features is , dynamically assign weights based on feature importance : ; where GAP(.) represents global average pooling, is a learnable parameter vector corresponding to the weights of K groups of channels, exp(.) represents the exponential function with the natural constant e as the base, and the same operation is performed on the extended feature map to generate a grouped feature map and dynamically assign weights to feature importance ; Calculate the hierarchical covariance matrix for each group of features separately and perform weighted fusion. The covariance matrix contains the spatial relationship of the feature maps, and the feature maps and The covariance matrices are respectively ; ; where and are the means of each group of features 、 respectively, 、 、 are the row, column, and channel grouping numbers of the feature map respectively, 、 are and the covariance matrices of the feature maps respectively; Define a difference metric based on the covariance matrix: ; Among them, represents the average covariance matrix.

6. The remote sensing image domain generalization semantic segmentation method according to claim 5, wherein, For the S23 feature decoupling, apply a differentiable mask to replace the hard binary mask with learnable parameters: ; Gumbel(0,1); where denotes a random coefficient following the Gumbel(0,1) distribution, is a learnable control coefficient that controls the sparsity of the mask.

7. The remote sensing image domain generalization semantic segmentation method according to claim 6, wherein The S23 feature decoupling also includes a mask joint constraint term: ; And in combination with the whitening loss function, construct a joint decoupling loss function: ; Finally, the loss function of feature decoupling and the cross-entropy loss function together constitute the joint loss function of the remote sensing image domain generalization semantic segmentation method: ; Among them is the output feature of the encoder in the fourth stage, is the label data of the source domain image.

8. The remote sensing image domain generalization semantic segmentation method according to claim 1, wherein The learning of domain-invariant features from the input of the remote sensing image data to the multi-scale instance encoding module to the domain-invariant instance decoupling module is one stage. Four stages of feature learning are repeated before the domain generalization semantic decoding. The features extracted in the second, third, and fourth stages are jointly used as the feature input for the domain generalization semantic decoding.

9. The remote sensing image domain generalization semantic segmentation method according to claim 8, wherein, The S30 domain generalization semantic decoding is specifically to aggregate and splice the features of the last three stages of the encoder, use a lightweight hamburger module to model the global context, and then use a multi-layer perceptron for per-pixel classification.

10. The remote sensing image domain generalization semantic segmentation method according to claim 1, wherein S10 - S30 is the training stage, and only the remote sensing images from the source domain are input into the framework. S40 is the inference stage, and the remote sensing images from the unseen target domain are input into the trained framework.