Remote sensing image domain generalization semantic segmentation method based on multi-scale instance decoupling

Through multi-scale instance coding and domain-invariant instance decoupling module, domain-invariant feature representation is learned, and global context modeling is combined with the Hamburg module and multi-layer perception machine, the problem of cross-domain generalization of remote sensing image segmentation model is solved, and efficient semantic segmentation on the target domain has been achieved.

CN119992289AActive Publication Date: 2025-05-13GUIZHOU SURVEY & DESIGN RES INST FOR WATER RESOURCES & HYDROPOWER

Patent Information

Application Number
CN202510484907.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-05-13
Estimated Expiration
2045-04-17

AI Technical Summary

Technical Problem

When facing the differences in the feature distribution of training data and inference data, it is difficult to achieve effective cross-domain generalization, resulting in a degradation in the segmentation performance of the model on the target domain that has not been seen before.

Method used

The domain generalization semantic segmentation method based on multi-scale instance decoupling is adopted. Through the multi-scale instance encoding (MSIE) and domain invariant instance decoupling (DII) modules, the domain invariant feature representation is learned, and the domain generalization semantic decoding module is used to model global context and phase-by-phase element classification using the Hamburg module and multi-layer perceptron.

Benefits of technology

Effective semantic segmentation on unseen target domains is achieved, which improves the cross-domain generalization capability of the model, reduces dependence on specific target domain data, and reduces computational load, making it possible to process larger data sets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992289A_ABST
    Figure CN119992289A_ABST
Patent Text Reader

Abstract

The invention relates to a remote sensing image domain generalization semantic segmentation method based on multi-scale instance decoupling, which relates to the technical field of remote sensing image interpretation and comprises the steps of multi-scale instance coding, domain invariant instance decoupling and domain generalization semantic decoding. Image features in a segmented backbone network are used as input, and a multi-scale instance coding module effectively senses feature responses of cross-domain instances through deep convolution, multi-scale cooperation and multi-path sensing and is not affected by scale differences, spatial deformation and noise; designing a domain invariant instance decoupling module to decouple the cross-domain style from the image representation; and finally, the decoupled features are input into a domain generalization semantic decoding module so as to predict a segmented image on a target domain which is not seen, and the segmentation method can generalize various remote sensing images which are not seen so as to obtain better segmentation capability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of remote sensing image interpretation, and in particular to a remote sensing image domain generalized semantic segmentation method based on multi-scale instance decoupling. Background Art

[0002] Remote sensing image interpretation is a basic task of earth observation. It plays an important role in many applications, such as land use classification, geospatial object extraction, and surface time series monitoring. In the remote sensing image interpretation task, remote sensing semantic segmentation occupies a unique position. It can realize the recognition and classification of each ground feature element. With the development of deep learning technology, remote sensing image interpretation methods based on deep learning have been widely studied, which has greatly promoted the development of remote sensing image semantic segmentation methods. However, most existing remote sensing image segmentation methods still assume that the training data and the inference data follow the same and independent statistical distribution. These segmentation methods are usually trained on a remote sensing training set of a specific benchmark and then verified on a test set of the same benchmark. The remote sensing images in the same benchmark usually have the same feature distribution. In actual applications, remote sensing image data often come from different observation platforms and sensors. Factors such as surface landscape, spatial resolution, lighting conditions, and sensor parameters will cause significant differences in the feature distribution of inference data (also known as the target domain) and training data (also known as the source domain). At this time, the segmentation model trained on the training data set will not be able to effectively achieve semantic segmentation of the target domain image. This deviation in feature distribution between training data and inference data is called inter-domain difference. The existence of inter-domain difference will greatly reduce the generalization ability of remote sensing image semantic segmentation model.

[0003] In the prior art, people have carried out various works to improve the cross-domain generalization ability of remote sensing image segmentation models. Most of these methods are based on the idea of ​​domain adaptation. By aligning the feature distribution of source domain images and target domain images during training, the cross-domain generalization ability of segmentation models is improved. Although this method can achieve significant segmentation performance improvement on target domain images, it is difficult to obtain the labeled data of target domain images at the same time in practical applications. In order to solve this weakness, unsupervised domain adaptive remote sensing image semantic segmentation methods are proposed. They no longer require the true value of target domain images during training, but reduce the domain deviation between source domain and target domain through methods such as data generation and low-cost functions. Since the target domain is involved in the training process, the segmentation models constructed by domain adaptation and unsupervised domain adaptation methods can only be applied to remote sensing data of a specific target domain. For other target domain data that are not involved in training, the generalization segmentation ability of the model will be affected. This weakness limits the versatility and practical application of domain adaptation methods and unsupervised domain adaptation methods. Ideally, remote sensing image segmentation models should be able to reason on various remote sensing images that have not been seen during training.

[0004] To solve this problem, the inventor team introduced a domain generalization method for remote sensing image segmentation, which eliminates the need for a target domain in training and can generalize to other target domain images. Summary of the invention

[0005] In order to solve the above problems, the purpose of the present invention is to provide a generalized semantic segmentation method in remote sensing image domain based on multi-scale instance decoupling, comprising the following steps: S10 Multi-scale Instance Encoding (MSIE): The multi-scale instance encoding module takes the image features in the segmentation backbone network as input, uses deep convolution, cross-scale interactive encoding, integrates multi-scale features, constructs deformation-frequency domain units, and perceives the feature response of cross-domain instances; S20 Domain-Invariant Instance Decoupling (DII): The domain-invariant instance decoupling module performs decoupling at the feature level based on mixed distribution feature reconstruction and covariance optimization to achieve learning of domain-invariant features; S30 domain generalization semantic decoding: The features after decoupling from the S20 domain invariant instances are input into the domain generalization semantic decoding module, and the hamburger module is used to model the global context; S40 Generalized semantic segmentation: Remote sensing images of unseen target domains are input into the trained framework to predict segmentation maps on unseen target domains.

[0006] Furthermore, the S10 multi-scale instance encoding (MSIE) specifically includes the following steps: S11 Input image features: Input the original image from the source domain, divide the original image into patches of size 16*16*3 (length*width*channel), and map each patch to a one-dimensional vector to form an original image feature sequence and input it into the network. The operation of generating a one-dimensional image patch sequence is referred to as Patch Embedding. S12 Cross-Scale Interaction Encoding (CSE): Its core operation is mathematically defined as, ; Where F represents the input image feature, To scale the operation to adapt to the characteristics of different receptive fields, represents element-wise multiplication, DW-Conv represents depth-wise convolution, It is a cross-scale interaction mechanism; this design ensures that information from different scales can be integrated synergistically, solving the problem of cross-domain scale variability. The output of the deep convolution undergoes cross-scale interaction operations, reducing the semantic fragmentation caused by independent processing of features at each scale. The scale-up operation is performed to accommodate multiple receptive fields, enabling the encoder to capture both local and global contextual information.

[0007] The cross-scale interaction mechanism expression is as follows: ; ; ; in, is a high-level scale feature. is a low-level scale feature, is bilinear upsampling, represents the sigmoid function, (.) is the aggregation and splicing operation; high-level scale features After upsampling and deep convolution operations, low-level scale features are Convolution operation, respectively generating features that can match the current scale Feature map and After the scale features of the three levels interact across scales, they can generate features that can replace the current scale features. Feature map The cross-scale interaction mechanism can establish topological dependencies between multi-scale features, high-level semantics can propagate from top to bottom to guide low-level features, and low-level details can correct high-level predictions from bottom to top, thus achieving two-way propagation of high- and low-level features.

[0008] S13 Multi-scale feature integration: The contextual attention mechanism is used to integrate feature information of different scales, dynamically assign importance to features from different scales, and further integrate feature information of different scales. Its mathematical expression is: ; ; in, is the result of cross-scale interactive encoding in step S12, is a learnable convolutional layer for generating attention weights, represents element-wise multiplication, represents the weight, Represents the attention-weighted feature map; by applying these attention weights, the encoder emphasizes the most informative features while suppressing irrelevant information, thereby achieving more effective aggregation of multi-scale features. These re-weighted features are summed and passed through an additional convolutional layer to produce the final encoded representation.

[0009] S14 Deformation-Frequency Domain Unit (DFSU) Design: Design a deformation-frequency domain heterogeneous convolution group to improve the robustness of cross-domain image segmentation through differentiated processing of multi-branch features.

[0010] S15 Scale Normalization: In cross-domain scenarios, differences in spatial resolution often lead to domain shifts. To alleviate this situation, the encoder introduces a scale normalization step to ensure the consistency of feature representation at different scales, which is expressed as: ; ; Among them, μ and σ are the mean and standard deviation of the features, γ and β are learnable parameters. The deformation-frequency domain unit of step S14 outputs a feature map; the features are standardized, the deviation of specific fields is reduced, and the ability of the encoder to generalize on aerial and satellite images is enhanced.

[0011] Furthermore, the S14 deformation-frequency domain unit (DFSU) is designed to design a deformation-frequency domain heterogeneous convolution group, which improves the robustness of cross-domain image segmentation through differentiated processing of multi-branch features, and uses the synergistic combination of deformation convolution and hole rate adaptation to improve the perception of geometric deformation of cross-domain features. Its expression is: ; in, is the learnable offset, is the hole rate adaptation, , is a multilayer perceptron with 2 hidden layers. Weighted feature map for attention The result after the synergistic combination of deformable convolution and hole rate adaptation; It solves the problem that the geometric morphology of image objects in different sensor platforms and different regions is significantly different, and is affected by the atmosphere and imaging hardware stability. In the cross-domain remote sensing image segmentation task, in addition to the influence of multi-scale features, there are also problems such as sensor noise and feature deformation.

[0012] The two-dimensional discrete cosine transform (DCT2D) operation is used to separate high-frequency textures from low-frequency semantics, suppress sensor noise and retain detail information. The core operation is: ; Among them, DCT2D is a two-dimensional discrete cosine transform, IDCT2D is its inverse transform, Conv is a standard convolution, and the attention weighted feature map After two-dimensional discrete cosine transform (DCT2D), the transformed feature map is generated , Transform feature map The result after 5x5 standard convolution is inverse transformed to generate a 2D discrete cosine transform weighted feature map .

[0013] Furthermore, the S20 domain invariant instance decoupling (DII) specifically includes the following steps: S21 Example feature trend calculation: given the output feature map of the encoder , calculate the central tendency of instance-level features and dispersion , ; ; reflects the overall strength of the feature map, while quantifies the variability between spatial elements. , , The height, width and number of channels of the feature map F are input to step S12 respectively; these indicators effectively capture the spatial complexity and domain-specific characteristics of remote sensing data.

[0014] S22 Constructing a remote sensing instance feature reconstruction model: Remote sensing data usually involves complex relationships between objects, textures, and environmental backgrounds. In order to eliminate bias in specific areas, and Rescale to the interval [0,1] and use the statistical distribution method to generate randomness; introduce the Bayesian probability framework, build a remote sensing instance feature reconstruction model based on the Beta-Student mixed distribution, and replace the original and ,The domain adaptability of cross-domain features can be improved to a certain extent through mixed distribution operations. Specifically, ; ; in, is the mixing weight, is the degrees of freedom of StudentT distribution, is the shape parameter of the Beta distribution, To truncate distributed, is the output feature map of a given encoder, Linear rectifier function; This feature reconstruction method eliminates the influence of domain-specific style information and forces the network to learn domain-independent features that focus on the underlying spatial structure and semantic content of remote sensing images. This is particularly important for remote sensing tasks because domain changes often obscure meaningful information.

[0015] The adjusted statistics are integrated back into the feature map to generate augmented feature embeddings : ; S23 Feature Decoupling: To further enhance feature generalization, the original feature embedding and the expanded feature embedding are decoupled through a shared encoder-decoder network. At each layer, the intermediate feature map And its expanded corresponding feature map Decouple the domain-invariant instances, where , , They are intermediate feature maps The height, width and number of channels.

[0016] Furthermore, for each feature map, hierarchical dynamic covariance decomposition is performed, and the channels of the feature map are divided into K groups through semantically guided channel grouping, such as edge, texture, shape and other semantic categories. Each group of features is , dynamically assign weights based on feature importance : ; Where GAP(.) represents global average pooling, is a learnable parameter vector corresponding to the weights of K groups of channels. exp(.) represents an exponential function with the natural constant e as the base, which is used to expand the feature map. Perform the same operation to generate group feature maps Dynamically assign weights based on feature importance ; The hierarchical covariance matrix is ​​calculated for each set of features and weighted fusion is performed. The feature map , The covariance matrices are: ; ; in and For each group of features , The mean of , , are the row, column and channel grouping numbers of the feature map, respectively. , The feature maps are and The covariance matrix of .

[0017] These covariance matrices contain the spatial relationships of feature maps, which is crucial for remote sensing applications that require detailed understanding of spatial distribution. The traditional way of global calculation of covariance matrices leads to the coupling of domain-dependent and domain-independent features. By applying semantically guided channels for hierarchical covariance calculation, the robustness of domain-invariant feature decoupling is improved and the computational complexity is reduced.

[0018] To improve variability, a variance measure is defined based on the covariance matrix: ; in represents the mean covariance matrix. This metric ensures that the disentangled features are consistent with domain-invariant properties while preserving spatial semantic consistency.

[0019] Furthermore, in order to improve the method's ability to select cross-domain features of remote sensing images, the Gumbel-Softmax differentiable mask is applied to replace the hard binary mask with a learnable parameter to improve the feature extraction capability in weak texture areas: ; Gumbel(0,1).

[0020] in represents the random coefficients that follow the Gumbel (0,1) distribution, Learnable control coefficients to control mask sparsity.

[0021] The collaborative design of dynamic grouped covariance decomposition and differentiable masking optimizes and improves computational efficiency, feature decoupling and domain generalization stability, enabling it to better extract and retain domain-invariant feature information that exists in different image domains.

[0022] Furthermore, a new mask joint constraint is added: ; Apply joint constraints to prevent uncontrollable mask sparsity, and combine with the whitening loss function to construct a joint decoupling loss function: .

[0023] Finally, the loss function for feature decoupling Together with the cross entropy loss function, it constitutes the joint loss function of the generalized semantic segmentation method in the remote sensing image domain: .

[0024] in is the output feature of the fourth stage encoder, is the label data of the source domain image.

[0025] By leveraging domain-invariant instance decoupling of the DII module, our approach effectively addresses the challenges posed by complex spatial relations and domain variability, enabling robust, domain-invariant feature extraction for remote sensing applications.

[0026] Furthermore, the learning of domain-invariant features of remote sensing image data from the input multi-scale instance encoding module to the domain-invariant instance decoupling module is one stage. The four stages of feature learning are repeated before domain generalization semantic decoding. The features extracted in the second, third, and fourth stages are used as the feature input of domain generalization semantic decoding.

[0027] Furthermore, the features of the last three stages of the encoder are concatenated and a lightweight hamburger module (Head) is used to model the global context, and then a multi-layer perceptron (MLP) with two hidden layers is used for phase-by-phase meta-classification. By focusing on the last three stages, the over-detailing of low-level information in the first stage is avoided, which is less relevant for high-level semantic tasks.

[0028] Furthermore, S10-S30 are the training phase, where only remote sensing images from the source domain are input into the framework, and S40 is the inference phase, where remote sensing images from the unseen target domain are input into the trained framework.

[0029] The beneficial effects of the present invention are: 1. Based on the general idea of ​​multi-scale extraction, a multi-scale feature interaction mechanism and deformation-frequency domain unit are constructed to learn stable instance representations at multiple scales, suppress the influence of noise, and maintain semantic consistency.

[0030] 2. A domain-invariant instance decoupling (DII) method based on mixed distribution feature reconstruction and covariance optimization was constructed to reduce the impact of domain-specific features in image appearance on the generalization ability of semantic segmentation.

[0031] 3. Realize domain-invariant feature representation and integrate features from multiple stages to capture local details and global context, which is crucial for understanding complex spatial relationships in remote sensing scenes. It does not require the presence of a target domain during training and can be generalized to other target domain images to predict segmented images on unseen target domains. At the same time, the reduction in computational load makes it possible to process larger data sets. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1: This is a network framework diagram of the present invention (Source Domain is the source domain, Unseen Domain is the unseen target domain, Patch Embedding is the operation of generating a one-dimensional image block sequence, MSIE Block 1, MSIE Block 2, MSIEBlock 3, and MSIE Block 4 are multi-scale instance encoding modules 1, 2, 3, and 4 of the four stages respectively, DIIdisentanglement is domain-invariant instance decoupling, Concatenate is a feature aggregation concatenation operation, Head is a lightweight hamburger module, MLP is a multi-layer perceptron with 2 hidden layers, Label is the label data of the source domain image, Loss is a joint loss function, Unseen Prediction is the image prediction segmentation result of the unseen target domain, and Stage1-4 are stages 1-4 respectively); Figure 2 It is a structural diagram of the multi-scale example encoding module of the present invention (GAT is a cross-scale interactive gate, DCT and IDCT are discrete cosine transform and its inverse transform); Figure 3 This is the domain-invariant instance decoupling module structure diagram of the present invention ( is the covariance ); Figure 4 The image segmentation results of various feature types on the ISPRS dataset in Example 1 of the present invention; Figure 5 This is the image segmentation result of building types on the WHU dataset in Example 1 of the present invention. DETAILED DESCRIPTION

[0033] The present invention will be described in detail below in conjunction with the accompanying drawings and specific embodiments. The exemplary embodiments and descriptions of the present invention are used to explain the present invention but are not intended to limit the present invention.

[0034] Example 1

[0035] A generalized semantic segmentation method for remote sensing image domain based on multi-scale instance decoupling includes the following steps: S10 Multi-scale Instance Encoding (MSIE): The multi-scale instance encoding module takes the image features in the segmentation backbone network as input, uses deep convolution, cross-scale interactive encoding, integrates multi-scale features, constructs deformation-frequency domain units, and perceives the feature response of cross-domain instances; S11 Input image features: Input the original image from the source domain, divide the original image into patches of size 16*16*3 (length*width*channel), and map each patch to a one-dimensional vector to form an original image feature sequence and input it into the network. The process of generating a one-dimensional image patch sequence is referred to as Patch Embedding. S12 Cross-Scale Interaction Encoding (CSE): Its core operation is mathematically defined as, ; Where F represents the input image feature, To scale the operation to adapt to the characteristics of different receptive fields, represents element-wise multiplication, DW-Conv represents depth-wise convolution, It is a cross-scale interaction mechanism; this design ensures that information from different scales can be integrated synergistically, solving the problem of cross-domain scale variability. The output of the deep convolution undergoes cross-scale interaction operations, reducing the semantic fragmentation caused by independent processing of features at each scale. The scale-up operation is performed to accommodate multiple receptive fields, enabling the encoder to capture both local and global contextual information.

[0036] The cross-scale interaction mechanism expression is as follows: ; ; ; is bilinear upsampling, represents the sigmoid function, (.) is the aggregation and splicing operation; high-level scale features After upsampling and deep convolution operations, low-level scale features are Convolution operation, respectively generating features that can match the current scale Feature map and After the scale features of the three levels interact across scales, they can generate features that can replace the current scale features. Feature map The cross-scale interaction mechanism can establish topological dependencies between multi-scale features, high-level semantics can propagate from top to bottom to guide low-level features, and low-level details can correct high-level predictions from bottom to top, thus achieving two-way propagation of high- and low-level features.

[0037] S13 Multi-scale feature integration: The contextual attention mechanism is used to integrate feature information of different scales, dynamically assign importance to features from different scales, and further integrate feature information of different scales. Its mathematical expression is: ; ; in, is the result of cross-scale interactive encoding in step S12, is a learnable convolutional layer for generating attention weights, represents element-wise multiplication, represents the weight, Represents the attention-weighted feature map; by applying these attention weights, the encoder emphasizes the most informative features while suppressing irrelevant information, thereby achieving more effective aggregation of multi-scale features. These re-weighted features are summed and passed through an additional convolutional layer to produce the final encoded representation.

[0038] S14 Deformation-Frequency Domain Unit (DFSU) Design: Design a deformation-frequency domain heterogeneous convolution group to improve the robustness of cross-domain image segmentation through differentiated processing of multi-branch features.

[0039] The deformable convolution and dilation rate adaptation are used in synergistic combination to improve the perception of geometric deformation of cross-domain features. The expression is: ; in, is the learnable offset, is the hole rate adaptation, , is a multilayer perceptron with 2 hidden layers. Weighted feature map for attention The result after the synergistic combination of deformable convolution and hole rate adaptation; It solves the problem that the geometric morphology of image objects in different sensor platforms and different regions is significantly different, and is affected by the atmosphere and imaging hardware stability. In the cross-domain remote sensing image segmentation task, in addition to the influence of multi-scale features, the problems of sensor noise and feature deformation are also affected.

[0040] The two-dimensional discrete cosine transform (DCT2D) operation is used to separate high-frequency textures from low-frequency semantics, suppress sensor noise and retain detail information. The core operation is: ; Among them, DCT2D is a two-dimensional discrete cosine transform, IDCT2D is its inverse transform, Conv is a standard convolution, and the attention weighted feature map After two-dimensional discrete cosine transform (DCT2D), the transformed feature map is generated , Transform feature map The result after 5x5 standard convolution is inverse transformed to generate a 2D discrete cosine transform weighted feature map .

[0041] S15 Scale Normalization: In cross-domain scenarios, differences in spatial resolution often lead to domain shifts. To alleviate this, the encoder introduces a scale normalization step to ensure the consistency of feature representations at different scales, which is expressed as: ; ; Among them, μ and σ are the mean and standard deviation of the features, γ and β are learnable parameters. The deformation-frequency domain unit of step S14 outputs a feature map; the features are standardized, the deviation of specific fields is reduced, and the ability of the encoder to generalize on aerial and satellite images is enhanced.

[0042] S20 Domain-Invariant Instance Decoupling (DII): The domain-invariant instance decoupling module performs decoupling at the feature level based on mixed distribution feature reconstruction and covariance optimization to achieve learning of domain-invariant features; S21 Example feature trend calculation: given the output feature map of the encoder , calculate the central tendency of instance-level features and dispersion , ; ; reflects the overall strength of the feature map, while quantifies the variability between spatial elements. , , The height, width and number of channels of the feature map F are input to step S12 respectively; these indicators effectively capture the spatial complexity and domain-specific characteristics of remote sensing data.

[0043] S22 Constructing a remote sensing instance feature reconstruction model: Remote sensing data usually involves complex relationships between objects, textures, and environmental backgrounds. In order to eliminate bias in specific areas, and Rescale to the interval [0,1] and use the statistical distribution method to generate randomness; introduce the Bayesian probability framework, build a remote sensing instance feature reconstruction model based on the Beta-Student mixed distribution, and replace the original and ,The domain adaptability of cross-domain features can be improved to a certain extent through mixed distribution operations. Specifically, ; ; in, is the mixing weight, is the degrees of freedom of StudentT distribution, is the shape parameter of the Beta distribution, To truncate distributed, is the output feature map of a given encoder, Linear rectifier function; This feature reconstruction method eliminates the influence of domain-specific style information and forces the network to learn domain-independent features that focus on the underlying spatial structure and semantic content of remote sensing images. This is particularly important for remote sensing tasks because domain changes often obscure meaningful information.

[0044] The adjusted statistics are integrated back into the feature map to generate augmented feature embeddings : ; S23 Feature Decoupling: To further enhance feature generalization, the original feature embedding and the expanded feature embedding are decoupled through a shared encoder-decoder network. At each layer, the intermediate feature map And its expanded corresponding feature map Decouple the domain-invariant instances, where , , They are intermediate feature maps The height, width and number of channels.

[0045] Furthermore, for each feature map, hierarchical dynamic covariance decomposition is performed, and the channels of the feature map are divided into K groups through semantically guided channel grouping, such as edge, texture, shape and other semantic categories. Each group of features is , dynamically assign weights based on feature importance : ; Where GAP(.) represents global average pooling, is a learnable parameter vector corresponding to the weights of K groups of channels. exp(.) represents an exponential function with the natural constant e as the base, which is used to expand the feature map. Perform the same operation to generate group feature maps Dynamically assign weights based on feature importance ; The hierarchical covariance matrix is ​​calculated for each set of features and weighted fusion is performed. The covariance matrix contains the spatial relationship of the feature map. , The covariance matrices are: ; ; in and For each group of features , The mean of , , are the row, column and channel grouping numbers of the feature map, respectively. , The feature maps are and The covariance matrix of .

[0046] These covariance matrices contain the spatial relationships of feature maps, which is crucial for remote sensing applications that require detailed understanding of spatial distribution. The traditional way of global calculation of covariance matrices leads to the coupling of domain-dependent and domain-independent features. By applying semantically guided channels for hierarchical covariance calculation, the robustness of domain-invariant feature decoupling is improved and the computational complexity is reduced.

[0047] To improve variability, a variance measure is defined based on the covariance matrix: ; in represents the mean covariance matrix. This metric ensures that the disentangled features are consistent with domain-invariant properties while preserving spatial semantic consistency.

[0048] Furthermore, in order to improve the method's ability to select cross-domain features of remote sensing images, the Gumbel-Softmax differentiable mask is applied to replace the hard binary mask with a learnable parameter to improve the feature extraction capability in weak texture areas: ; Gumbel(0,1).

[0049] in represents the random coefficients that follow the Gumbel (0,1) distribution, Learnable control coefficients to control mask sparsity.

[0050] The collaborative design of dynamic grouped covariance decomposition and differentiable masking optimizes and improves computational efficiency, feature decoupling and domain generalization stability, enabling it to better extract and retain domain-invariant feature information that exists in different image domains.

[0051] Furthermore, a new mask joint constraint is added: ; Apply joint constraints to prevent uncontrollable mask sparsity, and combine with the whitening loss function to construct a joint decoupling loss function: ; Finally, the loss function for feature decoupling Together with the cross entropy loss function, it constitutes the joint loss function of the generalized semantic segmentation method in the remote sensing image domain: .

[0052] in is the output feature of the fourth stage encoder, is the label data of the source domain image.

[0053] By leveraging domain-invariant instance decoupling of the DII module, our approach effectively addresses the challenges posed by complex spatial relations and domain variability, enabling robust, domain-invariant feature extraction for remote sensing applications.

[0054] The remote sensing image data is learned from the input multi-scale instance encoding module to the domain-invariant instance decoupling module to achieve domain-invariant features. The four stages of feature learning are repeated before domain generalization semantic decoding.

[0055] S30 domain generalization semantic decoding: The decoupled features of the S20 domain invariant instances are input into the domain generalization semantic decoding module, the features of the last three stages of the encoder are concatenated, and a lightweight hamburger module (Head) is used to model the global context, and then a multi-layer perceptron (MLP) with two hidden layers is used for phase-by-phase meta-classification. By focusing on the last three stages, the over-detailing of low-level information in the first stage is avoided, which is less relevant for high-level semantic tasks.

[0056] S40 Generalized semantic segmentation: Remote sensing images of unseen target domains are input into the trained framework to predict segmentation maps on unseen target domains.

[0057] S10-S30 are the training phase, where only remote sensing images from the source domain are input into the framework, and S40 is the inference phase, where remote sensing images from unseen target domains are input into the trained framework.

[0058] All experiments in this example are performed using Pytorch, and their implementation relies on the timm and mMSIEgmentation libraries to complete the segmentation task. The encoder in the segmentation model is pre-trained on the ImageNet-1K dataset.

[0059] These models were trained on a node equipped with 8 RTX 3090 graphics processing units (GPUs). In the segmentation experiments, common data augmentation techniques were applied, including random horizontal flipping, random scaling in the range of 0.5 to 2, and random cropping. The batch size was set to 8. The AdamW optimizer with an initial learning rate of 0.00006 and a polynomial learning rate decay strategy was used.

[0060] For the ISPRS dataset, the training consists of 160K iterations, and for the WHU dataset, the training consists of 80K iterations. For the WHU dataset, data augmentation is applied to each sub-dataset by randomly flipping and rotating the training samples; in contrast, for the ISPRS dataset, 10,000 images of size 512× 512 pixels are generated using traditional data augmentation techniques and randomly cropped from the augmented samples for training.

[0061] For evaluation, precision, recall, F1 score, and mean intersection over union (mIoU) are used to evaluate the performance of the model, specifically: IoU and mIoU are used as the main accuracy metrics to evaluate performance on the ISPRS and WHU datasets. Specifically, IoU evaluates the intersection-super-union of a single category, while mIoU represents the average IoU of all categories. For the WHU dataset, additional metrics such as precision, recall, and harmonic mean Fi-score are used to comprehensively evaluate the accuracy of building extraction.

[0062] Precision is defined as the ratio of true positives (TP) to the sum of true positives and false positives (FP): ; Recall measures the ratio of true positives to the sum of true positives and false negatives (FN): ; Fi-score provides a balanced measure by combining precision and recall, and is calculated as follows: ; IoU is a key metric in segmentation tasks, used to evaluate the overlap between the predicted region and the true region in a given category. It is calculated by dividing the number of true positive pixels by the total number of pixels in the union of the predicted region and the true region. For multi-category datasets, the mean IoU (mIoU) provides an overall segmentation accuracy measure by averaging the IoU values ​​of all categories. Here, N represents the total number of categories, as shown below: ; Among these indicators, TP and TN (True Negative) represent the number of samples correctly predicted as positive and negative, respectively, while FP and FN represent the number of incorrect predictions for negative and positive samples. Together, these indicators provide an understanding of the performance of the model under different datasets and evaluation criteria.

[0063] The proposed method is verified on four experimental sub-datasets of the ISPRS and WHU datasets respectively. The overall evaluation results are shown in Tables 1, 2 and 3.

[0064] Table 1 IoU of various feature categories in the ISPRS dataset ; As shown in Table 1, the images of Vaihingen and Postdam in the ISPRS dataset are used as the source domain and target domain respectively to evaluate the segmentation performance of six types of elements, including impervious surfaces, buildings, low vegetation, trees, cars, and background. The method has the best segmentation effect on three types of elements, including impervious surfaces, buildings, and cars. The classification accuracy on the tree category is relatively low, and the overall evaluation MIoU reaches 53.9 and 62.1 respectively. It can achieve cross-domain remote sensing image segmentation on the ISPRS dataset.

[0065] Table 2 IoU of various object categories in the WHU dataset ; As shown in Table 2, the Aerial and Satellite II datasets of the WHU dataset are used as the source domain and target domain respectively to evaluate the precision, recall, F1, and IoU of the proposed method for cross-domain classification of objects. The results show that even in the face of significant differences in spatial resolution, spectral characteristics, etc., the proposed method still has a high cross-domain segmentation performance of remote sensing images.

[0066] Table 3 Accuracy evaluation of each module in the model ablation experiment ; As shown in Table 3, through ablation experiments, the influence of each module on the segmentation performance in the proposed method is analyzed. The configuration containing all components (MSIE, DII and SD) has the best performance, with an IoU of 44.3%. This is a 4.0% improvement over the baseline model. The inclusion of MSIE enhances feature representation by capturing multi-scale information, while DII and SD further improve the model's ability to effectively separate and decode semantic features. When the semantic decoding (SD) component is excluded, the performance drops to 43.8% IoU, a decrease of 0.5%. This highlights the importance of SD in the final stage of the model, which helps accurately interpret features for segmentation. Excluding the domain-invariant instance disentanglement (DII) component causes a further drop in performance to 41.7% IoU, a significant drop of 2.1%. This shows that DII plays a key role in decoupling domain-specific features, which is crucial for improving performance.

[0067] Although the present invention has been described in detail above by general explanation and specific implementation methods, it is obvious to those skilled in the art that some modifications or improvements can be made on the basis of the present invention. Therefore, these modifications or improvements made on the basis of not departing from the spirit of the present invention all belong to the scope of protection claimed by the present invention.

Claims

1. A generalized semantic segmentation method for remote sensing image domain based on multi-scale instance decoupling, characterized in that: The following steps are involved: S10 Multi-scale Instance Coding: The multi-scale instance coding module takes the image features in the segmentation backbone network as input, uses deep convolution, cross-scale interactive coding, integrates multi-scale features, constructs deformation-frequency domain units, and perceives the feature response of cross-domain instances; S20 Domain-invariant instance decoupling: The domain-invariant instance decoupling module performs decoupling at the feature level based on mixed distribution feature reconstruction and covariance optimization to achieve learning of domain-invariant features. S30 domain generalization semantic decoding: The features after decoupling from the S20 domain invariant instances are input into the domain generalization semantic decoding module, and the hamburger module is used to model the global context; S40 Generalized semantic segmentation: Remote sensing images of unseen target domains are input into the trained framework to predict segmentation maps on unseen target domains.

2. The remote sensing image domain generalized semantic segmentation method according to claim 1, characterized in that: The S10 multi-scale instance encoding specifically includes the following steps: S11 Input image features: Input the original image from the source domain, divide the original image into blocks with a length*width*channel size of 16*16*3, and map each small block to a one-dimensional vector to form an original image feature sequence and input it into the network to generate a one-dimensional image block sequence; S12 Cross-scale Interaction Coding: Its core operation is mathematically defined as, ; Where F represents the input feature, To scale the operation to accommodate features of different receptive fields, DW-Conv stands for Deep Convolution. It is a cross-scale interaction mechanism; The cross-scale interaction mechanism expression is as follows: ; ; ; in, is a high-level scale feature. is a low-level scale feature. is bilinear upsampling, represents the sigmoid function, (.) is the aggregation and splicing operation; high-level scale features After upsampling and deep convolution operations, low-level scale features are Convolution operation, respectively generating features that can match the current scale Feature map and After the scale features of the three levels interact across scales, they can generate features that can replace the current scale features. Feature map ; S13 Multi-scale feature integration: The contextual attention mechanism is used to integrate feature information of different scales and dynamically assign importance to features from different scales. Its mathematical expression is: ; ; in, is the result of cross-scale interactive encoding in step S12, is a learnable convolutional layer for generating attention weights, represents element-wise multiplication, represents the weight, Representing the attention-weighted feature map, the encoder emphasizes the most informative features while suppressing irrelevant information. These re-weighted features are summed and passed through an additional convolutional layer to generate the final encoded representation; S14 Deformation-frequency domain unit design: Design deformation-frequency domain special-shaped convolution groups to differentially process multi-branch features; S15 scale normalization: Standardize the features, which is expressed as: ; ; Among them, μ and σ are the mean and standard deviation of the features, γ and β are learnable parameters. The deformation-frequency domain unit outputs a feature map in step S14.

3. The remote sensing image domain generalized semantic segmentation method according to claim 2, characterized in that: The S14 deformation-frequency domain unit design is specifically a synergistic combination of deformation convolution and hole rate adaptation, and its expression is: ; in, is the learnable offset, is the hole rate adaptation, , is a multilayer perceptron with 2 hidden layers. Weighted feature map for attention The result after the synergistic combination of deformable convolution and dilation rate adaptation; The two-dimensional discrete cosine transform operation is used to separate high-frequency textures from low-frequency semantics, suppress sensor noise and retain detail information. The core operations are: ; Among them, DCT2D is a two-dimensional discrete cosine transform, IDCT2D is its inverse transform, Conv is a standard convolution, and the attention weighted feature map After two-dimensional discrete cosine transform, the transformed feature map is generated , Transform feature map The result after 5x5 standard convolution is inverse transformed to generate a 2D discrete cosine transform weighted feature map .

4. The remote sensing image domain generalized semantic segmentation method according to claim 1, characterized in that: The S20 domain unchanged instance decoupling specifically includes the following steps: S21 Example feature trend calculation: given the output feature map of the encoder , calculate the central tendency of instance-level features and dispersion , ; ; reflects the overall strength of the feature map, while quantifies the variability between spatial elements. , , Input the height, width and number of channels of the feature map F into step S12 respectively; S22 builds a remote sensing instance feature reconstruction model: builds a remote sensing instance feature reconstruction model based on Beta-Student mixed distribution, replacing the original and , specifically, ; ; in, is the mixing weight, is the degrees of freedom of StudentT distribution, is the shape parameter of the Beta distribution, To truncate distributed, is the output feature map of a given encoder, Linear rectification function; The adjusted statistics are integrated back into the feature map to generate augmented feature embeddings : ; S23 Feature Decoupling: The original feature embedding and the expanded feature embedding are decoupled through a shared encoder-decoder network. And its expanded corresponding feature map Decouple the domain-invariant instances, where , , They are intermediate feature maps The height, width and number of channels.

5. The remote sensing image domain generalized semantic segmentation method according to claim 1, characterized in that: The S23 characteristic decoupling is specifically: For each feature map Perform hierarchical dynamic covariance decomposition and divide the channels of the feature map into K groups, each with the following features: , dynamically assign weights based on feature importance : ; Where GAP(.) represents global average pooling, is a learnable parameter vector corresponding to the weights of K groups of channels. exp(.) represents an exponential function with the natural constant e as the base, which is used to expand the feature map. Perform the same operation to generate group feature maps Dynamically assign weights based on feature importance ; The hierarchical covariance matrix is ​​calculated for each set of features and weighted fusion is performed. The covariance matrix contains the spatial relationship of the feature map. , The covariance matrices are, ; ; in and For each group of features , The mean of , , are the row, column and channel grouping numbers of the feature map, respectively. , The feature maps are and The covariance matrix of Define a variance measure based on the covariance matrix: in, represents the mean covariance matrix.

6. The remote sensing image domain generalized semantic segmentation method according to claim 5, characterized in that: The S23 feature decoupling applies a differentiable mask to replace the hard binary mask with a learnable parameter: ; Gumbel(0,1); in represents the random coefficients that follow the Gumbel (0,1) distribution, Learnable control coefficients to control mask sparsity.

7. The remote sensing image domain generalized semantic segmentation method according to claim 6, characterized in that: The S23 feature decoupling also includes a mask joint constraint: ; And combined with the whitening loss function, a joint decoupling loss function is constructed: ; Finally, the loss function for feature decoupling Together with the cross entropy loss function, it constitutes the joint loss function of the generalized semantic segmentation method in the remote sensing image domain: ; in is the output feature of the fourth stage encoder, is the label data of the source domain image.

8. The remote sensing image domain generalized semantic segmentation method according to claim 1, characterized in that: The remote sensing image data is input into a multi-scale instance encoding module and then into a domain-invariant instance decoupling module to realize the learning of domain-invariant features, which is a stage. Four stages of feature learning are repeated before domain generalization semantic decoding. The features extracted in the second, third, and fourth stages are used together as feature inputs for domain generalization semantic decoding.

9. The remote sensing image domain generalized semantic segmentation method according to claim 8, characterized in that: The S30 domain generalized semantic decoding is specifically to aggregate and concatenate the features of the last three stages of the encoder, use a lightweight hamburger module to model the global context, and then use a multi-layer perceptron to perform phase-by-phase meta-classification.

10. The remote sensing image domain generalized semantic segmentation method according to claim 1, characterized in that: S10-S30 are the training phase, where only remote sensing images from the source domain are input into the framework, and S40 is the inference phase, where remote sensing images from unseen target domains are input into the trained framework.

Citation Information

Patent Citations

  • Single domain generalization method for medical image segmentation

    CN116596832A

  • Cross-scene multi-domain fusion small sample remote sensing target robust identification method

    CN118918476A

  • Rice remote sensing image segmentation method based on multi-scale converter

    CN119741494A

  • Efficient segmentation of tumours from lung ct

    US20250061682A1

Cited By

  • Image recognition method and device, model training method and device, equipment and medium

    CN120375098A

  • Machine room energy-saving time sequence prediction method based on deep reinforcement learning

    CN120730686A

  • A data center energy-saving time sequence prediction method based on deep reinforcement learning

    CN120730686B

  • Intelligent low-altitude traffic unmanned aerial vehicle situation awareness method and system

    CN121617001A

  • Universal single-stage sperm morphology image segmentation method and system

    CN122265329A