Domain generalization method for remote sensing image segmentation task

Through feature decoupling and meta-learning technology, the domain-specific and domain-invariant information of remote sensing images are separated, which solves the problem of insufficient generalization of remote sensing image segmentation model in unknown domains, and achieves better generalization ability and semantic retention.

CN120259328APending Publication Date: 2025-07-04XIDIAN UNIV HANGZHOU RES INST +1
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510317241.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Existing remote sensing image segmentation models are insufficient in generalization when dealing with unknown domains, especially in the remote sensing field, lack effective domain-agnostic learning methods, which are difficult to adapt to the challenges of diversity and semantic preservation.

Method used

Feature decoupling, domain expansion and meta-learning technologies are adopted to decompose the feature map into domain-specific and domain-invariant components through feature decouplers, and combined with the variational autoencoder and meta-learning framework, model parameters are optimized to improve generalization capabilities.

Benefits of technology

It significantly improves the generalization ability of the remote sensing image segmentation model in unknown domains, can effectively adapt to multiple benchmark segmentation models, and handles multi-domain and single-domain generalization experimental settings, maintaining the diversity and original semantics of the training distribution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259328A_ABST
    Figure CN120259328A_ABST
Patent Text Reader

Abstract

The invention discloses a domain generalization method for a remote sensing image segmentation task, which relates to the technical field of image processing, and comprises the following steps of: firstly, extracting a feature map from a remote sensing image sample by utilizing a basic remote sensing image segmentation model, and decomposing the features into a domain specific component and a domain invariant component by adopting a feature decoupler; therefore, differences and generality among different domains can be captured more accurately. Thirdly, model parameters are updated based on the calculated cost loss, a meta-learning framework is used for optimizing a feature decoupler, and independence and semantic attributes of decoupling components are ensured; through the technical means of continuous value space modeling, vector orthogonal decoupling, generated data maintenance and the like, the remote sensing image segmentation accuracy is improved, and the generalization ability of the model in an unknown domain is remarkably enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and more particularly to a domain generalization method for remote sensing image segmentation tasks. Background Art

[0002] Due to the diverse data sources and variable acquisition scenarios of remote sensing images, the time-consuming and scarcity of high-quality pixel-level annotations, and the independent and identically distributed assumption followed by models under ideal experimental conditions, the high-precision remote sensing image segmentation models proposed by existing studies cannot effectively process out-of-distribution (OOD) data that widely exists and is difficult to annotate in actual tasks, and face huge generalization challenges. To this end, existing studies have carried out a lot of discussions around domain-agnostic learning (DAL), aiming to encourage artificial intelligence models to obtain discriminative general knowledge from training data (source domain) to better generalize to test data (target domain) with different data distributions. However, when dealing with the target domain with unknown data distribution in remote sensing image segmentation tasks, existing methods still face the following two major challenges: Most existing methods implement domain-agnostic learning based on domain labels, and this strategy lacks rationality and reliability in the remote sensing field; the training distributions used by existing methods either have limited diversity or have semantic failure problems, so it is difficult to obtain reliable performance on unknown remote sensing domains.

[0003] In the field of computer vision, DAL research first explored domain adaptation (DA). DA aims to directionally transfer artificial intelligence models from a source domain with abundant annotations to a specific target domain with scarce annotations. However, DA assumes that the data distribution of the target domain is known, which makes it difficult to effectively adapt to actual applications; to further process unknown target domains in the real world, many DAL studies have started to turn to a more challenging setting, namely domain generalization (DG). According to the specific settings of the source domain, DG can be divided into multiple domain generalization (MDG) and single domain generalization (SDG). MDG relies on multi-domain annotations that are difficult to obtain in actual applications, while SDG is more in line with the actual situation. It only uses single-source domain data and improves the diversity of the training distribution through domain extension techniques to improve the generalization ability of the model in unknown target domains.

[0004] The feature decoupling method aims to decompose a set of non-overlapping and interpretable component sets from existing signals. In recent years, this method has been extended to the fields of DA and DG, and related research can be divided into the reconstruction method and the regularization method. The heuristic gradient update strategy in MAML (Model-Agnostic Meta-Learning) has been widely introduced into the DG field. As the pioneering work of meta-learning in the DG field, MLDG (Meta-Learning with Differentiable Group Learning) obtains a domain-agnostic model by designing a meta-scenario that can simulate the real training-test domain shift and performing gradient backpropagation according to the task loss calculated from the meta-test data. In addition, many effective SDG methods have been developed in the field of computer vision, most of which need to improve the diversity of the training distribution through domain expansion to obtain effective domain-agnostic representations. According to the domain expansion method, the current SDG methods mainly include: parameter-free augmentation method, adversarial sample method, and model generation method. However, there are still the following problems in remote sensing domain-agnostic learning:

[0005] 1) Most feature decoupling methods and meta-learning methods need to perform domain-agnostic learning based on discrete domain labels. However, in the remote sensing field, discrete domain labels lack reliability and rationality, and these two types of methods usually cannot obtain effective performance in remote sensing image segmentation tasks;

[0006] 2) Although the regularization method in the feature decoupling method provides a solution without domain labels, the high-order statistics are not sufficient to simulate complex remote sensing domain shifts, and cannot effectively remove complex remote sensing domain shifts and extract excellent and general remote sensing domain-invariant information;

[0007] 3) Although some meta-learning methods combine domain expansion techniques and synthesize new domains in the meta-test stage, which enlarges the data distribution difference between the meta-test set and the meta-training set, the domain expansion adopted by the existing methods usually has difficulty in retaining the original semantics while improving the distribution diversity, and it is difficult to effectively improve the performance of the model on unknown target domains.

[0008] Therefore, how to improve the generalization of the model in the unknown domain in remote sensing image segmentation tasks is an urgent problem to be solved by those skilled in the art. Summary of the Invention

[0009] In view of this, the present invention provides a domain generalization method for remote sensing image segmentation tasks, which uses feature decoupling, domain expansion, and meta-learning techniques to solve the parameters of the basic model for remote sensing image segmentation tasks and the parameters of the feature decoupler, and improves the generalization of the model in the unknown domain in remote sensing image segmentation tasks.

[0010] To achieve the above object, the present invention adopts the following technical solutions:

[0011] A domain generalization method for remote sensing image segmentation tasks, comprising the following steps:

[0012] Step 1: Collect a number of remote sensing images, perform labeling processing on them to obtain remote sensing image samples, and aggregate all the remote sensing image samples into a source domain;

[0013] Step 2: Randomly extract a small batch of remote sensing image samples from the source domain to form a model training set, input the basic remote sensing image segmentation model to extract feature maps, and use a feature decoupler to decompose the feature maps into domain-specific components and domain-invariant components;

[0014] Step 3: The basic remote sensing image segmentation model processes the domain-specific components and domain-invariant components, calculates the overall cost loss of the model according to the processing results, and updates the model parameters;

[0015] Step 4: Randomly extract a small batch of remote sensing image samples from the source domain to form a meta-training set, use the meta-training set to train the feature decoupler, calculate the decoupling overall cost loss, and perform a first-order update on the feature decoupler;

[0016] Step 5: Randomly extract a small batch of remote sensing image samples from the source domain that are disjoint from the meta-training set to form a meta-test set, use the meta-test set to test the feature decoupler, calculate the meta-test objective loss, and perform a second-order update on the feature decoupler;

[0017] Step 6: Use the updated basic remote sensing image segmentation model and feature decoupler to perform remote sensing image segmentation tasks.

[0018] Preferably, the remote sensing image sample is represented as (x, y), where x represents the remote sensing image and y represents the corresponding label image, and H, W, and N c represent the length, width, and number of categories respectively; all the labeled remote sensing image samples (x, y) are aggregated to form the source domain wherein, is the k-th labeled data set, including a number of remote sensing image samples (x, y) with the same label.

[0019] Preferably, randomly extract a small batch of remote sensing image samples from the source domain to form a model training set input the feature extractor of the basic remote sensing image segmentation model to extract feature maps N d is the number of channels, and use a feature decoupler E fd to decompose the feature maps into domain-specific components and domain-invariant components

[0020] Preferably, the variational autoencoder VAE includes a probabilistic encoder E vaeand the probability decoder D vae .

[0021] Preferably, the specific steps of step 3 include:

[0022] Step 31: The domain-invariant component f di is input into the classifier of the basic remote sensing image segmentation model to perform remote sensing image segmentation and obtain a segmentation result According to the segmentation result calculate the cross-entropy loss

[0023] where δ(·) is an indicator function that returns 1 when the parameter y == c is true and 0 otherwise; y represents the label image corresponding to the remote sensing image sample; c represents the class index, which is a constant with a value range from 1 to N c ; θ f , θ c and θ fd are the parameters of the feature extractor classifier and the feature decoupler E fd respectively;

[0024] Step 32: The domain-specific component f ds is input into the variational autoencoder VAE of the basic remote sensing image segmentation model for continuous value modeling to obtain a reconstruction result According to the reconstruction result calculate the VAE cost loss which is expressed as:

[0025]

[0026] where KL() represents the KL divergence function; q(z ds |f ds ) represents the latent coding distribution; p(z ds ) represents the prior distribution, and a multivariate standard Gaussian distribution is used as the prior distribution; the domain-specific component f ds is input into the probability encoder E of the variational autoencoder VAE vae to obtain the latent coding distribution;

[0027] Step 33: Calculate the overall cost loss of the model according to the cross-entropy loss and the VAE cost loss, and update the basic remote sensing image segmentation model to update the model parameter θ base ; the model parameter θ base includes the parameter θ of the feature extractor f , the parameter θ of the variational autoencoder VAEvae , and the parameters θ of the classifier ; denoted as: c ; expressed as:

[0028]

[0029] Among them, α is the learning rate for updating the basic remote sensing image segmentation model; represents the mean value of the losses corresponding to all remote sensing image samples in the model training set; represents the derivative operation.

[0030] Preferably, step 4 specifically includes the following steps:

[0031] Step 41: Randomly extract a small batch of remote sensing image samples from the source domain to form a meta-training set and input them into the feature extractor of the updated basic remote sensing image segmentation model to extract the feature map f of the remote sensing image samples in the meta-training set and use the feature decoupler E fd to decompose the feature map f into domain-specific components f ds and domain-invariant components f di ;

[0032] Step 42: Perform vector orthogonal decoupling on the two components, and successively pass through the global average pooling layer and L2 regularization layer of the basic remote sensing image segmentation model to encode the domain-specific component f ds and the domain-invariant component f di into domain-specific embeddings and domain-invariant embeddings GAP(·) represents the global average pooling operation, represents the L2 regularization operation, and N d is the number of channels; according to the domain-specific embedding e ds and the domain-invariant embedding e di , calculate the orthogonal loss denoted as:

[0033]

[0034] Among them, θ fd is the decoupling parameter of the feature decoupler E fd ; |·| represents the absolute value operation;

[0035] Step 43: Perform semantic-guided decoupling on the two components, and input the feature map f, the domain-specific component f ds and the domain-invariant component f di into the classifier of the updated basic remote sensing image segmentation model Calculate the entropy values of the segmentation results of the three outputs using the entropy function, and calculate the semantic enhancement loss according to the entropy values and the semantic confusion loss According to the semantic enhancement loss and the semantic confusion loss Calculate the meta-training objective loss It is expressed as:

[0036]

[0037] Among them, represents the entropy function, are the segmentation results of the feature map f, the domain-specific component f ds and the domain-invariant component f di respectively; Softplus(·) represents a monotonically increasing function, Softplus(·) = ln(1 + exp(·)); θ c represents the parameters of the classifier ; θ fd represents the parameters of the feature decoupler E fd ;

[0038] Step 44: Calculate the decoupled overall cost loss according to the orthogonal loss and the meta-training objective loss Perform a first-order update on the feature decoupler E to update the decoupling parameters θ fd of the feature decoupler E fd ; It is expressed as:

[0039]

[0040] Among them, λ sg and λ vo are the hyperparameters of vector orthogonal decoupling and semantic-guided decoupling respectively; β is the learning rate during meta-training; θ base represents the model parameters; θ f represents the parameters of the feature extractor ; θ′ fd represents the decoupling parameters after the first-order update; fd represents the mean value of the losses corresponding to all remote sensing image samples in the meta-training set; represents the derivative operation.

[0041] Preferably, a meta-test set is composed of remote sensing image samples that are drawn in small batches from the source domain and are disjoint from the meta-training set Use the meta-test set to test the feature decoupler E fd and calculate the meta-test objective loss Perform a first-order update on the feature decoupler E fdPerform second-order update; the specific process is as follows:

[0042] Step 51: Randomly sample a latent vector from the prior distribution p(z ds ) and synthesize the domain-specific feature f′ through the probability decoder D in the variational autoencoder VAE of the updated basic remote sensing image segmentation model vae ds ,

[0043] Step 52: Synthesize the original feature by combining the domain-specific feature f′ ds and the domain-invariant component f corresponding to the remote sensing image samples in the meta-training set di

[0044] Step 53: Decouple the original feature using the updated feature decoupler E fd to obtain the domain-specific information and the domain-invariant information

[0045] Step 54: Calculate the decoupling consistency loss according to the domain-specific feature f′ ds , the domain-invariant component f corresponding to the meta-training set di , the domain-specific information and the domain-invariant information It is expressed as:

[0046]

[0047] where θ fd represents the decoupling parameter of the feature decoupler E fd ;

[0048] Step 55: Input the domain-invariant information into the classifier of the updated basic remote sensing image segmentation model for remote sensing image segmentation, and calculate the task loss according to the segmentation result It is expressed as:

[0049]

[0050] where represents the segmentation result; θ c represents the parameter of the classifier ; δ(·) represents the indicator function, which returns 1 if the parameter y == c is true, otherwise returns 0; y represents the label image corresponding to the remote sensing image sample, c represents a constant; N C represents the number of categories of image segmentation; ​​​​​​​

[0051] Step 56: According to the cross-entropy loss Task loss Orthogonal loss Semantic confusion loss Decoupled consistency loss Calculate the meta-test objective loss Perform a second-order update on the feature decoupler to update the decoupling parameter θ fd ; Expressed as:

[0052]

[0053] where, λ con is the weight hyperparameter of the decoupled consistency loss; represents the meta-test set which is the mean of the losses corresponding to all remote sensing image samples in it.

[0054] As can be seen from the above technical solutions, compared with the prior art, the present invention discloses a domain generalization method for remote sensing image segmentation tasks, which uses a continuous value space for remote sensing domain modeling, combines the denoising thinking of the diffusion model to decouple domain-invariant representations, and performs domain extension under the meta-learning framework by operating on domain-specific information. Among them, the present invention pioneers a new paradigm for continuous domain modeling to more accurately describe the remote sensing domain, thereby promoting more effective domain-agnostic learning. Compared with the DG method based on discrete domain modeling, the present invention has significant advantages in remote sensing image segmentation tasks, can effectively adapt to a variety of benchmark segmentation models, and can handle both MDG and SDG experimental settings simultaneously; referring to the denoising idea of the diffusion model, the mutual independence of the decoupled components is maintained through vector orthogonal decoupling, and the decoupled components are given corresponding semantic attributes based on semantic-guided decoupling. An effective feature decoupling is achieved through vector decomposition and semantic guidance, so as to successfully separate domain-invariant representations from domain-specific changes. Compared with the existing feature decoupling methods, the present invention can more effectively separate domain-specific and domain-invariant information in remote sensing image segmentation tasks; randomize domain-specific information in the meta-learning framework for effective domain extension, which can improve the diversity of the training distribution while fully retaining the original semantics, and use generated data to maintain decoupled consistency and construct a more realistic meta-scenario, which can effectively enhance the generalization ability of the model in unknown domains. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0056] Figure 1 Schematic diagram of the processing flow of a domain generalization method for remote sensing image segmentation tasks provided by the present invention;

[0057] Figure 2 Schematic diagram of the recognition results of different domain generalization methods under the MDG setting on the DAB dataset provided by the present invention;

[0058] Figure 3 Schematic diagram of the recognition results of different domain generalization methods under the SDG setting on the DAB dataset provided by the present invention;

[0059] Figure 4 Schematic diagram of the visualization results of the target domain features of the SDG and MDG experiments with the WHU-A dataset as the source domain and the target domain provided by the present invention. Specific implementation manners

[0060] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0061] The embodiments of the present invention disclose a domain generalization method for remote sensing image segmentation tasks, which is adapted to both the MDG and SDG settings at the same time. The process is as Figure 1 shown, and the specific implementation steps are as follows:

[0062] S1: Given a sample (x, y), where x represents a remote sensing image and y represents the corresponding label image, and H, W, and N c respectively represent the length, width, and number of categories;

[0063] S2: In the training stage, all labeled samples are aggregated to form the source domain where is the k-th labeled dataset, including several remotely sensed image samples (x, y) with the same label. In each iteration, a small batch is randomly sampled from for updating the basic model; the randomly selected sample (x, y) is sent to the feature extractor of the basic model to obtain the feature map where N d is the number of channels; through the feature decoupler E fd for vector orthogonal decoupling, the feature map f is decomposed into two independent and complementary components, namely the domain-specific component and the domain-invariant component

[0064] S3: Feed the domain-invariant component f obtained in S1 di into the classifier of the base model Perform remote sensing image segmentation to obtain the segmentation result; Feed the domain-specific component f obtained in S1 ds into the variational autoencoder VAE composed of the probability encoder E vae and the probability decoder D vae in the base model for continuous value modeling to obtain the reconstruction result Calculate the cross-entropy loss based on the domain-invariant component f di as shown in Equation 1; and calculate the VAE cost loss function based on f as shown in Equation 2; ds

[0065]

[0066] where represents the segmentation result; δ(·) is the indicator function that returns 1 when the parameter y == c is true and 0 otherwise; y represents the label image corresponding to the remote sensing image sample; c represents the class index, which is a constant with a value range from 1 to N c ; θ f , θ c and θ fd are the learnable parameters of the feature extractor classifier and the feature decoupler E fd respectively;

[0067]

[0068] where θ vae represents the learnable parameter of the variational autoencoder VAE; KL(·) represents the KL divergence calculation function; the first term is the reconstruction loss, and the second term can be regarded as a regularizer in the latent space, which constrains the KL divergence between the latent coding distribution q(z ds |f ds ) and the prior distribution p(z ds ), and the multivariate standard Gaussian distribution most commonly used in existing research is used as the prior distribution; the domain-specific component f ds is input into the probability encoder E of the variational autoencoder VAE vae to obtain the latent coding distribution;

[0069] S4: Update the base model based on the two losses calculated in S3; the parameters θ of the base model base include the parameters θ of the feature extractor f, the parameters θ of the variational autoencoder VAE vae , and the parameters θ of the classifier ; c The overall cost function of the model As shown in Equation 3, the parameter update can be formalized as Equation 4;

[0070]

[0071] where α is the learning rate for updating the base model; represents the mean of the losses corresponding to all remote sensing image samples in the model training set; represents the derivative operation;

[0072] S5: Randomly sample a mini-batch of samples from as the meta-training set Meta-training set Input the samples (x, y) in the meta-training set into the updated parameter θ obtained in S4 base and extract the feature maps f of the remote sensing image samples in the meta-training set from the base model after and use the feature decoupler E fd to divide it into domain-invariant information f di and domain-specific information f ds ; To effectively decompose the two components, perform vector orthogonal decoupling and semantic-guided decoupling;

[0073] In the vector orthogonal decoupling step, use the global average pooling layer and the L2 regularization layer that sequentially pass through the base remote sensing image segmentation model to encode the domain-specific information f ds and the domain-invariant information f di into the domain-specific embedding and the domain-invariant embedding That is and GAP(·) represents the global average pooling operation, represents the L2 regularization operation, N d is the number of channels, and then calculate the orthogonal loss according to Equation 5 By minimizing make f ds and f di orthogonal to ensure their mutual independence;

[0074]

[0075] where, θ fd is the learnable parameter of the feature decoupler E fd ; |·| represents the absolute value operation;

[0076] In the semantic-guided decoupling step, the feature map f and the two components f di and fds Input classifier Adopt the entropy function to model f, f di and f ds 's semantic sorting relationship, calculate their respective entropy values; calculate the semantic enhancement loss according to Equation 6 using the entropy values, and calculate the semantic confusion loss according to Equation 7;

[0077]

[0078] where Softplus(·) = ln(1 + exp(·)) is a monotonically increasing function, aiming to reduce the optimization difficulty caused by negative values;

[0079]

[0080] The training objective loss function of this step can be calculated by Equation 8:

[0081]

[0082] S6: Randomly sample a mini-batch of samples from as the meta-training set According to the loss calculated in S5 and perform a first-order update on the parameters θ fd of the feature decoupler E fd ; The overall decoupling cost loss function in the meta-training stage is as in Equation 9, and the first-order update process of θ fd can be formalized as Equation 10;

[0083]

[0084] where λ sg and λ vo are hyperparameters for vector orthogonal decoupling and semantic-guided decoupling; represents the mean of the losses corresponding to all remote sensing image samples in the meta-training set;

[0085]

[0086] where β is the learning rate in the meta-learning process;

[0087] S7: Sample a mini-batch of samples that is disjoint from the meta-training set from as the meta-test set First, randomly sample a latent vector from the prior distribution p(z ds ) and synthesize the domain-specific feature f′ through the probability decoder D vae ​ds , that is using the generated domain-specific feature f′ ds and the true domain-invariant component f decomposed from the meta-training set di to synthesize the original feature and then calculate the task loss based on , maintain the decoupling consistency constraint based on VAE, and finally combine the two decoupling losses calculated in S5 and the first-order update result θ′ obtained in S6 fd to perform a second-order update on the feature decoupler E fd to obtain the final meta-optimization result;

[0088] Use the feature decoupler E fd to perform decoupling to obtain the corresponding domain-specific information and domain-invariant information During this process, and should be consistent with f′ ds and f di respectively, so the decoupling consistency loss is obtained as shown in Equation 11;

[0089]

[0090] According to the domain-invariant features decoupled by perform the segmentation task, and the corresponding task loss is as shown in Equation 12;

[0091]

[0092] where

[0093] the meta-test target loss in the meta-test phase is calculated based on the parameters θ′ of the feature decoupler after the first-order update fd , and the meta-test target loss includes the task loss, vector orthogonality loss, semantic guidance loss, and decoupling consistency loss, as shown in Equation 13. Use the meta-test target loss to perform meta-update on the parameters of the feature decoupler, as shown in Equation 14;

[0094]

[0095] where λ con is the weight hyperparameter of the decoupling consistency function; represents the mean value of the losses corresponding to all remote sensing image samples in the meta-test set.

[0096] On the other hand, in a specific embodiment, to verify the effectiveness of the method of the present invention, a Domain Agnostic Building (DAB) building dataset was constructed based on the existing publicly available building remote sensing image segmentation dataset. This dataset consists of the Massachusetts Building (MASS-B) dataset, 3 subsets of the Wuhan University Building (WHU-B) dataset, and 5 subsets of the Inria Aerial Image Labeling (IAIL) dataset. In the domain generalization experiment, the MASS-B, WHU-A, WHU-S1, WHU-S2, IAIL-A, IAIL-C, IAIL-K, IAIL-T, and IAIL-V datasets constitute 9 pre-defined domains, which have significant differences in terms of geographical location, data source, spatial resolution, etc., as shown in Table 1.

[0097] Table 1 Details of the publicly available datasets used to construct the DAB dataset

[0098]

[0099] During the experiment, each pre-defined domain was divided into a training set and a test set. Specifically, the MASS-B, WHU-A, WHU-S1, and WHU-S2 datasets used their original training data as the training set, and the remaining data as the test set. In addition, for the convenience of model training, the samples of the MASS-B dataset were randomly cut into a size of 512×512. For the 5 subsets of the IAIL dataset, after being cut into 512×512 tiles, they were divided into a training set and a test set according to a ratio of 3:2. It should be noted that in the MDG experiment, these 9 datasets took turns as the target domain, and the remaining datasets served as the source domain. In the SDG experiment, WHU-A served as the source domain, and the remaining datasets served as the target domain. In the domain generalization training paradigm, the training set and test set of the source domain were used for model training and validation respectively, while the test set of the target domain was used for model testing.

[0100] The UNet was used as the basic remote sensing image segmentation model for comparative research on the DAB dataset. In addition to the basic remote sensing image segmentation model based on Empirical Risk Minimization (ERM), 10 state-of-the-art MDG methods and 8 state-of-the-art SDG methods were also compared. The accuracy results are shown in Tables 2 and 3, and the segmentation results are as Figure 2 And Figure 3 shown; Figure 2Shows the recognition results of different domain generalization methods under the MDG setting on the DAB dataset, where Figure 2 in (a) is the original image, Figure 2 in (b) is the label image, Figure 2 in (c)-(n) are the recognition results of ERM, CrossGrad, UFDN, GLIE, SagNet, RobustNet, TACIT, MLDG, M3L, PtM, SuA-SpML, and the method of the present invention respectively; Figure 3 Shows the recognition results of different domain generalization methods under the SDG setting on the DAB dataset, where Figure 2 in (a) is the original image, Figure 3 in (b) is the label image, Figure 3 in (c)-(l) are the recognition results of ERM, AugMix, MixStyle, RCKL, CCDR, ADA, M-ADA, L2D, PDEN, and the method of the present invention respectively. The experimental results show that the present invention has significant advantages in the remote sensing image segmentation task compared with other MDG methods and SDG methods.

[0101] Table 2 Precision comparison of different domain generalization methods under the MDG setting on the DAB dataset

[0102]

[0103]

[0104] Table 3 Precision comparison of different domain generalization methods under the SDG setting on the DAB dataset

[0105]

[0106]

[0107]

[0108] To visually explore the improvement effect of the present invention on the generalization of the remote sensing image segmentation model, the t-SNE tool is used to visualize the features used for classification by the same remote sensing image segmentation model before and after using this method. Before using this method, the remote sensing image segmentation model uses the output features of the feature extractor (denoted as f b ) for classification. After using this method, the remote sensing image segmentation model uses the decoupled domain-invariant features (denoted as ) for classification. In addition, to better illustrate the feature decoupling effect of this method, the output features of the feature extractor in the remote sensing image segmentation model after using this method are further obtained (denoted as f a) t-SNE visualization results, as shown in Figure 4 which shows the visualization results of the target domain features in the SDG and MDG experiments with the WHU-A dataset as the source domain and target domain respectively. The red dots are the background and the green dots are the buildings. Obviously, in the MDG and SDG experiments, f b has difficulty in distinguishing the two categories, as shown in Figure 4 where (a) in b is f Figure 4 in the MDG experiment, and (d) in b is f a in the SDG experiment; f a has a significant improvement in both intra-class aggregation and inter-class separation, as shown in Figure 4 where (b) in Figure 4 is f a in the MDG experiment, and (e) in is f a in the SDG experiment; and after further eliminating domain-specific information, Figure 4 the corresponding class confusion range in the t-SNE of Figure 4

[0109] Each embodiment in this specification is described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, refer to the description in the method part.

[0110] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A domain generalization method for remote sensing image segmentation tasks, characterized in that It includes the following steps: Step 1: Collect a number of remote sensing images, perform marking processing on them to obtain remote sensing image samples, and aggregate all remote sensing image samples into the source domain; Step 2: Randomly extract a small batch of remote sensing image samples from the source domain to form a model training set, input the basic remote sensing image segmentation model to extract feature maps, and use a feature decoupler to decompose the feature maps into domain-specific components and domain-invariant components; Step 3: The basic remote sensing image segmentation model processes the domain-specific components and domain-invariant components, calculates the overall cost loss of the model according to the processing results, and updates the model parameters; Step 4: Randomly extract a small batch of remote sensing image samples from the source domain to form a meta-training set, use the meta-training set to train the feature decoupler, calculate the overall decoupling cost loss, and perform a first-order update on the feature decoupler; Step 5: Randomly extract a small batch of remote sensing image samples from the source domain that do not intersect with the meta-training set to form a meta-test set, use the meta-test set to test the feature decoupler, calculate the meta-test objective loss, and perform a second-order update on the feature decoupler; Step 6: Use the updated basic remote sensing image segmentation model and feature decoupler to perform remote sensing image segmentation tasks.

2. The domain generalization method for remote sensing image segmentation task according to claim 1, wherein The remote sensing image sample is represented as (x, y), where x represents the remote sensing image and y represents the corresponding label image. and H, W, and N c represent the length, width, and number of classes, respectively.

3. A domain generalization method for remote sensing image segmentation tasks according to claim 2, characterized in that, Randomly extract a small batch of remote sensing image samples from the source domain to form a model training set Input the feature extractor of the basic remote sensing image segmentation model Extract the feature map N d Is the number of channels, and use the feature decoupler E fd Decompose the feature map into domain-specific components And domain-invariant components 4. A domain generalization method for remote sensing image segmentation tasks according to claim 2, characterized in that, The specific process of Step 3 includes: Step 31: Domain-invariant component f di Input it into the classifier of the basic remote sensing image segmentation model Perform remote sensing image segmentation to obtain a segmentation result According to the segmentation result Calculate the cross-entropy loss where, δ(·) is the indicator function, y represents the label image corresponding to the remote sensing image sample, c represents the class index; θ f , θ c and θ fd are the parameters of the feature extractor classifier and the feature decoupler E fd respectively; Step 32: Domain-specific component f ds Input it into the variational autoencoder (VAE) of the basic remote sensing image segmentation model for continuous value modeling to obtain the reconstruction result According to the reconstruction result Calculate the VAE cost loss It is expressed as: where KL() represents the KL divergence function; q(z ds |f ds ) represents the latent coding distribution output by the probability encoder of the variational autoencoder VAE; p(z ds ) represents the prior distribution, and a multivariate standard Gaussian distribution is adopted Step 33: Calculate the overall cost loss of the model based on the cross-entropy loss and the VAE cost loss, and update the basic remote sensing image segmentation model to update the model parameters θ base ; The model parameters θ base include the parameters θ of the feature extractor f , the parameters θ vae of the variational autoencoder VAE , and the parameters θ c of the classifier; expressed as: Among them, α is the learning rate for updating the basic remote sensing image segmentation model; represents the mean value of the losses corresponding to all remote sensing image samples in the model training set; represents the derivative operation.

5. A domain generalization method for remote sensing image segmentation tasks according to claim 4, characterized in that, Step 4 specifically includes the following steps: Step 41: Randomly sample a small batch of remote sensing image samples from the source domain to form a meta-training set and input them into the feature extractor of the updated basic remote sensing image segmentation model to extract the feature maps f of the remote sensing image samples in the meta-training set and use the feature decoupler E fd to decompose the feature map f into domain-specific components f ds and domain-invariant components f di ; Step 42: Perform vector orthogonal decoupling on the domain-specific component f ds and the domain-invariant component f di to calculate the orthogonal loss Step 43: Perform semantic-guided decoupling on the domain-specific component f ds and the domain-invariant component f di to calculate the meta-training objective loss Step 44: Calculate the decoupled overall cost loss according to the orthogonal loss and the meta-training objective loss ; Perform a first-order update on the feature decoupler E fd to update the decoupling parameter θ fd of the feature decoupler E fd ; which is expressed as: Among them, λ sg and λ vo are the hyperparameters of vector orthogonal decoupling and semantic-guided decoupling respectively; β is the learning rate in the meta-training process; θ base represents the model parameters; θ f represents the parameters of the feature extractor ; θ′ fd represents the decoupling parameters after the first-order update; represents the mean value of the losses corresponding to all remote sensing image samples in the meta-training set.

6. A domain generalization method for remote sensing image segmentation tasks according to claim 5, characterized in that The process of vector orthogonal decoupling is: Step 421: The domain-specific component f ds and the domain-invariant component f di are respectively encoded into the domain-specific embedding and the domain-invariant embedding GAP(·) represents the global average pooling operation, represents the L2 regularization operation, and N d is the number of channels; Step 422: Calculate the orthogonal loss based on the domain-specific embedding e ds and the domain-invariant embedding e di , which is expressed as: as follows: where θ fd is the decoupling parameter of the feature decoupler E fd ; |·| represents the absolute value operation.

7. A domain generalization method for remote sensing image segmentation tasks according to claim 5, characterized in that The process of semantic-guided decoupling is: Step 431: Input the feature map f, the domain-specific component f ds and the domain-invariant component f di into the classifier of the updated basic remote sensing image segmentation model respectively to obtain three segmentation results; Step 432: Use the entropy function to calculate the entropy values of the three segmentation results respectively; Step 433: Calculate the semantic enhancement loss according to the entropy value and the semantic confusion loss which is expressed as: Step 434: Calculate the meta-training objective loss according to the semantic enhancement loss and the semantic confusion loss which is expressed as: It is expressed as: Among them, represents the entropy function, are the segmentation results of the feature map f, the domain-specific component f ds and the domain-invariant component f di respectively; Softplus(·) represents a monotonically increasing function, softplus(·) = ln(1 + exp(·)); θ c represents the parameter of the classifier ; θ fd represents the parameter of the feature decoupler E fd .

8. A domain generalization method for remote sensing image segmentation tasks according to claim 5, characterized in that, Randomly extract a small batch of remote sensing image samples from the source domain that are disjoint from the meta-training set to form the meta-test set For the feature decoupler E fd The specific process of the second-order update is as follows: Step 51: Randomly sample a latent vector from the prior distribution p(z ds ) and synthesize the domain-specific feature f′ through the probability decoder of the VAE in the variational autoencoder of the updated basic remote sensing image segmentation model ds , Step 52: Combine the domain-specific feature f′ ds and the meta-training set with the domain-invariant component f corresponding to the remote sensing image samples di to synthesize the original feature Step 53: Use the updated feature decoupler E fd to decouple the original features and obtain domain-specific information and domain-invariant information Step 54: Calculate the decoupled consistency loss based on the domain-specific feature f′ ds , the meta-training set , the corresponding domain-invariant component f di , the domain-specific information , and the domain-invariant information which is expressed as: as follows: Among them, θ fd represents the decoupling parameter of the feature decoupler E fd ; Step 55: Input the domain-invariant information into the classifier of the updated basic remote sensing image segmentation model for remote sensing image segmentation, and calculate the task loss according to the segmentation result which is expressed as: Among them, represents the segmentation result; θ c represents the parameter of classifier c; Step 56: According to the cross-entropy loss Task loss Orthogonal loss Semantic confusion loss Decoupled consistency loss Calculate the meta-test objective loss Perform a second-order update on the feature decoupler to update the decoupling parameter θ fd ; Expressed as: Among them, λ con is the weight hyperparameter for decoupled consistency loss; represents the mean of the losses corresponding to all remote sensing image samples in the meta-test set.

Citation Information

Cited By

  • Design task intelligent decomposition method based on deep learning

    CN120909802A

  • A gas pipeline construction quality information management system

    CN122509792A