Self-adaption cross-domain remote sensing image semantic segmentation method based on unsupervised domain

Through the unsupervised domain adaptive method, the problem of domain offset in semantic segmentation of remote sensing images is solved, and the high-precision segmentation effect is achieved under the lack of labeled data.

CN120279273APending Publication Date: 2025-07-08SOUTH CHINA UNIV OF TECH
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510437737.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The prior art has degraded performance due to domain offset problems in remote sensing image semantic segmentation, especially in the absence of labeled data in the target domain, it is difficult to effectively align the feature distribution of the source domain and the target domain, resulting in poor segmentation performance.

Method used

Unsupervised domain adaptive method is adopted to decouple and represent enhancement of target domain features, and features are extracted using domain-invariant encoder and target domain-specific encoder, and features are aligned through orthogonal constraint loss, reconstruction loss and adversarial methods, combining multi-level feature alignment and self-learning to generate target domain fusion features to improve segmentation performance.

Benefits of technology

In the absence of labeled data, the segmentation accuracy of remote sensing images and the ability to adapt to complex scenarios are significantly improved, the dependence on target domain labeled data is reduced, and the cross-domain segmentation performance and robustness are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279273A_ABST
    Figure CN120279273A_ABST
Patent Text Reader

Abstract

The invention discloses a self-adaptive cross-domain remote sensing image semantic segmentation method based on an unsupervised domain, and the method comprises the steps: pre-training an encoder-decoder segmentation skeleton through a source domain image and a semantic tag of the source domain image; the difference between the domain invariant feature and the specific feature of the target domain is measured through the difference loss, and the reconstruction loss is applied to prevent information loss; the domain invariant high-level features of the source domain and the target domain are aligned based on an antagonism method; respectively processing high-layer and shallow-layer invariant features of a target domain through a main decoder and an auxiliary decoder of the decoders, and ensuring that prediction results of different levels of features are consistent by utilizing consistency loss; and fusing the domain-invariant high-level features and the specific features of the target domain to generate target domain fusion features, and optimizing the cross-domain adaptability of the decoder by aligning the prediction result of the target domain fusion features with the prediction result of the domain-invariant features. According to the method, the segmentation performance on the target domain is remarkably improved through feature decoupling and feature representation enhancement in combination with target domain self-learning based on pseudo labels.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the technical field of computer vision and deep learning, and particularly to a cross-domain remote sensing image semantic segmentation method based on unsupervised domain adaptation. Background Art

[0002] In recent years, deep neural networks have made remarkable progress in semantic segmentation. However, the excellent performance of deep learning technology depends on a large number of training samples and expensive pixel-level annotations. In practical remote sensing applications, due to the changes in remote sensing sensors and the obvious differences in landscapes at different geographical locations, the data distributions of training images and test images are different, resulting in a sharp decline in performance, that is, there is a domain shift problem.

[0003] Although existing studies have tried to alleviate the domain shift problem through methods such as data augmentation and transfer learning, most of these methods rely on labeled data in the source domain or cannot effectively align the feature distributions of the source domain and the target domain. Feature-based domain adaptation methods transfer knowledge by aligning the feature spaces of the source domain and the target domain. However, this method ignores the individual characteristics of the target domain.

[0004] Currently, decomposing features into shared features and dissimilar features for domain adaptation is a widely used method. Many decoupling methods will fuse the domain-specific information of the target domain and the domain-invariant information of the source domain to generate a new source domain image with the same style as the target domain, train the segmentation network using the new source domain image and labels, and then directly apply it to the target domain image. However, these methods usually have the problem of semantic loss during the process of generating new images, thus introducing segmentation errors. Some decoupled representation methods usually focus on more precisely extracting domain-invariant information while discarding domain-specific features. However, ignoring the domain-specific features of the target domain during the transfer process and solely relying on the domain-invariant features of the source domain may cause overfitting to the source domain data, weaken its generalization ability to the target domain data, and thus lead to a decline in performance in the target domain.

[0005] Therefore, how to effectively align the feature distributions of the source domain and the target domain without labeled data in the target domain and improve the segmentation performance in the target domain has become an urgent problem to be solved. Summary of the Invention

[0006] The purpose of the present invention is to overcome the deficiencies of the prior art and provide a cross-domain remote sensing image semantic segmentation method based on unsupervised domain adaptation. By performing personalized learning on the private features in the target domain, it can effectively alleviate the performance decline problem caused by data distribution differences in cross-domain segmentation tasks, especially in the case where the target domain data lacks annotations, and significantly improve the adaptability to complex scenes and segmentation accuracy.

[0007] To achieve the above object, the technical solution provided by the present invention is: a cross-domain remote sensing image semantic segmentation method based on unsupervised domain adaptation, comprising the following steps:

[0008] S1: Obtain remote sensing image data, and preprocess the remote sensing image data to generate source domain images and target domain images of a unified size; the source domain images refer to image data with known semantic labels for training and learning, and the target domain images are image data to be semantically segmented and lacking annotation information;

[0009] S2: Based on the source domain images obtained in step S1 and their corresponding semantic labels, pre-train the encoder-decoder segmentation framework for semantic segmentation, where the encoder-decoder segmentation framework includes an encoder and a decoder; the encoder includes a domain-invariant encoder and a target domain-specific encoder, and the decoder includes a main decoder and an auxiliary decoder; through pre-training, an initialized semantic segmentation framework is obtained. The domain-invariant encoder is used to extract source domain domain-invariant features and target domain domain-invariant features, where the source domain domain-invariant features include source domain domain-invariant high-level features and source domain domain-invariant low-level features, and the target domain domain-invariant features include target domain domain-invariant high-level features and target domain domain-invariant low-level features. The target domain-specific encoder extracts target domain-specific features;

[0010] S3: Based on the source domain domain-invariant high-level features, target domain domain-invariant high-level features, and target domain-specific features obtained in step S2, perform feature decoupling: First, use an orthogonal constraint loss to measure the difference between the target domain domain-invariant high-level features and the target domain-specific features, and apply a reconstruction loss to avoid information loss. Then, align the source domain domain-invariant high-level features and the target domain domain-invariant high-level features through an adversarial method to improve the cross-domain transfer ability, and obtain the decoupled and aligned target domain domain-invariant high-level features and target domain-specific features;

[0011] S4: Based on the decoupled target domain domain-invariant high-level features in step S3 and the target domain domain-invariant low-level features obtained in step S2, perform multi-level feature alignment to promote the representation enhancement of the target domain domain-invariant high-level features. The main decoder and the auxiliary decoder are used to process the target domain domain-invariant high-level features and the target domain domain-invariant low-level features respectively, and the consistency loss constraint is used to achieve the consistency of the prediction results of the same category features at different levels, and obtain the representation-enhanced target domain domain-invariant high-level features;

[0012] S5: Based on the target domain-specific features after decoupling in step S3 and the enhanced target domain domain-invariant high-level features obtained in step S4, perform target domain self-learning, fuse the two features to obtain the target domain fusion features, and improve the adaptability of the decoder to the target domain by aligning the prediction results of the target domain fusion features and the target domain domain-invariant high-level features, and obtain the target domain fusion features that fully represent the target domain information;

[0013] S6: Based on the target domain fusion features that fully represent the target domain information obtained in step S5, parse the target domain fusion features through the main decoder to generate the segmentation prediction results of the target domain image.

[0014] Furthermore, the specific operation steps of step S2 are as follows:

[0015] S21: Use the domain-invariant encoder E di and the target domain-specific encoder E ds to extract image features. Among them, the domain-invariant encoder E di extracts the domain-invariant features from the source domain image S and the target domain image T respectively, denoted as F s di = E di (S) and F t di = E di (T), and and where and are the source domain domain-invariant shallow features and the target domain domain-invariant shallow features, coming from the middle-level output of the domain-invariant encoder, and are the source domain domain-invariant high-level features and the target domain domain-invariant high-level features, coming from the last layer output of the domain-invariant encoder. The subscripts t and s represent the target domain and the source domain respectively. The target domain-specific encoder E ds only acts on the target domain image to extract the target domain-specific features F t ds = E ds (T); The encoder selects ResNet50 as the backbone network;

[0016] S22: The main decoder G M receives the source domain domain-invariant high-level features to generate the source domain main segmentation probability map P s di , G M integrates a dual attention module, and a full convolutional segmentation head is connected after this module; The dual attention module combines the global attention mechanism and the channel attention mechanism, and can effectively capture the global context information while retaining local details; Use P sdi and the pixel-level cross-entropy loss L of the source domain label L obtained based on step S1 s to optimize the segmentation skeleton: task In the formula, H and W are the height and width of P

[0017]

[0018] ; C is the number of categories of the source domain label L s di ; h and w are the height position index and width position index calculated currently, and c is the category calculated currently; s

[0019] S23: The auxiliary decoder G Aux is composed of fully convolutional layers. G Aux receives the source domain domain-invariant shallow features to generate the source domain auxiliary segmentation probability map P s Aux ; the auxiliary segmentation loss function L of G Aux is defined as the pixel-level cross-entropy loss of P aux_task and the source domain label L: s Aux In the formula, H and W are the height and width of P s ; the width, height, and number of categories of P

[0020]

[0021] and P s Aux are all the same. s Aux s di di

[0022] Furthermore, the specific operation steps of step S3 are as follows:

[0023] S31: Based on the target domain specific feature F t ds obtained in step S2, the target domain domain-invariant high-level feature and the target domain domain-invariant shallow feature first use to pass through the auxiliary decoder G Aux to output the target domain auxiliary probability map P t Aux ; then, based on P t Aux calculate the category center feature, and the formula is as follows:

[0024]

[0025] In the formula, is the result of normalizing the probability of P by the Softmax function t Aux in the spatial position, is the class feature center of the target-domain specific features, is the class feature center of the target-domain domain-invariant high-level features;

[0026] S32: To achieve feature decoupling, normalize and to obtain the class feature centers of the normalized target-domain specific features and target-domain domain-invariant high-level features:

[0027]

[0028] In the formula, ||·||2 represents the L2 norm, and the constant ε = 10 -6 is used to prevent division-by-zero errors;

[0029] Then calculate the cosine similarity and construct the orthogonal constraint loss L diff , forcing the domain-specific and domain-invariant features to be orthogonally distributed in the feature space and eliminating redundant semantic interference. The formula is as follows:

[0030]

[0031] S33: To preserve the information integrity during the feature decoupling process, add F t ds and pixel by pixel and input the result into the reconstruction decoder R to generate the reconstructed image where R consists of 3 convolutional layers with channel numbers {256, 64, 3}; define the reconstruction loss L recon based on the mean squared error loss. The formula is as follows:

[0032]

[0033] In the formula, 3 represents the number of image channels, is the product of the height and width of the input image, which is the total number of pixels per channel; ch is the currently calculated channel number, i is the currently calculated pixel index, is the original image pixel feature calculated currently, is the reconstructed image pixel feature calculated currently;

[0034] S34: Introduce a fully convolutional domain discriminator D. D consists of 4 convolutional layers with a kernel size of 3×3, a stride of 1, and channel numbers {256, 128, 64, 2}. Except for the last layer, each convolutional layer is followed by a LeakyReLU activation function with a parameter of 0.2; D is applied to the source-domain domain-invariant high-level features Domain - invariant high - level features in the target domain Perform domain classification in the spatial dimension to obtain the source domain classification prediction result P s D And the target domain classification prediction result P t D , the adversarial loss function L of the discriminator D adv Is defined as:

[0035]

[0036] The domain - invariant encoder implicitly optimizes feature generation through the gradient reversal strategy, making the discriminator unable to distinguish the feature source and achieving feature alignment.

[0037] Furthermore, the specific operation steps of step S4 are as follows:

[0038] S41: Based on the target domain domain - invariant shallow features obtained in step S2 Via the auxiliary decoder G Aux Obtain the target domain auxiliary segmentation probability map P t Aux , based on the decoupled target domain domain - invariant high - level features obtained in step S3 Via the main decoder G M Obtain the target domain main segmentation probability map P t di ;

[0039] S42: Generate pseudo - labels based on P t di For each pixel in P Determine the highest probability value and its corresponding class index, and at the same time calculate the second - highest probability value; if the difference between the highest probability value and the second - highest probability value exceeds a predefined threshold, then assign the label of the pixel to the class index corresponding to the highest probability value; otherwise, set the pixel label to - 1, indicating that the pixel is ignored and does not participate in the loss calculation. t di

[0040] S43: Use the pixel - level cross - entropy loss L between P t Aux And the pseudo - labels To optimize the domain - invariant encoder, the main decoder, and the auxiliary decoder: MFAM

[0041]

[0042] Finally, obtain the enhanced target domain domain - invariant high - level features

[0043] ​​Further, the specific operation steps of step S5 are as follows:

[0044] S51: Based on the decoupled target domain-specific feature F t ds obtained in step S3 and the enhanced target domain domain-invariant high-level feature obtained in step S4, generate a more representative target domain fusion feature by pixel-wise summation

[0045] S52: Input F t di+ds into the main decoder G M to generate the target domain fusion segmentation probability map P t di+ds = G M (F t di +ds ), and use the pseudo-label as the supervision signal to optimize the adaptability to the target domain data;

[0046] S53: Use P t Aux and the pixel-wise cross-entropy loss L based on the pseudo-label TEM obtained in step S4 to optimize the domain-invariant encoder, domain-specific encoder, and main decoder:

[0047]

[0048] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0049] 1. The present invention reduces the dependence on labeled data in the target domain and lowers the cost of data annotation through an unsupervised domain adaptation method. Existing methods usually require a large amount of labeled data for training, while the present invention can achieve effective cross-domain segmentation without labeled data through feature alignment.

[0050] 2. By decoupling and enhancing the representation of the target domain features, the present invention can effectively align the feature distributions of the source domain and the target domain, and solve the domain shift problem. The performance of existing methods drops significantly when the data distributions of the source domain and the target domain are quite different, while the present invention significantly improves the segmentation performance in the target domain through feature decoupling and enhancement.

[0051] 3. The present invention enhances the adaptability to the target domain data by self-learning the target domain data and making full use of the individual characteristics of the target domain. Existing methods often ignore the individual characteristics of the target domain, resulting in insufficient generalization ability, while the present invention further improves the segmentation accuracy and robustness by fusing the common features and private features of the target domain. Brief Description of the Drawings

[0052] Figure 1 is the framework diagram of the method of the present invention.

[0053] Figure 2 is the schematic diagram of the decoupling process.

[0054] Figure 3 is the schematic diagram of the multi-level feature alignment process.

[0055] Figure 4 is the schematic diagram of the target domain self-learning. Detailed Description of the Preferred Embodiments

[0056] The present invention will be further described in detail below with reference to the embodiments and the drawings, but the embodiments of the present invention are not limited thereto.

[0057] As Figures 1 to 4 shown, this embodiment discloses a cross-domain remote sensing image semantic segmentation method based on unsupervised domain adaptation, which is a cross-domain remote sensing image semantic segmentation method developed using the Python language and can run on multiple platforms, and includes the following steps:

[0058] 1) Obtain data.

[0059] Obtain remote sensing image data, and preprocess the remote sensing image data to generate source domain images and target domain images of a unified size; the source domain images refer to image data with known semantic labels for training and learning, and the target domain images are image data to be subjected to semantic segmentation and lacking annotation information;

[0060] 2) Use the source domain images and their semantic labels to pre-train the encoder-decoder segmentation framework.

[0061] 21) Use the domain-invariant encoder E di and the target domain-specific encoder E ds to extract image features, wherein the domain-invariant encoder E di extracts domain-invariant features from the source domain image S and the target domain image T respectively, denoted as F s di = E di (S) and F t di = E di (T), and and wherein and are the source domain domain-invariant shallow features and the target domain domain-invariant shallow features, from the middle layer output of the domain-invariant encoder, and are the domain - invariant high - level features of the source domain and the target domain, which are the outputs of the last layer of the domain - invariant encoder. The subscripts t and s represent the target domain and the source domain respectively; the target - domain - specific encoder E ds only acts on the target - domain images to extract the target - domain - specific features F t ds βE ds (T); The encoder selects ResNet50 as the backbone network;

[0062] 22) The main decoder G M receives the domain - invariant high - level features of the source domain to generate the source - domain main segmentation probability map P s di , G M integrates a dual - attention module, which is followed by a fully - convolutional segmentation head; The dual - attention module combines the global attention mechanism and the channel attention mechanism, which can effectively capture the global context information while retaining local details; Using P s di and the source - domain label L obtained in step 1) s the pixel - level cross - entropy loss L task to optimize the segmentation framework:

[0063]

[0064] where H and W are the height and width of P s di , C is the number of classes of the source - domain label L s , h and w are the height - position index and width - position index of the current calculation, and c is the class of the current calculation;

[0065] 23) The auxiliary decoder G Aux is composed of fully - convolutional layers, G Aux receives the domain - invariant shallow - level features of the source domain to generate the source - domain auxiliary segmentation probability map P s Aux , G Aux the auxiliary segmentation loss function L of aux_task is defined as the pixel - level cross - entropy loss of P s Aux and the source - domain label L s :

[0066]

[0067] where H and W are the height and width of P s Aux , P s Aux and P s diThe width, height, and number of categories are all the same.

[0068] 3) Measure the difference between the domain-invariant high-level features of the target domain and the target domain-specific features through the difference loss, and combine the reconstruction process to prevent information loss; align the domain-invariant high-level features of the source domain and the domain-invariant high-level features of the target domain based on the adversarial method.

[0069] 31) Based on the target domain-specific features F obtained in step 2) t ds , the domain-invariant high-level features of the target domain and the domain-invariant shallow-level features of the target domain First, use Pass through the auxiliary decoder G Aux to output the target domain auxiliary probability map P t Aux , and then based on P t Aux Calculate the category center features, and the formula is as follows:

[0070]

[0071] In the formula, is the result of normalizing the probability of P t Aux at the spatial position, is the category feature center of the target domain-specific features, is the category feature center of the domain-invariant high-level features of the target domain;

[0072] 32) To achieve feature decoupling, normalize and to obtain the normalized category feature centers of the target domain-specific features and the domain-invariant high-level features of the target domain:

[0073]

[0074]

[0075] In the formula, ||·||2 represents the L2 norm, and the constant ε = 10 -6 is used to prevent division-by-zero errors;

[0076] Then calculate the cosine similarity and construct the orthogonal constraint loss L diff , forcing the domain-specific and domain-invariant features to be orthogonally distributed in the feature space and eliminating redundant semantic interference. The formula is as follows:

[0077]

[0078] 33) To preserve the information integrity in the feature decoupling process, Ft ds and After adding pixel by pixel, the result is input into the reconstruction decoder R to generate a reconstructed image where R consists of 3 convolutional layers with channel numbers {256, 64, 3}; the reconstruction loss L is defined based on the mean squared error loss recon , and the formula is as follows:

[0079]

[0080] In the formula, 3 represents the number of image channels, is the product of the height and width of the input image, which is the total number of pixels per channel; ch is the channel number being currently calculated, i is the pixel index being currently calculated, is the pixel feature of the original image being currently calculated, is the pixel feature of the reconstructed image being currently calculated;

[0081] 34) Introduce a fully convolutional domain discriminator D, which consists of 4 convolutional layers with a kernel size of 3×3, a stride of 1, and channel numbers {256, 128, 64, 2}. Except for the last layer, each convolutional layer is followed by a LeakyReLU activation function with a parameter of 0.2; D performs domain classification in the spatial dimension on the source domain domain-invariant high-level features and the target domain domain-invariant high-level features to obtain the source domain domain classification prediction result P s D and the target domain domain classification prediction result P t D , and the adversarial loss function L of the discriminator D adv is defined as:

[0082]

[0083] The domain-invariant encoder implicitly optimizes feature generation through the gradient reversal strategy, making the discriminator unable to distinguish the feature source and achieving feature alignment.

[0084] 4) Process the target domain domain-invariant high-level features and the target domain domain-invariant low-level features through the main decoder and the auxiliary decoder respectively, and use the consistency loss to ensure that the prediction results of different-level features are consistent.

[0085] 41) Based on the target domain domain-invariant low-level features obtained in step 2 through the auxiliary decoder G Aux obtain the target domain auxiliary segmentation probability map P t Aux , and based on the decoupled target domain domain-invariant high-level features obtained in step 3 through the main decoder G MObtain the target domain main segmentation probability map P t di ;

[0086] 42) Based on P t di Generate pseudo - labels For each pixel in P t di Determine the highest probability value and its corresponding class index, and at the same time calculate the second - highest probability value; if the difference between the highest probability value and the second - highest probability value exceeds a predefined threshold, assign the label of the pixel to the class index corresponding to the highest probability value; otherwise, set the pixel label to - 1, indicating that the pixel is ignored and does not participate in the loss calculation.

[0087] 43) Use P t Aux And the pseudo - labels The pixel - level cross - entropy loss L MFAM Optimize the domain - invariant encoder, main decoder, and auxiliary decoder:

[0088]

[0089] Finally, obtain the enhanced target domain domain - invariant high - level features

[0090] 5) Fuse the target domain domain - invariant high - level features and target domain - specific features to generate target domain fused features, and perform target domain self - learning by aligning its prediction results with the prediction results of the domain - invariant features to optimize the cross - domain adaptability of the decoder.

[0091] 51) Based on the decoupled target domain - specific features F t ds′ Obtained in step 3) and the enhanced target domain domain - invariant high - level features Obtained in step 4), generate more representative target domain fused features through pixel - by - pixel summation

[0092] 52) Input F t di+ds Into the main decoder G M To generate the target domain fused segmentation probability map P t di+ds = G M (F t di+ds ), use the pseudo - labels As the supervision signal to optimize the adaptability to the target domain data;

[0093] 53) Use P t AuxWith the pseudo-labels obtained in step 4) of the pixel-level cross-entropy loss L TEM Optimize the domain-invariant encoder, domain-specific encoder, and main decoder:

[0094]

[0095] 6) Obtain the semantic segmentation result of the target domain image: Based on the target domain fusion features obtained in step 5), the main decoder is used to parse the target domain fusion features to generate the segmentation prediction result of the target domain image.

[0096] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.

Claims

1. A cross-domain remote sensing image semantic segmentation method based on unsupervised domain adaptation, characterized in that, It includes the following steps: S1: Obtain remote sensing image data, and preprocess the remote sensing image data to generate source domain images and target domain images of a unified size; the source domain images refer to image data with known semantic labels for training and learning, and the target domain images are image data to be semantically segmented and lacking annotation information; S2: Based on the source domain images and their corresponding semantic labels obtained in step S1, pre-train the encoder-decoder segmentation framework for semantic segmentation, where the encoder-decoder segmentation framework includes an encoder and a decoder; the encoder includes a domain-invariant encoder and a target domain-specific encoder, and the decoder includes a main decoder and an auxiliary decoder; through pre-training, an initialized semantic segmentation framework is obtained. The domain-invariant encoder is used to extract source domain domain-invariant features and target domain domain-invariant features, where the source domain domain-invariant features include source domain domain-invariant high-level features and source domain domain-invariant low-level features, and the target domain domain-invariant features include target domain domain-invariant high-level features and target domain domain-invariant low-level features. The target domain-specific encoder extracts target domain-specific features; S3: Based on the source domain domain-invariant high-level features, target domain domain-invariant high-level features, and target domain-specific features obtained in step S2, perform feature decoupling: First, use the orthogonal constraint loss to measure the difference between the target domain domain-invariant high-level features and the target domain-specific features, and apply the reconstruction loss to avoid information loss. Then, align the source domain domain-invariant high-level features and the target domain domain-invariant high-level features through an adversarial method to improve the cross-domain transfer ability, and obtain the decoupled and aligned target domain domain-invariant high-level features and target domain-specific features; S4: Based on the decoupled target domain domain-invariant high-level features in step S3 and the target domain domain-invariant low-level features obtained in step S2, perform multi-level feature alignment to promote the representation enhancement of the target domain domain-invariant high-level features. The main decoder and the auxiliary decoder are used to process the target domain domain-invariant high-level features and the target domain domain-invariant low-level features respectively, and the consistency loss constraint is used to achieve the consistency of the prediction results of the same category features at different levels, and obtain the representation-enhanced target domain domain-invariant high-level features; S5: Based on the decoupled target domain-specific features in step S3 and the representation-enhanced target domain domain-invariant high-level features obtained in step S4, perform target domain self-learning, fuse the two features to obtain the target domain fusion features, and improve the adaptability of the decoder to the target domain by aligning the prediction results of the target domain fusion features and the target domain domain-invariant high-level features, and obtain the target domain fusion features that fully represent the target domain information; S6: Based on the target domain fusion features that fully represent the target domain information obtained in step S5, parse the target domain fusion features through the main decoder to generate the segmentation prediction results of the target domain images.

2. The cross-domain remote sensing image semantic segmentation method based on unsupervised domain adaptation according to claim 1, characterized in that The specific operation steps of step S2 are: S21: Use the domain-invariant encoder E di and the target-domain specific encoder E ds to extract image features. Among them, the domain-invariant encoder E di extracts the domain-invariant features from the source-domain image S and the target-domain image T, denoted as F s di = E di (S) and F t di = E di (T), and and where and are the source-domain domain-invariant shallow features and the target-domain domain-invariant shallow features, from the middle-layer output of the domain-invariant encoder, and are the source-domain domain-invariant high-level features and the target-domain domain-invariant high-level features, from the last-layer output of the domain-invariant encoder. The subscripts t and s represent the target domain and the source domain respectively. The target-domain specific encoder E ds only acts on the target-domain image to extract the target-domain specific feature F t ds = E ds (T); The encoder selects ResNet50 as the backbone network; S22: Main decoder G M Receives the source domain domain-invariant high-level features Generates the source domain main segmentation probability map P s di , G M Integrates a dual attention module followed by a fully convolutional segmentation head; the dual attention module combines the global attention mechanism and the channel attention mechanism, which can effectively capture global context information while retaining local details; uses P s di And the source domain label L obtained based on step S1 s The pixel-level cross-entropy loss L task Optimize the segmentation backbone: where H and W are the height and width of P s di and C is the number of classes of the source domain label L s ; h and w are the height position index and width position index calculated currently, and c is the class calculated currently; S23: Auxiliary decoder G Aux which is composed of fully convolutional layers, G Aux receives the source domain domain-invariant shallow features and generates the source domain auxiliary segmentation probability map P s Aux , G Aux The auxiliary segmentation loss function L of aux_task is defined as the pixel-level cross-entropy loss between P s Aux and the source domain label L s : Wherein, H and W are the height and width of P s Aux , P s Aux and P s di have the same width, height, and number of categories.

3. A cross-domain remote sensing image semantic segmentation method based on unsupervised domain adaptation according to claim 2, characterized in that The specific operation steps of step S3 are: S31: Based on the target domain specific feature F obtained in step S2 t ds , the target domain domain-invariant high-level feature and the target domain domain-invariant low-level feature First use to output the target domain auxiliary probability map P through the auxiliary decoder G Aux t Aux , and then calculate the class center feature based on P t Aux The formula is as follows:​ In the formula, is the result of normalizing the probability of P t Aux in the spatial position, is the class feature center of the target domain specific feature, is the class feature center of the target domain domain-invariant high-level feature; S32: To achieve feature decoupling, and are normalized to obtain the class feature centers of the normalized target domain-specific features and the target domain domain-invariant high-level features: where ||·||2 represents the L2 norm, and the constant ε = 10 -6 is used to prevent division-by-zero errors; Then calculate the cosine similarity and construct the orthogonal constraint loss L diff , which forces the domain-specific and domain-invariant features to be orthogonally distributed in the feature space and eliminates redundant semantic interference. The formula is as follows: S33: To preserve the information integrity during the feature decoupling process, add F t ds and pixel by pixel and input the result into the reconstruction decoder R to generate the reconstructed image where R consists of 3 convolutional layers with the number of channels being {256, 64, 3}; define the reconstruction loss L based on the mean squared error loss recon , and the formula is as follows: In the formula, 3 represents the number of image channels, is the product of the height and width of the input image, which is the total number of pixels per channel; ch is the channel number currently being calculated, and i is the pixel index currently being calculated. is the pixel feature of the original image currently being calculated, is the pixel feature of the reconstructed image currently being calculated; S34: Introduce a fully convolutional domain discriminator D, which consists of 4 convolutional layers with a kernel size of 3×3, a stride of 1, and the number of channels being {256, 128, 64, 2} respectively. Except for the last layer, each convolutional layer is followed by a LeakyReLU activation function with a parameter of 0.2; D performs domain classification in the spatial dimension on the domain-invariant high-level features of the source domain and the domain-invariant high-level features of the target domain to obtain the source domain classification prediction result P s D and the target domain classification prediction result P t D , and the adversarial loss function L adv of the discriminator D is defined as: The domain-invariant encoder implicitly optimizes feature generation through the gradient reversal strategy, making the discriminator unable to distinguish the feature source and achieving feature alignment.

4. A cross-domain remote sensing image semantic segmentation method based on unsupervised domain adaptation according to claim 3, characterized in that The specific operation steps of step S4 are: S41: Based on the target domain domain-invariant shallow features obtained in step S2 through the auxiliary decoder G Aux obtain the target domain auxiliary segmentation probability map P t Aux , based on the decoupled target domain domain-invariant high-level features obtained in step S3 through the main decoder G M obtain the target domain main segmentation probability map P t di ; S42: Based on P t di Generate pseudo labels For each pixel in P t di Determine the highest probability value and its corresponding class index, and at the same time calculate the second highest probability value; if the difference between the highest probability value and the second highest probability value exceeds a predefined threshold, assign the label of the pixel to the class index corresponding to the highest probability value; otherwise, set the pixel label to -1, indicating that the pixel is ignored and does not participate in the loss calculation. S43: Use P t Aux with the pseudo-label of the pixel-level cross-entropy loss L MFAM Optimize the domain-invariant encoder, the main decoder, and the auxiliary decoder: Finally, the enhanced target-domain domain-invariant high-level features are obtained 5. A cross-domain remote sensing image semantic segmentation method based on unsupervised domain adaptation according to claim 4, characterized in that, The specific operation steps of step S5 are as follows: S51: Based on the decoupled target-domain specific feature F obtained in step S3 t ds′ and the enhanced target-domain domain-invariant high-level feature obtained in step S4 generate a more representative target-domain fusion feature through pixel-wise summation S52: Input F t di+ds into the main decoder G M to generate the target domain fusion segmentation probability map P t di+ds = G M (F t di+ds ), and use the pseudo-label as the supervision signal to optimize the adaptability to the target domain data; S53: Use P t Aux Optimize the domain-invariant encoder, domain-specific encoder, and main decoder with the pixel-level cross-entropy loss L based on the pseudo-labels obtained in step S4 TEM :

Citation Information

Cited By

  • Cluster robot cross-view collaborative sensing method and system for aircraft manufacturing

    CN120495643A

  • View angle robust traffic accident detection method based on space-time attention domain adaptation

    CN121191046A

  • Unsupervised domain adaptive medical image segmentation method, device and equipment

    CN121259338A

  • An unsupervised domain adaptation medical image segmentation method, device and equipment

    CN121259338B