RetNet-based unsupervised domain adaptive remote sensing semantic segmentation method
By introducing attention mechanism with spatial attenuation matrix and TokenMix data enhancement, combined with self-training and teacher network, the problem of the difference in data distribution between source and target domains in semantic segmentation of remote sensing images is solved, and high-precision remote sensing image segmentation is achieved, especially in the identification of building categories in urban road scenarios.
Patent Information
- Application Number
- CN202510523667.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-08-15
AI Technical Summary
The existing remote sensing image semantic segmentation method is difficult to achieve high-precision segmentation when facing the differences in data distribution of source and target domains caused by factors such as lighting conditions and seasonal changes in different scenarios, and relying on a large amount of labeled data leads to high cost.
The unsupervised domain adaptive remote sensing semantic segmentation method based on RetNet is adopted, and the attention mechanism with spatial attenuation matrix and TokenMix data enhancement is introduced, combining self-training and teacher networks to reduce inter-domain differences and improve the segmentation accuracy of the model for remote sensing images.
It effectively improves the model's understanding of complex backgrounds, reduces domain gaps, and realizes high-precision semantic segmentation of remote sensing images, especially in the identification of different building categories in urban road scenarios.
Smart Images

Figure CN120495653A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image processing, and in particular to an unsupervised domain adaptive remote sensing semantic segmentation method based on RetNet. Background Art
[0002] With the continuous development of remote sensing technology, remote sensing images are increasingly being used in fields such as urban planning, environmental monitoring, and traffic management. Semantic segmentation of remote sensing images of urban road scenes is crucial for applications such as urban management and autonomous driving. However, due to factors such as lighting conditions, seasonal variations, and camera angles in different scenarios, data distribution differs between the source and target domains, posing a challenge to semantic segmentation of remote sensing images.
[0003] Traditional semantic segmentation methods for remote sensing images typically rely on large amounts of labeled data and trained through supervised learning. However, in practical applications, obtaining large amounts of labeled data is not only time-consuming and labor-intensive, but also costly. Furthermore, traditional methods often struggle to achieve ideal segmentation results when faced with disparate data distributions between the source and target domains.
[0004] In recent years, deep learning technology has made significant progress in the field of image semantic segmentation. Transformer-based models have garnered widespread attention due to their ability to effectively capture long-range dependencies in images. However, directly applying Transformers to semantic segmentation of remote sensing images still faces several challenges, such as insufficient understanding of complex backgrounds and a large domain gap between the source and target domains.
[0005] To address these issues, this paper proposes a RetNet-based unsupervised domain-adaptive remote sensing semantic segmentation method. By introducing a RetNet-based segmentation network and a TokenMix hybrid module for an unsupervised domain-adaptive framework, this method effectively improves the model's ability to understand the complex background of remote sensing images, while reducing the domain gap between the source and target domains, achieving high-precision semantic segmentation of remote sensing images. Summary of the Invention
[0006] Purpose of the invention: The present invention aims to provide an unsupervised domain-adaptive remote sensing semantic segmentation method based on RetNet, which effectively reduces the distribution difference between the source domain and the target domain by introducing TokenMix data enhancement and a dual-teacher network alternating training mechanism, and improves the segmentation accuracy of the model for remote sensing images.
[0007] Technical solution: This paper proposes an unsupervised domain adaptive remote sensing semantic segmentation method based on RetNet, which includes the following steps:
[0008] Step 1: Obtain remote sensing image data of urban road scenes, manually annotate them according to the categories of different buildings in the scene, build a training dataset, and preprocess the dataset images;
[0009] Step 2: Build a backbone network for remote sensing image segmentation. The backbone network is based on RetNet. In the encoder stage, an attention mechanism with a spatial attenuation matrix is used to replace the traditional attention mechanism. It extracts multi-level features and generates hierarchical features of the input image.
[0010] Step 3: Construct a decoder, using bilinear upsampling and convolution to embed the multi-level feature maps extracted in the encoding phase, and use convolution kernels with different expansion rates combined with point-by-point convolution to achieve context-aware fusion features;
[0011] Step 4: Add an intermediate domain between the source and target domains; use the TokenMix method to alternately mix the cross-domain data, introduce dynamic mixing ratios and introduce feature consistency constraints; use self-training and teacher networks to form a complete RetNet-based unsupervised domain-adaptive remote sensing semantic segmentation model;
[0012] Step 5: Iteratively train the remote sensing semantic segmentation model to obtain the optimal training weights and finally obtain the optimized remote sensing image semantic segmentation model.
[0013] Furthermore, in step 1, the label image and the original image of the remote sensing dataset are scaled and cropped; then, the data is further enhanced using color jittering and Gaussian blurring, and the enhanced dataset is used to generate a training set, a validation set, and a test set.
[0014] Furthermore, in step 2, in the encoder, not only is the traditional attention mechanism replaced by a self-attention based on the attenuation matrix, but the one-dimensional attenuation is also extended to a two-dimensional attenuation; since the observed unidirectional and one-dimensional temporal attenuation is converted into a bidirectional and two-dimensional spatial attenuation, the spatial attenuation introduces the explicit spatial prior related to the Manhattan distance into the visual backbone, thereby forming a new self-attention based on the spatial attenuation matrix.
[0015] The attenuation matrix calculation formula is:
[0016]
[0017] in, represents the bidirectional spatial attenuation weight between pixel points n and m, γ is a hyperparameter that controls the attenuation strength, and x n ,y n Represents the spatial coordinates of pixel n, x m ,y mRepresents the spatial coordinates of pixel point m. By introducing a bidirectional spatial attenuation mechanism, the model's ability to understand the complex background of remote sensing images is enhanced.
[0018] Furthermore, in step 2, the specific implementation formula of the improved attention mechanism MaSA is:
[0019] MaSA(X)=(Softmax(QK T ⊙D 2d ))V
[0020] Among them, Q, K, and V are the query matrix, key matrix, and value matrix obtained by linear transformation of input X, respectively. 2d is the bidirectional spatial attenuation matrix.
[0021] Furthermore, in step 2, in order to reduce the computational complexity of global information modeling, the attention based on the spatial attenuation matrix can be converted into a decomposable form, which can model global information with linear complexity while maintaining the spatial attenuation matrix. The attention scores in the horizontal and vertical directions of the image are calculated separately, and then a one-dimensional bidirectional attenuation matrix is applied to these attention weights. The decomposed matrix is:
[0022]
[0023] Furthermore, in step 2, a decomposable form of self-attention is characterized in that the horizontal and vertical distances between markers are represented by a one-dimensional decay matrix. The decomposed attention is:
[0024]
[0025] MaSA(X)=MaSA H (MaSA W V) T
[0026] MaSA H Indicates the attention in the vertical direction, D H Represents the spatial attenuation matrix decomposed in the vertical direction, MaSA W Indicates attention in the horizontal direction, D W Represents the spatial attenuation matrix decomposed in the vertical direction. The source and target domain images are decomposed into multiple overlapping image blocks, which are analyzed using the local attention mechanism MaSA and fed back to the RetNet encoders at different stages;
[0027] Furthermore, in step 2, after layer normalization and MaSA operation in RetNetBlock, it goes through a layer normalization and feedforward network, and the calculation formula is:
[0028]
[0029] MLP is a multi-layer perceptron, W is the weight matrix and bias vector, which are used to adjust the features. MaSA is used to capture the relationship between different parts of the input sequence. When calculating, the output Z of the previous layer is L-1 Perform layer normalization and pass the result to the MaSA function for output. The result is the same as the output Z of the previous layer. L-1 The sum is subtracted from the bias vector to obtain the layered output.
[0030] Furthermore, in step 3, before feature fusion, each input feature map Fi is embedded into the same number of channels C through a 1×1 convolution, and all feature maps are adjusted to the same spatial size as the highest resolution feature map using bilinear upsampling, and the adjusted feature maps are spliced in the channel direction.
[0031] Furthermore, in step 3, context-aware feature fusion is performed using multiple parallel 3×3 depthwise separable convolutions with different dilation rates, similar to the ASPP structure. Unlike traditional ASPP, global average pooling is not used. Instead, a final 1×1 convolution is used to fuse multi-scale features into a single output.
[0032] Furthermore, in step 4, the input image is divided into blocks, and the input image x is divided into non-overlapping patches, where Patches are converted into visual tokens through linear projection to form a token sequence that is adapted to RetNet;
[0033] Furthermore, in step 4, a dynamic mixing ratio λ and a random mask M are generated at the Token level. t , where M t ∈R H / P×W / P , where λ is dynamically calculated from the attention graph:
[0034]
[0035] in and The attention maps for source and target domain samples are generated by spatially normalized activation maps before the last classification head of the pre-trained network. The coverage ratio of the source domain token in the mask is determined based on λ, and a mask with non-uniform occlusion in multiple regions is generated.
[0036] Furthermore, in step 4, according to the mask M t , the source domain image x S and the target domain image x TTokens are mixed to generate new samples
[0037]
[0038] At the same time, combined with the source domain true label y s and target domain pseudo labels Generate mixed labels
[0039]
[0040] in Generate prediction results for target domain samples through the dual-teacher network, and select highly reliable pseudo labels based on the confidence threshold;
[0041] Furthermore, in step 4, feature consistency constraints are introduced at the same time, by maximizing the hybrid feature f m and the source domain feature f s , target domain features f t The cosine similarity of can be used to optimize the model parameters during the training iteration:
[0042] L sim =λ·cos(f m ,f s )+(1-λ)·cos(f m ,f t ),
[0043] Furthermore, in step 4, the mixed sample With label After inputting the segmentation network, it will be jointly trained with the cross entropy loss and feature consistency loss to gradually align the distribution differences between domains.
[0044] Where cos() represents the cosine similarity calculation, which constrains the consistency of the mixed sample in the feature space with the source domain and the target domain through back propagation;
[0045] Furthermore, in step 4, the target domain and the intermediate domain are alternately blended. First, a blended image is generated by combining the source domain and the intermediate domain, and a blended label is generated by combining the source domain true label and the intermediate domain pseudo label. Then, a blended image is generated by combining the source domain true label and the target domain pseudo label, and a blended label is generated by combining the source domain true label and the target domain pseudo label. These two blending steps are alternated in each iteration to train the model to adapt to data from different domains.
[0046] Furthermore, in step 4, a self-training method is used to first obtain a set of annotated images K from the source domain. S , expressed as:
[0047]
[0048] in is the i-th image in the source domain, yes The corresponding annotations or labels, Ns is the total number of images; secondly, the sample annotations in the target domain are:
[0049]
[0050] in is the i-th image in the target domain, yes The corresponding annotation or label, N T is the total number of images; finally, a large amount of unlabeled data is obtained from the target domain, namely:
[0051]
[0052] in is the i-th unlabeled image in the target domain, N M is the total number of unlabeled target domain images.
[0053] Furthermore, in step 4, the student model g is firstly performed on the source domain. θ The model uses the cross entropy loss function to measure the difference between the model prediction and the true annotation, which is expressed as follows:
[0054]
[0055] in, represents the loss function on the source domain, θ is the model parameter, represents the true label of the i-th source image and model predictions The cross entropy between .
[0056] Furthermore, in step 4, the dual teacher networks hφ1 and hφ2 are used for alternating training in the self-training process to generate pseudo labels for the target domain data, and the formula is:
[0057]
[0058] in, Represents the pseudo label probability of the jth pixel corresponding to the i-th target domain image in category c. This probability is based on Teacher Network The prediction of the jth pixel of the image over all possible categories c′ is determined by selecting the category c with the maximum probability as the pseudo label for the pixel.
[0059] Furthermore, in step 4, a complete RetNet-based unsupervised domain-adaptive remote sensing semantic segmentation model is constructed. First, the segmentation model is trained using labeled data from the source domain. This allows the segmentation network to learn basic segmentation capabilities and obtain gradient information through backpropagation to optimize model parameters. Next, the TokenMix method is used to perform data mixing between the source domain and the newly introduced intermediate domain. During this process, the segmentation network continues to learn and backpropagates gradients based on the new mixed samples to further adjust its parameters. Then, the same TokenMix method is used to perform mixed training on data from the source and target domains. This mixed training helps the segmentation network better adapt to the data distribution of the target domain and reduce the differences between the source and target domains. This process allows the segmentation network to obtain more gradient information, thereby improving its generalization ability on unseen data. Finally, the model adopts a dual-teacher network strategy, in which the two teacher networks are trained alternately with the segmentation network. The predictions of the teacher networks are used to generate pseudo-labels for the target domain data, and the segmentation network learns based on these pseudo-labels. In addition, the exponential moving average (EMA) technique is used to smooth the model's prediction results, that is, the weight of the teacher network is an exponentially weighted average of the weight of the segmentation network. This approach can help stabilize the training process and promote the effective transfer of knowledge from the teacher network to the segmentation network.
[0060] Beneficial effects:
[0061] The present invention provides an unsupervised domain adaptive remote sensing semantic segmentation method based on RetNet, which has the following advantages:
[0062] 1. We introduce an attention mechanism with a spatial decay matrix that decays attention scores based on the distance between pixels, extracts multi-level features, and generates hierarchical features for the input image. This effectively extracts and generates hierarchical features, enhancing the model's ability to understand complex backgrounds in remote sensing images and making it suitable for semantic segmentation of different building categories in urban road scenes.
[0063] 2. The method provided by this invention utilizes multi-level features, applies them to bottleneck features, and fuses all stacked multi-level features, utilizing low-level features at different resolutions and semantically rich high-level features. Depthwise separable convolution uses fewer parameters than ordinary convolution, reducing the risk of overfitting to the source domain. Multi-scale feature fusion and depthwise separable convolution efficiently integrate high- and low-level features while avoiding overfitting. This enables the model to better capture detailed information and contextual relationships in semantic segmentation tasks.
[0064] 3. This paper adds an intermediate domain between the source and target domains; uses the TokenMix method to alternately mix cross-domain data, introduces dynamic mixing ratios, and introduces feature consistency constraints; and utilizes self-training and a teacher network to form a complete RetNet-based unsupervised domain-adaptive remote sensing semantic segmentation model.
[0065] 4. Experiments show that this method rapidly increases the accuracy of the model from 40% to 70% in the early stages of training, demonstrating rapid learning. The training loss plot shows a rapid decrease from 0.7 to 0.5. RetNet performs better in the middle and late stages, with the loss decreasing even faster, reaching approximately 0.4. On the Potsdam→Vaihingen dataset, the RetNet model achieves high F1 scores for object recognition, including clutter, car, and tree. It achieves a particularly high F1 score of 54.79% for clutter, the highest among all models. For building recognition, the RetNet model achieves the highest F1 score, reaching 94.62%, demonstrating its advantages in building segmentation. Overall, the RetNet model achieves the highest mIoU and mF1 scores, at 58.49% and 70.25%, respectively. This demonstrates that it provides the most balanced and highest performance across all considered categories. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] Figure 1 :Flowchart of the unsupervised domain adaptive remote sensing semantic segmentation method based on RetNet;
[0067] Figure 2 :The complete RetNet-based unsupervised domain adaptive remote sensing semantic segmentation model in this invention;
[0068] Figure 3 : encoder of segmentation network RetNet;
[0069] Figure 4 : RetNetBlock of the segmentation network RetNet in the encoding stage;
[0070] Figure 5 : Decoder of segmentation network RetNet;
[0071] Figure 6 : TokenMix cross-domain mixing process;
[0072] Figure 7 : A self-training framework for unsupervised domain adaptation;
[0073] Figure 8 : Training accuracy curve;
[0074] Figure 9 : training loss rate curve;
[0075] Figure 10 : IoU performance comparison histogram;
[0076] Figure 11 : F1 performance comparison bar chart.
[0077] Specific implementation examples
[0078] For a better understanding of the present invention, the present invention is further described below with reference to the accompanying drawings in the examples of the present invention, but is not intended to limit the present invention. Various modifications and improvements made to the technical solution of the present invention by ordinary persons in the art without departing from the design concept of the present invention should fall within the scope of protection of the present invention.
[0079] like Figure 1-11 As shown in Figure 1, an unsupervised domain adaptive remote sensing semantic segmentation method based on RetNet is as follows:
[0080] This paper describes in detail the implementation of an unsupervised domain-adaptive remote sensing semantic segmentation method based on RetNet using the Potsdam→Vaihingen dataset. This implementation includes a complete flow of data processing, feature extraction, fusion, and prediction output, strictly adhering to the technical solutions of claims 1-6. This is for illustrative purposes only and does not limit the scope of protection of this invention.
[0081] Step 1: Obtain remote sensing image data of urban road scenes, manually annotate them according to the categories of different buildings in the scene, construct a training dataset, and preprocess the dataset images; scale and crop the labeled images and original images of the remote sensing dataset to a size of 512×512; then further enhance the data using color jittering and Gaussian blurring, and use the enhanced dataset to generate training, validation, and test sets.
[0082] Step 2: Build a backbone network for remote sensing image segmentation. The backbone network is based on RetNet. In the encoder stage, an attention mechanism with a spatial attenuation matrix is used to replace the traditional attention mechanism. It extracts multi-level features and generates hierarchical features of the input image.
[0083] (1) In the encoder, not only is the traditional attention mechanism replaced by a self-attention based on the attenuation matrix, but the one-dimensional attenuation is also extended to a two-dimensional attenuation; since the observed unidirectional and one-dimensional temporal attenuation is converted to a bidirectional and two-dimensional spatial attenuation, the spatial attenuation introduces the explicit spatial prior related to the Manhattan distance into the visual backbone, thus forming a new self-attention based on the spatial attenuation matrix.
[0084] The attenuation matrix calculation formula is:
[0085]
[0086] in, represents the bidirectional spatial attenuation weight between pixel points n and m, γ is a hyperparameter that controls the attenuation strength, and x n ,y n Represents the spatial coordinates of pixel n, x m ,y m Represents the spatial coordinates of pixel point m. By introducing a bidirectional spatial attenuation mechanism, the model's ability to understand the complex background of remote sensing images is enhanced.
[0087] (2) Improved attention mechanism MaSA:
[0088] MaSA(X)=(Softmax(QK T ⊙D 2d ))V
[0089] Among them, Q, K, and V are the query matrix, key matrix, and value matrix obtained by linear transformation of input X, respectively. 2d is the bidirectional spatial attenuation matrix.
[0090] (3) To reduce the computational complexity of global information modeling, the attention based on the spatial attenuation matrix can be converted into a decomposable form, which can model global information with linear complexity while maintaining the spatial attenuation matrix. The attention scores in the horizontal and vertical directions of the image are calculated separately, and then a one-dimensional bidirectional attenuation matrix is applied to these attention weights. The decomposed matrix is:
[0091]
[0092] (4) Decomposable form of self-attention, using a one-dimensional decay matrix to represent the horizontal and vertical distances between tokens. The decomposed attention is:
[0093]
[0094] MaSA(X)=MaSA H (MaSA W V) T
[0095] MaSA H Indicates the attention in the vertical direction, D H Represents the spatial attenuation matrix decomposed in the vertical direction, MaSA W Indicates attention in the horizontal direction, D WRepresents the spatial attenuation matrix decomposed in the vertical direction. The source and target domain images are decomposed into multiple overlapping image blocks, which are analyzed using the local attention mechanism MaSA and fed back to the RetNet encoders at different stages;
[0096] (5) After layer normalization and MaSA operation in RetNetBlock, it goes through a layer normalization and feedforward network, and the calculation formula is:
[0097]
[0098] MLP is a multi-layer perceptron, W is the weight matrix and bias vector, which are used to adjust the features. MaSA is used to capture the relationship between different parts of the input sequence. When calculating, the output Z of the previous layer is L-1 Perform layer normalization and pass the result to the MaSA function for output. The result is the same as the output Z of the previous layer. L-1 The sum is subtracted from the bias vector to obtain the layered output.
[0099] Step 3: Construct a decoder, using bilinear upsampling and convolution to embed the multi-level feature maps extracted in the encoding phase, and use convolution kernels with different expansion rates combined with point-by-point convolution to achieve context-aware fusion features;
[0100] (1) Before feature fusion, each input feature map Fi is embedded into the same number of channels C through a 1×1 convolution, all feature maps are adjusted to the same spatial size as the highest resolution feature map using bilinear upsampling, and the adjusted feature maps are spliced along the channel direction.
[0101] (2) Context-aware feature fusion is performed using multiple parallel 3×3 depthwise separable convolutions with different dilation rates, similar to the ASPP structure. Unlike traditional ASPP, global average pooling is not used. Instead, a 1×1 convolution is used to fuse multi-scale features into a single output.
[0102] Step 4: Add an intermediate domain between the source and target domains; use the TokenMix method to alternately mix the cross-domain data, introduce dynamic mixing ratios and introduce feature consistency constraints; use self-training and teacher networks to form a complete RetNet-based unsupervised domain-adaptive remote sensing semantic segmentation model;
[0103] (1) The input image is divided into blocks, and the input image x is divided into non-overlapping patches, where Patches are converted into visual tokens through linear projection to form a token sequence that is adapted to RetNet;
[0104] (2) Generate dynamic mixing ratio λ and random mask M at the Token level t , where M t ∈R H / P×W / P , where λ is dynamically calculated from the attention graph:
[0105]
[0106] in and The attention maps for source and target domain samples are generated by spatially normalized activation maps before the last classification head of the pre-trained network. The coverage ratio of the source domain token in the mask is determined based on λ, and a mask with non-uniform occlusion in multiple regions is generated.
[0107] (3) According to the mask M t , the source domain image x S and the target domain image x T Tokens are mixed to generate new samples
[0108]
[0109] At the same time, combined with the source domain true label y s and target domain pseudo labels Generate mixed labels
[0110]
[0111] in Generate prediction results for target domain samples through the dual-teacher network, and select highly reliable pseudo labels based on the confidence threshold;
[0112] (4) At the same time, feature consistency constraints are introduced to maximize the hybrid feature f m and the source domain feature f s , target domain features f t The cosine similarity of can be used to optimize the model parameters during the training iteration:
[0113] L sim =λ·cos(f m ,f s )+(1-λ)·cos(f m ,f t ),
[0114] (5) Mixed samples With label After inputting the segmentation network, it will be jointly trained with the cross entropy loss and feature consistency loss to gradually align the distribution differences between domains.
[0115] Where cos() represents the cosine similarity calculation, which constrains the consistency of the mixed sample in the feature space with the source domain and the target domain through back propagation;
[0116] (6) Alternately mix the target domain and the intermediate domain. First, generate a mixed image by combining the source domain and the intermediate domain. The source domain's true label and the intermediate domain's pseudo label generate a mixed label. Then, generate a mixed image by combining the source domain's true label and the target domain's pseudo label generate a mixed label. These two mixing steps are performed alternately in each iteration to train the model to adapt to data from different domains.
[0117] Step 5: Iteratively train the remote sensing semantic segmentation model to obtain the optimal training weights and finally obtain the optimized remote sensing image semantic segmentation model.
[0118] (1) Using self-training method, first obtain a set of annotated images K from the source domain S , expressed as:
[0119]
[0120] in is the i-th image in the source domain, yes The corresponding annotations or labels, Ns is the total number of images; secondly, the sample annotations in the target domain are:
[0121]
[0122] in is the i-th image in the target domain, yes The corresponding annotation or label, N T is the total number of images; finally, a large amount of unlabeled data is obtained from the target domain, namely:
[0123]
[0124] in is the i-th unlabeled image in the target domain, N M is the total number of unlabeled target domain images.
[0125] (2) First, perform the student model g on the source domain θ The model uses the cross entropy loss function to measure the difference between the model prediction and the true annotation, which is expressed as follows:
[0126]
[0127] in, represents the loss function on the source domain, θ is the model parameter, represents the true label of the i-th source image and model predictions The cross entropy between .
[0128] (3) In the self-training process, the dual teacher networks hφ1 and hφ2 are used for alternating training to generate pseudo labels for the target domain data. The formula is:
[0129]
[0130] in, Represents the pseudo label probability of the jth pixel corresponding to the i-th target domain image in category c. This probability is based on Teacher Network The prediction of the jth pixel of the image over all possible categories c′ is determined by selecting the category c with the maximum probability as the pseudo label for the pixel.
[0131] (4) Construct a complete RetNet-based unsupervised domain-adaptive remote sensing semantic segmentation model. First, the segmentation model is trained using labeled data from the source domain. The purpose is to enable the segmentation network to learn basic segmentation capabilities and obtain gradient information through backpropagation to optimize the model parameters. Next, the TokenMix method is used to perform data mixing between the source domain and the newly introduced intermediate domain. In this process, the segmentation network continues to learn and returns gradients based on the new mixed samples to further adjust its parameters. Then, the same TokenMix method is used to perform mixed training on the source domain and target domain data. This mixed training aims to help the segmentation network better adapt to the data distribution of the target domain and reduce the difference between the source domain and the target domain. Through this process, the segmentation network can obtain more gradient information, thereby improving its generalization ability on unseen data. Finally, the model adopts a dual teacher network strategy. The two teacher networks are trained alternately with the segmentation network. The prediction results of the teacher network are used to generate pseudo labels for the target domain data, and the segmentation network learns based on these pseudo labels. In addition, the exponential moving average (EMA) technique is used to smooth the model's prediction results, that is, the weight of the teacher network is an exponentially weighted average of the weight of the segmentation network. This approach can help stabilize the training process and promote the effective transfer of knowledge from the teacher network to the segmentation network.
[0132] This section compares the proposed algorithm with various domain adaptation algorithms, including AdaptSegNet, Advent, CLAN, MUCSS, ProCA, and DAFormer, a total of 8 algorithms. Experiments are conducted on the Potsdam→Vaihingen dataset domain adaptation dataset configuration. The results are shown in Table 1.
[0133] Table 1 Potsdam → Vaihingen performance comparison experimental data
[0134]
[0135] On the Potsdam→Vaihingen dataset, the RetNet model achieved high F1 scores for clutter, car, and tree recognition. In particular, it achieved a 54.79% F1 score for clutter, the highest among all models. For building recognition, the RetNet model achieved the highest F1 score, reaching 94.62%, demonstrating its advantages in building segmentation. Overall, the RetNet model achieved the highest mIoU and mF1 scores, at 58.49% and 70.25%, respectively. This demonstrates that it provides the most balanced and highest performance across all considered categories. Compared to the original CLAN model, it achieved significant improvements in every category. For example, on the road surface category, CLAN achieved an F1 score of 75.68%, while RetNet achieved 80.30%.
[0136] The above embodiments are intended only to illustrate the technical concepts and features of the present invention. Their purpose is to enable those skilled in the art to understand the contents of the present invention and implement them accordingly. They are not intended to limit the scope of protection of the present invention. Any equivalent changes or modifications made in accordance with the spirit of the present invention shall be covered by the present invention.
Claims
1. An unsupervised domain adaptive remote sensing semantic segmentation method based on RetNet, characterized by: The following steps are involved: Step 1: Obtain remote sensing image data of urban road scenes and preprocess the dataset images; Step 2: Build a backbone network for remote sensing image segmentation. The backbone network is based on RetNet and uses an attention mechanism with a spatial attenuation matrix in the encoder stage to extract multi-level features and generate hierarchical features of the input image. Step 3: Construct a decoder, using bilinear upsampling and convolution to embed the multi-level feature maps extracted in the encoding phase. It also uses convolution kernels with different expansion rates combined with point-by-point convolution to achieve context-aware fusion features. Step 4: Add an intermediate domain between the source and target domains, use the TokenMix method to alternately mix the cross-domain data, introduce dynamic mixing ratios and feature consistency constraints, and use self-training and teacher networks to form a complete RetNet-based unsupervised domain-adaptive remote sensing semantic segmentation model; Step 5: Iteratively train the remote sensing semantic segmentation model to obtain the optimal training weights and finally obtain the optimized remote sensing image semantic segmentation model.
2. The unsupervised domain adaptive remote sensing semantic segmentation method based on RetNet according to claim 1, characterized in that The step 1 includes scaling and cropping the label image and the original image of the remote sensing dataset; then further enhancing the data using color jittering and Gaussian blurring, and generating a training set, a validation set, and a test set from the enhanced dataset.
3. The remote sensing semantic segmentation method based on RetNet according to claim 1, characterized in that The segmentation network in step 2 includes: (2.1) In the encoder, the one-dimensional attenuation is extended to a two-dimensional attenuation; since the observed unidirectional and one-dimensional temporal attenuation is converted to a bidirectional and two-dimensional spatial attenuation, the spatial attenuation introduces an explicit spatial prior related to the Manhattan distance into the visual backbone, thus forming a new self-attention based on the spatial attenuation matrix, The attenuation matrix calculation formula is: in, represents the bidirectional spatial attenuation weight between pixel points n and m, γ is a hyperparameter that controls the attenuation strength, and x n ,y n Represents the spatial coordinates of pixel n, x m ,y m Represents the spatial coordinates of pixel point m, (2.2) The specific implementation formula of the improved attention mechanism MaSA is: MaSA(X)=(Softmax(QK T ⊙D 2d ))V Among them, Q, K, and V are the query matrix, key matrix, and value matrix obtained by linear transformation of input X, respectively. 2d is the bidirectional spatial attenuation matrix, (2.3) Converting the attention based on the spatial attenuation matrix into a decomposable form, it is possible to model global information with linear complexity while maintaining the spatial attenuation matrix. The attention scores in the horizontal and vertical directions of the image are calculated separately, and then the one-dimensional bidirectional attenuation matrix is applied to these attention weights. The decomposed matrix is: (2.4) The decomposable form of self-attention uses a one-dimensional decay matrix to represent the horizontal and vertical distances between tags. The decomposed attention is: Time(X)=Time H (Time) W V) T MaSA H Indicates the attention in the vertical direction, D H Represents the spatial attenuation matrix decomposed in the vertical direction, MaSA W Indicates attention in the horizontal direction, D W Represents the spatial attenuation matrix decomposed in the vertical direction, decomposing the images of the source and target domains into multiple overlapping image blocks, using the local attention mechanism MaSA to analyze these image blocks and feed them back to the RetNet encoders at different stages; (2.5) After layer normalization and MaSA operation in RetNetBlock, it goes through a layer normalization and feedforward network, and the calculation formula is: MLP is a multi-layer perceptron, W is a weight matrix and a bias vector, which is used to adjust the features. MaSA is used to capture the relationship between different parts of the input sequence. When calculating, the output Z of the previous layer is L-1 Perform layer normalization and pass the result to the MaSA function for output. The result is the same as the output Z of the previous layer. L-1 Add them together, subtract the summed result from the bias vector, and finally get the layered output.
4. The unsupervised domain adaptive remote sensing semantic segmentation method based on RetNet according to claim 1, characterized in that The specific implementation of the decoder in step 3 includes the following steps: (3.1) Before feature fusion, each input feature map Fi is embedded into the same number of channels C through a 1×1 convolution, and all feature maps are adjusted to the same spatial size as the highest resolution feature map using bilinear upsampling, and the adjusted feature maps are spliced along the channel direction; (3.2) Perform context-aware feature fusion by using multiple parallel 3×3 depth-wise separable convolutions with different dilation rates, and finally fuse the multi-scale features into a single output through a 1×1 convolution.
5. The unsupervised domain adaptive remote sensing semantic segmentation method based on RetNet according to claim 1, characterized in that The specific process of the TokenMix hybrid method for unsupervised domain adaptation in step 4 is as follows: (4.1) The input image is divided into blocks. The input image x is divided into non-overlapping patches, where Patches are converted into visual tokens through linear projection to form a token sequence that is adapted to RetNet; (4.2) Generate dynamic mixing ratio λ and random mask M at the Token level t , where M t ∈R H / P×W / P , where λ is dynamically calculated from the attention graph: in and The attention maps for source and target domain samples are generated by spatially normalized activation maps before the last classification head of the pre-trained network. The coverage ratio of the source domain token in the mask is determined based on λ, and a mask with non-uniform occlusion in multiple regions is generated. (4.3) According to the mask M t , the source domain image x S and the target domain image x T Tokens are mixed to generate new samples At the same time, combined with the source domain true label y s and target domain pseudo labels Generate mixed labels in Generate prediction results for target domain samples through the dual-teacher network, and select highly reliable pseudo labels based on the confidence threshold; (4.4) At the same time, feature consistency constraints are introduced by maximizing the mixed feature f m and the source domain feature f s , target domain features f t The cosine similarity of can be used to optimize the model parameters during the training iteration: L sim =λ·cos(f m ,f s )+(1-λ)·cos(f m ,f t ), (4.5) Mixed samples With label After inputting the segmentation network, it will be jointly trained with the cross entropy loss and feature consistency loss to gradually align the distribution differences between domains. Where cos() represents the cosine similarity calculation, which constrains the consistency of the mixed sample in the feature space with the source domain and the target domain through back propagation; (4.6) Alternately mix the target domain and the intermediate domain. First, generate a mixed image by combining the source domain and the intermediate domain. The source domain true label and the intermediate domain pseudo label generate a mixed label. Then, generate a mixed image by combining the source domain and the target domain true label and the target domain pseudo label generate a mixed label. Alternate these two mixing steps in each iteration to train the model to adapt to data from different domains.
6. The unsupervised domain adaptive remote sensing semantic segmentation method based on RetNet according to claim 1, characterized in that The specific implementation process of the unsupervised domain adaptation framework training in step 5 is as follows: (5.1) Using self-training method, first obtain a set of annotated images K from the source domain S , expressed as: in is the i-th image in the source domain, yes The corresponding annotations or labels, Ns is the total number of images; secondly, the sample annotations in the target domain are: in is the i-th image in the target domain, yes The corresponding annotation or label, N T is the total number of images; finally, a large amount of unlabeled data is obtained from the target domain, namely: in is the i-th unlabeled image in the target domain, N M is the total number of unlabeled target domain images; (5.2) First, perform the student model g on the source domain θ The model uses the cross entropy loss function to measure the difference between the model prediction and the true annotation, which is expressed as follows: in, represents the loss function on the source domain, θ is the model parameter, represents the true label of the i-th source image and model predictions The cross entropy between (5.3) In the self-training process, the dual teacher networks hφ1 and hφ2 are trained alternately to generate pseudo labels for the target domain data. The formula is: in, Represents the pseudo label probability of the jth pixel corresponding to the i-th target domain image in category c. This probability is based on Teacher Network The prediction of the jth pixel of the image over all possible categories c′ is determined by selecting the category c with the maximum probability as the pseudo label of the pixel; (5.4) Construct a complete RetNet-based unsupervised domain-adaptive remote sensing semantic segmentation model, use the labeled data in the source domain to train the segmentation model, use the TokenMix method to perform data mixing between the source domain and the newly introduced intermediate domain, and use the same TokenMix method to perform mixed training on the source domain and target domain data. The model adopts a dual-teacher network strategy, and the two teacher networks are trained alternately with the segmentation network. The prediction results of the teacher network are used to generate pseudo labels for the target domain data, and the segmentation network is learned based on these pseudo labels. The exponential moving average (EMA) technique is used to smooth the model's prediction results.
Citation Information
Cited By
Unmanned aerial vehicle cluster collaborative fuming scheduling method oriented to maximized smoke coverage rate
CN122111084A
Unmanned aerial vehicle cluster cooperative smoke dispatching method for maximizing smoke coverage
CN122111084B