Unsupervised domain adaptation method for semantic segmentation of high-resolution remote sensing image
By introducing a bidirectional recursive attention pyramid and multi-discriminator module in the unsupervised domain adaptation method, the fine-grained feature interaction between the source domain and the target domain in high-resolution remote sensing images is achieved, and the problem of insufficient domain offset and feature alignment is solved, and the segmentation performance of the model is improved.
Patent Information
- Application Number
- CN202510233806.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-28
AI Technical Summary
Existing unsupervised domain adaptation methods face domain offset problems and insufficient feature alignment granularity when processing high-resolution remote sensing images, resulting in a degradation in the performance of the model on the target domain.
The unsupervised domain adaptation method based on the bidirectional recursive attention pyramid (BRAP) and multi-discriminator module (MDM) is adopted to achieve fine-grained feature interaction between the source domain and the target domain through multi-scale pyramid decomposition, recursive bidirectional attention bridge and dynamic gated fusion device, and the multi-discriminator module is used to improve the discriminant ability of features.
The domain offset problem is effectively solved, the semantic segmentation performance of the model in the target domain is improved, and the adaptability and segmentation accuracy of complex scenarios are enhanced.
Smart Images

Figure CN120070896A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of computer vision and remote sensing image processing, and particularly relates to an unsupervised domain adaptation method based on a bidirectional recursive attention pyramid and a multi-discriminator module for semantic segmentation of high-resolution remote sensing images. Background Art
[0002] With the rapid development of unmanned aerial vehicle and remote sensing technologies, high-resolution remote sensing images have been widely used in fields such as infrastructure planning, land management, and urban change detection. Semantic segmentation, as a key technology in remote sensing image processing, aims to divide each pixel in the image into predefined categories. However, due to the complexity and diversity of remote sensing images, traditional supervised learning methods require a large amount of labeled data and are difficult to handle the differences between different domains in practical applications.
[0003] Unsupervised domain adaptation - hereinafter referred to as the UDA method reduces the dependence on labeled data in the target domain by using the labeled data in the source domain to adapt to the unlabeled data in the target domain. However, the existing UDA methods face the following challenges when processing remote sensing images:
[0004] Domain shift problem: The feature distribution difference between the source domain and the target domain is large, resulting in a decline in the performance of the model in the target domain.
[0005] Insufficient granularity of feature alignment: Existing methods often ignore the fine-grained semantic information during feature alignment, resulting in insufficient adaptability of the model to complex scenes.
[0006] To solve the above problems, the present invention proposes an unsupervised domain adaptation method based on a bidirectional recursive attention pyramid and a multi-discriminator module, which realizes fine-grained feature interaction between the source domain and the target domain through multi-scale pyramid decomposition, recursive bidirectional attention bridge, and dynamic gating fusion, and further improves the discriminative ability of features by using the multi-discriminator module, thereby effectively solving the domain shift problem.
[0007] Compared with natural images, remote sensing images exhibit greater cross-domain differences in terms of color, texture, spatial resolution, and scene content. These differences are attributed to the following two facts: First, the spectral heterogeneity of remote sensing images is higher than that of natural images acquired in the RGB band. In addition, remote sensing images have richer structural information, and the target scale changes more significantly. Therefore, when the performance of these UDA methods is not satisfactory, directly applying them to large variations, such as the camera angle, image resolution, especially the number and variety of remote sensing sensors have increased significantly in the past decade, and the availability of higher-quality VHR data has further expanded the differences between different domains. Recently, some DL-based UDA methods have been used for remote sensing tasks. For example, GANAI is the first GAN-based UDA algorithm designed for semantic segmentation of aerial images, while TriADA introduced a class-aware self-training technique to enhance its recognizability. WTIC further studied the invariant semantic features in VHR remote sensing images and proposed a dynamic training strategy by incorporating multiple weak supervision constraints. Recently, subspace alignment equipped with a CNN-based shallow feature extractor was designed to explore and align rich semantic vector subspaces, where complex image content can be further represented and analyzed. In particular, CCAGAN adopts global and class-level adversarial losses to enhance local semantic consistency by employing negative transfer during domain adaptation and a class attention module for processing deep features. Although these pioneering works perform well in cross-domain semantic segmentation, they still face two main problems when dealing with remote sensing images. First, existing methods are insufficient in bidirectional information interaction between domains and fail to fully exploit the potential associations between the source domain and the target domain, resulting in limited feature alignment effects. This is mainly because the complexity of remote sensing images, including their spectral heterogeneity, rich structural information, and significant target scale changes, makes it difficult for simple feature alignment strategies to capture the complex semantic relationships between the source domain and the target domain. Second, the granularity of feature alignment is relatively coarse and cannot accurately capture fine-grained semantic information, thus affecting the model's adaptability to complex scenes and segmentation accuracy. Summary of the Invention
[0008] An object of the present invention is to provide an unsupervised domain adaptation method based on a bidirectional recursive attention pyramid - hereinafter denoted as BRAP and a multi-discriminator module - hereinafter denoted as MDM for semantic segmentation of high-resolution remote sensing images. This method realizes fine-grained feature interaction between the source domain and the target domain through the BRAP module and further enhances the discriminative ability of features by using the MDM, thereby effectively solving the domain shift problem.
[0009] To achieve the above object, an embodiment of the present invention provides an unsupervised domain adaptation method based on a bidirectional recursive attention pyramid and a multi-discriminator module, including the following steps:
[0010] S1: Obtain the labeled source domain dataset and the unlabeled target domain dataset;
[0011] S3: Extract multi-level features from the source domain and target domain data through a feature extractor;
[0012] S3: Input the extracted features into a bidirectional recursive attention pyramid module, and complete dual-domain feature interaction through multi-scale pyramid decomposition, recursive bidirectional attention bridge, and dynamic gating fuser;
[0013] The bidirectional recursive attention pyramid module includes multi-scale pyramid decomposition, recursive bidirectional attention bridge, and dynamic gating fuser. Among them, multi-scale pyramid decomposition decomposes the feature map into multi-scale feature maps with sizes of 1×1, 3×3, 6×6, and the original size at each level. The recursive bidirectional attention bridge realizes cross-domain closed-loop interaction through forward and backward attention transfer between levels. The dynamic gating fuser aggregates features of each scale through dynamic gating weights;
[0014] S4: Discriminate the interacted features through a multi-discriminator module;
[0015] The multi-discriminator module includes a feature discriminator, an adversarial discriminator, and a class discriminator. Among them, the feature discriminator is used to confuse dual-domain features, the adversarial discriminator is used to realize adversarial learning of overall features, and the class discriminator completes feature confrontation from the class level to further improve the discrimination ability of features;
[0016] S5: Perform backpropagation calculation through the defined segmentation loss, global feature difference loss, domain confusion loss, and adversarial loss;
[0017] S6: Deploy the trained model to the target domain to achieve semantic segmentation of high-resolution remote sensing images in the target domain.
[0018] Furthermore, the bidirectional recursive attention pyramid module consists of a multi-scale pyramid decomposition unit, a recursive bidirectional attention interaction unit, and a dynamic gating fusion unit. Among them, the multi-scale pyramid decomposition unit is used to decompose the source domain and target domain features into hierarchical representations of different granularities. The recursive bidirectional attention interaction unit realizes deep coupling of dual-domain semantics and details through cross-level closed-loop transfer. The dynamic gating fusion unit aggregates the aligned features of each level through adaptive weights to output a unified cross-domain adaptation result.
[0019] Furthermore, the multi-scale pyramid decomposition unit decomposes the input feature map into four layers of feature representations with different resolutions through a spatial pyramid pooling strategy, including a 1×1 global semantic descriptor, 3×3 medium local structure features, 6×6 fine-grained detail features, and high-resolution features of the original size, covering the scale differences from the macroscopic scene to the microscopic target in the remote sensing image, while retaining the original resolution features to avoid loss of small target information.
[0020] Furthermore, the recursive bidirectional attention interaction unit realizes cross-level closed-loop interaction in three iterations, specifically including two paths: forward transmission and backward transmission. In forward transmission, the query and key of the low-level features in the source domain act on the value matrix of the high-level features in the target domain, injecting global semantics into the detailed features of the target domain. In backward transmission, the query and key of the high-level features in the target domain act on the value matrix of the low-level features in the source domain, feeding back the adapted local information to the global features in the source domain. During this process, the value matrix in the target domain and the query matrix in the source domain are mixed through cross-domain residual connections to retain cross-domain commonalities and differences, ensuring the stability of feature alignment.
[0021] Furthermore, after upsampling the aligned features at each level, the dynamic gating fusion unit dynamically assigns weights based on feature similarity, preferentially retaining the hierarchical features with high alignment quality, and finally generating the adapted high-resolution features through weighted fusion. This unit generates anchor features through global average pooling as the benchmark for weight calculation, and adaptively adjusts the contributions of each level in combination with the Softmax function, thereby balancing the importance differences of multi-scale features in cross-domain adaptation and enhancing the robustness of the model to complex scenarios.
[0022] Furthermore, the global feature difference loss Lglobal is obtained by calculating the difference between the high-level aligned feature maps of the source domain and the target domain; the domain confusion loss Ldc confuses the domain feature distributions of the two domains by using the true labels of the source domain as the true labels of the target domain, and calculates the difference between the prediction results of the target domain and the true labels of the source domain; the segmentation loss Lseg is calculated by obtaining the difference between the source domain features and the true boundaries; the class confusion loss Lmc compares the mean differences of the class features in the two domains in the high-dimensional reproducing kernel Hilbert space.
[0023] Furthermore, the feature adversarial loss Ladv realizes overall feature confrontation by training the discriminator's ability to distinguish whether the features come from the source domain or the target domain; the class adversarial loss Lamc realizes fine-grained feature alignment by judging the feature source at the class level.
[0024] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0025] The present invention solves the deficiencies of existing unsupervised domain adaptation methods in remote sensing image semantic segmentation through the bidirectional recursive attention pyramid interaction module and the multi-discriminator module. BRAP captures the global-local dependence relationship between the source domain and the target domain through multi-scale pyramid bidirectional closed-loop interaction between the two domains, realizing fine-grained feature interaction. MDM further improves the discriminative ability of features by discriminating features from different perspectives through multiple discriminators, thereby effectively solving the domain shift problem and improving the segmentation performance of the model in the target domain. Brief Description of the Drawings
[0026] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments:
[0027] Figure 1 : It is a framework flowchart of an unsupervised domain adaptation method for high-resolution remote sensing image semantic segmentation provided by the first embodiment of the present invention;
[0028] Figure 2 : It is a schematic structural diagram of a bidirectional recursive attention pyramid module provided by the first embodiment of the present invention;
[0029] Figure 3 : It is a schematic structural diagram of a multi-discriminator module provided by the first embodiment of the present invention;
[0030] Figure 4 : It is a flowchart of the steps of an unsupervised domain adaptation method for high-resolution remote sensing image semantic segmentation provided by the first embodiment of the present invention; Specific Embodiments
[0031] The specific embodiments of the present invention are as follows, which describe in detail the whole process from data acquisition to model deployment, ensuring that every step and technical detail are fully explained.
[0032] As Figure 1 shown, a framework flowchart of an unsupervised domain adaptation method for high-resolution remote sensing image semantic segmentation provided by the first embodiment of the present invention includes:
[0033] S1: Obtain the dataset
[0034] Obtain the labeled source domain dataset Ds = {(xs, ys)} and the unlabeled target domain dataset Dt = {xt}, where xs and ys are the samples and their corresponding labels in the source domain respectively, and xt is the sample in the target domain;
[0035] S2: Feature extraction
[0036] Use the DeepLab-v2 framework based on ResNet-101 as the feature extractor, pre-trained on the ImageNet dataset. The feature extractor extracts multi-level features from the source domain and target domain data, and outputs high-level feature maps Fs and Ft with abstract semantic information, with a size of 8 / H × 8 / W × D, where D is the feature dimension, D = 2048, and H and W are the height and width of the image respectively;
[0037] S3: Feature interaction
[0038] Input the extracted features Fs and Ft into the bidirectional recursive attention pyramid module, and realize fine-grained feature interaction through multi-scale pyramid decomposition, recursive bidirectional attention bridge and dynamic gating fuser. Specifically as follows:
[0039] Multi-scale Pyramid Decomposition: The feature maps of the source domain and the target domain are decomposed into four feature maps with different resolutions through spatial pyramid pooling. The sizes of the feature maps at each level are 1×1, 3×3, 6×6, and the original size. Layer L1: Extract the global semantic descriptor through global average pooling; Layer L2: Divide it into 3×3 sub-regions, and perform average pooling on each region to capture medium local structures; Layer L3: Divide it into 6×6 sub-regions, and retain fine-grained local details after pooling; Layer L4: Directly retain the high-resolution feature map to avoid information loss caused by downsampling.
[0040] Recursive Bidirectional Attention Bridge: The recursive design breaks through the traditional unidirectional attention limit, enabling the deep coupling of dual-domain features at the spatial and semantic levels. Each recursive iteration contains two paths, forward and backward, and a total of 3 iterations are performed to gradually deepen the cross-domain interaction. In the three iterations, the feature of the i-th layer in the source domain acts on the feature of the i+1-th layer in the target domain through the attention mechanism, and the feature of the i-th layer in the target domain is fed back to the i-1-th layer in the source domain through reverse attention, forming a cross-level closed-loop interaction;
[0041] Forward Pass: The query matrix Qs (i) and the key matrix Ks (i) of the i-th layer in the source domain are cross-domain residual connection matrix Vmixedt (i+1) with the value of the i+1-th layer in the target domain for attention calculation. The formula is:
[0042]
[0043] Backward Pass: The query matrix Qt (i) and the key matrix Kt (i) of the i-th layer in the target domain are cross-domain residual connection matrix Vmixeds (i-1) with the value of the i-1-th layer in the source domain for attention calculation. The formula is:
[0044]
[0045] The recursive bidirectional attention bridge introduces cross-domain residual connections in each iteration. During the forward propagation process, the value matrix Vt in the target domain and the query matrix Qs in the source domain are mixed through skip connections:
[0046] V mixedt = V t + γMLP(Q S )
[0047] During the backward propagation process, the value matrix Vs in the source domain and the query matrix Qt in the target domain are mixed through skip connections:
[0048] V mixeds = V s + γMLP(Q t )
[0049] where γ is a learnable scaling factor and MLP is used for dimension matching.
[0050] The dynamic gating fuser upsamples the features of each layer to the original size at the top layer of the pyramid and performs weighted summation through the similarity weight β k The formula is:
[0051]
[0052] where Fanchor is the anchored feature after global average pooling, which is used to measure the importance of the features at each level. The final output is:
[0053]
[0054] As Figure 2 shown, a bidirectional recursive attention pyramid module is provided in the present invention for interacting with local information and global information in the dual domain. Pyramid decomposition effectively addresses the problem of large object scale differences in high-resolution remote sensing images, while the recursive mechanism enhances cross-level information flow; the recursive design breaks through the traditional one-way attention limitation, enabling deep coupling of dual-domain features at the spatial and semantic levels; the gating mechanism avoids manually setting hierarchical weights and improves the model's adaptability to complex scenarios.
[0055] S4: Discriminate the aligned features through a multi-discriminator module;
[0056] The multi-discriminator module includes a feature discriminator, an adversarial discriminator, and a class discriminator. The feature discriminator is used to process high-level aligned features, the adversarial discriminator is used to implement adversarial learning of the overall features, and the class discriminator completes feature adversarial learning from the class level to further improve the discriminative ability of the features;
[0057] S5: Perform backpropagation calculation through the defined segmentation loss, global feature difference loss, and adversarial loss;
[0058] As Figure 3 、 Figure 4 shown, a multi-discriminator module is provided in the present invention, including a feature discriminator, an adversarial discriminator, and a class discriminator. The feature discriminator and the adversarial discriminator are composed of 4 self-attention modules to enhance the discriminator's attention to semantic information; the class discriminator is composed of a decoder and an encoder. The label information from the source domain is accessed to the encoded features during the encoding stage and passes through a self-attention module to achieve adversarial learning from the class level.
[0059] Global feature difference loss Lglobal: This loss function is used to align the high-level feature maps of the source domain and the target domain and reduce the distribution difference between the two domains. The maximum mean difference is used to calculate this loss:
[0060]
[0061] and are the high - level feature maps of the source domain and the target domain processed by the bidirectional recursive attention pyramid module respectively. P represents the average pooling operation on the feature map of each channel to extract spatial information. λg is a weight coefficient used to balance the contribution of this loss in the overall optimization.
[0062] Boundary loss Lseg: This loss function is used to ensure that the generator can accurately perform semantic segmentation on the source domain images. It is calculated based on the cross - entropy loss, and the specific formula is:
[0063]
[0064] Ps and are the predicted probability maps of the classifier for the source domain images. Ys is the ground - truth label of the source domain images.
[0065] Domain confusion loss Ldc: This loss function is used to deceive the feature discriminator Df into misclassifying the feature map of the target domain as that of the source domain. The specific formula is:
[0066]
[0067] is the prediction result of the feature discriminator for the feature map of the target domain. λf is the weight coefficient.
[0068] Class - mixing loss Lmc: Comparing the mean differences of the dual - domain class features in the high - dimensional reproducing kernel Hilbert space, the loss function uses the maximum mean discrepancy, and the specific formula is:
[0069]
[0070] φ(·): The kernel function (such as Gaussian kernel, polynomial kernel) maps to the high - dimensional space. The Hilbert space corresponding to the kernel function.
[0071] These four loss functions are designed to train the generator. At this time, the discriminator parameters are fixed, and the discriminator parameters will not be updated during the training of the generator.
[0072] Feature adversarial loss Ladv: This loss function is used to train the feature discriminator Df to enable it to distinguish the feature maps of the source domain and the target domain. The specific formula is as follows:
[0073]
[0074] and are the prediction results of the feature discriminator for the feature maps of the source domain and the target domain respectively.
[0075] Category adversarial loss Lamc: The prediction result incorporates the true class information from the source domain, and fine-grained feature alignment is achieved by judging the feature source at the class level. The specific formula is as follows:
[0076]
[0077] and are the prediction results of the feature discriminator for each class of the source domain and target domain feature maps respectively. λk is the weight parameter of the proportion of the loss calculation result of each class in the entire loss function.
[0078] These two loss functions are designed to train the discriminator, improve the discrimination performance of the discriminator module, and enhance the effect of adversarial learning. During the training of the discriminator, the parameters of the generation part are frozen and not updated. The generator and the discriminator are alternately trained during the training phase.
[0079] S6: Model deployment
[0080] Deploy the trained model to the target domain to achieve semantic segmentation of high-resolution remote sensing images in the target domain.
[0081] The above are only the preferred embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
Claims
1. An unsupervised domain adaptation method for semantic segmentation of high-resolution remote sensing images, characterized in that: At least the following steps are included: S1: Obtain a labeled source domain dataset and an unlabeled target domain dataset; S2: Extract multi-level features from source and target domain data through feature extractor; S3: The extracted features are input into the bidirectional recursive attention pyramid module, and the dual-domain feature interaction is completed through multi-scale pyramid decomposition, recursive bidirectional attention bridge and dynamic gated fuser; The bidirectional recursive attention pyramid module includes a multi-scale pyramid decomposition, a recursive bidirectional attention bridge, and a dynamic gated fuser, wherein the multi-scale pyramid decomposition decomposes the feature map into a multi-scale feature map, and the sizes of each level are 1×1, 3×3, 6×6, and the original size. The recursive bidirectional attention bridge realizes cross-domain closed-loop interaction through forward and reverse attention transfer between levels, and the dynamic gated fuser aggregates features of each scale through dynamic gated weights; S4: discriminate the interacted features through a multi-discriminator module; The multi-discriminator module includes a feature discriminator, an adversarial discriminator and a category discriminator, wherein the feature discriminator is used to confuse dual-domain features, the adversarial discriminator is used to achieve adversarial learning of overall features, and the category discriminator completes feature adversarial learning from the category level to further improve the feature discrimination ability; S5: Back propagation calculation is performed through the defined segmentation loss, global feature difference loss, domain confusion loss and adversarial loss; S6: Deploy the trained model to the target domain to achieve semantic segmentation of high-resolution remote sensing images in the target domain.
2. The unsupervised domain adaptation method for semantic segmentation of high-resolution remote sensing images according to claim 1, characterized in that: The bidirectional recursive attention pyramid module achieves cross-domain interaction through the following steps: Multi-scale pyramid decomposition: The dual-domain features are decomposed into four layers of feature maps with different resolutions through spatial pyramid pooling, with each level having a size of 1×1, 3×3, 6×6, and the original size; Recursive bidirectional attention bridge: In three iterations, the i-th layer features of the source domain act on the i+1-th layer features of the target domain through the attention mechanism, and the i-th layer features of the target domain are fed back to the i-1-th layer features of the source domain through reverse attention, forming a cross-level closed-loop interaction; Dynamic gated fuser: After upsampling the features of each layer at the top layer of the pyramid, the output is adaptively aggregated through similarity weights.
3. The unsupervised domain adaptation method for semantic segmentation of high-resolution remote sensing images according to claim 1, characterized in that: The category discriminator first decodes the features to the original input image size, and then inputs them into the encoder together with the category information of the source domain. After the category information undergoes multi-level feature extraction, the high-level semantic information obtained is fused with the features to complete the category information access.
4. The unsupervised domain adaptation method for semantic segmentation of high-resolution remote sensing images according to claim 2, characterized in that: The specific calculation method of the recursive bidirectional attention bridge includes: Forward pass: query matrix Qs of the i-th layer in the source domain (i) and bond matrix Ks (i) The cross-domain residual connection matrix Vmixedt with the value of the target domain i+1 layer (i+1) To calculate attention, the formula is: Backward pass: query matrix Qt of the target domain i-th layer (i) and key matrix Kt (i) The cross-domain residual connection matrix Vmixeds with the value of the target domain i-1 layer (i-1) To calculate attention, the formula is:
5. The unsupervised domain adaptation method for semantic segmentation of high-resolution remote sensing images according to claim 2, characterized in that: The dynamic gated fusion upsamples the features of each layer to the original size at the top layer of the pyramid and uses the similarity weight β k Weighted summation, the formula is: Among them, Fanchor is the anchor feature after global average pooling, which is used to measure the importance of features at each level.
6. The unsupervised domain adaptation method for semantic segmentation of high-resolution remote sensing images according to claim 4, characterized in that: The recursive bidirectional attention bridge introduces a cross-domain residual connection in each iteration, and mixes the value matrix Vt of the target domain with the query matrix Qs of the source domain through a skip connection during the forward propagation process: V mixedt =V t +γMLP(Q S ) In the reverse propagation process, the value matrix Vs of the source domain is mixed with the query matrix Qt of the target domain through skip connections: V mixeds =V s +γMLP(Q t ) Where γ is a learnable scaling factor and MLP is used for dimensionality matching.
7. The unsupervised domain adaptation method for semantic segmentation of high-resolution remote sensing images according to claim 3, characterized in that: The feature discriminator in the multi-discriminator module confuses the dual-domain feature distribution by taking the source domain truth label as the target domain truth label; The adversarial discriminator trains the discriminator by discriminating dual-domain features, enhancing the discriminant ability and improving the model's adversarial performance.
Citation Information
Patent Citations
Unsupervised remote sensing image semantic segmentation method based on super-resolution and domain self-adaption
CN113160234A
Feature adaptive alignment unsupervised domain adaptive remote sensing image semantic segmentation method
CN113378906A
PCR liquid drop image detection technology system and use method thereof
CN113469103A
Unsupervised remote sensing image semantic segmentation based on spatial resolution domain adaptation
CN113850813A
Unsupervised domain adaptive semantic segmentation method and system
CN115631337A