Unsupervised Domain Adaptation Remote Sensing Image Semantic Segmentation Method Based on Parallel Network Structure
The local and global features of remote sensing images are extracted by combining CNN and Transformer through a parallel network structure, and feature weighting is used to minimize distribution differences, which solves the lack of feature alignment and global modeling in semantic segmentation of remote sensing images, achieving high-precision semantic segmentation effect.
Patent Information
- Application Number
- CN202510414052.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-04-03
AI Technical Summary
The existing unsupervised domain adaptation methods have problems such as insufficient feature alignment, global modeling and incomplete local detail expression in semantic segmentation of remote sensing images, resulting in poor segmentation effect.
The parallel network structure is adopted, and local and global features are extracted in combination with CNN and Transformer branches, and feature weighting is performed through the channel attention mechanism to minimize distribution differences and achieve feature alignment between the source domain and the target domain.
It significantly improves the feature expression ability of remote sensing images, improves segmentation accuracy and model adaptability in the target domain, and achieves high-precision semantic segmentation.
Smart Images

Figure CN119942127B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image segmentation, and in particular, to an unsupervised domain adaptation remote sensing image semantic segmentation method based on a parallel network structure. Background Art
[0002] Remote sensing image semantic segmentation is a core task in the field of image analysis, aiming to assign each pixel in a remote sensing image to a specific land cover class, such as buildings, water bodies, and vegetation. This technology has important application values in fields such as land use classification, disaster monitoring, urban planning, and ecological environment assessment.
[0003] However, remote sensing image semantic segmentation faces complex technical challenges. Firstly, remote sensing images usually have complex scene structures and multi-scale characteristics, and there are significant morphological differences among different land cover classes. For example, buildings usually have obvious geometric features, while vegetation may have irregular shapes. In addition, the annotation cost of high-resolution remote sensing images is extremely high. To annotate remote sensing images, not only accurate spatial resolution is required, but also the professional knowledge of domain experts. This annotation process consumes a large amount of manpower and time and is easily affected by human subjectivity. Secondly, the differences in imaging conditions, sensor types, time seasons, etc. between different datasets result in the difficulty of directly applying a model trained on one dataset to another dataset. These phenomena greatly limit the generality of remote sensing image semantic segmentation technology.
[0004] To address these problems, unsupervised domain adaptation technology has gradually become a research hotspot. The core idea of this technology is to reduce the distribution difference between the source domain (labeled data) and the target domain (unlabeled data) so that the model can achieve good segmentation performance on the target domain without the support of labeled data for the target domain. This method can effectively reduce the dependence on labeled data and also provides a solution to the problem of domain distribution difference in remote sensing semantic segmentation. Current research mainly focuses on using generative adversarial networks to achieve the alignment of inter-domain features. By designing a generator and a discriminator, the distributions of the source domain and the target domain in the feature space gradually tend to be consistent. In recent years, self-training methods have also gradually become a research hotspot. By generating pseudo-labels for the target domain and performing iterative optimization, the adaptability of the model to the target domain is enhanced, effectively alleviating the problem of insufficient labeled data.
[0005] However, the current unsupervised domain adaptation methods still have some limitations. First of all, most domain adaptation methods based on generative adversarial networks rely too much on the adversarial training of the generator and discriminator. However, due to the complex scenes and blurred class boundaries of remote sensing images, the generated features often lack the ability to express details, making it difficult to accurately align the distributions of the source domain and the target domain. Secondly, although the self-training method can iteratively optimize the segmentation performance of the target domain through pseudo-labels, the quality of the pseudo-labels themselves highly depends on the accuracy of the initial model. When the data distribution of the target domain is complex, it is easy to introduce noise, resulting in a decline in the stability and performance of model training. In addition, in the process of feature alignment, most current semantic segmentation methods often ignore the synergistic effect between local details and global semantics, and only rely on a certain type of network (such as CNN or Transformer) alone for feature modeling. This limitation may lead to incomplete feature expression, making it difficult to comprehensively capture the spatial and semantic information of remote sensing images. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to solve the deficiencies of the existing unsupervised domain adaptation methods in feature alignment, global modeling, and local detail expression, and to achieve high-precision semantic segmentation. To overcome the defects of the above-mentioned prior art (or related technologies), the present invention provides an unsupervised domain adaptation remote sensing image semantic segmentation method based on a parallel network structure.
[0007] The present invention provides an unsupervised domain adaptation remote sensing image semantic segmentation method based on a parallel network structure, including:
[0008] Step S1, obtaining a labeled source domain dataset and an unlabeled target domain dataset;
[0009] Step S2, constructing a parallel network model;
[0010] Step S3, extracting source domain features from the source domain dataset through the parallel network model and generating a first prediction result, and constructing a source domain cross-entropy loss function based on the first prediction result and the source domain dataset;
[0011] Step S4, predicting the target domain dataset through the parallel network model to generate pseudo-labels to obtain an enhanced target domain dataset, and extracting target domain features to generate a second prediction result, and constructing a target domain cross-entropy loss function based on the second prediction result and the enhanced target domain dataset;
[0012] Step S5, constructing a channel attention weighting module to weight the source domain features and the target domain features to obtain a shared attention weight, and using the shared attention weight to perform channel-by-channel weighting on the source domain features and the target domain features respectively to generate aligned source domain features and aligned target domain features;
[0013] Step S6: Obtain the minimized distribution difference based on the aligned source domain features and the aligned target domain features, and obtain the total loss function based on the distribution difference, the source domain cross-entropy loss function, and the target domain cross-entropy loss function to optimize the model;
[0014] Step S7: Test the parallel network model to obtain the mean intersection over union as the semantic segmentation evaluation metric.
[0015] Compared with the prior art, the unsupervised domain adaptation remote sensing image semantic segmentation method based on the parallel network structure of the present invention has the following advantages:
[0016] In the present invention, the parallel network model is used to simultaneously obtain global features and local features, effectively combining the local detail capture ability and the global semantic modeling ability, significantly improving the feature expression ability of the model for complex remote sensing images. At the same time, the channel attention mechanism is introduced to dynamically weight the features, highlighting the key features shared between domains and suppressing the interference information unique to the domain; by minimizing the distribution difference, the feature distributions of the source domain and the target domain are aligned, thereby enhancing the segmentation effect, improving the segmentation accuracy and adaptation ability of the parallel network model in the target domain, and solving the deficiencies of the existing unsupervised domain adaptation methods in feature alignment, global modeling, and local detail expression, and achieving high-precision semantic segmentation.
[0017] In a possible implementation manner, the parallel network model includes a parallel encoder for feature extraction and a decoder for result prediction. The parallel encoder includes a CNN branch encoder for extracting local features and a Transformer branch encoder for extracting global features. In step S2, the local features and the global features are obtained through the following calculation formula:
[0018] ;
[0019] where
[0020] represents the local features;
[0021] represents the CNN branch encoder;
[0022] represents the input image of the parallel encoder;
[0023] represents the global features;
[0024] represents the Transformer branch encoder.
[0025] In a possible implementation manner, the source domain data set includes multiple source domain images with labels. In step S3, the source domain cross-entropy loss function is constructed through the following calculation formula:
[0026] ;
[0027] represents the source domain cross-entropy loss function;
[0028] represents the label of each source domain image;
[0029] represents the first prediction result;
[0030] represents the length of the source domain image;
[0031] represents the width of the source domain image;
[0032] represents the number of categories of each source domain image.
[0033] In a possible implementation manner, the target domain data set includes multiple unlabeled target domain images. In step S4, the target domain cross-entropy loss function is constructed through the following calculation formula:
[0034] ;
[0035] ;
[0036] where,
[0037] represents the target domain cross-entropy loss function;
[0038] represents the pseudo-label;
[0039] represents the pixel prediction confidence of the th row and th column in the target domain image, , ;
[0040] represents the second prediction result;
[0041] represents the third prediction result generated by the parallel network model for predicting the target domain data set;
[0042] Represents the length of the target domain image;
[0043] Represents the width of the target domain image;
[0044] Represents the number of categories of each of the target domain images.
[0045] In a possible implementation manner, in the step S5, the shared attention weight is obtained through the following calculation formula:
[0046] ;
[0047] ;
[0048] Wherein,
[0049] Represents the global average pooling result of the source domain features;
[0050] Represents the height of the feature map;
[0051] Represents the width of the feature map;
[0052] Represents the feature value of the c-th channel of the source domain feature map at the position (i, j);
[0053] Represents the global average pooling result of the target domain features;
[0054] Represents the feature value of the c-th channel of the target domain feature map at the position (i, j);
[0055] Represents the shared attention weight;
[0056] Represents the Sigmoid activation function;
[0057] And Represents the fully connected layer weight matrix;
[0058] Represents the ReLU activation function.
[0059] In a possible implementation manner, in the step S5, the aligned source domain features and the aligned target domain features are obtained through the following calculation formula:
[0060] ;
[0061] Wherein,
[0062] represent the aligned source domain features;
[0063] represent the source domain features;
[0064] represent the shared attention weights;
[0065] represent the aligned target domain features;
[0066] represent the target domain features.
[0067] In a possible implementation, in step S6, the distribution difference is obtained through the following calculation formula:
[0068] ;
[0069] where
[0070] represents the distribution difference;
[0071] represents the similarity of samples within the aligned source domain features;
[0072] represents the similarity of samples within the aligned target domain features;
[0073] represents the Gaussian kernel function;
[0074] represents the similarity between the aligned source domain features and the aligned target domain features.
[0075] In a possible implementation, in step S6, the total loss function is obtained through the following calculation formula:
[0076] ;
[0077] where
[0078] represents the total loss function;
[0079] represents the source domain cross-entropy loss function;
[0080] represents the target domain cross-entropy loss function;
[0081] represents a preset parameter;
[0082] represents the said distribution difference.
[0083] In a possible implementation manner, in the said step S7, the average intersection over union is obtained through the following calculation formula:
[0084] ;
[0085] wherein,
[0086] represents the intersection over union;
[0087] represents the number of positive examples correctly predicted during the test process of the said parallel network model;
[0088] represents the number of positive examples wrongly predicted during the test process of the said parallel network model;
[0089] represents the number of negative examples wrongly predicted during the test process of the said parallel network model;
[0090] represents the said average intersection over union;
[0091] represents the th category of the said intersection over union, . BRIEF DESCRIPTION OF THE DRAWINGS
[0092] Figure 1 is the step flow chart of the present invention;
[0093] Figure 2 is the schematic diagram of the semantic segmentation framework of the present invention;
[0094] Figure 3 is the schematic diagram of the specific design of the parallel encoder of the present invention;
[0095] Figure 4 is the schematic diagram of the specific design of the feature fusion module of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0096] First of all, those skilled in the art should understand that these embodiments are only used to explain the technical principles of the embodiments of the present invention, and are not intended to limit the protection scope of the embodiments of the present invention. Those skilled in the art can make adjustments according to needs to adapt to specific application scenarios.
[0097] The present invention will be further described in detail below with reference to the drawings and specific embodiments.
[0098] SeeFigures 1-4 , an embodiment of the present invention discloses an unsupervised domain adaptation remote sensing image semantic segmentation method based on a parallel network structure, including:
[0099] Step S1, collect a labeled source domain dataset and an unlabeled target domain dataset , where represents the source domain image in the source domain dataset, is the label corresponding to the source domain image , represents the target domain image in the target domain dataset, and the source domain dataset and the target domain dataset share the same label space;
[0100] Step S2, construct a parallel network model M, including a parallel encoder E and a decoder D. The parallel encoder E is composed of a CNN branch encoder and a Transformer branch encoder. The CNN branch encoder is used to extract local features , and the Transformer branch encoder is used to extract global features , and the decoder D is used to receive the features output by the parallel encoder E and generate a prediction result;
[0101] Step S3, train the parallel network model M on the source domain dataset , extract the source domain features through the parallel encoder E , and generate a prediction result through the decoder D , construct a source domain cross-entropy loss function , and optimize the parallel network model M in a supervised manner; the source domain cross-entropy loss function is:
[0102] ;
[0103] In the formula, is the prediction result, is the label of the source domain image; h and w are the length and width of the source domain image respectively, and c is the number of classes of the source domain image;
[0104] Step S4, use the parallel network model M to predict the target domain dataset and generate pseudo-labels :
[0105] ;
[0106] where is the pixel prediction confidence of the -th row and -th column in the target domain image, , ; then for the target domain image set after data augmentation Extract target domain features , and generate prediction results through decoder D , calculate the cross-entropy loss function of the target domain , and optimize the prediction ability of the target domain in an unsupervised manner. The cross-entropy loss function of the target domain is as follows:
[0107] ;
[0108] Step S5, construct a channel attention weighting module to weight the source domain features and target domain features , and obtain the shared attention weight through global average pooling and a fully connected network :
[0109] ;
[0110] ;
[0111] Among them, and are the weight matrices of the fully connected layers, represents the ReLU activation function, represents the Sigmoid activation function;
[0112] Utilize this shared attention weight to weight the source domain features and target domain features channel by channel respectively, and generate the aligned features and :
[0113] ;
[0114] Step S6, achieve the alignment of the weighted source domain features and weighted target domain features from a statistical perspective by minimizing the distribution difference :
[0115] ;
[0116] Among them, represents the Gaussian kernel function, which is used to measure the similarity of features; represents the similarity of samples within the aligned source domain features, represents the similarity of samples within the aligned target domain features, represents the similarity between the aligned source domain features and the aligned target domain features;
[0117] The total loss function in this specific example is:
[0118] ;
[0119] In the formula, is a constant used to control the loss ratio; Step S7, test on the test set of the target domain dataset, and use the mean intersection over union (mIoU) as the evaluation metric to evaluate the semantic segmentation performance of the parallel network model M on the target domain. The formula for the mean intersection over union mIoU is:
[0120] ;
[0121] In the formula, n is the number of target domain label categories, represents the intersection over union of the th category, .
[0122] The source domain dataset in Step S1 and the target domain dataset contain two datasets, Potsdam and Vaihingen. Among them, the Potsdam dataset covers the urban area of Potsdam, Germany, and contains 38 high-resolution aerial images of a fixed size (6000×6000 pixels) with a resolution of 5 cm / pixel. Each image consists of four bands: red (R), green (G), blue (B), and near-infrared (IR), and provides detailed ground truth annotations, which are divided into six categories: impervious surfaces, buildings, low vegetation, trees, cars, and background; the Vaihingen dataset covers the small town of Vaihingen, Germany, and contains 33 high-resolution aerial images of different sizes with a resolution of 9 cm / pixel. The images mainly consist of three bands: near-infrared (IR), red (R), and green (G), and the ground truth annotations are the same as those of the Potsdam dataset; the images of Potsdam and Vaihingen are respectively cropped to 896×896 pixels and 512×512 pixels. Therefore, Potsdam has a total of 1764 images, and Vaihingen has a total of 1696 images; in terms of dataset division, the Potsdam dataset is divided into a training set and a test set. The training set contains 1323 images, and the test set contains 441 images; similarly, the Vaihingen dataset is also divided into a training set and a test set, where the training set has 1256 images and the test set has 440 images; in this specific embodiment, 4 source domain images and 4 target domain images are randomly selected from the dataset each time.
[0123] The CNN branch encoder in Step S2 uses ResNet as the backbone network to capture the detailed information of the image through layer-by-layer convolution operations. The Transformer branch encoder uses the MiT network to model the long-range dependencies in the image through the multi-head self-attention mechanism to obtain global semantic information. The parallel encoder E converts the input image X into a multi-scale and multi-level feature representation, where:
[0124] ;
[0125] The decoder D adopts a SegFormer decoder structure to achieve efficient feature decoding and accurate semantic segmentation.
[0126] Continue to refer to Figure 4 , after performing step S2, it further includes:
[0127] Construct a feature fusion module in the parallel encoder E to fuse the CNN features and the Transformer features to obtain the fused features , and the design of the feature fusion module is as follows: First, extract the global information of through global average pooling to generate a channel description vector; next, input this vector into two consecutive 1×1 convolutional layers for non-linear transformation and channel compression and expansion; then, generate a dynamic weight vector through the Sigmoid activation function; finally, generate the fused features:
[0128] .
[0129] In the description of the present invention, the descriptions referring to terms such as "one embodiment", "some embodiments", "in this embodiment", "specific examples", or "some examples", etc. mean that the specific features, mechanisms, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic descriptions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, mechanisms, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without conflict, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0130] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. An unsupervised domain adaptation remote sensing image semantic segmentation method based on a parallel network structure, characterized in that: The following steps are involved: Step S1, obtaining a labeled source domain dataset and an unlabeled target domain dataset; Step S2, constructing a parallel network model; Step S3, extracting source domain features from the source domain data set through the parallel network model and generating a first prediction result, and constructing a source domain cross entropy loss function based on the first prediction result and the source domain data set; Step S4, predicting the target domain data set through the parallel network model to generate pseudo labels to obtain an enhanced target domain data set, extracting target domain features to generate a second prediction result, and constructing a target domain cross entropy loss function based on the second prediction result and the enhanced target domain data set; Step S5, constructing a channel attention weighting module to perform feature weighting on the source domain features and the target domain features to obtain shared attention weights, and using the shared attention weights to perform channel-by-channel weighting on the source domain features and the target domain features to generate aligned source domain features and aligned target domain features; Step S6, obtaining a minimized distribution difference according to the aligned source domain features and the aligned target domain features, and obtaining a total loss function based on the distribution difference, the source domain cross entropy loss function and the target domain cross entropy loss function for model optimization; Step S7, testing the parallel network model to obtain an average intersection-over-union ratio as a semantic segmentation evaluation index.
2. The unsupervised domain adaptation remote sensing image semantic segmentation method based on a parallel network structure according to claim 1, characterized in that: The parallel network model includes a parallel encoder for feature extraction and a decoder for result prediction, and the parallel encoder includes a CNN branch encoder for extracting local features and a Transformer branch encoder for extracting global features. In step S2, the local features and the global features are obtained by the following calculation formula: ; in, representing the local features; represents the CNN branch encoder; represents an input image of the parallel encoder; representing the global feature; Represents the Transformer branch encoder.
3. The unsupervised domain adaptation remote sensing image semantic segmentation method based on a parallel network structure according to claim 1, characterized in that: The source domain dataset includes a plurality of source domain images with labels. In step S3, the source domain cross entropy loss function is constructed by the following calculation formula: ; represents the source domain cross entropy loss function; Representing the label of each of the source domain images; represents the first prediction result; represents the length of the source domain image; represents the width of the source domain image; Represents the number of categories of each source domain image.
4. The unsupervised domain adaptation remote sensing image semantic segmentation method based on a parallel network structure according to claim 1, characterized in that: The target domain dataset includes a plurality of unlabeled target domain images. In step S4, the target domain cross entropy loss function is constructed by the following calculation formula: ; ; in, represents the target domain cross entropy loss function; represents the pseudo label; represents the first Line The pixel prediction confidence of the column, , ; represents the second prediction result; A third prediction result representing the target domain data set generated by the parallel network model; represents the length of the target domain image; represents the width of the target domain image; Represents the number of categories of each target domain image.
5. The unsupervised domain adaptation remote sensing image semantic segmentation method based on a parallel network structure according to claim 1, characterized in that: In step S5, the shared attention weight is obtained by the following calculation formula: ; ; in, Represents the global average pooling result of source domain features; Indicates the height of the feature map; Indicates the width of the feature map; Represents the feature value of the cth channel at position (i, j) in the source domain feature map; Represents the global average pooling result of the target domain features; Represents the feature value of the cth channel at position (i, j) in the target domain feature map; represents the shared attention weight; Represents the Sigmoid activation function; and represents the weight matrix of the fully connected layer; Represents the ReLU activation function.
6. The unsupervised domain adaptation remote sensing image semantic segmentation method based on a parallel network structure according to claim 1, characterized in that: In step S5, the aligned source domain features and the aligned target domain features are obtained by the following calculation formula: ; in, represents the source domain features after alignment; represents the source domain feature; represents the shared attention weight; represents the aligned target domain features; Represents the target domain features.
7. The unsupervised domain adaptation remote sensing image semantic segmentation method based on a parallel network structure according to claim 1, characterized in that: In step S6, the distribution difference is obtained by the following calculation formula: ; in, represents said distribution difference; Represents the similarity of samples within the source domain features after the alignment; Represents the similarity of samples within the aligned target domain features; represents the Gaussian kernel function; Represents the similarity between the aligned source domain features and the aligned target domain features.
8. The unsupervised domain adaptation remote sensing image semantic segmentation method based on a parallel network structure according to claim 1, characterized in that: In step S6, the total loss function is obtained by the following calculation formula: ; in, represents the total loss function; represents the source domain cross entropy loss function; represents the target domain cross entropy loss function; Indicates preset parameters; represents the distribution difference.
9. The unsupervised domain adaptation remote sensing image semantic segmentation method based on a parallel network structure according to claim 1, characterized in that: In step S7, the average intersection-over-combination ratio is obtained by the following calculation formula: ; in, represents intersection and union ratio; represents the number of positive examples correctly predicted during the test of the parallel network model; Indicates the number of positive examples that are incorrectly predicted during the testing of the parallel network model; represents the number of negative examples that are incorrectly predicted during the testing of the parallel network model; represents the average intersection-over-union ratio; Indicates The intersection-over-union ratio of the categories is .
Citation Information
Patent Citations
Unsupervised domain adaptive remote sensing road semantic segmentation method based on GAN network
CN113888547A
Semi-supervised domain adaptive image semantic segmentation method, system and device and storage medium
CN116229080A