Parallel network structure-based unsupervised domain adaptive remote sensing image semantic segmentation method
By adopting the parallel network structure and channel attention mechanism in the semantic segmentation of remote sensing images, the shortcomings of the existing unsupervised domain adaptation methods in feature alignment and detail expression are solved, and high-precision semantic segmentation of remote sensing images are achieved.
Patent Information
- Application Number
- CN202510414052.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-03
AI Technical Summary
The existing unsupervised domain adaptation methods have shortcomings in feature alignment, global modeling and local detail expression, and it is difficult to achieve high-precision semantic segmentation of remote sensing images.
Using a method based on parallel network structure, global and local features are simultaneously extracted through parallel encoder, and a channel attention mechanism is introduced to weight the features to minimize distribution differences to achieve the alignment of feature distribution between the source domain and the target domain.
It significantly improves the model's feature expression ability on complex remote sensing images, enhances the segmentation effect, and improves the segmentation accuracy and adaptability of the parallel network model on the target domain.
Smart Images

Figure CN119942127A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image segmentation, and in particular to an unsupervised domain adaptation remote sensing image semantic segmentation method based on a parallel network structure. Background Art
[0002] Remote sensing image semantic segmentation is a core task in the field of image analysis. It aims to assign each pixel in the remote sensing image to a specific land object category, such as buildings, water bodies, and vegetation. This technology has important application value in land use classification, disaster monitoring, urban planning, and ecological environmental assessment.
[0003] However, remote sensing image semantic segmentation faces complex technical challenges. First, remote sensing images usually have complex scene structures and multi-scale characteristics, and different types of objects have significant differences in morphology. For example, buildings usually have obvious geometric features, while vegetation may have irregular shapes. In addition, the annotation cost of high-resolution remote sensing images is extremely high. Annotating remote sensing images requires not only accurate spatial resolution, but also the expertise of domain experts. This annotation process consumes a lot of manpower and time, and is easily affected by human subjectivity. Secondly, the differences in imaging conditions, sensor types, time and seasons between different data sets make it difficult for a model trained on one data set to be directly applied to another data set. These phenomena have greatly limited the versatility of remote sensing image semantic segmentation technology.
[0004] In response to these problems, unsupervised domain adaptation technology has gradually become a research hotspot. The core idea of this technology is to reduce the distribution difference between the source domain (labeled data) and the target domain (unlabeled data), so that the model can achieve good segmentation performance in the target domain without the support of labeled data in the target domain. This method can effectively reduce the dependence on labeled data and also provide a solution to the problem of domain distribution differences in remote sensing semantic segmentation. Current research mainly focuses on using generative adversarial networks to achieve inter-domain feature alignment. By designing generators and discriminators, the distribution of the source domain and the target domain in the feature space gradually converges. In recent years, self-training methods have also gradually become a research hotspot. By generating pseudo-labels for the target domain and performing iterative optimization, the target domain adaptability of the model is enhanced, effectively alleviating the problem of insufficient labeled data.
[0005] However, the current unsupervised domain adaptation methods still have some limitations. First, most domain adaptation methods based on generative adversarial networks rely too much on adversarial training of generators and discriminators. However, due to the complex scenes and fuzzy category boundaries of remote sensing images, the generated features often lack the ability to express details and it is difficult to accurately align the distribution of the source domain and the target domain. Secondly, although the self-training method can iteratively optimize the segmentation performance of the target domain through pseudo-labels, the quality of the pseudo-labels themselves is highly dependent on the accuracy of the initial model. When the target domain data distribution is complex, it is easy to introduce noise, resulting in the stability and performance of model training. In addition, most current semantic segmentation methods often ignore the synergy between local details and global semantics during feature alignment, and only rely on a certain type of network (such as CNN or Transformer) for feature modeling. This limitation may lead to incomplete feature expression, making it difficult to fully capture the spatial and semantic information of remote sensing images. Summary of the invention
[0006] The technical problem to be solved by the present invention is to solve the deficiencies of existing unsupervised domain adaptation methods in feature alignment, global modeling and local detail expression, and to achieve high-precision semantic segmentation. In order to overcome the defects of the above-mentioned prior art (or related technology), the present invention provides an unsupervised domain adaptation remote sensing image semantic segmentation method based on a parallel network structure.
[0007] The present invention provides an unsupervised domain adaptation remote sensing image semantic segmentation method based on a parallel network structure, comprising: Step S1, obtaining a labeled source domain dataset and an unlabeled target domain dataset; Step S2, constructing a parallel network model; Step S3, extracting source domain features from the source domain data set through the parallel network model and generating a first prediction result, and constructing a source domain cross entropy loss function based on the first prediction result and the source domain data set; Step S4, predicting the target domain data set through the parallel network model to generate pseudo labels to obtain an enhanced target domain data set, extracting target domain features to generate a second prediction result, and constructing a target domain cross entropy loss function based on the second prediction result and the enhanced target domain data set; Step S5, constructing a channel attention weighting module to perform feature weighting on the source domain features and the target domain features to obtain shared attention weights, and using the shared attention weights to perform channel-by-channel weighting on the source domain features and the target domain features to generate aligned source domain features and aligned target domain features; Step S6, obtaining a minimized distribution difference according to the aligned source domain features and the aligned target domain features, and obtaining a total loss function based on the distribution difference, the source domain cross entropy loss function and the target domain cross entropy loss function for model optimization; Step S7, testing the parallel network model to obtain an average intersection-over-union ratio as a semantic segmentation evaluation index.
[0008] Compared with the prior art, the unsupervised domain adaptation remote sensing image semantic segmentation method based on the parallel network structure of the present invention has the following advantages: The present invention utilizes a parallel network model to simultaneously acquire global features and local features, effectively combining the local detail capture capability and the global semantic modeling capability, significantly improving the model's feature expression capability for complex remote sensing images. At the same time, a channel attention mechanism is introduced to dynamically weight features, highlight key features shared between domains, and suppress domain-specific interference information. The feature distribution of the source domain and the target domain are aligned by minimizing the distribution difference, thereby enhancing the segmentation effect and improving the segmentation accuracy and adaptability of the parallel network model in the target domain. This solves the shortcomings of existing unsupervised domain adaptation methods in feature alignment, global modeling, and local detail expression, and achieves high-precision semantic segmentation.
[0009] In a possible implementation, the parallel network model includes a parallel encoder for feature extraction and a decoder for result prediction, the parallel encoder includes a CNN branch encoder for extracting local features and a Transformer branch encoder for extracting global features, and in step S2, the local features and the global features are obtained by the following calculation formula: ; in, representing the local features; represents the CNN branch encoder; represents an input image of the parallel encoder; representing the global feature; Represents the Transformer branch encoder.
[0010] In a possible implementation, the source domain dataset includes a plurality of source domain images with labels. In step S3, the source domain cross entropy loss function is constructed by the following calculation formula: ; represents the source domain cross entropy loss function; Representing the label of each of the source domain images; represents the first prediction result; represents the length of the source domain image; represents the width of the source domain image; Represents the number of categories of each source domain image.
[0011] In a possible implementation, the target domain dataset includes a plurality of unlabeled target domain images. In step S4, the target domain cross entropy loss function is constructed by the following calculation formula: ; ; in, represents the target domain cross entropy loss function; represents the pseudo label; represents the first Line The pixel prediction confidence of the column, , ; represents the second prediction result; A third prediction result representing the target domain data set generated by the parallel network model; represents the length of the target domain image; represents the width of the target domain image; Represents the number of categories of each target domain image.
[0012] In a possible implementation, in step S5, the shared attention weight is obtained by the following calculation formula: ; ; in, Represents the global average pooling result of source domain features; Indicates the height of the feature map; Indicates the width of the feature map; Represents the feature value of the cth channel at position (i, j) in the source domain feature map; Represents the global average pooling result of the target domain features; Represents the feature value of the cth channel at position (i, j) in the target domain feature map; represents the shared attention weight; Represents the Sigmoid activation function; and represents the weight matrix of the fully connected layer; Represents the ReLU activation function.
[0013] In a possible implementation, in step S5, the aligned source domain features and the aligned target domain features are obtained by the following calculation formula: ; in, represents the source domain features after alignment; represents the source domain feature; represents the shared attention weight; represents the aligned target domain features; Represents the target domain features.
[0014] In a possible implementation, in step S6, the distribution difference is obtained by the following calculation formula: ; in, represents said distribution difference; Represents the similarity of samples within the source domain features after the alignment; Represents the similarity of samples within the aligned target domain features; represents the Gaussian kernel function; Represents the similarity between the aligned source domain features and the aligned target domain features.
[0015] In a possible implementation, in step S6, the total loss function is obtained by the following calculation formula: ; in, represents the total loss function; represents the source domain cross entropy loss function; represents the target domain cross entropy loss function; Indicates preset parameters; represents the distribution difference.
[0016] In a possible implementation, in step S7, the average intersection-over-union ratio is obtained by the following calculation formula: ; in, represents intersection and union ratio; represents the number of positive examples correctly predicted during the test of the parallel network model; Indicates the number of positive examples that are incorrectly predicted during the testing of the parallel network model; represents the number of negative examples that are incorrectly predicted during the testing of the parallel network model; represents the average intersection-over-union ratio; Indicates The intersection-over-union ratio of the categories is . BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 is a flow chart of the steps of the present invention; Figure 2 It is a schematic diagram of the semantic segmentation framework of the present invention; Figure 3 A schematic diagram of a specific design of a parallel encoder of the present invention; Figure 4 This is a schematic diagram of the specific design of the feature fusion module of the present invention. DETAILED DESCRIPTION
[0018] First, those skilled in the art should understand that these implementations are only used to explain the technical principles of the embodiments of the present invention, and are not intended to limit the protection scope of the embodiments of the present invention. Those skilled in the art can make adjustments to them as needed to adapt to specific application scenarios.
[0019] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0020] See also Figure 1-4 The embodiment of the present invention discloses an unsupervised domain adaptation remote sensing image semantic segmentation method based on a parallel network structure, comprising: Step S1: Collect labeled source domain datasets and unlabeled target domain dataset ,in, represents the source domain images in the source domain dataset, is the source domain image The corresponding label, Representing the target domain images in the target domain dataset, the source domain dataset and the target domain dataset share the same label space; Step S2, construct a parallel network model M, including a parallel encoder E and a decoder D. The parallel encoder E consists of a CNN branch encoder and a Transformer branch encoder, where the CNN branch encoder is used to extract local features. , Transformer branch encoder is used to extract global features , the decoder D is used to receive the features output by the parallel encoder E and generate prediction results; Step S3, in the source domain dataset The parallel network model M is trained and the source domain features are extracted through the parallel encoder E. , and generates the prediction result through the decoder D , construct the source domain cross entropy loss function , optimize the parallel network model M in a supervised manner; the source domain cross entropy loss function is: ; In the formula, is the prediction result, is the label of the source domain image; h and w are the length and width of the source domain image respectively, and c is the number of categories of the source domain image; Step S4: Use the parallel network model M to train the target domain dataset Make predictions and generate pseudo labels : ; in, is the first Line The pixel prediction confidence of the column, , ; Then the target domain image set after data enhancement Extracting target domain features , and generates prediction results through decoder D , calculate the target domain cross entropy loss function , optimize the prediction ability of the target domain in an unsupervised manner, and the cross entropy loss function of the target domain is: ; Step S5: construct a channel attention weighting module to and target domain features Perform feature weighting and obtain shared attention weights through global average pooling and fully connected networks : ; ; in, and is the weight matrix of the fully connected layer, represents the ReLU activation function, Represents the Sigmoid activation function; Using this shared attention weight The source domain features and target domain features are weighted channel by channel to generate aligned features and : ; Step S6, by minimizing the distribution difference , from a statistical perspective, the alignment of the weighted features of the source domain and the weighted features of the target domain is achieved: ; in, Represents the Gaussian kernel function, which is used to measure the similarity of features; Represents the similarity of samples within the source domain features after alignment, Represents the similarity of samples within the target domain features after alignment, Represents the similarity between the aligned source domain features and the aligned target domain features; The total loss function in this specific example is: ; In the formula, is a constant used to control the loss ratio; step S7, testing is performed on the test set of the target domain dataset, and the mean intersection over union (mIoU) is used as the evaluation index to evaluate the semantic segmentation performance of the parallel network model M in the target domain. The formula of the mean intersection over union mIoU is: ; Where n is the number of target domain label categories, Indicates The intersection-over-union ratio of the categories, .
[0021] Source domain dataset in step S1 and target domain dataset The package contains two datasets, Potsdam and Vaihingen. The Potsdam dataset covers the urban area of Potsdam, Germany, and contains 38 high-resolution aerial images of a fixed size (6000×6000 pixels) with a resolution of 5 cm / pixel. Each image consists of four bands: red (R), green (G), blue (B), and near infrared (IR). It provides detailed ground truth annotations, which are divided into six categories: opaque surfaces, buildings, low vegetation, trees, cars, and background; the Vaihingen dataset covers the urban area of Germany. The town of Vaihingen contains 33 high-resolution aerial images of different sizes with a resolution of 9 cm / pixel. The images are mainly composed of three bands: near infrared (IR), red (R) and green (G). The ground truth annotations are consistent with the Potsdam dataset. The images of Potsdam and Vaihingen are cropped to 896×896 pixels and 512×512 pixels, respectively. Therefore, there are 1764 images in Potsdam and 1696 images in Vaihingen. In terms of dataset division, the Potsdam dataset is divided into a training set and a test set, with 1323 images in the training set and 441 images in the test set. Similarly, the Vaihingen dataset is also divided into a training set and a test set, with 1256 images in the training set and 440 images in the test set. In this specific embodiment, 4 source domain images and 4 target domain images are randomly extracted from the dataset each time.
[0022] The CNN branch encoder in step S2 uses ResNet as the backbone network to capture the detailed information of the image through layer-by-layer convolution operations. The Transformer branch encoder uses the MiT network to model the long-distance dependencies in the image through a multi-head self-attention mechanism to obtain global semantic information. The parallel encoder E converts the input image X into a multi-scale, multi-level feature representation, where: ; The decoder D adopts the SegFormer decoder structure to achieve efficient feature decoding and accurate semantic segmentation.
[0023] Continue to see Figure 4 , after executing step S2, further comprising: Construct a feature fusion module in the parallel encoder E to fusion the CNN features and Transformer features Fusion is performed to obtain fusion features ,The design of the feature fusion module is as follows: First, extract The global information of the network is used to generate a channel description vector. Next, the vector is input into two consecutive 1×1 convolutional layers for nonlinear transformation and channel compression and expansion. Then, the dynamic weight vector is generated through the Sigmoid activation function. ;Finally, generate fusion features: .
[0024] In the description of the present invention, the description with reference to the terms "one embodiment", "some embodiments", "in the present embodiment", "specific example", or "some examples" etc. means that the specific features, mechanisms, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, mechanisms, materials or characteristics described may be combined in a suitable manner in any one or more embodiments or examples. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.
[0025] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed by the present invention should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. An unsupervised domain adaptation remote sensing image semantic segmentation method based on a parallel network structure, characterized in that: The following steps are involved: Step S1, obtaining a labeled source domain dataset and an unlabeled target domain dataset; Step S2, constructing a parallel network model; Step S3, extracting source domain features from the source domain data set through the parallel network model and generating a first prediction result, and constructing a source domain cross entropy loss function based on the first prediction result and the source domain data set; Step S4, predicting the target domain data set through the parallel network model to generate pseudo labels to obtain an enhanced target domain data set, extracting target domain features to generate a second prediction result, and constructing a target domain cross entropy loss function based on the second prediction result and the enhanced target domain data set; Step S5, constructing a channel attention weighting module to perform feature weighting on the source domain features and the target domain features to obtain shared attention weights, and using the shared attention weights to perform channel-by-channel weighting on the source domain features and the target domain features to generate aligned source domain features and aligned target domain features; Step S6, obtaining a minimized distribution difference according to the aligned source domain features and the aligned target domain features, and obtaining a total loss function based on the distribution difference, the source domain cross entropy loss function and the target domain cross entropy loss function for model optimization; Step S7, testing the parallel network model to obtain an average intersection-over-union ratio as a semantic segmentation evaluation index.
2. The unsupervised domain adaptation remote sensing image semantic segmentation method based on a parallel network structure according to claim 1, characterized in that: The parallel network model includes a parallel encoder for feature extraction and a decoder for result prediction, and the parallel encoder includes a CNN branch encoder for extracting local features and a Transformer branch encoder for extracting global features. In step S2, the local features and the global features are obtained by the following calculation formula: ; in, representing the local features; represents the CNN branch encoder; represents an input image of the parallel encoder; representing the global feature; Represents the Transformer branch encoder.
3. The unsupervised domain adaptation remote sensing image semantic segmentation method based on a parallel network structure according to claim 1, characterized in that: The source domain dataset includes a plurality of source domain images with labels. In step S3, the source domain cross entropy loss function is constructed by the following calculation formula: ; represents the source domain cross entropy loss function; Representing the label of each of the source domain images; represents the first prediction result; represents the length of the source domain image; represents the width of the source domain image; Represents the number of categories of each source domain image.
4. The unsupervised domain adaptation remote sensing image semantic segmentation method based on a parallel network structure according to claim 1, characterized in that: The target domain dataset includes a plurality of unlabeled target domain images. In step S4, the target domain cross entropy loss function is constructed by the following calculation formula: ; ; in, represents the target domain cross entropy loss function; represents the pseudo label; represents the first Line The pixel prediction confidence of the column, , ; represents the second prediction result; A third prediction result representing the target domain data set generated by the parallel network model; represents the length of the target domain image; represents the width of the target domain image; Represents the number of categories of each target domain image.
5. The unsupervised domain adaptation remote sensing image semantic segmentation method based on a parallel network structure according to claim 1, characterized in that: In step S5, the shared attention weight is obtained by the following calculation formula: ; ; in, Represents the global average pooling result of source domain features; Indicates the height of the feature map; Indicates the width of the feature map; Represents the feature value of the cth channel at position (i, j) in the source domain feature map; Represents the global average pooling result of the target domain features; Represents the feature value of the cth channel at position (i, j) in the target domain feature map; represents the shared attention weight; Represents the Sigmoid activation function; and represents the weight matrix of the fully connected layer; Represents the ReLU activation function.
6. The unsupervised domain adaptation remote sensing image semantic segmentation method based on a parallel network structure according to claim 1, characterized in that: In step S5, the aligned source domain features and the aligned target domain features are obtained by the following calculation formula: ; in, represents the source domain features after alignment; represents the source domain feature; represents the shared attention weight; represents the aligned target domain features; Represents the target domain features.
7. The unsupervised domain adaptation remote sensing image semantic segmentation method based on a parallel network structure according to claim 1, characterized in that: In step S6, the distribution difference is obtained by the following calculation formula: ; in, represents said distribution difference; Represents the similarity of samples within the source domain features after the alignment; Represents the similarity of samples within the aligned target domain features; represents the Gaussian kernel function; Represents the similarity between the aligned source domain features and the aligned target domain features.
8. The unsupervised domain adaptation remote sensing image semantic segmentation method based on a parallel network structure according to claim 1, characterized in that: In step S6, the total loss function is obtained by the following calculation formula: ; in, represents the total loss function; represents the source domain cross entropy loss function; represents the target domain cross entropy loss function; Indicates preset parameters; represents the distribution difference.
9. The unsupervised domain adaptation remote sensing image semantic segmentation method based on a parallel network structure according to claim 1, characterized in that: In step S7, the average intersection-over-combination ratio is obtained by the following calculation formula: ; in, represents intersection and union ratio; represents the number of positive examples correctly predicted during the test of the parallel network model; represents the number of positive examples that are incorrectly predicted during the testing of the parallel network model; represents the number of negative examples that are incorrectly predicted during the testing of the parallel network model; represents the average intersection-over-union ratio; Indicates The intersection-over-union ratio of the categories is .
Citation Information
Patent Citations
Unsupervised domain adaptive remote sensing road semantic segmentation method based on GAN network
CN113888547A
Semi-supervised domain adaptive image semantic segmentation method, system and device and storage medium
CN116229080A
Event-based unsupervised domain adaptive semantic segmentation network training method
CN116935047A
Remote sensing image cross-domain semantic segmentation method based on domain adaptation
CN118691824A
Domain adaptation semantic segmentation method based on cross pseudo supervision
CN118968062A