Cross-domain remote sensing image semantic segmentation method based on context orthogonal decoupling

By combining the Segformer framework and the decoupling module, using self-training to generate pseudo labels and contrastive learning to optimize model training, the problems of redundant features and insufficient category distinction performance in cross-domain remote sensing image semantic segmentation are solved, and the adaptability and generalization ability of the model are improved.

CN120689878APending Publication Date: 2025-09-23CHINA UNIV OF MINING & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510781995.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Existing cross-domain remote sensing image semantic segmentation models lack adaptability and generalization capabilities between different domains, and are particularly inhibited by redundant features and category discrimination performance in remote sensing images, resulting in a decline in model effectiveness.

Method used

A cross-domain remote sensing image semantic segmentation model based on the Segformer framework is adopted. Pseudo labels are generated through self-training for image mixing. The decoupling module is used to decouple and reorganize the features. The model training process is optimized by combining contrastive learning and category confusion matrix to improve the model's semantic feature recognition and domain feature differentiation capabilities.

Benefits of technology

It effectively improves the model's ability to distinguish redundant features, improves the model's adaptability between different domains and its ability to distinguish semantic features, and enhances the model's generalization performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120689878A_ABST
    Figure CN120689878A_ABST
Patent Text Reader

Abstract

The invention discloses a cross-domain remote sensing image semantic segmentation method based on context orthogonal decoupling, and the method comprises the steps: building a model based on a Segformer frame, generating a pseudo tag of a target domain image, and mixing a source domain image and the target domain image to obtain a mixed image; decoupling the source domain features and the target domain features by using a decoupling module to obtain a semantic feature weight matrix; respectively separating the source domain features and the target domain features to obtain domain features and semantic features; drawing close two groups of semantic features of the same category in the domain features and the semantic features by utilizing comparative learning, drawing the distance between the two groups of semantic features of different categories, distinguishing two groups of center domain features by utilizing cosine similarity, and recombining to obtain recombined features; sending the recombined features into a classifier to obtain confidence, and calculating the difference between the recombined features and the source domain features by using a category confusion matrix; training the model; and inputting a to-be-identified cross-domain remote sensing image into the model to obtain an image pixel-level segmentation result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a cross-domain remote sensing image semantic segmentation method based on context orthogonal decoupling, and belongs to the technical field of cross-domain remote sensing image semantic segmentation methods based on context orthogonal decoupling. Background Art

[0002] At this stage, the application scenarios of remote sensing imagery are constantly expanding, and their importance is becoming increasingly prominent. Semantic segmentation, as a key foundational technology, has gradually become a research hotspot. For example, tasks such as accurately extracting building outlines from remote sensing images, calculating forest cover, and performing road segmentation are all of great practical value. Deep learning has achieved remarkable success in the field of computer vision, and semantic segmentation technology is constantly advancing and improving. However, when a model trained on source domain images is directly applied to target domain images, the performance often deteriorates significantly, making cross-domain semantic segmentation a current research focus.

[0003] In the field of remote sensing images, the cross-domain phenomenon is particularly prominent. Images acquired from different regions and time points, or even from different sensors, may have significant differences in color, texture, illumination, and many other aspects. In order to train an effective semantic segmentation model, a large amount of labeled data is usually required, and this data must cover the category labels corresponding to each pixel. However, obtaining and labeling this data often requires a lot of time and money. Therefore, how to make the model trained on the source domain image have stronger adaptability to cope with the differences between images and maintain good generalization ability has become an urgent problem to be solved. The main challenges currently faced by the existing cross-domain remote sensing image semantic segmentation are as follows: (1) Remote sensing images are mixed with too many external interference factors, which have an inhibitory effect on the training model, making it difficult for the model to extract generalization features. (2) Remote sensing images contain richer content, and high-altitude shooting causes some similar categories to lose their unique features, which further reduces the model's ability to distinguish. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to overcome the defects of the existing technology, provide a cross-domain remote sensing image semantic segmentation method based on context orthogonal decoupling, and propose corresponding solutions to the problems of redundant feature factors and category differentiation performance.

[0005] Preferably, the present invention provides a cross-domain remote sensing image semantic segmentation method based on context orthogonal decoupling, comprising:

[0006] Step 1: Data preprocessing is performed on the previously acquired cross-domain remote sensing image semantic segmentation dataset ISPRS Potsdam and cross-domain remote sensing image semantic segmentation dataset ISPRS Vaihingen. ISPRS Potsdam is the source domain image, and ISPRS Vaihingen is the target domain image.

[0007] Step 2: Build a cross-domain remote sensing image semantic segmentation model based on the Segformer framework, generate pseudo labels for target domain images using a self-training method, and mix the source domain image and the target domain image to obtain a mixed image for training;

[0008] Step 3: Use the decoupling module to decouple the input source domain features and target domain features to obtain the semantic feature weight matrix;

[0009] Step 4: Separate the input source domain features and target domain features to obtain domain features and semantic features. Use contrastive learning to bring two groups of semantic features of the same category closer, and to increase the distance between two groups of semantic features of different categories. Use cosine similarity to distinguish the two groups of central domain features.

[0010] Step 5: Recombining the source domain semantic features and the target domain domain features to obtain recombined features; feeding the recombined features into the classifier to obtain the confidence level, and using the category confusion matrix to calculate the difference between the recombined features and the source domain features;

[0011] Step 6: Based on the loss function and the mixed image, the cross-domain remote sensing image semantic segmentation model is trained;

[0012] The cross-domain remote sensing image to be identified is input into the trained cross-domain remote sensing image semantic segmentation model to predict the pixel-level segmentation result of the image.

[0013] Preferably, step 1 comprises:

[0014] (1.1) The ISPRS Potsdam dataset contains cross-domain remote sensing images and includes IR-RG, RGB, and IR-R-GB band modes. The ISPRS Vaihingen dataset contains cross-domain remote sensing images and includes IR-RG band mode.

[0015] (1.2) Cropping the cross-domain remote sensing image to cut the cross-domain remote sensing image into images of a preset size;

[0016] (1.3) The cropped cross-domain remote sensing image is subjected to data preprocessing including random rotation, flipping, and random photometric distortion to obtain a preprocessed cross-domain remote sensing image.

[0017] Preferably, step 2 comprises:

[0018] (2.1) Build the Segformer framework and obtain the student model;

[0019] (2.2) Copy the student model to the teacher model. The cross-domain remote sensing image semantic segmentation model includes a student model and a teacher model. The parameter initialization of the teacher model is consistent with that of the student model. The subsequent parameters of the teacher model are updated by using the exponential moving average through the student model:

[0020]

[0021] Among them, t represents the current iteration round, α represents the update weight, represents the parameters of the teacher model for the tth iteration, represents the parameters of the student model at the tth iteration;

[0022] (2.3) Using the teacher model with updated parameter values, generate online pseudo labels for the input target domain image, and take the category with the highest score as the pseudo label of the target domain image;

[0023] (2.4) Calculate the ratio of the number of pixels with a score greater than the threshold τ in all pixels of the target domain image, and obtain the weight of the target domain cross entropy loss:

[0024]

[0025] Where H and W are the height and width of the input target domain image, respectively, q represents the confidence score corresponding to the pseudo label; [·] represents Iverson brackets, indicating that the value is 1 when the condition in the brackets is true;

[0026] (2.5) Use the ClassMix mixing method to generate the mask matrix M and mix the source domain image and the target domain image:

[0027] x mix =M·x S +(1-M)·x T (3)

[0028] Among them, x mix is the mixed image, x S 、x T Represent the source domain image and target domain image respectively;

[0029] (2.6) The mixed labels of the mixed image are mixed with the mask matrix M, and the loss weight of the source domain pixel is 1 and the target domain pixel is w T ;

[0030] (2.7) The mixed images and mixed labels are fed into the student model for training.

[0031] Preferably, step 3 comprises:

[0032] (3.1) Perform linear transformation on the source domain features and target domain features using nn.Linear(inp, inp*3), where inp is the channel dimension of the input tensor. The linearly transformed source domain features and target domain features are evenly separated into three parts according to the channel dimension, and the corresponding linearly transformed features Q, K, and V are obtained.

[0033] (3.2) The feature K of the source domain feature and the feature K of the target domain feature are averaged according to the channel dimension to obtain Calculate K and The square of the difference between

[0034] (3.3) Yes Sum by channel, using Divide by The sum of and normalize to get the source domain normalization parameter K S , target domain normalization parameter K T ;

[0035] (3.4) Using the source domain normalization parameter K S Perform matrix multiplication with the corresponding feature Q to obtain the source domain relationship matrix R between channels S ;Use the target domain normalization parameter K T Perform matrix multiplication with the corresponding feature Q to obtain the target domain relationship matrix R between channels T ;

[0036] (3.5) The source domain relationship matrix R S and the target domain relationship matrix R T After adding, divide by 2 and transform using the softmax activation function to obtain the unified channel relationship matrix R C ;

[0037] (3.6) The unified channel relationship matrix R C Perform matrix multiplication with the feature V of the source domain feature to obtain the attention matrix A of the target domain T , the unified channel relationship matrix R C Perform matrix multiplication with the feature V of the target domain feature to obtain the attention matrix A of the target domain of the source domainT , attention matrix A S And the target domain's attention matrix A T Contains batch dimension, channel dimension and spatial dimension;

[0038] (3.7) Focus the source domain’s attention matrix A S , the target domain's attention matrix A T Splicing is performed according to the spatial dimension, and nonlinear transformation is performed using convolution nn.Conv2d. mip is the middle channel dimension, which is 1 / 8 of the input channel dimension of inp, and is processed through batch layers and Relu functions.

[0039] (3.8) Separate the attention matrix A of the source domain after processing in step 3.7 S , the target domain's attention matrix A T , and processed with convolution nn.Conv2d and sigmoid activation function to obtain the semantic feature weight matrix w of the source domain s , the semantic feature weight matrix w of the target domain t .

[0040] Preferably, step 4 comprises:

[0041] (4.1) Semantic feature weight matrix w based on the source domain s , the semantic feature weight matrix w of the target domain t , we can calculate:

[0042]

[0043] where f S 、f T 、 They are source domain features, target domain features, source domain semantic features, source domain field features, target domain semantic features and target domain field features;

[0044] (4.2) Based on the true source domain labels and the target domain pseudo labels, obtain the mask matrix M for each category. Use the mask matrix to calculate the category center features of the source domain semantic features and the category center features of the target domain semantic features. Calculate the dot product between the source domain semantic center features and the target domain semantic center features of different categories, and use contrastive learning to construct the loss function:

[0045]

[0046] in, are the central features of the source domain semantic category and the central features of the target domain semantic category, respectively; K1 and K2 are the number of source domain categories and the number of target domain categories, respectively;

[0047] (4.3) Calculate the central domain features of the source domain and the central domain features of the target domain:

[0048]

[0049] Among them, L d is the central domain feature, and cos represents the calculation of cosine similarity.

[0050] Preferably, step 5 comprises:

[0051] (5.1) Add the source domain semantic features and target domain features to form a new feature tensor f u :

[0052]

[0053] (5.2) The new feature tensor f u Send it to the classifier for prediction and get the category confidence C u ; Confidence C for the category u Transpose the confidence matrix composed of Transpose the matrix Perform matrix multiplication to obtain the category confusion matrix y of each pixel of the recombined feature u :

[0054]

[0055] in, represents matrix multiplication;

[0056] (5.3) Calculate the mean L1 of the sum of the off-diagonal elements of the confusion matrix for each pixel as the class confusion loss:

[0057]

[0058] in, Represents the confusion matrix y u The value of row i and column j, N represents the number of all categories in the source domain;

[0059] (5.4) Calculate the category confusion matrix y of the source domain features s , subtract the two category confusion matrices and get the square of the difference y between the category confusion matrices d :

[0060] y d =(y u -y s ) 2 (13)

[0061] (5.5) Sum the diagonal differences between the category confusion matrix of the recombined features and the confusion matrix of the source domain features to obtain the sum L2:

[0062]

[0063] Where, represents the square y d The value in the i-th row and j-th column of the matrix.

[0064] Preferably, step 6 comprises:

[0065] (6.1) Calculate the cross entropy loss L for semantic segmentation of source domain images S :

[0066]

[0067] in, is the one-hot label of the source domain image, g s is the student model, H is the height of the source domain image, W is the width of the source domain image, C is the number of semantic categories of the source domain image, x s is the source domain image;

[0068] (6.2) Calculate the semantic segmentation cross entropy loss L of the mixed image mix :

[0069]

[0070] Among them, w is the mixing weight, is the one-hot label of the mixed image, x mix It is a mixed image;

[0071] (6.3) Calculate the decoupling contrast loss L DC :

[0072] L DC =λ1L s +λ2L d (17)

[0073] Among them, λ1 and λ2 are the semantic contrast loss weight and domain contrast loss weight respectively;

[0074] (6.4) Calculate the category confusion loss L con :

[0075] L con =λ3(L1+L2) (18)

[0076] Among them, λ3 is the category confusion loss weight;

[0077] (6.5) Using cross entropy loss LS , semantic segmentation cross entropy loss L mix , decoupled contrast loss L DC and category confusion loss L con The cross-domain remote sensing image semantic segmentation model is trained until the number of iterations meets the preset iteration number requirement.

[0078] Preferably, the present invention provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of any one of the methods described in the first aspect when executing the program.

[0079] Preferably, the present invention provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of any one of the methods described in the first aspect when executed by a processor.

[0080] The beneficial effects achieved by the present invention are:

[0081] First, the existing technology cannot deal well with the impact of redundant features on model training, while the context orthogonal decoupling provided by the present invention can help the model to distinguish pure semantic features from irrelevant domain features, and then use contrastive learning to compare the similarity of semantic features and domain features. Compared with directly using encoding features for contrastive learning, this is more conducive to the model's ability to learn features.

[0082] Second, the present invention reconstructs features and feeds them into the classifier, helping the model focus on key factors and reduce attention to irrelevant features. The present invention also utilizes a class confusion matrix to better demonstrate the connections between categories, further improving the model's ability to distinguish between similar categories. BRIEF DESCRIPTION OF THE DRAWINGS

[0083] In order to more clearly illustrate the technical solution of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0084] Figure 1 is a flow chart of some embodiments of the present application;

[0085] Figure 2 This is a diagram of the overall structure of the network in some embodiments of the present application;

[0086] Figure 3 It is a specific flow chart of the decoupling module in some embodiments of the present application. DETAILED DESCRIPTION

[0087] See also Figure 1The present application discloses a cross-domain remote sensing image semantic segmentation method based on context orthogonal decoupling, which is characterized by comprising the following steps:

[0088] Step 1: Data preprocessing is performed on the previously acquired cross-domain remote sensing image semantic segmentation dataset ISPRS Potsdam and cross-domain remote sensing image semantic segmentation dataset ISPRS Vaihingen. ISPRS Potsdam is the source domain image, and ISPRS Vaihingen is the target domain image.

[0089] Step 2: Build a cross-domain remote sensing image semantic segmentation model based on the Segformer framework, generate pseudo labels for target domain images using a self-training method, and mix the source domain image and the target domain image to obtain a mixed image for training;

[0090] Step 3: Use the decoupling module to decouple the input source domain features and target domain features to obtain the semantic feature weight matrix;

[0091] Step 4: Separate the input source domain features and target domain features to obtain domain features and semantic features. Use contrastive learning to bring two groups of semantic features of the same category closer, and to increase the distance between two groups of semantic features of different categories. Use cosine similarity to distinguish the two groups of central domain features.

[0092] Step 5: Recombining the source domain semantic features and the target domain domain features to obtain recombined features. The recombined features are fed into the classifier to obtain confidence scores. The difference between the recombined features and the source domain features is calculated using the category confusion matrix to suppress the impact of domain features on model performance.

[0093] In step 6, based on the loss function and the mixed image, the cross-domain remote sensing image semantic segmentation model is trained and tested on the target domain; the cross-domain remote sensing image to be identified is input into the trained cross-domain remote sensing image semantic segmentation model to predict the image pixel-level segmentation result.

[0094] Furthermore, step 1 includes:

[0095] (1.1) Obtain the cross-domain remote sensing image semantic segmentation dataset ISPRS Potsdam and the cross-domain remote sensing image semantic segmentation dataset ISPRS Vaihingen. The cross-domain remote sensing image semantic segmentation dataset ISPRS Potsdam contains 38 cross-domain remote sensing images with a size of 6000*6000 and a spatial resolution of 5 cm. The cross-domain remote sensing image semantic segmentation dataset ISPRS Potsdam provides band modes including IR-RG, RGB, and IR-R-GB; the cross-domain remote sensing image semantic segmentation dataset ISPRS Vaihingen contains 33 cross-domain remote sensing images with a size of 2000*2000 and a spatial resolution of 9 cm. The cross-domain remote sensing image semantic segmentation dataset ISPRS Vaihingen includes the band mode IR-RG.

[0096] (1.2) Cropping the cross-domain remote sensing image into an image of a preset size of 512*512;

[0097] (1.3) The cropped cross-domain remote sensing image is subjected to data preprocessing including random rotation, flipping, and random photometric distortion to obtain a preprocessed cross-domain remote sensing image.

[0098] Furthermore, step 2 includes:

[0099] (2.1) Build the Segformer framework and obtain the student model;

[0100] (2.2) Copy the student model to the teacher model. The cross-domain remote sensing image semantic segmentation model includes a student model and a teacher model. The parameter initialization of the teacher model is consistent with that of the student model. The subsequent parameters of the teacher model are updated by using the exponential moving average through the student model:

[0101]

[0102] Among them, t represents the current iteration round, α represents the update weight, represents the parameters of the teacher model for the tth iteration, represents the parameters of the student model at the tth iteration;

[0103] (2.3) Using the teacher model with updated parameter values, generate online pseudo labels for the input target domain image, and take the category with the highest score as the pseudo label of the target domain image;

[0104] (2.4) However, it is inappropriate to directly use this label to train the student model. The target domain pseudo-label is screened by using a threshold τ. The ratio of the number of pixels with a score greater than the threshold τ in all pixels of the target domain image is calculated to obtain the weight of the target domain cross entropy loss. The calculation expression is as follows:

[0105]

[0106] Where H and W are the height and width of the input target domain image, respectively, q represents the confidence score corresponding to the pseudo label; [·] represents Iverson brackets, indicating that the value is 1 when the condition in the brackets is true;

[0107] (2.5) In addition, the ClassMix mixing method is used to generate the mask matrix M to mix the source domain image and the target domain image. The formula is as follows:

[0108] x mix =M·x S +(1-M)·x T (3)

[0109] Among them, x mix is the mixed image, x S 、x T Represent the source domain image and target domain image respectively;

[0110] (2.6) Similarly, the mixed labels of the mixed image are also mixed using the mask matrix M, with the loss weights of the source domain pixels being 1 and the target domain pixels being w T ;

[0111] (2.7) The mixed images and mixed labels are fed into the student model for training.

[0112] Furthermore, in step 3, the decoupling module is used to decouple the input source domain features and target domain features to obtain a semantic feature weight matrix, including:

[0113] (3.1) Figure 2 As shown, the source domain features and target domain features are linearly transformed using nn.Linear(inp, inp*3), where inp is the channel dimension of the input tensor. The linearly transformed source domain features and target domain features are evenly separated into three parts according to the channel dimension, and the corresponding linearly transformed features Q, K, and V are obtained.

[0114] (3.2) The feature K of the source domain feature and the feature K of the target domain feature are transformed respectively. The transformation process is to obtain the average value according to the channel dimension. Calculate K and The square of the difference between

[0115] (3.3) Yes Sum according to the channel and then use Divide by The sum of and normalize to get the source domain normalization parameter KS , target domain normalization parameter K T ;

[0116] (3.4) Using the source domain normalization parameter K S , target domain normalization parameter K T Perform matrix multiplication with the corresponding feature Q respectively to obtain the source domain relationship matrix R between channels S , target domain relationship matrix R T ;

[0117] (3.5) The source domain relationship matrix R S and the target domain relationship matrix R T After adding, divide by 2 and transform using the softmax activation function to obtain the unified channel relationship matrix R C ;

[0118] (3.6) The unified channel relationship matrix R C Perform matrix multiplication with the feature V of the source domain and the feature V of the target domain respectively to obtain the source domain attention matrix A S , the target domain's attention matrix A T , attention matrix A S And the target domain's attention matrix A T Contains batch dimension, channel dimension and spatial dimension;

[0119] (3.7) Focus the source domain’s attention matrix A S , the target domain's attention matrix A T The concatenation is performed along the last dimension, i.e., the spatial dimension, and nonlinear transformation is performed using convolution nn.Conv2d(inp, mip, kernel_size = 1, stride = 1, padding = 0). Mip is the intermediate channel dimension, which is 1 / 8 of the inp input channel dimension. Kernel_size is the convolution kernel size, stride is the convolution kernel stride, and padding is the number of padded pixels. The convolution is then normalized by a batch layer and activated by the ReLU function to improve robustness.

[0120] (3.8) Separate the attention matrix A of the source domain processed in step 3.7 S , the target domain's attention matrix A T And use convolution nn.Conv2d (mip, inp, kernel_size = 1, stride = 1, padding = 0) and sigmoid activation function to obtain the semantic feature weight matrix w of the source domain. s , the semantic feature weight matrix w of the target domain t .

[0121] Furthermore, step 4 includes:

[0122] (4.1) Semantic feature weight matrix w based on the source domain s , the semantic feature weight matrix w of the target domain t , we can calculate:

[0123]

[0124] where f S 、f T 、 They are source domain features, target domain features, source domain semantic features, source domain field features, target domain semantic features, and target domain field features;

[0125] (4.2) Based on the true source domain labels and the target domain pseudo labels, obtain the mask matrix M for each category. Use the mask matrix to calculate the category center features of the source domain semantic features and the category center features of the target domain semantic features. Calculate the dot product between the source domain semantic center features and the target domain semantic center features of different categories, and use contrastive learning to construct a loss function. The loss function formula is as follows:

[0126]

[0127] in, are the central features of the source domain semantic category and the central features of the target domain semantic category, respectively; K1 and K2 are the number of source domain categories and the number of target domain categories, respectively;

[0128] (4.3) Calculate the central domain features of the source domain and the central domain features of the target domain; use cosine similarity to measure the similarity of the two domain features; force the similarity of the two domain features to tend to 0, forming orthogonal independence, thereby helping the model learn:

[0129]

[0130] Among them, L d is the central domain feature, and cos represents the calculation of cosine similarity.

[0131] Furthermore, step 5 includes:

[0132] (5.1) Recombining the separated semantic features and domain features. Specifically, the source domain semantic features and target domain features are added together to form a new feature tensor f u , the formula is as follows:

[0133]

[0134] (5.2) The new feature tensor f uSend it to the classifier for prediction and get the category confidence C u ; For confidence C u Transpose the confidence matrix composed of Transpose the matrix Perform matrix multiplication to obtain the category confusion matrix y of each pixel of the recombined feature u , the formula is as follows:

[0135]

[0136] in, stands for matrix multiplication;

[0137] (5.3) Calculate the mean L1 of the sum of the off-diagonal elements of the confusion matrix of each pixel as the category confusion loss. Because the decoupling matrix uses context information to extract semantic weights, the mean of the sum of the off-diagonal elements of the confusion matrix of each pixel can reflect the association between context information and different categories. The calculation formula is as follows:

[0138]

[0139] in, Represents the confusion matrix y u The value of row i and column j, N represents the number of all categories in the source domain;

[0140] (5.4) Calculate the category confusion matrix y of the source domain features s , subtract the two category confusion matrices and get the square of the difference y between the category confusion matrices d , the calculation process is as follows:

[0141] y d =(y u -y s ) 2 (13)

[0142] (5.5) For square y d , mainly calculates the mean of the off-diagonal sum; because the previous off-diagonal sum has ensured that the elements will tend to be in the same category when discriminating, but it cannot guarantee whether the judgment is accurate; the diagonal difference between the category confusion matrix of the reconstructed features and the confusion matrix of the source domain features itself is summed to obtain the sum value L2, so as to ensure that the reconstructed features can still be accurately judged. The calculation formula is as follows:

[0143]

[0144] Where, represents the square y d The value in the i-th row and j-th column of the matrix.

[0145] Furthermore, step 6 includes:

[0146] (6.1) Calculate the cross entropy loss L for semantic segmentation of source domain images S , the formula is as follows:

[0147]

[0148] in, is the one-hot label of the source domain image, g s is the student model, H is the height of the source domain image, W is the width of the source domain image, C is the number of semantic categories of the source domain image, x s is the source domain image;

[0149] (6.2) Calculate the semantic segmentation cross entropy loss L of the mixed image mix :

[0150]

[0151] Among them, w is the mixing weight, is the one-hot label of the mixed image, x mix It is a mixed image;

[0152] (6.3) Calculate the decoupling contrast loss L DC :

[0153] L DC =λ1L s +λ2L d (17)

[0154] Among them, λ1 and λ2 are the semantic contrast loss weight and domain contrast loss weight respectively;

[0155] (6.4) Calculate the category confusion loss L con :

[0156] L con =λ3(L1+L2) (18)

[0157] Among them, λ3 is the category confusion loss weight;

[0158] (6.5) Using cross entropy loss L S , semantic segmentation cross entropy loss L mix , decoupled contrast loss L DC and category confusion loss L con The cross-domain remote sensing image semantic segmentation model is trained until the number of iterations meets the preset iteration number requirement.

[0159] In an embodiment of the present application, the present invention provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of any of the above methods when executing the program.

[0160] In an embodiment of the present application, the present invention provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of any of the above methods when executed by a processor.

[0161] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments.

[0162] Those skilled in the art will readily appreciate other embodiments of the present invention after considering the specification and practicing the invention as disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not invented herein, and the description and examples are to be considered merely as exemplary.

[0163] The above specific implementation methods further illustrate the purpose, technical solutions and beneficial effects of this application in detail. It should be understood that the above are only specific implementation methods of this application and are not intended to limit the scope of protection of this application. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of this application should be included in the scope of protection of this application.

Claims

1. A cross-domain remote sensing image semantic segmentation method based on context orthogonal decoupling, characterized by: include: Step 1: Data preprocessing is performed on the previously acquired cross-domain remote sensing image semantic segmentation dataset ISPRS Potsdam and cross-domain remote sensing image semantic segmentation dataset ISPRS Vaihingen. ISPRS Potsdam is the source domain image, and ISPRS Vaihingen is the target domain image. Step 2: Build a cross-domain remote sensing image semantic segmentation model based on the Segformer framework, generate pseudo labels for target domain images using a self-training method, and mix the source domain image and the target domain image to obtain a mixed image for training; Step 3: Use the decoupling module to decouple the input source domain features and target domain features to obtain the semantic feature weight matrix; Step 4: Separate the input source domain features and target domain features to obtain domain features and semantic features. Use contrastive learning to bring two groups of semantic features of the same category closer, and to increase the distance between two groups of semantic features of different categories. Use cosine similarity to distinguish the two groups of central domain features. Step 5: Recombining the source domain semantic features and the target domain domain features to obtain recombined features; feeding the recombined features into the classifier to obtain the confidence level, and using the category confusion matrix to calculate the difference between the recombined features and the source domain features; Step 6: Based on the loss function and the mixed image, the cross-domain remote sensing image semantic segmentation model is trained; The cross-domain remote sensing image to be identified is input into the trained cross-domain remote sensing image semantic segmentation model to predict the pixel-level segmentation result of the image.

2. The cross-domain remote sensing image semantic segmentation method based on context orthogonal decoupling according to claim 1 is characterized in that: Step 1 includes: (1.1) The ISPRS Potsdam dataset contains cross-domain remote sensing images and includes IR-RG, RGB, and IR-R-GB band modes. The ISPRS Vaihingen dataset contains cross-domain remote sensing images and includes IR-RG band mode. (1.2) Cropping the cross-domain remote sensing image to cut the cross-domain remote sensing image into images of a preset size; (1.3) The cropped cross-domain remote sensing image is subjected to data preprocessing including random rotation, flipping, and random photometric distortion to obtain a preprocessed cross-domain remote sensing image.

3. The cross-domain remote sensing image semantic segmentation method based on context orthogonal decoupling according to claim 1 is characterized in that: Step 2 includes: (2.1) Build the Segformer framework and obtain the student model; (2.2) Copy the student model to the teacher model. The cross-domain remote sensing image semantic segmentation model includes a student model and a teacher model. The parameter initialization of the teacher model is consistent with that of the student model. The subsequent parameters of the teacher model are updated by using the exponential moving average through the student model: Among them, t represents the current iteration round, α represents the update weight, represents the parameters of the teacher model for the tth iteration, represents the parameters of the student model at the tth iteration; (2.3) Using the teacher model with updated parameter values, generate online pseudo labels for the input target domain image, and take the category with the highest score as the pseudo label of the target domain image; (2.4) Calculate the ratio of the number of pixels with a score greater than the threshold τ in all pixels of the target domain image, and obtain the weight of the target domain cross entropy loss: Where H and W are the height and width of the input target domain image, respectively, q represents the confidence score corresponding to the pseudo label; [·] represents Iverson brackets, indicating that the value is 1 when the condition in the brackets is true; (2.5) Use the ClassMix mixing method to generate the mask matrix M and mix the source domain image and the target domain image: x mix =M·x S +(1-M)·x T (3) Among them, x mix is the mixed image, x S 、x T Represent the source domain image and target domain image respectively; (2.6) The mixed labels of the mixed image are mixed with the mask matrix M, and the loss weight of the source domain pixel is 1 and the target domain pixel is w T ; (2.7) The mixed images and mixed labels are fed into the student model for training.

4. The cross-domain remote sensing image semantic segmentation method based on context orthogonal decoupling according to claim 1 is characterized in that: Step 3 includes: (3.1) Perform linear transformation on the source domain features and target domain features using nn.Linear(inp, inp*3), where inp is the channel dimension of the input tensor. The linearly transformed source domain features and target domain features are evenly separated into three parts according to the channel dimension, and the corresponding linearly transformed features Q, K, and V are obtained. (3.2) The feature K of the source domain feature and the feature K of the target domain feature are averaged according to the channel dimension to obtain Calculate K and The square of the difference between (3.3) Yes Sum by channel, using Divide by The sum of and normalize to get the source domain normalization parameter K S , target domain normalization parameter K T ; (3.4) Using the source domain normalization parameter K S Perform matrix multiplication with the corresponding feature Q to obtain the source domain relationship matrix R between channels S ;Use the target domain normalization parameter K T Perform matrix multiplication with the corresponding feature Q to obtain the target domain relationship matrix R between channels T ; (3.5) The source domain relationship matrix R S and the target domain relationship matrix R T After adding, divide by 2 and transform using the softmax activation function to obtain the unified channel relationship matrix R C ; (3.6) The unified channel relationship matrix R C Perform matrix multiplication with the feature V of the source domain feature to obtain the attention matrix A of the target domain T , the unified channel relationship matrix R C Perform matrix multiplication with the feature V of the target domain feature to obtain the attention matrix A of the target domain of the source domain T , attention matrix A S And the target domain's attention matrix A T Contains batch dimension, channel dimension and spatial dimension; (3.7) Focus the source domain’s attention matrix A S , the target domain's attention matrix A T Splicing is performed according to the spatial dimension, and nonlinear transformation is performed using convolution nn.Conv2d. mip is the middle channel dimension, which is 1 / 8 of the input channel dimension of inp, and is processed through batch layers and Relu functions. (3.8) Separate the attention matrix A of the source domain after step 3.7 S , the target domain's attention matrix A T , and processed with convolution nn.Conv2d and sigmoid activation function to obtain the semantic feature weight matrix ws of the source domain and the semantic feature weight matrix wt ​​of the target domain.

5. The cross-domain remote sensing image semantic segmentation method based on context orthogonal decoupling according to claim 4 is characterized in that: Step 4 includes: (4.1) Based on the semantic feature weight matrix ws of the source domain and the semantic feature weight matrix wt ​​of the target domain, we can calculate: where f S 、f T 、 They are source domain features, target domain features, source domain semantic features, source domain field features, target domain semantic features and target domain field features; (4.2) Based on the true source domain labels and the target domain pseudo labels, obtain the mask matrix M for each category. Use the mask matrix to calculate the category center features of the source domain semantic features and the category center features of the target domain semantic features. Calculate the dot product between the source domain semantic center features and the target domain semantic center features of different categories, and use contrastive learning to construct the loss function: Among them, L ss is the loss value, are the central features of the source domain semantic category and the central features of the target domain semantic category, respectively; K1 and K2 are the number of source domain categories and the number of target domain categories, respectively; (4.3) Calculate the central domain features of the source domain and the central domain features of the target domain: Among them, L d is the central domain feature, and cos represents the calculation of cosine similarity.

6. The cross-domain remote sensing image semantic segmentation method based on context orthogonal decoupling according to claim 5 is characterized in that: Step 5 includes: (5.1) Add the source domain semantic features and target domain features to form a new feature tensor f u : (5.2) The new feature tensor fu is sent to the classifier for prediction to obtain the category confidence Cu; the confidence matrix composed of the category confidence Cu is transposed to obtain the transposed matrix Transpose the matrix Perform matrix multiplication to obtain the category confusion matrix yu for each pixel of the recombined feature: in, represents matrix multiplication; (5.3) Calculate the mean L1 of the sum of the off-diagonal elements of the confusion matrix for each pixel as the class confusion loss: in, Represents the value of the i-th row and j-th column of the confusion matrix yu, and N represents the number of all categories in the source domain; (5.4) Calculate the category confusion matrix ys of the source domain features, subtract the two category confusion matrices, and obtain the square of the difference between the category confusion matrices y d : and d =(and u -and s ) 2 (13) (5.5) Sum the diagonal differences between the category confusion matrix of the recombined features and the confusion matrix of the source domain features to obtain the sum L2: Where, It represents the value of the i-th row and j-th column of the square yd matrix.

7. The cross-domain remote sensing image semantic segmentation method based on context orthogonal decoupling according to claim 6 is characterized in that: Step 6 includes: (6.1) Calculate the cross entropy loss L for semantic segmentation of source domain images S : in, is the one-hot label of the source domain image, gs is the student model, H is the height of the source domain image, W is the width of the source domain image, C is the number of semantic categories of the source domain image, x s is the source domain image; (6.2) Calculate the semantic segmentation cross entropy loss L of the mixed image mix : Where w is the mixing weight, is the one-hot label of the mixed image, x mix It is a mixed image; (6.3) Calculate the decoupling contrast loss L DC : L DC =λ1L s +λ2L d (17) Among them, λ1 and λ2 are the semantic contrast loss weight and domain contrast loss weight respectively; (6.4) Calculate the category confusion loss L con : <h2 style=";text-align:left;direction:ltr">L<h2 style=";text-align:left;direction:ltr"> con <h2 style=";text-align:left;direction:ltr"> =λ3(L1+L2) (18) Among them, λ3 is the category confusion loss weight; (6.5) Using cross entropy loss L S , semantic segmentation cross entropy loss L mix , decoupled contrast loss L DC and category confusion loss L con The cross-domain remote sensing image semantic segmentation model is trained until the number of iterations meets the preset iteration number requirement.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the method according to any one of claims 1 to 7 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.