Remote sensing image semantic segmentation method based on causal style invariance
Through the causal invariant characterization learning network ICRNet, combined with the style intervention module and the identity regularization module, the generalization problem caused by the space-time heterogeneity of remote sensing images is solved, and the invariant content representation learning and better generalization performance of remote sensing images are achieved.
Patent Information
- Application Number
- CN202510117706.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-13
AI Technical Summary
The existing remote sensing image intelligent understanding technology shows poor generalization performance when facing the differences in visual representations at different times in the same region, and it is difficult to learn invariant characteristics to solve the generalization problem caused by spatial and temporal heterogeneity.
A semantic segmentation method for remote sensing images based on causal style invariance is proposed. By constructing a causal invariant characterization learning network ICRNet, including a style intervention module and an identity regularization module. The style intervention module interferes with the image style by simulating the seasonal changes of remote sensing images, and the identity regularization module forces the model to learn unchanged content features through a comparative learning framework.
Effectively learning the constant content representation of remote sensing images improves the model's generalization ability when facing changes in the style of remote sensing images, and significantly improves the performance and accuracy of semantic segmentation of remote sensing images.
Smart Images

Figure CN119992095A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the field of electronic information technology, and in particular to a remote sensing image semantic segmentation method based on causal style invariance. Background Art
[0002] Remote sensing image understanding tasks are based on image representation learning methods, which automatically learn representation information from remote sensing images and are used in the prediction of various remote sensing vision tasks. The current development process of image representation learning can be summarized as: supervised representation learning stage, self-supervised representation learning stage, and causal invariant representation learning stage.
[0003] Supervised representation learning: Early feature extraction mainly relied on manual features designed by experts based on domain knowledge. Researchers manually extracted features from images through various algorithms and used these features in machine learning models for classification and recognition tasks, such as scale-invariant feature transform (SIFT), histogram-oriented gradients (HOG), and visual word bag models. With the development of deep neural networks, representation learning methods dominated by convolutional neural networks (CNN) automatically learn useful features from images and use them in the prediction of various visual tasks, greatly reducing labor costs. At this time, the representation learning method is mainly a supervised learning paradigm, which adds supervision signals to the model through labeled data to guide the model to automatically learn useful representation information. Typical networks include AlexNet, VGG, ResNet, etc., which have achieved good results in many visual tasks.
[0004] Self-supervised representation learning: The success of supervised learning is inseparable from a large amount of labeled data. However, the acquisition of a large amount of labeled data is often time-consuming and labor-intensive, especially in remote sensing images with wide coverage and rich information. Therefore, self-supervised representation learning methods are widely used in various visual tasks because they do not require a large amount of labeled data and can learn good representation information of images. Self-supervised representation learning methods design a proxy task so that the model can learn useful representations without relying on supervisory signals. The Inpainting method blocks part of the image and predicts its content, so that the model can learn the global structure and semantic information of the image without a large amount of labeled data. SimCLR adopts a contrastive learning framework, generates positive and negative sample pairs through random data enhancement, and uses contrastive loss to maximize the similarity of positive sample pairs and minimize the similarity of negative sample pairs, thereby learning effective image representations without a large amount of labeled data. The MoCo series proposes a dynamic dictionary queue mechanism, using a momentum-updated encoder to maintain a large negative sample dictionary, which solves the problem of insufficient number of negative samples in contrastive learning, allowing the number of negative samples to be greatly increased, and improving the effect of representation learning.
[0005] In the field of remote sensing, some researchers have conducted research based on self-supervised representation learning methods. Li et al. proposed the GLCNet network, which adds a global style contrast learning module and a local feature matching contrast learning module to the traditional contrast learning framework, which can effectively learn the global and local features of remote sensing images. Zhang et al. proposed a contrast learning method based on the gradient-guided sampling strategy, Grass, which improves the representation learning effect of contrast learning by dynamically selecting positive and negative samples using gradient information during the training process, thereby significantly improving the performance and accuracy of semantic segmentation of remote sensing images. Muhtar et al. proposed a new self-supervised dense representation learning method for semantic segmentation of remote sensing images, IndexNet, which learns the pixel-level representation of remote sensing images by tracking the position of objects, and learns spatiotemporal invariant features by combining image-level contrast and pixel-level contrast. Li et al. proposed a self-supervised multi-task representation learning method, SSLR, which effectively captures the visual representation of remote sensing images by designing three different proxy tasks and a triple-joint network to simultaneously learn high-level and low-level image features.
[0006] Causal invariant representation learning: Existing representation learning methods have made a lot of achievements, but there are still many unresolved problems. Invariant representation learning is a difficult point in the field of image understanding. Invariant representation is a relatively stable and robust representation information in visual tasks. Learning invariant representation can effectively improve the generalization performance of the model in visual tasks. Now many researchers have conducted research on invariant representation learning based on causal mechanisms. In order to solve the generalization problem caused by the domain difference between artificial data and real data, Gilhyun et al. proposed a causal invariant learning method GCISG. Using the contrastive learning framework, causal invariant loss was designed to learn style-invariant representation. RELIC explains the contrastive learning method from the perspective of causal representation learning, and designs regularization constraints to force the model to learn invariant representation. The CIRL algorithm is used to learn causal factors in images. First, the non-causal factors are intervened by the causal intervention module, and then the representations of the original image and the enhanced image are sent to the factorization module to force the representation to be separated from the non-causal factors and to be independent of each other. Finally, the adversarial mask module is used to learn the invariant representation information. Summary of the invention
[0007] With the development of remote sensing technology, it is becoming easier and easier to obtain remote sensing images. Remote sensing image intelligent understanding technology based on deep neural networks, with its powerful data processing and feature extraction capabilities, can automatically identify and extract information from remote sensing images, and is widely used in urban planning, ecological monitoring, land use and other fields.
[0008] Due to the interference of various offset factors, the visual representation of remote sensing images of the same area may be different at different times. For example, the color of vegetation and ground objects may be quite different in winter and autumn. In the face of these visual differences, existing remote sensing image intelligent understanding technologies show poor generalization performance. Therefore, one of the main challenges in the current remote sensing image understanding task is how to learn invariant features from remote sensing images to solve the generalization problem caused by the temporal and spatial heterogeneity of remote sensing images.
[0009] Technical solution:
[0010] A remote sensing image semantic segmentation method based on causal style invariance is implemented by constructing a causal invariant representation learning network ICRNet, wherein the causal invariant representation learning network ICRNet includes a style intervention module and an identity regularization module;
[0011] in:
[0012] Style intervention module: intervenes in the style of images by simulating seasonal changes in remote sensing images;
[0013] Identity Regularization Module: Learning Invariant Content Features from Images with Changing Styles.
[0014] Preferably, the style intervention module includes two stages:
[0015] (1) Based on the DiffusionCLIP network, we add seasonal text descriptions as guiding conditions through CLIP, and use remote sensing images to fine-tune the reverse denoising process of the diffusion model, so that the generated remote sensing images have the characteristics of seasonal changes;
[0016] (2) Set the change rules for different objects and use the labels of the dataset to generate the areas that need to be edited in the corresponding images.
[0017] Preferably, in the reverse denoising process of the diffusion model, the generated results of the same time step in the forward process and the reverse process are multiplied by mask and used as the initial value of the next denoising process.
[0018] Specifically, the initial value formula for the next denoising process is as follows:
[0019]
[0020] In the formula, and Indicates the generated results when the time step is T in the reverse process and the forward process; represents the generated result when the time step is T-1 in the denoising process, and m represents the binary mask generated by the label of the dataset.
[0021] Preferably, the changing rules of the different land features are as follows:
[0022]
[0023] Preferably, in the identity regularization module, the following formula is used to make the model predict the same remote sensing objects with different styles but the same content:
[0024]
[0025] Among them, d o (S=s1) means that seasonal offset simulation is performed on various objects in remote sensing images according to the change rules to intervene in their style characteristics.
[0026] Preferably, in the identity regularization module, the symmetric JS divergence is used as the similarity measurement method, and the design style intervention identity regularization R SIIR :
[0027]
[0028] Among them, JS() represents xxxx.
[0029] Preferably, in the framework of self-supervised contrastive learning, R SIIR It is expressed as:
[0030]
[0031] in, It represents the feature information of a pair of positive samples with the same content but different styles extracted by the encoder and prediction head structure. Sim represents the cosine similarity and τ is the temperature coefficient.
[0032] Preferably, the loss function is defined as:
[0033] L=L C +α·R SIIR
[0034] Among them, α is the weight parameter, L C is the contrast loss, and the calculation formula is as follows:
[0035]
[0036] in, It represents the feature information of a pair of positive samples with the same content but different styles extracted by the encoder and prediction head structure. Sim represents the cosine similarity and τ is the temperature coefficient.
[0037] Beneficial effects of the present invention
[0038] This paper proposes a causal invariant representation learning network, ICRNet, which is used to learn invariant content representations from remote sensing images to solve the generalization problem caused by the spatiotemporal heterogeneity of remote sensing images. Based on the assumption of causal mechanism, style intervention identity regularization (SIIR) is designed under the framework of contrastive learning. We use a diffusion model to simulate the seasonal shift of remote sensing images, intervene in the image style, and then regard images with the same content but different styles as positive sample pairs, and force the model to learn invariant content representations through SIIR. Experimental results show that compared with other contrastive learning methods, our proposed method can better learn the invariant content representation of remote sensing images and show better generalization ability when facing style changes of remote sensing images. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 The ideal cause-effect diagram in the specific implementation method is
[0040] Figure 2 The actual task cause-effect diagram in the specific implementation method
[0041] Figure 3 Schematic diagram of ICRNet structure in a specific implementation method
[0042] Figure 4 Schematic diagram of the style intervention fine-tuning stage in a specific implementation method
[0043] Figure 5 Schematic diagram of the style intervention generation stage in a specific implementation method
[0044] Figure 6 The result diagram of the seasonal offset simulation in the embodiment
[0045] Figure 7 This is a visualization of the semantic segmentation results of the autumn data in the embodiment.
[0046] Figure 8 This is the visualization result of ICRNet's attention to different styles of images in the embodiment. DETAILED DESCRIPTION
[0047] The present invention looks at the remote sensing image understanding task from a causal perspective: a remote sensing image contains different ground object entities Ei, and the features of the ground object entities are composed of content features C and style features S. Ideally, the content features are the direct cause variables of the target task and directly determine the results of the target task. The style features are non-causal variables of the target task and will not affect the target task. The causal diagram in the ideal case is as follows: Figure 1 shown.
[0048] In actual tasks, the model is generally used as a predictor of the target task. When learning the feature information of remote sensing images, the model will learn the content features and style features at the same time, making the style features an indirect cause of the target task. The offset factor affects the model's prediction results for the target task by interfering with the style features, resulting in poor generalization performance of the model. Figure 2 shown.
[0049] Therefore, we believe that content features are the invariant representations that need to be learned in remote sensing image understanding tasks, while style is the interference feature that causes the generalization performance of the model to deteriorate. This paper proposes a causal invariant representation learning network, ICRNet, which aims to learn invariant content features from remote sensing images of different styles. The overall structure of the ICRNet network is as follows: Figure 3 shown.
[0050] The ICRNet network mainly consists of two modules:
[0051] (a) The style intervention module mainly intervenes in the style of remote sensing images by simulating seasonal changes in remote sensing images. Based on the powerful generation capability of the diffusion model, we add text to guide the denoising generation process of the diffusion model in combination with CLIP, so that the objects in the final remote sensing images meet the laws of seasonal changes.
[0052] (b) The identity regularization module is mainly used to force the model to learn invariant content features from images with changing styles. Since we cannot directly decouple content features and style features, we use the contrastive learning framework to indirectly make the model learn invariant content features. We use remote sensing images with different styles but the same content as a pair of positive samples, extract deep feature information of the image through the encoder and prediction head structure, and add style intervention identity regularization on the basis of contrast loss to force the model to learn invariant content features.
[0053] 2. Style Intervention Module
[0054] Given that seasonal changes are common in remote sensing image migration and the degree of change is large, we decided to simulate the seasonal changes of remote sensing images through the style intervention module and intervene in the style of the images. The implementation of the style intervention module mainly includes two stages.
[0055] (1) In the first stage, based on the DiffusionCLIP network, we added seasonal text descriptions as guiding conditions through CLIP, and used remote sensing images to fine-tune the reverse denoising process of the diffusion model, so that the generated remote sensing images have the characteristics of seasonal changes. The process is as follows Figure 4 shown.
[0056] (2) In the second stage, considering that different objects in remote sensing images change differently with the seasons, we set the change rules for different objects, as shown in Table 1. Then, with the help of the labels of the dataset, we generate the areas that need to be edited in the corresponding images. In the reverse denoising of the diffusion model, we multiply the generated results of the same time step in the forward process and the reverse process by the mask, and use them as the initial value of the next denoising process. The calculation formula is as follows. After T times of denoising, the local change simulation of the mask area is finally achieved. The process is as follows Figure 5 shown.
[0057]
[0058] In the formula, and Indicates the generated results when the time step is T in the reverse process and the forward process, represents the generated result when the time step is T-1 in the denoising process, and m represents the binary mask generated by the label of the dataset.
[0059] Table 1 Seasonal offset simulation rules
[0060]
[0061] (III) Identity Regularization Module
[0062] In order to make the model learn the content features C of the image as much as possible and ignore the style features S, we set the following identity goal: to make the model predict the same results for remote sensing objects with different styles and the same content. Specifically, it can be expressed as:
[0063]
[0064] Among them, d o (S=s1) indicates that seasonal offset simulation is performed on various objects in the remote sensing image according to Rule Table 1 to intervene in their style characteristics.
[0065] We hope to achieve the goal of formula (2) by adding regularization constraints. Considering that the similarity of a pair of positive samples should be consistent and will not change with the change of calculation order, we abandon the traditional KL divergence and use the symmetric JS divergence as the similarity measurement method, and design Style Intervention Identity Regularization (SIIR):
[0066]
[0067] In the framework of self-supervised contrastive learning, the above formula can be formalized as:
[0068]
[0069] In formula (4), It represents the feature information of a pair of positive samples with the same content but different styles extracted by the encoder and prediction head structure. Sim represents the cosine similarity and τ is the temperature coefficient.
[0070] Combined with the loss of contrastive learning, the final loss function is defined as:
[0071] L=L C +α·R SIIR (5)
[0072] In formula (5), α is a weight parameter, which is generally set to 0.6. C The calculation formula is as follows:
[0073]
[0074] By minimizing the loss function (5), the model is forced to learn invariant content representations from images with varying styles.
[0075] The present invention will be further described below in conjunction with embodiments, but the protection scope of the present invention is not limited thereto:
[0076] (I) Dataset
[0077] We verify the performance of the ICRNet network on three semantic segmentation datasets. The Potsdam dataset is a commonly used dataset in remote sensing image semantic segmentation tasks. Most of the areas in the LoveDA Rural and LandCover datasets are vegetation objects, which are greatly affected by seasonal changes and are suitable for seasonal simulation.
[0078] (1) ISPRS Potsdam Dataset: The Potsdam dataset is a high-resolution remote sensing image dataset provided by ISPRS. It contains 38 aerial images of 6000×6000 pixels with a resolution of 0.05m and six categories: buildings, trees, low vegetation, vehicles, opaque surfaces, and others. We use 24 images as training data and 14 as test data, and crop each image into a small block of 256×256 pixels, finally obtaining 13,824 images for pre-training and 8,064 images for testing. In the fine-tuning stage, we select 1% of the pre-training data, that is, 138 labeled data for fine-tuning.
[0079] (2) LoveDA Rural: The loveDA dataset is a dataset on land cover classification provided by RSIDEA of Wuhan University. It includes two areas: urban and rural areas. Considering that there are more vegetation objects in rural areas and the degree of seasonal changes is large, the data of rural areas are selected for semantic segmentation experimental testing. The image data resolution of rural areas is 0.3m, including farmland, forests, shrubs, roads, buildings and water bodies. After processing, 21,856 training data and 15,872 test data are finally obtained, with a size of 256×256. Similarly, this paper chooses to pre-train 1% size, that is, 218 labeled data are fine-tuned on the semantic segmentation task.
[0080] (3) Deep Globe Land Cover Classification Dataset: LandCover is a remote sensing image dataset for land use classification released by CVPR in 2018. It has a resolution of 0.5m and contains seven categories. After processing, we divided 32,100 training data and 40,200 test data with a size of 256×256. This paper chooses to pre-train 1% size, that is, 321 annotated data for fine-tuning on the semantic segmentation task.
[0081] Table 2 Main information of the dataset
[0082]
[0083] (II) Experimental setup
[0084] (1) Baseline: To verify the ability of ICRNet to learn invariant representations, we compared five contrastive learning methods, two of which are typical contrastive learning methods, and the other three have different improvements in the remote sensing field. In addition, we randomly initialize the model weights and use the same amount of data for fine-tuning as the baseline for the semantic segmentation task.
[0085] 1. Random baseline: Without pre-training, the model weights are randomly initialized, and fine-tuned on the downstream semantic segmentation task using the same amount of data as the other methods.
[0086] 2. Typical contrastive learning method: SimCLR is based on a typical contrastive learning framework, which learns representation information by forcing positive samples to be similar and negative samples to be dissimilar. MoCo v2 adds dynamic queues and momentum encoders to the classic contrastive learning framework, effectively solving the problems of negative sample selection and feature representation stability in contrastive learning.
[0087] 3. Contrastive learning methods in remote sensing: GLCNet adds a global style contrast learning module and a local feature matching contrast learning module to the traditional contrast learning framework, which can effectively learn the global and local features of remote sensing images and perform well in remote sensing image semantic tasks. On the one hand, IndexNet learns the pixel-level representation of remote sensing images by tracking the position of objects, and on the other hand, learns spatiotemporal invariant features by combining image-level contrast and pixel-level contrast. Grass improves the effect of contrast learning by dynamically selecting positive and negative samples using gradient information during training, thereby significantly improving the performance and accuracy of semantic segmentation of remote sensing images.
[0088] (2) Evaluation indicators
[0089] This paper cites OA, Kappa and performance loss rate as evaluation indicators for semantic segmentation tasks.
[0090] OA represents the ratio of the number of pixels correctly predicted as positive examples to the total number of pixels, as shown in the formula:
[0091]
[0092] In the formula, N represents the total number of pixels.
[0093] The Kappa coefficient is used to measure the accuracy of semantic segmentation. The calculation formula is as shown below:
[0094]
[0095] In the formula, p e The calculation formula is shown in formula (9):
[0096]
[0097] In the formula, a c represents the actual number of pixels of class c, b c Represents the number of pixels predicted for class c.
[0098] In order to further reflect the model's ability to learn invariant representation information, this paper uses the performance loss rate as an evaluation indicator to calculate the decrease in semantic segmentation accuracy from the original test set to the test set of different styles. The smaller the value, the less affected the model is by seasonal changes, and the stronger the ability to learn invariant representation is. The calculation formula is shown as follows:
[0099]
[0100] (3) Experimental details
[0101] In order to more fairly reflect the ability of different methods to learn representations, we use the same ResNet50 network as the backbone in the self-supervised pre-training process and pre-train according to each method. Then in the fine-tuning process, we uniformly use the network architecture of Deeplab v3+, load the pre-trained backbone network weights of each method in the Encoder, and use random initialization weights for the Random baseline. Finally, we use the same amount of labeled data to fine-tune on the semantic segmentation task. We conducted experiments on RTX4070Ti, using Adam optimizer to optimize parameters, setting the initial learning rate to 0.01, pre-training for 100 epochs, fine-tuning for 50 epochs, and a batch size of 16.
[0102] For the ICRNet model, we simulated data from three seasons, winter, spring, and summer, on the Potsdam dataset and used them together with the original images as data of different styles for pre-training. On the LoveDA Rural and LandCover datasets, since there are many green vegetation objects in the dataset and the changes are not obvious in spring and summer, we only simulated winter data and used them together with the original images as data of different styles for pre-training. In the seasonal shift simulation, we set the time step of the forward diffusion and reverse diffusion of the diffusion model to 40. Depending on the season, the range of the time step is set to [0,500] or [0,700], and the learning rate is set to 8e-6 during fine-tuning.
[0103] (1) Seasonal shift simulation: We simulated the seasonal changes of the three datasets through the style intervention module to intervene in the style of remote sensing images. In the Potsdam dataset, we simulated data from four seasons: spring, summer, autumn, and winter. In the LoveDARural and LandCover datasets, since there are many green vegetation objects in the datasets and the changes in spring and summer are not obvious, we only simulated winter and autumn data. The results of seasonal simulation are shown in Figure 2. Figure 6 shown.
[0104] (2) Comparison of various models on the autumn style test set: We first generated autumn style test data for the three datasets through seasonal offset simulation. For all models, the autumn style test set is new style data that has not been observed. We test the generalization ability of the model on the autumn style dataset as an indirect evaluation of the model's learning of invariant representations. From the results in Table 1, we can see that compared with other models, our ICRNet network has achieved the highest semantic segmentation accuracy on the three datasets, indicating that the ICRNet network performs better when facing new style data. The semantic segmentation visualization results of the autumn style data are shown in Figure 1. Figure 7As shown in the figure, compared with other models, the segmentation results of the ICRNet network are closer to the labeled data, that is, the model performs better on the autumn dataset. In order to intuitively reflect the generalization performance of the model when the style changes, we compare the performance loss rate of the model from the original test set to the autumn style test set. From the results in Table 2, we can see that when the style of the data changes, the ICRNet network loses the least performance, which shows that ICRNet has the best generalization performance and has the ability to learn invariant content representation.
[0105] Table 3 Kappa and OA accuracy comparison of autumn style validation set
[0106]
[0107] Table 4 Comparison of performance loss rate from the original validation set to the autumn style validation set
[0108]
[0109] (3) Ablation experiment: In order to verify whether the style intervention identity regularization (SIIR) in ICRNet helps the model learn invariant representations, we conducted an ablation experiment to compare the segmentation accuracy of the model with and without SIIR on the autumn style test set. As shown in Table 3, after adding SIIR, ICRNet shows higher segmentation accuracy when facing autumn style data, indicating that SIIR can improve the model's ability to learn invariant representations.
[0110] Table 5 Ablation experiment results of SIIR
[0111]
[0112] (4) Attention maps in different styles: To further verify the ability of ICRNet to learn invariant representations, we visualize the feature attention maps output by ICRNet Backbone in different styles. Taking the trees in the Potsdam dataset as an example, the attention visualization results are shown in Figure 2. Figure 8 From the results of the attention map, we can see that in the three different styles, the upper right corner area where the trees are located has achieved a high degree of attention, indicating that the model has indeed learned the invariant representation information.
[0113] The experimental results show that when faced with new style data, the semantic segmentation accuracy of ICRNet is higher than other self-supervised representation learning methods, and the model has the least performance loss from the original data to the autumn style data. The experiment shows that ICRNet can learn invariant content representation and show good generalization ability.
[0114] The specific embodiments described herein are merely examples of the spirit of the present invention. Those skilled in the art may make various modifications or additions to the specific embodiments described or replace them in similar ways, but they will not deviate from the spirit of the present invention or exceed the scope defined by the appended claims.
Claims
1. A remote sensing image semantic segmentation method based on causal style invariance, characterized by It is achieved by constructing a causal invariant representation learning network ICRNet, which includes a style intervention module and an identity regularization module; wherein: Style intervention module: intervenes in the style of images by simulating seasonal changes in remote sensing images; Identity Regularization Module: Learning Invariant Content Features from Images with Changing Styles.
2. The method according to claim 1, characterized in that The style intervention module consists of two stages: (1) Based on the DiffusionCLIP network, we add seasonal text descriptions as guiding conditions through CLIP, and use remote sensing images to fine-tune the reverse denoising process of the diffusion model, so that the generated remote sensing images have the characteristics of seasonal changes; (2) Set the change rules for different objects and use the labels of the dataset to generate the areas that need to be edited in the corresponding images.
3. The method according to claim 2, characterized in that In the reverse denoising process of the diffusion model, the generated results of the same time step in the forward process and the reverse process are multiplied by the mask and used as the initial value of the next denoising process.
4. The method according to claim 3, characterized in that The initial value formula for the next step of the denoising process is as follows: In the formula, and Indicates the generated results when the time step is T in the reverse process and the forward process; represents the generated result when the time step is T-1 in the denoising process, and m represents the binary mask generated by the label of the dataset.
5. The method according to claim 2, characterized in that The specific change rules of different landforms are as follows:
6. The method according to claim 1, characterized in that In the identity regularization module, the following formula is used to make the model predict the same remote sensing objects with different styles but the same content: Among them, do(S=s) means to simulate the seasonal shift of the objects in the remote sensing images according to the change rules and intervene in their style characteristics, s1 and s2 represent two different seasons; Indicates the degree of influence of the terrain features in season s on the target task.
7. The method according to claim 6, characterized in that In the identity regularization module, the symmetric JS divergence is used as the similarity measurement method, and the design style intervention identity regularization R SIIR : Among them, JS() represents JS divergence, and the JS divergence of distributions P1 and P2 can be expressed as: Where KL() represents KL divergence.
8. The method according to claim 7, characterized in that In the framework of self-supervised contrastive learning, R SIIR It is expressed as: Among them, M represents the number of all negative samples, It represents the feature information of a pair of positive samples with the same content but different styles extracted by the encoder and prediction head structure. Sim represents the cosine similarity and τ is the temperature coefficient.
9. The method according to claim 6, characterized in that The loss function is defined as: L=L C +α·R SIIR Among them, α is the weight parameter, L C is the contrast loss, and the calculation formula is as follows: Among them, M represents the number of all negative samples, It represents the feature information of a pair of positive samples with the same content but different styles extracted by the encoder and prediction head structure. Sim represents the cosine similarity and τ is the temperature coefficient.