A Remote Sensing Image Change Detection Method and System Based on Large Model Domain Adaptation

The proposed method addresses domain disparities and feature integration issues in SAM2-based remote sensing change detection by using a layer-wise low-rank adaptive encoder and difference adaptive enhancement module, enhancing sensitivity and accuracy in change detection.

CN119625540BActive Publication Date: 2025-07-15XIAN UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510146940.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-07-15
Estimated Expiration
2045-02-11

AI Technical Summary

Technical Problem

When the SAM2 model is used for remote sensing image change detection in the prior art, there are problems of domain differences and boundary displacements, and it is difficult to effectively distinguish the slight changes in the remote sensing image and integrate the different spatial levels and semantic particle size characteristics of the bi-time phase image.

Method used

The hierarchical low-rank adaptive encoder, differential adaptive enhancement module and residual convolution decoder are adopted to adaptively adjust the feature extraction and fusion of the SAM2 model through low-rank matrix adaptive adjustment, and differential information is captured in combination with global and local details to generate high-quality change detection maps.

Benefits of technology

It realizes efficient and accurate detection of remote sensing image changes, improves the sensitivity and detection accuracy of the model in remote sensing scenarios, and reduces boundary displacement and in-class inconsistencies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119625540B_ABST
    Figure CN119625540B_ABST
Patent Text Reader

Abstract

The present application discloses a remote sensing image change detection method and system based on large model domain adaptation. The method includes the following steps: S1: Input a pair of dual-temporal remote sensing images to be detected into a trained remote sensing image change detection network; S2: Output a change detection map. The present application belongs to the technical field of remote sensing image change detection. The present application solves the problems of domain difference and boundary displacement existing when the SAM2 model is currently used in the change detection task. Through a hierarchical low-rank adaptation strategy, the present application overcomes the knowledge difference between SAM2 and the change detection task. By introducing a low-rank matrix into the key layer of SAM2, the model is guided to learn remote sensing domain knowledge, realizing the domain adaptation of SAM2 to remote sensing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of remote sensing image change detection, and particularly relates to a remote sensing image change detection method and system based on large model domain adaptation. Background Art

[0002] Remote sensing image change detection aims to obtain the change information between two images taken at the same geographical location at different times, and plays an important role in many fields such as disaster monitoring, urban planning, and environmental investigation. It has become an important research hotspot in remote sensing earth observation technology. Traditional change detection methods rely on manual feature extraction and are difficult to achieve efficient and accurate change recognition in complex scenarios.

[0003] Recently, vision foundation models have driven a transformation in the change detection paradigm due to their task-agnostic advantages in different downstream tasks. Different from other models, vision foundation models, driven by their good generalization performance and few-shot learning ability, make it increasingly attractive to use vision foundation models to process change detection tasks. The Segment Anything Model 2 (SAM2) is a segmentation model recently developed by Meta, and its core idea is to predict the segmentation mask of the target according to the prompts provided by the user. After high-quality training on a large-scale natural dataset, SAM2 has demonstrated strong zero-shot generalization and feature representation capabilities in various vision tasks. However, when applying the SAM2 model to the change detection task, the following problems exist: (1) Domain difference: At the data level, there are significant differences between remote sensing images and natural images in terms of spatial resolution, imaging conditions, and object scales. Remote sensing images usually have lower resolutions, and there is a large spectral overlap between ground objects, which makes it difficult for the model to distinguish foreground and background. When directly applying SAM2 to the remote sensing change detection task, the network often relies on its inherent knowledge in natural scenes, which limits its sensitivity to subtle changes in remote sensing scenes. (2) Boundary displacement: At the feature level, the SAM2 model is difficult to integrate different spatial levels and semantic granularity features of dual-temporal images. Simple feature fusion methods (such as splicing) fail to fully exploit the temporal differences of dual-temporal features. Summary of the Invention

[0004] The purpose of the present invention is to provide a remote sensing image change detection method and system based on large model domain adaptation, which solves the problems of domain difference and boundary displacement existing when the SAM2 model is used in the change detection task in the prior art.

[0005] The present application provides a technical solution: A remote sensing image change detection method based on large model domain adaptation, comprising the following steps:

[0006] S1: Input a pair of dual-temporal remote sensing images to be detected into a trained remote sensing image change detection network;

[0007] The change detection network includes the following steps:

[0008] S11: Hierarchical low-rank adaptive encoder. First, introduce a low-rank matrix into the self-attention layer and the linear layer of the SAM2 image encoder, and then extract multi-scale features of a pair of multi-temporal remote sensing images ; Extract multi-scale features of a pair of multi-temporal remote sensing images through the hierarchical low-rank adaptive encoder , where i = 1, 2, 3, 4.

[0009] S12: Difference adaptive enhancement module. First, upsample the multi-scale features of the S11 multi-temporal remote sensing images , i = 2, 3, 4 to the size of , and then adaptively fuse each group of to generate 4 groups of difference features; Through the difference adaptive enhancement module, fuse the multi-scale features of a pair of multi-temporal remote sensing images in S11 , i = 1, 2, 3, 4.

[0010] S13: Residual convolutional decoder. Decode the 4 groups of difference features generated by S12 and output a change map; The output change map contains two channels. The first channel corresponds to the probability of the non-change class, and the second channel corresponds to the probability of the change class; Use the argmax operation along the channel dimension to obtain a binary change map.

[0011] S2: Output the change detection map.

[0012] Preferably, the size of each group of difference features is the same as .

[0013] Preferably, the hierarchical low-rank adaptive encoder in S11 specifically includes the following steps:

[0014] Select the SAM2 image encoder as the feature extractor of the multi-temporal remote sensing images.

[0015] In each two Transformer modules of the SAM2 image encoder, introduce a low-rank adaptive strategy into its self-attention layer and multi-layer perceptron layer.

[0016] Preferably, the low-rank adaptive strategy includes:

[0017] Given the weight matrix of a certain layer of the Transformer module , and represents the dimension of the weight matrix ; Add a branch on one side of , and this branch is composed of two low-rank matrices decomposed by , consists of, where r represents the weight matrix and rank, and ; During training, freeze the weights of , , replace with the update; Given the input feature of this layer as , the output sequence , this process is expressed as:

[0018] (1)

[0019] (2)

[0020] In Equation 1 and Equation 2, is the weight matrix after introducing LoRA for this layer, is the weight matrix that replaces during training for update.

[0021] Preferably, the S12 differential adaptive enhancement module includes the following steps:

[0022] The differential adaptive enhancement module captures differential information through two branches: global differential perception and local detail optimization. The specific steps of the two branches are as follows:

[0023] In the global differential perception branch, first add the in S12 with the element to get , then perform global average pooling operation on each channel feature element. The size of changes from [C, H, W] to [C, 1, 1], where C represents the number of feature channels, and H and W represent the height and width of the feature map respectively. The global attention feature is calculated as:

[0024] (3)

[0025] (4)

[0026] In Equation 3, represents element-wise addition. In Equation 4, GAP represents the global average pooling operation, represents 1×1 convolution operation, and BN represents batch normalization operation, ; represents the global attention feature, represents the input feature of the global differential perception branch;

[0027] In the local detail optimization branch, first, subtract the from the element and take the absolute value to obtain . The max pooling operation takes the maximum value in the local receptive field, and the features processed in each layer have the same size as the . The local attention feature is calculated as:

[0028] (5)

[0029] (6)

[0030] In Equation 5, represents element-wise subtraction, represents the absolute value operation, and in Equation 6, MAP represents the max pooling operation; ;

[0031] Then, add to the feature and send it to the Sigmoid function to obtain the attention weight matrices and ;

[0032] Finally, and serve as the measurement coefficients of global attention and local attention.

[0033] Preferably, the attention weight matrices and are calculated specifically as follows:

[0034] (7)

[0035] (8)

[0036] The extracted spatial detail features and class discriminant features are multiplied by their respective relevant weights at the pixel level, and then weighted summation is performed at the pixel level. The enhanced feature of each layer is calculated by the following formula:

[0037] (9)

[0038]

[0039] In Equation 9, represents the feature aggregation strategy of adaptive weights, represents element-wise multiplication;

[0040] Divide the samples containing labels into a training set, a validation set, and a test set; at the beginning of training, the parameters of the SAM2 image encoder itself are kept frozen and do not participate in the update during training; Are initialized to 0 and 1 respectively, Represent the low-rank matrix decomposed by And Represents the weight matrix of a certain layer in the transformer. And The rank value of is set to 16; the parameters of the differential adaptive enhancement module and the residual convolution decoder are randomly initialized; during training, data augmentation strategies of random flipping and random rotation are adopted.

[0041] Preferably, the optimization strategy of the network is as follows:

[0042] During training, optimize the performance of the network by minimizing the cross-entropy loss; for the binary change detection problem, the mathematical expression of the cross-entropy loss function is:

[0043] (10)

[0044] In Equation (10), Represents the class label of the true sample, Represents the predicted class label, and N represents the total number of pixels of each sample; Is an index to measure the difference between the predicted value and the true value of the model. The smaller the value of the loss function, the closer the predicted value of the model is to the true value, and the better the performance of the model.

[0045] The present invention also provides another technical solution:

[0046] A remote sensing image change detection system based on large model domain adaptation, the system includes: an image input module, a network composition and processing module, and an image output module.

[0047] Image input module: Input the dual-temporal remote sensing image into the change detection network.

[0048] Network composition and processing module: Include: a hierarchical low-rank adaptive encoder, a differential adaptive enhancement module, and a residual convolution decoder.

[0049] The hierarchical low-rank adaptive encoder extracts multi-scale features (i = 1, 2, 3, 4) by introducing a low-rank matrix into the self-attention and linear layers of the SAM2 image encoder; the differential adaptive enhancement module generates 4 groups of differential features by upsampling and adaptively fusing the multi-scale features (i = 2, 3, 4); the residual convolution decoder decodes the differential features to obtain a change map containing two types of probabilities, and converts the change map into a binary change map using the argmax operation.

[0050] The image output module is used to output the change detection map.

[0051] The beneficial effects of the present invention are as follows: Through the collaborative work of three main modules, namely the hierarchical low-rank adaptive encoder, the differential adaptive enhancement module, and the residual convolution decoder, the remote sensing image change detection network of the present invention can achieve the change detection of dual-temporal remote sensing images. By using technologies such as multi-scale feature extraction, feature fusion, and residual convolution decoding, it can effectively extract and visualize the change information of the images. This method combines low-rank matrices, adaptive enhancement, and residual structures, and has certain innovation and practicality in the field of remote sensing image analysis. Description of the Drawings

[0052] Figure 1 It is a schematic flow framework diagram of the remote sensing image change detection method based on large model domain adaptation of the present application;

[0053] Figure 2 It is a schematic flow framework diagram of the hierarchical low-rank adaptive strategy of the present application;

[0054] Figure 3 It is a schematic flow framework diagram of the differential adaptive enhancement module of the present application;

[0055] Figure 4 It is a detection result map obtained by the remote sensing image change detection method based on large model domain adaptation of the present application. Detailed Embodiments

[0056] Next, the technical solutions in the embodiments of the present application will be described clearly and completely. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0057] As Figure 1 shown, a remote sensing image change detection method based on large model domain adaptation of the present application includes the following steps:

[0058] S1: Input a pair of dual-temporal remote sensing images to be detected into the trained remote sensing image change detection network.

[0059] S2: Output the change detection map.

[0060] Among them, the change detection network specifically includes:

[0061] S11. The hierarchical low-rank adaptive encoder, by extracting multi-scale features of a pair of dual-temporal remote sensing images (i = 1, 2, 3, 4); pre represents the previous moment, and post represents the subsequent moment; first, introduce the low-rank matrix into the self-attention layer and linear layer of the SAM2 image encoder, and then perform feature extraction.

[0062] S12. The differential adaptive enhancement module, by fusing the multi-scale features of a pair of dual-temporal remote sensing images (i = 1, 2, 3, 4); first (i = 2, 3, 4) are upsampled to the size of, and then each group of is adaptively fused to generate 4 groups of differential features with rich semantic information. The size of each group of differential features is the same as the same.

[0063] S13. The residual convolutional decoder, by decoding 4 groups of differential features and outputting the change map. Among them, the first channel corresponds to the probability of the non-change class, and the second channel corresponds to the probability of the change class. The binary change map is obtained by using the argmax operation along the channel dimension.

[0064] The training process of the change detection network is as follows:

[0065] Divide the samples containing labels into training set, validation set, and test set.

[0066] At the beginning of training, the parameters of the SAM2 image encoder itself are kept frozen and do not participate in the update during the training process; are initialized to 0 and 1 respectively, represents the weight matrix of a certain layer in the transformer, represents the low-rank matrix decomposed from , and the rank of is set to 16. The parameters of the differential adaptive enhancement module and the residual convolutional decoder are randomly initialized. During the training process, data augmentation strategies of random flipping and random rotation are adopted.

[0067] Embodiment

[0068] The remote sensing image change detection method based on large model domain adaptation in this application specifically includes the following steps:

[0069] Step 1: Construct the training samples of the change detection network, including the training set, validation set, and test set. Before training, normalize the dual-temporal images and labels of the training set to between 0 and 1.

[0070] Step 2: Select the SAM2 image encoder as the feature extractor of the dual-temporal images.

[0071] Step 3: Construct a remote sensing image change detection network based on large model domain adaptation, including a hierarchical low-rank adaptation encoder, a difference adaptive enhancement module, and a residual convolutional decoder. The processing procedures corresponding to the hierarchical low-rank adaptation encoder, the difference adaptive enhancement module, and the residual convolutional decoder are as follows:

[0072] Select the SYSU-CD, WHU-CD, and LEVIR-CD datasets, and provide change labels (i.e., annotation maps for the changed area and the unchanged area) for each pair of dual-temporal images. Divide the remote sensing image dataset into three parts: a training set, a validation set, and a test set.

[0073] The training set is used for model training, the validation set is used to adjust hyperparameters and prevent overfitting. The test set is used to evaluate the performance of the final model. Normalize the dual-temporal images and their labels in the training set, validation set, and test set. Convert the pixel values of the dual-temporal images from the original numerical range to values between -0.5 and 0.5 to improve the training stability and convergence speed.

[0074] Construct a hierarchical low-rank adaptation encoder. After high-quality learning on a large-scale natural dataset, SAM2 has excellently completed various downstream tasks. However, the domain difference between natural images and remote sensing images limits its performance in change detection tasks. To extract the inductive bias of remote sensing images and further enhance the generalization ability of SMA2, this application sets a hierarchical low-rank adaptation strategy based on the parameter-efficient transfer learning method to fine-tune the encoder network of SAM2.

[0075] As Figure 2 shown, this application introduces low-rank adaptation in the self-attention layer and the perceptron layer of the Transformer module in the SAM2 image encoder: in the self-attention layer, adjust the weight matrix through low-rank to optimize the adaptive weight allocation ability of the attention mechanism, so that the model can more accurately capture the global relationship between input features; in the multi-layer perceptron layer, perform non-linear transformation on the features. After introducing low-rank adaptation, it can more effectively adjust the feature expression during the transformation process.

[0076] Given the weight matrix of a certain layer of the Transformer module, R represents the feature space, and represent the dimensions of the weight matrix . Add a branch on one side of , and this branch is composed of two low-rank matrices decomposed by , . Let r represent the ranks of the weight matrices and , and ; during the training process, freeze the weights of , , Replacement update; given the input features of this layer as , the output sequence is , and this process is expressed as:

[0077] (1)

[0078] (2)

[0079] In Equations 1 to 2, is the weight matrix after introducing LoRA for this layer, is the weight matrix that replaces during training. In the multi-head self-attention mechanism of the Transformer module, the cosine similarity between different positions is calculated to determine which features should be weighted. This application applies the low-rank adaptation strategy to the projection layers of the query ( ) and value ( ), affecting the attention scores calculated in the self-attention mechanism, thereby adjusting the model's attention to different regions. Therefore, the attention calculation process after introducing the low-rank adaptation strategy is as follows:

[0080] (3)

[0081] (4)

[0082] (5)

[0083] (6)

[0084] In Equations 3 to 6, , and are the query, key, and value matrices respectively. , , are the frozen projection layers of SAM2 , weight matrices, , , , are trainable low-rank matrices, and F is the multi-scale context features extracted by the SAM2 image encoder.

[0085] In the MLP layer of the Transformer module, the nonlinear transformation allows the model to learn more complex feature representations. This application introduces low-rank adaptation in the first hidden layer of the MLP to capture the complex relationships between input features. The calculation process of the linear layer after introducing LoRA is as follows:

[0086] (7)

[0087] (8)

[0088] Among them, is the input sequence of the MLP layer, represents the weight matrix after introducing a low-rank matrix in the first linear layer of the MLP layer, , and , are respectively the weight matrices of linear layer 1 and linear layer 2 in the MLP layer, are the decomposed low-rank matrices, , represent the bias terms.

[0089] Construct a differential adaptive enhancement module: This application proposes a differential adaptive enhancement module. Different from the single-branch feature fusion strategy, the differential adaptive enhancement module realizes the accurate capture of dual-temporal difference information through two branches of global difference perception and local detail optimization.

[0090] The two branches of the differential adaptive enhancement module respectively use global attention and local attention to capture category discriminative features and spatial detail features, output their attention weights and adaptively fuse them.

[0091] As Figure 3 shown, the multi-scale features extracted by the encoder are respectively denoted as (i = 1, 2, 3, 4), and the differential features enhanced by the differential adaptive enhancement module are denoted as (i = 1, 2, 3, 4).

[0092] The operation steps of the differential adaptive enhancement module are specifically as follows:

[0093] Add and element-wise to obtain , subtract from element-wise to obtain . Then, input and into the global difference perception and local detail optimization branches respectively, considering both local and global context information.

[0094] In the global difference perception branch, first perform global average pooling operation on each channel feature element to average, The size changes from [C, H, W] to [C, 1, 1], where C represents the number of feature channels, and H and W represent the height and width of the feature map respectively. The global attention feature is calculated as:

[0095] (9)

[0096] (10)

[0097] In Equations (9) and (10), denotes element-wise addition, GAP represents the global average pooling operation, denotes the 1×1 convolution operation, and BN represents the batch normalization operation, , represents the global attention feature. represents the input feature of the global difference perception branch.

[0098] In the local detail optimization branch, first, the in S12 is subtracted from the element by element, and the absolute value is taken to obtain . The maximum pooling operation takes the maximum value of the local receptive field, and the feature processed in each layer has the same size as . The local attention feature

[0099] (11)

[0100] (12)

[0101] In Equations (11) and (12), denotes element-wise subtraction, denotes the absolute value operation, MAP represents the maximum pooling operation, represents the input feature of the local detail optimization branch.

[0102] Then, is added to the feature, and the result is fed into the Sigmoid function to obtain the attention weight matrix and . Specifically as follows:

[0103] (13)

[0104] (14)

[0105] Finally, is As the measurement coefficients of global attention and local attention. Specifically, the extracted spatial detail features and category discrimination features are multiplied by their respective relevant weights at the pixel level, and then weighted summation is performed at the pixel level. The enhanced features of each layer can be calculated by the following formula:

[0106] (15)

[0107]

[0108] where represents the feature aggregation strategy of adaptive weights represents element-wise multiplication.

[0109] The optimization strategy of the network is as follows:

[0110] During the training process, the performance of the network is optimized by minimizing the cross-entropy loss. For the binary change detection problem, the mathematical expression of the cross-entropy loss function is:

[0111] (16)

[0112] where is the class label of the true sample, is the predicted class label, and N is the total number of pixels of each sample. is an index to measure the difference between the predicted value and the true value of the model. The smaller the value of the loss function, the closer the predicted value of the model is to the true value, and the better the performance of the model.

[0113] The details during the training process are as follows:

[0114] The SAM2-large image encoder is used as the feature extractor, and multi-scale features are output at the 2nd, 8th, 44th, and 48th transformer modules to ensure the integrity of the feature space and semantic information. To obtain the best performance of the model, hierarchical low-rank adaptation is introduced once every two Transformer Blocks in the encoder, and the rank of the decomposition matrix is set to 16. Data augmentation strategies of random flipping and random rotation are used during the training process, and the backbone part of the network is kept frozen. The network is trained on a single Nvidia 4090 GPU, the low-rank matrices are initialized to 0 and 1 respectively, and the remaining trainable parameters are randomly initialized. According to experience, the learning rate is set to 2.1e-4, and the model is trained for 150 epochs. The learning rate is linearly decayed until the last epoch, and the AdamW optimizer with a weight decay of 0.01 and beta values of (0.9, 0.999) is used.

[0115] To more intuitively illustrate the effectiveness of the remote sensing image change detection method based on large model domain adaptation provided in this application, a fully supervised binary change detection experiment was conducted. The evaluation metrics include: F1 score (F1), precision (Pre.), recall (Rec.), overall accuracy (OA), and intersection over union (IoU).

[0116] Among them, true positive (TP) represents the number of pixels where the change is correctly detected, false positive (FP) represents the number of pixels that actually have no change but are wrongly detected as having changed, true negative (TN) represents the number of pixels that have no change and are correctly detected as not having changed, and false negative (FN) represents the number of pixels that actually have changed but are wrongly detected as not having changed.

[0117] (17)

[0118] (18)

[0119] (19)

[0120] (20)

[0121] (21)

[0122] As shown in Table 1 below, a performance comparison of the remote sensing image change detection method based on large model domain adaptation in this application with other detection methods on the SYSU-CD dataset is given. Higher IoU and F1 indicate better detection effects, and the best results in each column are in bold font.

[0123] Table 1 Quantitative comparison of the remote sensing image change detection method based on large model domain adaptation in this application with other methods on the SYSU-CD and WHU-CD datasets

[0124]

[0125] The remote sensing image change detection method based on large model domain adaptation in this application performs best on the SYSU-CD dataset. Compared with BAN, recall, IoU, F1 score, and OA are improved by 7.73%, 5.50%, 3.80%, and 0.91% respectively. The above experimental results show that the remote sensing image change detection method based on large model domain adaptation provided in this application has the best performance compared with other change detection networks, verifying that the detection method provided in this application is the most effective.

[0126] To further verify the superiority of the remote sensing image change detection method based on large model domain adaptation in this application, the following was conducted Figure 4From the visual analysis shown, it can be observed that most methods have boundary offsets and intra-class inconsistencies in detecting large-scale objects, while the remote sensing image change detection method based on large model domain adaptation provided in this application shows integrity inside and smooth boundaries. In areas where change objects are dense, the results detected by the remote sensing image change detection method based on large model domain adaptation in this application are more complete than other methods, and there is basically no missed detection or false detection.

[0127] This application provides a remote sensing image change detection system based on large model domain adaptation, which includes: an image input module, a network composition and processing module, and an image output module.

[0128] Image input module: By inputting dual-temporal remote sensing images into the change detection network.

[0129] Network composition and processing module: Includes: a hierarchical low-rank adaptive encoder, a differential adaptive enhancement module, and a residual convolutional decoder.

[0130] The hierarchical low-rank adaptive encoder extracts multi-scale features (i = 1, 2, 3, 4) by introducing low-rank matrices into the self-attention and linear layers of the SAM2 image encoder; the differential adaptive enhancement module generates 4 groups of differential features by upsampling and adaptively fusing the multi-scale features (i = 2, 3, 4); the residual convolutional decoder decodes the differential features to obtain a change map containing two types of probabilities, and converts the change map into a binary change map using the argmax operation.

[0131] The image output module is used to output the change detection map.

[0132] The hierarchical low-rank adaptive strategy of this application overcomes the knowledge difference between SAM2 and the change detection task. By introducing low-rank matrices into the key layers of SAM2, it guides the model to learn remote sensing domain knowledge and realizes the domain adaptation of SAM2 to remote sensing.

[0133] The differential adaptive enhancement module of this application effectively coordinates the spatial detail information and the class discrimination information. By generating the attention weights of the two, it adaptively enhances the differential information and suppresses the boundary displacement and intra-class inconsistency phenomena.

[0134] Experimental results on multiple benchmark datasets show that the remote sensing image change detection method based on large model domain adaptation in this application is superior to other state-of-the-art change detection networks in multiple metrics. In addition, the dual-branch structure of the remote sensing image change detection method based on large model domain adaptation in this application can be easily extended to visual tasks such as RGB-T / D semantic segmentation and MRI-CT image fusion.

[0135] Although the content of the present application has been described in detail through the above preferred embodiments, it should be recognized that the above description should not be considered as a limitation of the present application. After those skilled in the art have read the above content, various modifications and alternatives to the present application will be obvious. Therefore, the scope of protection of the present application should be defined by the appended claims.

Claims

1. A remote sensing image change detection method based on large model domain adaptation, characterized in that It includes the following steps: S1: Input a pair of double-time remote sensing images to be detected into a trained remote sensing image change detection network; The change detection network includes the following steps: S11: Hierarchical Low-Rank Adaptive Encoder. First, introduce a low-rank matrix into the self-attention layer and the linear layer of the SAM2 image encoder, and then extract multi-scale features of a pair of dual-temporal remote sensing images and ; extract multi-scale features of a pair of dual-temporal remote sensing images through the Hierarchical Low-Rank Adaptive Encoder and , where i = 1, 2, 3, 4; S12: Differential Adaptive Enhancement Module. First, the multi-scale features of the dual-temporal remote sensing images and , where i = 2, 3, 4, are upsampled to and in terms of size. Then, for each group of and , they are adaptively fused to generate 4 groups of differential features. Through the Differential Adaptive Enhancement Module, the multi-scale features of a pair of dual-temporal remote sensing images in S11 and , where i = 1, 2, 3, 4, are fused. The S12 difference adaptive enhancement module includes the following steps: The difference adaptive enhancement module captures difference information through two branches of global difference perception and local detail optimization. The specific steps of the two branches are as follows: In the global difference perception branch, first add the and elements in S12 to obtain . Then, perform global average pooling operation on each channel feature element. The size of changes from [C, H, W] to [C, 1, 1], where C represents the number of feature channels, and H and W represent the height and width of the feature map respectively; the global attention feature is calculated as: (3) (4) In formula (3), represents element-wise addition. In formula (4), GAP represents the global average pooling operation, represents the 1×1 convolution operation, and BN represents the batch normalization operation, ; represents the global attention feature, represents the input feature of the global difference perception branch; In the local detail optimization branch, first subtract the from the element and take the absolute value to obtain . The max pooling operation takes the maximum value of the local receptive field, and the features processed in each layer have the same size as . The local attention feature is calculated as: (5) (6) In formula (5) represents element-wise subtraction, represents the absolute value operation, and MAP in formula (6) represents the max pooling operation; ; Then, add and features, and send the result into the Sigmoid function to obtain the attention weight matrix and ; Finally, and serve as the measurement coefficients for global attention and local attention; S13: A residual convolutional decoder decodes the 4 groups of difference features generated by S12 and outputs a change map; the output change map contains two channels. The first channel corresponds to the probability of the non-change class, and the second channel corresponds to the probability of the change class; Use the argmax operation to obtain a binary change map along the channel dimension; S2: Output a change detection map.

2. The remote sensing image change detection method based on large model domain adaptation according to claim 1, characterized in that The dimensions of each set of the four sets of differential features are the same as and identical.

3. The remote sensing image change detection method based on large model domain adaptation according to claim 1, wherein The S11 hierarchical low-rank adaptive encoder specifically includes the following steps: Select the SAM2 image encoder as the feature extractor for the double-time remote sensing images; In the Transformer modules of every two SAM2 image encoders, introduce a low-rank adaptive strategy into its self-attention layer and multi-layer perceptron layer.

4. The method for remote sensing image change detection based on large model domain adaptation according to claim 3, wherein, The low-rank adaptive strategy includes: Given the weight matrix of a certain layer of the Transformer module , representing the real number space, and denoting the dimension of the weight matrix ; adding a branch on one side, and this branch consists of two low-rank matrices decomposed by , , where r represents the rank of the weight matrix and ; during the training process, freeze the weights of , and , replace with their updates; given the input feature of this layer as , and the output sequence , this process is expressed as: (1) (2) In Equations (1) and (2), is the weight matrix after introducing LoRA to this layer, is the weight matrix that replaces during training for updating.

5. The remote sensing image change detection method based on large model domain adaptation according to claim 1, characterized in that The attention weight matrix and is calculated as follows: (7) (8) The extracted spatial detail features and class discrimination features are respectively multiplied by their relevant weights at the pixel level, and then weighted summation is performed at the pixel level. The enhanced features of each layer are calculated by the following formula: (9) In formula (9), represents the feature aggregation strategy of the adaptive weight, represents the element-wise multiplication; Divide the samples containing tags into a training set, a validation set, and a test set; at the start of training, the parameters of the SAM2 image encoder itself are kept frozen and do not participate in the update during the training process; and are initialized to 1 and 0 respectively, and the rank value is set to 16; the parameters of the differential adaptive enhancement module and the residual convolutional decoder are randomly initialized; during the training process, data augmentation strategies of random flipping and random rotation are adopted.

6. The remote sensing image change detection method based on large model domain adaptation according to claim 5, characterized in that, The optimization strategy of the change detection network is as follows: During the training process, optimize the performance of the network by minimizing the cross-entropy loss; for the binary change detection problem, the mathematical expression of the cross-entropy loss function is: (10) In Equation (10), represents the class label of the real sample, represents the predicted class label, and N represents the total number of pixels of each sample; is an index to measure the difference between the predicted value and the true value of the model. The smaller the value of the loss function, the closer the predicted value of the model is to the true value, and the better the performance of the model.

7. A system for applying the remote sensing image change detection method based on large model domain adaptation according to any one of claims 1-6, characterized in that, The system includes: an image input module, a network composition and processing module, and an image output module; The image input module: Input the double-time remote sensing images into the change detection network; The network composition and processing module: includes: a hierarchical low-rank adaptive encoder, a difference adaptive enhancement module, and a residual convolutional decoder; The hierarchical low-rank adaptive encoder extracts multi-scale features by introducing a low-rank matrix into the self-attention and linear layers of the SAM2 image encoder; the difference adaptive enhancement module generates 4 groups of difference features by upsampling and adaptively fusing the multi-scale features; the residual convolutional decoder decodes the difference features to obtain a change map containing two types of probabilities, and uses the argmax operation to convert the change map into a binary change map; The image output module is used to output the change detection map.

Citation Information

Patent Citations

  • Remote sensing image change detection method and system based on semantic fusion

    CN119068351A

  • Remote sensing image segmentation method and device

    CN119152213A