Remote sensing image change detection method and system based on measurement calibration adaptive SAM2 network

By introducing adapters and semantic metric calibration mechanisms into the SAM2 network, the problem of non-semantic change interference in complex scenarios is solved, and the accuracy and reliability of high-resolution remote sensing image change detection is achieved.

CN120236177APending Publication Date: 2025-07-01Chinese People's Liberation Army Cyberspace Force Information Engineering University
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510301449.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The existing methods of remote sensing image change detection using SAM2 are difficult to effectively identify non-semantic changes in complex scenarios, resulting in more pseudo-changes in the change detection results.

Method used

The remote sensing image change detection method based on the metric calibration adaptive SAM2 network is adopted, and the accuracy of change detection is improved by introducing the first and second adapters to adaptive adjustments between the encoder and the decoder, and semantic measurements and change calibrations are performed in the change perceptron.

Benefits of technology

In complex scenarios, the accuracy of change detection is significantly improved, and the pseudo-change in the change detection results is reduced. It is suitable for high-resolution remote sensing image change detection tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236177A_ABST
    Figure CN120236177A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of remote sensing image processing, in particular to a remote sensing image change detection method and system based on a metric calibration adaptive SAM2 network, and the method comprises the steps: obtaining a change result in a dual-time-phase remote sensing image of a target region through a remote sensing image change detection model, the model comprises an encoder used for extracting multi-scale features of a dual-time-phase image block, the decoder is used for recovering, decoding and outputting the multi-scale features of the double-time-phase image blocks layer by layer; the change sensor is used for carrying out semantic measurement and change calibration on the decoded and output features of the double-time-phase image blocks and outputting a semantic measurement result and a change mapping result through up-sampling; the encoder is internally provided with a first adapter for gradually optimizing and adjusting the semantic encoding process of the image blocks in the extraction of the multi-scale features, and a second adapter for adaptively adjusting the change detection task and carrying out feature dimension reduction processing is arranged between the encoder and the decoder. According to the invention, the change area in the complex scene can be accurately captured, and false change in the change detection result is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of remote sensing image processing, and particularly to a remote sensing image change detection method and system based on a metric calibration adaptive SAM2 network. Background Art

[0002] The change detection technology in high-resolution remote sensing images plays a crucial role in identifying the changes in land cover types in different periods, such as the changes in building areas, garden green spaces, forests, and grasslands. The achievements of this technology have been widely penetrated into many fields such as ecological environment monitoring, agricultural production management, and military defense. The leap of contemporary high-resolution remote sensing technology has made it possible to capture the delicate details of land cover, providing valuable information resources for change detection tasks. However, this has also brought a series of new challenges. In remote sensing images, there are a wide variety of ground objects with their own characteristics. In addition, imaging factors such as lighting conditions, shadow effects, and sensor types have a significant impact on data quality. These factors cause objects with the same semantics to exhibit very different spectral characteristics at different spatial positions and time dimensions.

[0003] In recent years, deep learning technology has made significant progress in the field of change detection. Related research combines an encoder and a decoder to design a siamese network architecture, and by introducing diverse backbone networks, such as convolutional neural network series, convolutional neural network series, Transformer series, graph convolutional network series, and the recently popular Mamba network, the semantic expression ability of deep features has been significantly improved. In addition, some research attempts to introduce metric learning strategies to further enhance the network's ability to distinguish between ground object categories. Although these methods have achieved remarkable success, their ability to identify non-semantic changes still needs to be improved. The powerful universality and adaptability of recent basic models have been firmly established. It builds a pre-trained model based on large-scale unlabeled data and auxiliary tasks (Pretext), and demonstrates strong generalization ability in various downstream tasks of vision and language understanding. In this context, Meta AI launched SAM2 based on SAM, aiming to achieve fast visual segmentation in videos and images. SAM2 pre-trains the Hiera image encoder using MAE and integrates hierarchical features to achieve higher segmentation accuracy. However, directly applying SAM2 to the remote sensing image change detection task will inevitably bring the problem of domain differences. Summary of the Invention

[0004] Therefore, the present invention provides a remote sensing image change detection method and system based on a metric calibration adaptive SAM2 network to solve the problem of interference from non-semantic changes in complex scenes existing in the existing remote sensing image change detection using SAM2.

[0005] According to the design solution provided by the present invention, on the one hand, a remote sensing image change detection method based on a metric calibration adaptive SAM2 network is provided, including:

[0006] Obtain dual-temporal remote sensing images of the target area, and preprocess the dual-temporal remote sensing images to obtain corresponding dual-temporal target image patches;

[0007] Input the dual-temporal target image patches into a pre-trained remote sensing image change detection model, and use the remote sensing image change detection model to obtain the change results in the dual-temporal remote sensing images of the target area. Among them, the remote sensing image change detection model includes: an encoder for extracting multi-scale features of the dual-temporal image patches, a decoder for gradually restoring and decoding the output of the multi-scale features of the dual-temporal image patches layer by layer, and a change perceptron for performing semantic metric and change calibration on the decoded output features of the dual-temporal image patches and outputting the semantic metric results and change mapping results through upsampling. And a first adapter for gradually optimizing and adjusting the semantic encoding process of the image patches during the extraction of multi-scale features is provided inside the encoder, and a second adapter for adaptively adjusting the change detection task and performing dimensionality reduction processing on the features is provided between the encoder and the decoder.

[0008] As the remote sensing image change detection method based on the metric calibration adaptive SAM2 network of the present invention, further, the remote sensing image change detection model includes: a plurality of encoders stacked, and each encoder is composed of a first adapter and a Hiera encoder for extracting hierarchical features of the image patches.

[0009] As the remote sensing image change detection method based on the metric calibration adaptive SAM2 network of the present invention, further, the first adapter includes a downsampling fully connected layer, a first GeLU activation function, an upsampling fully connected layer, and a second GeLU activation function.

[0010] As the remote sensing image change detection method based on the metric calibration adaptive SAM2 network of the present invention, further, the remote sensing image change detection model includes: a number of decoders corresponding to and stacked with each encoder, and each decoder includes: a transposed convolutional layer for feature space alignment and a decoding layer for feature decoding, and the decoding layer is composed of a convolutional layer, a normalization layer, and a ReLU activation function.

[0011] As the remote sensing image change detection method based on the metric calibration adaptive SAM2 network of the present invention, further, the second adapter includes: a convolutional layer, a batch normalization layer, and a ReLU activation function.

[0012] As the remote sensing image change detection method based on the metric calibration adaptive SAM2 network of the present invention, further, the change sensor includes: a change enhancement unit for cross-mixing the dual-temporal features along the channels and capturing the change features between the dual-temporal phases, a semantic metric unit for performing feature transformation on the dual-temporal features and measuring the semantic similarity between the features, a feature calibration unit for using the semantic similarity metric features to guide the change features to focus on the semantically inconsistent regions and output, and an output unit for performing upsampling operations on the semantic similarity metric features and the change features and outputting the semantic metric result and the change mapping result.

[0013] As the remote sensing image change detection method based on the metric calibration adaptive SAM2 network of the present invention, further, the training process of the remote sensing image change detection model includes:

[0014] Obtain a remote sensing data set, crop the images in the remote sensing data set into fixed-size and non-overlapping image patches and label the ground object targets in the image patches to construct a change detection sample set;

[0015] Divide the change detection sample set into a training data set, a validation data set, and a test data set according to a specified ratio;

[0016] Construct a model training loss function using the semantic metric loss and the binary cross-entropy loss, train the remote sensing image change detection model based on the model training loss function using the training data set, evaluate and optimize the trained remote sensing image change detection model using the validation data set and the test data set, and use the final remote sensing image change detection model as the target model for performing the remote sensing change detection task in the target area.

[0017] In another aspect, the present invention also provides a remote sensing image change detection system based on the metric calibration adaptive SAM2 network, including: an image acquisition module and a change detection module, wherein,

[0018] The image acquisition module is used to acquire the dual-temporal remote sensing images of the target area and preprocess the dual-temporal remote sensing images to obtain the corresponding dual-temporal target image patches;

[0019] A change detection module is used to input dual-temporal target image patches into a pre-trained remote sensing image change detection model, and utilize the remote sensing image change detection model to obtain the change results in the dual-temporal remote sensing images of the target area. Among them, the remote sensing image change detection model includes: an encoder for extracting multi-scale features of dual-temporal image patches, a decoder for gradually restoring and decoding the output of the multi-scale features of dual-temporal image patches layer by layer, and a change sensor for performing semantic measurement and change calibration on the decoded output features of dual-temporal image patches and outputting the semantic measurement results and change mapping results through upsampling. And a first adapter for gradually optimizing and adjusting the semantic encoding process of image patches during the extraction of multi-scale features is arranged inside the encoder, and a second adapter for adaptively adjusting the change detection task and performing dimensionality reduction processing on the features is arranged between the encoder and the decoder.

[0020] Advantages of the present invention:

[0021] The present invention uses the encoder of frozen SAM2 to extract general visual representations in remote sensing images, and embeds an internal adapter and an external adapter inside and outside the encoder respectively to achieve the adaptive adjustment from natural image processing tasks to remote sensing image change detection tasks. On the basis of obtaining enhanced change information, metric learning between dual-temporal semantic features is introduced, and the change information is calibrated through semantic measurement, further improving the separability of land cover feature types in complex scenes. And further, a comparison is implemented on a remote sensing dataset. The experimental results show that the change detection of the present solution can accurately capture the changed areas in complex scenes, effectively reduce the false changes in the change detection results, and is applicable to the change detection task of high-resolution remote sensing images. Description of the drawings

[0022] Figure 1 Schematic diagram of the remote sensing image change detection process based on the metric calibration adaptive SAM2 network in the embodiment;

[0023] Figure 2 Example of non-semantic change interference caused by imaging conditions in the embodiment;

[0024] Figure 3 Schematic diagram of the MSAM2CD architecture of the remote sensing image change detection model in the embodiment. Detailed implementation manners

[0025] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with the drawings and technical solutions.

[0026] Implementing change detection using high-resolution remote sensing images is of great significance for deeply understanding the dynamic evolution of the earth's surface. However, the land cover in remote sensing images has significant intra-class differences and inter-class similarities, which limits the performance of change detection in complex scenarios. Therefore, in the embodiments of the present invention, see Figure 1 as shown, a remote sensing image change detection method based on a metric calibration adaptive SAM2 network is provided, including:

[0027] S101. Obtain dual-temporal remote sensing images of the target area, and preprocess the dual-temporal remote sensing images to obtain corresponding dual-temporal target image patches.

[0028] S102. Input the dual-temporal target image patches into a pre-trained remote sensing image change detection model, and use the remote sensing image change detection model to obtain the change results in the dual-temporal remote sensing images of the target area. Among them, the remote sensing image change detection model includes: an encoder for extracting multi-scale features of the dual-temporal image patches, a decoder for gradually restoring and decoding the output of the multi-scale features of the dual-temporal image patches layer by layer, and a change sensor for performing semantic measurement and change calibration on the decoded output features of the dual-temporal image patches and outputting the semantic measurement results and change mapping results through upsampling. And a first adapter for gradually optimizing and adjusting the semantic encoding process of the image patches during the extraction of multi-scale features is arranged inside the encoder, and a second adapter for adaptively adjusting the change detection task and performing dimensionality reduction processing on the features is arranged between the encoder and the decoder.

[0029] As Figure 2 shown, in the change detection of ground objects in complex scenarios, there are significant differences in the shape and appearance between buildings a and b with the same semantics; while between buildings a and c, b and d, there are obvious differences in color and shadow due to factors such as lighting. Therefore, in order to accurately identify changes in complex and variable scenarios, the model needs to comprehensively consider the intra-class difference degree and inter-class similarity between ground objects, and effectively eliminate the interference of non-semantic changes caused by imaging conditions. Therefore, in this embodiment, see Figure 3 shown, for the given dual-temporal remote sensing images of the target area, the pre-trained SAM2 is used as a frozen encoder to extract high-level semantic information from the remote sensing images; in order to efficiently adapt to the CD task of change detection, trainable adapters are introduced inside and outside SAM2 respectively; and the multi-level encoded features are refined and fused in the U-shaped decoder; in the change perception of metric calibration, semantic measurement, change enhancement and calibration perception are performed between the dual-temporal decoded features, and the measurement results and change mapping are output through upsampling.

[0030] Among them, multiple encoders that can be stacked are adopted. Each encoder consists of a first adapter and a Hiera encoder for extracting hierarchical features of image patches. The first adapter includes a downsampling fully connected layer, a first GeLU activation function, an upsampling fully connected layer, and a second GeLU activation function. A number of decoders corresponding to and stacked with each encoder are included. Each decoder contains: a transposed convolutional layer for feature space alignment and a decoding layer for feature decoding. The decoding layer consists of a convolutional layer, a normalization layer, and a ReLU activation function. The second adapter includes: a convolutional layer, a batch normalization layer, and a ReLU activation function. The change perception unit can be set to include: a change enhancement unit for cross-mixing dual-temporal features along channels and capturing change features between dual-temporal phases, a semantic metric unit for performing feature transformation on dual-temporal features and measuring the semantic similarity between features, a feature calibration unit for using the semantic similarity metric features to guide the change features to focus on semantically inconsistent regions and output, and an output unit for performing upsampling operations on the semantic similarity metric features and the change features and outputting the semantic metric result and the change mapping result.

[0031] In this embodiment, as Figure 3 shown, Hiera is used as the encoder. Different from the ViT encoder adopted by SAM, the hierarchical structure of Hiera can effectively capture multi-scale features. Specifically, given dual-temporal remote sensing images Hiera outputs four hierarchical features respectively where H, W, and C i represent the length, width, and number of channels of the feature respectively.

[0032] In view of the inherent domain differences between natural images and remote sensing images, in order to fully exploit the potential of the SAM2 model in the change detection task, different types of adapters are ingeniously embedded inside and outside the SAM2 encoder. Specifically, in order to gradually adjust and optimize the semantic encoding process, adapters (denoted as AdapterI) are carefully inserted between the encoding blocks of Hiera. Each AdapterI sequentially includes a downsampling fully connected layer, a GeLU activation function, an upsampling fully connected layer, and a GeLU activation function again. Furthermore, in order to achieve task adaptability and feature dimensionality reduction after each encoding block, another type of adapter (denoted as AdapterII) is connected. Each AdapterII consists of a 1×1 convolutional layer, a batch normalization layer, and a ReLU activation function. The introduction of the two types of adapters helps to further refine and optimize the feature representation, providing more accurate and effective information support for the change detection task.

[0033] On the basis of obtaining multi-level semantic features, the decoder uses a refined U-shaped structure to gradually restore semantic information layer by layer. Its input is multi-level encoded features Output the decoded features D at the top layer through three decoding blocks {D3, D2, D1} t . Specifically, in each decoding block, spatial alignment of features is achieved through transposed convolution. After feature concatenation, feature decoding is achieved through 3×3 convolution, BN, and ReLU activation function. The calculation formula for this stage can be expressed as follows:

[0034]

[0035] where UpConv() and Conc() represent transposed convolution and feature concatenation operations respectively.

[0036] To achieve an effective transformation from dual-temporal semantic features to change information, a change perception module with semantic metric information calibration is designed. As Figure 3 shown, MCCPM is mainly composed of a change enhancement unit (CEU) and a semantic metric unit (SMU). In CEU, different from element subtraction and feature concatenation, etc., the dual-temporal features are cross-mixed along the channels, and grouped convolution is used to capture more detailed change features between the dual-temporal phases. Specifically, the input dual-temporal decoded features obtain the mixed features after cross-mixing of the features. The grouped convolution uses a convolution kernel of C groups of 3×3. During the convolution process, each group of convolution kernels can perform convolution on the dual-temporal features in both the spatial dimension and the time dimension simultaneously. The specific calculation of CEU is as follows:

[0037]

[0038] F CEU = ReLU(BN(GroupConv 3×3 (F mix ))) (4)

[0039] where represents the output change features. In SMU, the input dual-temporal decoded features first use two convolution blocks composed of a 1×1 convolution layer, a BN layer, and a ReLU activation function to perform feature transformation on the dual-temporal decoded features, and correspondingly obtain the intermediate features To minimize the intra-class difference and maximize the inter-class difference between the dual-temporal features, the Euclidean distance is used to measure the semantic similarity between the features, and the metric features are obtained. The calculation formula for this stage is as follows:

[0040] F t = Convblock 1×1 (Convblock 1×1 (D t )) (5)

[0041] F SMU = |norm(F 1 ) - norm(F 2 )|² (6)

[0042] where norm() represents normalizing the feature vector, and ||² represents calculating the L2 norm. Further, the metric feature is used to guide the change feature to focus on the regions with semantic inconsistency, obtaining the calibrated feature The specific calculation can be expressed as follows:

[0043]

[0044] where Sigmoid() represents the Sigmoid activation function, represents multiplying features.

[0045] Among them, the training process of the remote sensing image change detection model can be designed to include:

[0046] Obtain the remote sensing dataset, crop the images in the remote sensing dataset into non-overlapping image patches of a fixed size and label the ground object targets in the image patches to construct a change detection sample set;

[0047] Divide the change detection sample set into a training dataset, a validation dataset, and a test dataset according to a specified ratio;

[0048] Construct a model training loss function using the semantic metric loss and the binary cross-entropy loss. Based on the model training loss function, train the remote sensing image change detection model using the training dataset, evaluate and optimize the trained remote sensing image change detection model using the validation dataset and the test dataset, and use the final remote sensing image change detection model as the target model for performing the remote sensing change detection task in the target area.

[0049] As Figure 3 shown, at the very end of MSAM2CD, the metric feature and the change feature are upsampled to align with the size of the ground truth label. During the network training process, using the metric loss and the binary cross-entropy loss as the optimization objectives, the specific calculation formula of the model loss can be expressed as follows:

[0050]

[0051] where represents the total loss function, n represents the number of feature vectors, l ij and l c represent the ground truth label, f ij and y c represent the metric result and the change result output by MSAM2CD respectively.

[0052] Furthermore, based on the above method, an embodiment of the present invention further provides a remote sensing image change detection system based on a metric calibration adaptive SAM2 network, including: an image acquisition module and a change detection module, where,

[0053] The image acquisition module is used to acquire dual-temporal remote sensing images of the target area, and preprocess the dual-temporal remote sensing images to obtain corresponding dual-temporal target image patches;

[0054] The change detection module is used to input the dual-temporal target image patches into a pre-trained remote sensing image change detection model, and use the remote sensing image change detection model to obtain the change results in the dual-temporal remote sensing images of the target area. Among them, the remote sensing image change detection model includes: an encoder for extracting multi-scale features of the dual-temporal image patches, a decoder for gradually restoring and decoding the output of the multi-scale features of the dual-temporal image patches layer by layer, and a change perceptron for performing semantic metric and change calibration on the decoded output features of the dual-temporal image patches and outputting the semantic metric results and change mapping results through upsampling. And a first adapter for gradually optimizing and adjusting the semantic encoding process of the image patches during the extraction of multi-scale features is provided inside the encoder, and a second adapter for adaptively adjusting the change detection task and performing dimensionality reduction processing on the features is provided between the encoder and the decoder.

[0055] To verify the effectiveness of this solution, the following further explains with experimental data:

[0056] Experiments were carried out on two publicly available change detection datasets, LEVIR-CD and WHU-CD. The LEVIR-CD dataset contains 637 pairs of remote sensing images, with a size of 1024×1024 and a spatial resolution of 0.5m. In the experiment, these images were cropped into non-overlapping patches with a size of 256×256 and divided according to their official standards. Therefore, 7120 / 1024 / 2048 pairs of patches were obtained for training / validation / testing respectively. The WHU-CD dataset contains a pair of high-resolution images, with a size of 32507×15354 and a spatial resolution of 0.75m. Due to memory limitations and the absence of official data segmentation, this pair of images was cropped into non-overlapping patches with a size of 256×256 and randomly divided into training / validation / test sets with samples of 6096 / 762 / 762.

[0057] The experiment was conducted on a laptop equipped with an Intel Xeon Gold 5220 CPU and a Tesla V100 PCIE GPU. This solution was mainly implemented using the PyTorch toolbox and the Python language. During the model training process, mixed-precision training was adopted to improve the training efficiency, and the batch size and the number of iterations were set to 32 and 100 respectively. The SGD optimizer was used to optimize the network weights, where the weight decay and the initial learning rate (denoted as lr initial ) were set to 0.0005 and 0.1 respectively. The learning rate changed with the number of iterations according to the formula lr = lr initial ×(1 - iterations / total_iterations) 3 . Six widely used evaluation metrics were adopted to measure the performance of different methods, namely the overall classification accuracy (OA), F1-score (F1), intersection over union (IoU), recall (Rec), precision (Pre), and Kappa coefficient (Kap).

[0058] This solution MSAM2CD was compared with a series of state-of-the-art change detection (CD) methods, including FC-Siam_D, BIT, TFIGR, USSFCNet, P2VCD, SAMCD, and EATDer. The comparison results are shown in Tables 1 and 2. Compared with other state-of-the-art methods, this solution MSAM2CD achieved better performance. Specifically, on the LEVIR-CD dataset, this solution obtained the highest values in terms of OA, F1, IoU, and Kap. On the WHU-CD dataset, this solution performed best in all metrics except the Pre metric. In particular, MSAM2CD could ensure an effective balance between a high Rec and a high Pre, indicating that it could accurately capture the changed regions and effectively reduce the false changes in the CD results. The effectiveness of this solution was demonstrated by both qualitative evaluation and quantitative results.

[0059] Table 1 Quantitative comparison of evaluation metrics of different methods on the LEVIR-CD dataset

[0060]

[0061]

[0062] Table 2 Quantitative comparison of evaluation metrics of different methods on the WHU-CD dataset

[0063]

[0064] To fully verify the effectiveness of each component in the model of this solution, ablation experiments were carried out on two datasets, as shown in Table 3.

[0065] Table 3 Ablation Experiment Results

[0066]

[0067] When removing CEU and SMU separately or simultaneously, the performance decreases, indicating the effectiveness of using semantic metrics to calibrate change information. In addition, when removing AdapterI and AdapterII separately or simultaneously, the performance also decreases, verifying the feasibility of fine-tuning the vision base model for downstream tasks. Furthermore, after replacing the SAM2 model with ResNet18, the change performance drops sharply, confirming the great potential of SAM2 for change detection tasks. The above observations show that each component in SAM2CD in this solution plays an irreplaceable role in complex scene change detection tasks.

[0068] The above experimental data shows that this solution can be applied to the change detection task of high-resolution remote sensing images and has good application prospects in the field of ground object change detection in complex scenes.

[0069] Unless otherwise specifically stated, the relative steps, numerical expressions, and values of the components and steps set forth in these embodiments do not limit the scope of the present invention.

[0070] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0071] The units and method steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those of ordinary skill in the art can use different methods to implement the described functions for each specific application, but such implementation is not considered to exceed the scope of the present invention.

[0072] Those of ordinary skill in the art can understand that all or part of the steps in the above methods can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium, such as a read-only memory, a magnetic disk, or an optical disc, etc. Optionally, all or part of the steps of the above embodiments can also be implemented using one or more integrated circuits. Correspondingly, each module / unit in the above embodiments can be implemented in the form of hardware or in the form of a software function module. The present invention is not limited to any specific form of combination of hardware and software.

[0073] Finally, it should be noted that the above-described embodiments are only specific embodiments of the present invention, used to illustrate the technical solutions of the present invention, rather than limiting it. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions recorded in the foregoing embodiments or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be determined by the protection scope of the claims.

Claims

1. A remote sensing image change detection method based on a metrically calibrated adaptive SAM2 network, characterized in that: Include: Acquire a dual-temporal remote sensing image of the target area, and preprocess the dual-temporal remote sensing image to obtain a corresponding dual-temporal target image block; The dual-phase target image block is input into a pre-trained remote sensing image change detection model, and the remote sensing image change detection model is used to obtain the change results in the dual-phase remote sensing image of the target area, wherein the remote sensing image change detection model includes: an encoder for extracting multi-scale features of the dual-phase image block, a decoder for recovering and decoding the multi-scale features of the dual-phase image block layer by layer, and a change sensor for performing semantic measurement and change calibration on the decoded output features of the dual-phase image block and outputting the semantic measurement results and change mapping results through upsampling, and a first adapter for gradually optimizing and adjusting the semantic encoding process of the image block in extracting multi-scale features is arranged inside the encoder, and a second adapter for adaptively adjusting the change detection task and performing dimensionality reduction processing on the features is arranged between the encoder and the decoder.

2. The remote sensing image change detection method based on the metric calibration adaptive SAM2 network according to claim 1 is characterized in that: The remote sensing image change detection model comprises: a plurality of encoders arranged in a stacked manner, each encoder being composed of a first adapter and a Hiera encoder for extracting hierarchical features of image blocks.

3. The remote sensing image change detection method based on the metric calibration adaptive SAM2 network according to claim 1 or 2 is characterized in that: The first adapter includes a downsampling fully connected layer, a first GeLU activation function, an upsampling fully connected layer and a second GeLU activation function.

4. The remote sensing image change detection method based on the metric calibration adaptive SAM2 network according to claim 2 is characterized in that: The remote sensing image change detection model includes: a plurality of decoders corresponding to each encoder and stacked, each decoder includes: a transposed convolution layer for feature space alignment and a decoding layer for feature decoding, and the decoding layer is composed of a convolution layer, a normalization layer and a ReLU activation function.

5. The remote sensing image change detection method based on metric calibration adaptive SAM2 network according to claim 1 is characterized in that: The second adapter includes: a convolutional layer, a batch normalization layer and a ReLU activation function.

6. The remote sensing image change detection method based on metric calibration adaptive SAM2 network according to claim 1 is characterized in that: The change sensor includes: a change enhancement unit for cross-mixing bi-phase features along channels and capturing bi-phase change features, a semantic measurement unit for performing feature conversion on bi-phase features and measuring semantic similarity between features, a feature calibration unit for guiding change features to focus on semantically inconsistent areas using semantic similarity measurement features and outputting them, and an output unit for upsampling semantic similarity measurement features and change features and outputting semantic measurement results and change mapping results.

7. The remote sensing image change detection method based on metric calibration adaptive SAM2 network according to claim 1 is characterized in that: The remote sensing image change detection model training process includes: Acquire a remote sensing data set, crop the images in the remote sensing data set into fixed-size and non-overlapping image blocks, and mark the ground objects in the image blocks to construct a change detection sample set; Divide the change detection sample set into a training data set, a validation data set, and a test data set according to a specified ratio; The semantic metric loss and binary cross entropy loss are used to construct the model training loss function. The remote sensing image change detection model is trained based on the model training loss function and the training dataset. The trained remote sensing image change detection model is evaluated and optimized using the validation dataset and the test dataset. According to the evaluation and optimization results, the final remote sensing image change detection model is used as the target model for performing the remote sensing change detection task in the target area.

8. A remote sensing image change detection system based on a metrically calibrated adaptive SAM2 network, characterized in that: It includes: image acquisition module and change detection module, among which, An image acquisition module is used to acquire a dual-phase remote sensing image of a target area and preprocess the dual-phase remote sensing image to obtain a corresponding dual-phase target image block; A change detection module is used to input a dual-phase target image block into a pre-trained remote sensing image change detection model, and use the remote sensing image change detection model to obtain change results in the dual-phase remote sensing image of the target area, wherein the remote sensing image change detection model includes: an encoder for extracting multi-scale features of the dual-phase image block, a decoder for recovering and decoding the multi-scale features of the dual-phase image block layer by layer, and a change sensor for performing semantic measurement and change calibration on the decoded output features of the dual-phase image block and outputting the semantic measurement results and change mapping results through upsampling, and a first adapter for gradually optimizing and adjusting the semantic encoding process of the image block in extracting multi-scale features is arranged inside the encoder, and a second adapter for adaptively adjusting the change detection task and performing dimensionality reduction processing on the features is arranged between the encoder and the decoder.

9. An electronic device, characterized in that: include: at least one processor, and a memory coupled to the at least one processor; The memory stores a computer program, and the computer program can be executed by the at least one processor to implement the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed, the method according to any one of claims 1 to 7 can be implemented.