Remote sensing image change detection method and system
Patent Information
- Application Number
- CN202611005908.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-07
- Publication Date
- 2026-09-29
AI Technical Summary
当前遥感影像变化检测方法的核心共性问题为:过度依赖编码器-解码器架构,未能打破解码器的固有设计,同时特征融合策略设计不合理,多尺度特征冗余与噪声抑制不足,导致模型难以兼顾检测性能与计算效率
[0016]本申请实施例提供了一种遥感影像变化检测方法、遥感影像变化检测系统。该方法先获取待检测的双时相遥感影像数据,再将该双时相遥感影像数据输入预训练的目标影像变化检测模型进行检测。该检测模型包括依次连接的茎层融合模块、编码器模块、多尺度校准融合模块和变化检测头。模型的检测过程包括:通过茎层融合模块对双时相遥感影像数据进行特征提取与融合处理,得到时序融合特征;通过编码器模块对时序融合特征进行多尺度特征提取处理,得到多个不同尺度的特征;通过多尺度校准融合模块对多个不同尺度的特征进行校准融合处理,得到多通道融合特征;通过变化检测头根据多通道融合特征进行变化检测,得到二值变化掩码;最后根据二值变化掩码得到影像变化检测结果。本申请实施例通过茎层融合模块对影像数据进行特征提取和融合,能够有效替代传统的通道拼接方式,提升融合特征质量;通过多尺度校准融合模块对特征进行校准融合,能够有效解决特征融合效果差、多尺度冗余与噪声干扰的问题;通过彻底摒弃传统检测模型的解码器结构,消除解码器带来的所有计算开销和参数冗余,能够从架构层面实现模型的轻量化,解决当前模型结构冗余、计算复杂度高的问题。综上,本申请实施例能够在保证高检测精度的同时,实现模型的轻量化和计算效率的有效提升。
Smart Images

Figure CN122841952A_ABST
Abstract
Description
Technical Field
[0001] This application relates to, but is not limited to, the field of image recognition technology, and in particular to a method and system for detecting changes in remote sensing images. Background Technology
[0002] With the rapid development of remote sensing technology, the ability to acquire high-resolution and ultra-high-resolution remote sensing images has been continuously improved. These images contain richer details and semantic information about ground features, but this also places higher demands on change detection algorithms. On the one hand, they need to ensure detection accuracy, accurately identifying surface changes of different scales and types; on the other hand, they need to balance computational efficiency and lightweight models to meet the practical application needs of large-scale batch processing of remote sensing images and real-time detection by edge devices. Before the application of deep learning technology to change detection in remote sensing images, traditional methods mainly relied on pixel-level grayscale and spectral information analysis, such as change vector analysis, support vector machine classification, and principal component analysis. These methods can only capture low-level pixel differences, making it difficult to model complex semantic change features. They are highly sensitive to interference factors such as illumination changes, sensor noise, and terrain undulations, resulting in limited detection accuracy and robustness, which can no longer meet the change detection needs of high-resolution remote sensing images. The rise of deep learning has driven the rapid development of change detection technology for remote sensing images, improving the accuracy and robustness of change detection.
[0003] Despite advancements in deep learning-based remote sensing image change detection methods, core shortcomings remain regarding the balance between lightweight models, computational efficiency, and detection performance. Current common problems in remote sensing image change detection methods include: over-reliance on the encoder-decoder architecture, failure to break free from the inherent design of the decoder, and unreasonable feature fusion strategies, resulting in insufficient multi-scale feature redundancy and noise suppression. This makes it difficult for models to balance detection performance and computational efficiency. Late-stage fusion offers high accuracy but incurs high computational overhead and structural redundancy, while early-stage fusion attempts to improve efficiency but lacks accuracy and remains constrained by the decoder. Summary of the Invention
[0004] The following is an overview of the subject matter described in detail herein. This overview is not intended to limit the scope of the claims.
[0005] This application provides a method and system for detecting changes in remote sensing images, which can achieve lightweighting of the model and effective improvement of computational efficiency while ensuring high detection accuracy.
[0006] In a first aspect, embodiments of this application provide a method for detecting changes in remote sensing images, comprising: acquiring dual-temporal remote sensing image data to be detected; inputting the dual-temporal remote sensing image data into a pre-trained target image change detection model, the target image change detection model comprising a stem-layer fusion module, an encoder module, a multi-scale calibration fusion module, and a change detection head connected in sequence; performing feature extraction and fusion processing on the dual-temporal remote sensing image data through the stem-layer fusion module to obtain temporal fusion features; performing multi-scale feature extraction processing on the temporal fusion features through the encoder module to obtain features at multiple different scales; performing calibration fusion processing on the features at multiple different scales through the multi-scale calibration fusion module to obtain multi-channel fusion features; performing change detection on the multi-channel fusion features through the change detection head to obtain a binary change mask; and obtaining an image change detection result based on the binary change mask.
[0007] In conjunction with the first aspect, in one embodiment of this application, the step of calibrating and fusing features at multiple different scales through the multi-scale calibration and fusion module to obtain multi-channel fused features includes: performing bilinear interpolation on the features at multiple different scales; performing feature fusion enhancement on the interpolated features at each scale to obtain preliminary multi-scale fused features; performing feature splitting on the preliminary multi-scale fused features to obtain a first feature branch, a second feature branch, and a third feature branch; determining feature relationship weights based on the first feature branch and the second feature branch; and obtaining multi-channel fused features based on the feature relationship weights and the third feature branch.
[0008] In conjunction with the first aspect, in one embodiment of this application, the multiple features at different scales include a first feature, a second feature, a third feature, and a fourth feature; the feature fusion enhancement processing of the interpolated features at each scale to obtain preliminary multi-scale fusion features includes: performing average pooling along the channel dimension on the interpolated fourth feature to obtain a pooled fourth feature, wherein the number of channels of the pooled fourth feature is the same as that of the third feature; performing mean and hyperbolic tangent processing on the third feature to obtain a first attention weight; weighting the pooled fourth feature according to the first attention weight, and fusing the weighted pooled fourth feature with the third feature to obtain a fusion-enhanced deep feature; performing channel-dimensional average pooling on the fusion-enhanced deep feature to obtain a pooled feature; adding the pooled feature and the second feature and then performing max pooling to obtain a fusion-enhanced feature; and obtaining preliminary multi-scale fusion features based on the fusion-enhanced feature, the first feature, and the fourth feature.
[0009] In conjunction with the first aspect, in one embodiment of this application, the training steps of the target image change detection model include: acquiring raw image detection data, the raw image detection data including dual-temporal remote sensing image data samples and corresponding binary change mask annotations; dividing the raw image detection data into a training set and a validation set; constructing an initial image change detection model, the initial image change detection model including an initial stem-layer fusion module, an initial encoder module, an initial multi-scale calibration fusion module, and an initial change detection head; acquiring a late-stage fusion change detection network; inputting the training set into the late-stage fusion change detection network for forward propagation to obtain a first prediction probability map; inputting the training set into the initial image change detection model for forward propagation to obtain intermediate fusion features and a second prediction. The process involves: using a probability map; determining a total loss function based on the first predicted probability map, the intermediate fusion features, the second predicted probability map, and the binary change mask annotations corresponding to the dual-temporal remote sensing image data samples in the training set, according to a preset loss function; determining the parameter gradient of the initial image change detection model based on the total loss function; updating the parameters of each module in the initial image change detection model based on the parameter gradient to obtain an updated image change detection model; evaluating the detection performance of the updated image change detection model based on the validation set; determining the model target weight based on the performance detection results; and obtaining the target image change detection model based on the model target weight when the model training reaches a preset number of training rounds or the performance of the validation set detection no longer improves after multiple rounds.
[0010] In conjunction with the first aspect, in one embodiment of this application, before dividing the raw image detection data into a training set and a validation set, the method further includes: performing spatial registration, resolution unification, and pixel value normalization on the dual-temporal remote sensing image data samples to obtain standard dual-temporal remote sensing image data samples; and performing data augmentation on the standard dual-temporal remote sensing image data samples to obtain processed raw image detection data.
[0011] In conjunction with the first aspect, in one embodiment of this application, the preset loss function includes at least one of the following: cross-entropy loss function, binary cross-entropy loss function, and mean absolute error loss function.
[0012] In conjunction with the first aspect, in one embodiment of this application, the loss function includes the cross-entropy loss function, the binary cross-entropy loss function, and the mean absolute error loss function; the step of determining the total loss function according to the first prediction probability map, the intermediate fusion feature, the second prediction probability map, and the binary change mask annotations corresponding to the dual-temporal remote sensing image data samples in the training set, according to the preset loss function, includes: calculating a first loss value using the cross-entropy loss function based on the second prediction probability map and the binary change mask annotations corresponding to the dual-temporal remote sensing image data samples; calculating a second loss value using the binary cross-entropy loss function based on the intermediate fusion feature; calculating a third loss value using the mean absolute error loss function based on the first prediction probability map and the second prediction probability map; and performing weighted fusion based on the first loss value, the second loss value, and the third loss value to obtain the total loss function.
[0013] In conjunction with the first aspect, in one embodiment of this application, obtaining the image change detection result based on the binary change mask includes: if the dual-temporal remote sensing image data is segmented image data, stitching together the binary change masks of all segmented image data to obtain the change mask of the original image; visualizing the change mask of the original image to obtain the image change detection result.
[0014] Secondly, embodiments of this application provide a remote sensing image change detection system, applied to the remote sensing image change detection method described above. The system includes: a data input module for acquiring dual-temporal remote sensing image data to be detected; a change detection module for inputting the dual-temporal remote sensing image data into a pre-trained target image change detection model, the target image change detection model including a stem-layer fusion module, an encoder module, a multi-scale calibration fusion module, and a change detection head connected in sequence; the stem-layer fusion module performs feature extraction and fusion processing on the dual-temporal remote sensing image data to obtain temporal fusion features; the encoder module performs multi-scale feature extraction processing on the temporal fusion features to obtain features at multiple different scales; the multi-scale calibration fusion module performs calibration fusion processing on the features at multiple different scales to obtain multi-channel fusion features; the change detection head performs change detection based on the multi-channel fusion features to obtain a binary change mask; and a result output module for obtaining image change detection results based on the binary change mask.
[0015] In conjunction with the second aspect, in one embodiment of this application, the system further includes a data preprocessing module for converting the dual-temporal remote sensing image data into standardized input data.
[0016] This application provides a method and system for detecting changes in remote sensing images. The method first acquires dual-temporal remote sensing image data to be detected, and then inputs this data into a pre-trained target image change detection model for detection. The detection model includes a stem-layer fusion module, an encoder module, a multi-scale calibration fusion module, and a change detection head, connected sequentially. The model's detection process includes: extracting and fusing features from the dual-temporal remote sensing image data using the stem-layer fusion module to obtain temporal fusion features; extracting multi-scale features from the temporal fusion features using the encoder module to obtain features at multiple different scales; calibrating and fusing the features at multiple different scales using the multi-scale calibration fusion module to obtain multi-channel fusion features; detecting changes based on the multi-channel fusion features using the change detection head to obtain a binary change mask; and finally, obtaining the image change detection result based on the binary change mask. This application's embodiments utilize a stem-layer fusion module to extract and fuse features from image data, effectively replacing traditional channel stitching and improving the quality of fused features. A multi-scale calibration fusion module calibrates and fuses features, effectively addressing issues of poor feature fusion performance, multi-scale redundancy, and noise interference. By completely abandoning the decoder structure of traditional detection models, eliminating all computational overhead and parameter redundancy introduced by the decoder, the model achieves lightweighting at the architectural level, resolving the problems of redundant structure and high computational complexity in current models. In summary, this application's embodiments can achieve both high detection accuracy and a lightweight model with significantly improved computational efficiency. Attached Figure Description
[0017] Figure 1 This is a flowchart of a remote sensing image change detection method provided in one embodiment of this application; Figure 2 This is a flowchart illustrating the training process of a target image change detection model provided in one embodiment of this application. Figure 3 This is an overall framework diagram of an image change detection model provided in one embodiment of this application; Figure 4 This is a flowchart illustrating the determination of the total loss function according to an embodiment of this application; Figure 5 This is provided in one embodiment of the present application. Figure 1 The detailed flowchart of step 150; Figure 6 This is a structural diagram of a multi-scale calibration fusion module provided in one embodiment of this application; Figure 7 This is an overall flowchart of a remote sensing image change detection method provided in a specific embodiment of this application; Figure 8 This is an architecture diagram of a remote sensing image change detection system provided in one embodiment of this application. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0019] It should be noted that although the flowchart shows a logical order, in some cases, the steps shown or described may be performed in a different order than that shown in the flowchart. The terms "first," "second," etc., used in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the structures, proportions, sizes, etc., depicted in the drawings are only used to complement the content disclosed in the specification for those skilled in the art to understand and read, and are not intended to limit the implementation conditions of this application. Therefore, they have no substantial technical significance. Any modifications to the structure, changes in proportions, or adjustments to size, without affecting the effects and purposes achieved by this application, should still fall within the scope of the technical content disclosed in this application. Similarly, the terms such as "upper," "lower," "left," "right," "middle," and "one" used in this specification are only for clarity of description and are not used to limit the scope of implementation of this application. Changes or adjustments in their relative relationships, without substantially altering the technical content, should also be considered within the scope of implementation of this application.
[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0021] Remote sensing image change detection is a core research direction at the intersection of remote sensing information processing and computer vision. Through intelligent analysis of dual-temporal / multi-temporal remote sensing images, it automatically extracts surface change information, providing key data support and decision-making basis for fields such as land and resources management, urban planning, ecological environment protection, disaster emergency response, and agricultural production monitoring.
[0022] With the rapid development of remote sensing technology, the ability to acquire high-resolution and ultra-high-resolution remote sensing images has been continuously improved. These images contain richer details and semantic information about ground features, but this also places higher demands on change detection algorithms. On the one hand, they need to ensure detection accuracy, accurately identifying surface changes of different scales and types; on the other hand, they need to balance computational efficiency and lightweight models to meet the practical application needs of large-scale batch processing of remote sensing images and real-time detection by edge devices. Before the application of deep learning technology to change detection in remote sensing images, traditional methods mainly relied on pixel-level grayscale and spectral information analysis, such as change vector analysis, support vector machine classification, and principal component analysis. These methods can only capture low-level pixel differences, making it difficult to model complex semantic change features. They are highly sensitive to interference factors such as illumination changes, sensor noise, and terrain undulations, resulting in limited detection accuracy and robustness, which can no longer meet the change detection needs of high-resolution remote sensing images. The rise of deep learning has driven the rapid development of change detection technology for remote sensing images. Methods based on deep learning architectures such as convolutional neural networks (CNN), Transformer, and Mamba have become the mainstream. These methods extract high-level semantic features of images through deep networks, significantly improving the accuracy and robustness of change detection.
[0023] Current deep learning-based remote sensing image change detection methods, whether for early or late fusion, all use an encoder-decoder architecture. The core difference lies in the timing of the fusion of dual temporal features. Both approaches suffer from inherent limitations in balancing detection performance and computational efficiency, as detailed below: Late-stage fusion is currently the most widely used and relatively superior solution for change detection in remote sensing images. Representative methods include ChangeFormer, BIT, RSMamba, and Convformer-CD. The core architecture of this solution consists of a twin encoder, a temporal fusion module, a decoder, and a change detection head. The specific process is as follows: The bi-temporal remote sensing images (Ipre and Ipost) are input into two weight-shared twin encoder branches, independently extracting deep semantic features from the bi-temporal images; temporal fusion of the bi-temporal features is performed through operations such as subtraction, stitching, multiplication, and attention weighting to obtain a preliminary change feature representation; the change features are input into a complex decoder structure, and operations such as upsampling, skip connections, and feature fusion are used to restore the spatial resolution of the features, compensating for the detail loss caused by encoder downsampling; finally, the change detection head outputs a binary change mask with the same resolution as the input images. The advantage of late-stage fusion is that it can fully preserve the feature information of each bi-temporal image, and the independently extracted features can better characterize the semantics of the bi-temporal images, thus achieving good detection accuracy in various change detection scenarios. However, this approach has significant inherent drawbacks: First, the twin encoder requires feature extraction from both temporal images separately, which introduces additional computational overhead and significantly increases the model's floating-point operations and inference latency. Second, in order to improve detection accuracy, late-stage fusion schemes typically design complex decoder structures containing a large number of convolutional, upsampling, and skip connection layers, further increasing the model's parameter count, computational complexity, and memory footprint. Third, the overall model structure is redundant, making it difficult to deploy on edge devices or embedded devices, and it also fails to meet the real-time, batch detection requirements of large-scale remote sensing images.
[0024] Early fusion schemes are improved solutions proposed to address the high computational complexity of late-stage fusion. Representative methods include dual-temporal image fusion detection methods based on FCN and UNet++. The core architecture of this scheme is a single encoder-decoder-change detection head. The specific process is as follows: dual-temporal remote sensing images are stitched together to form a multi-channel fusion input; the fusion input is directly input into a single encoder-decoder network. The encoder extracts fusion features, the decoder restores the feature resolution and refines the features, and finally, the change detection head outputs a binary change mask. The advantage of early fusion schemes is that they eliminate the need for a twin encoder, extracting fusion features through a single encoder, effectively reducing the double computational overhead of the twin encoder; at the same time, the dual-temporal features interact directly in the early stages of the network, enabling better capture of subtle differences in the dual-temporal images. However, existing early fusion schemes still have key unresolved issues: First, they still rely on traditional complex decoder structures to restore feature resolution and refine features. The presence of the decoder still increases the complexity and computational cost of the model, failing to achieve true lightweighting. Second, the use of simple interpolation and stitching fusion methods for multi-scale features extracted by the encoder makes it difficult to fully explore the semantic change information in multi-scale features, resulting in poor feature fusion effects. Consequently, their detection performance is generally lower than that of later fusion schemes, failing to fully leverage the computational advantages of early fusion. Third, the feature fusion design of existing early fusion schemes is relatively crude, achieving the fusion of two-temporal images only through simple channel stitching, without considering the refined extraction and interaction of primary features in the two-temporal phases, and the quality of fused features needs to be improved. Fourth, direct stitching of multi-scale features easily introduces redundant responses, makes it impossible to explicitly model the spatial-spectral dependencies of features, and makes it difficult to suppress noise interference.
[0025] In summary, the core common problems of current remote sensing image change detection methods are: over-reliance on the encoder-decoder architecture, failure to break the inherent design of the decoder, unreasonable feature fusion strategy design, insufficient multi-scale feature redundancy and noise suppression, making it difficult for the model to balance detection performance and computational efficiency. Late-stage fusion has high accuracy but high computational cost and structural redundancy, while early-stage fusion attempts to improve efficiency but lacks accuracy and is still constrained by the decoder.
[0026] In view of this, embodiments of this application provide a method and system for detecting changes in remote sensing images. The method first acquires the dual-temporal remote sensing image data to be detected, and then inputs the dual-temporal remote sensing image data into a pre-trained target image change detection model for detection. The detection model includes a stem-layer fusion module, an encoder module, a multi-scale calibration fusion module, and a change detection head, connected sequentially. The model's detection process includes: performing feature extraction and fusion processing on the dual-temporal remote sensing image data through the stem-layer fusion module to obtain temporal fusion features; performing multi-scale feature extraction processing on the temporal fusion features through the encoder module to obtain features at multiple different scales; performing calibration fusion processing on the features at multiple different scales through the multi-scale calibration fusion module to obtain multi-channel fusion features; performing change detection based on the multi-channel fusion features using the change detection head to obtain a binary change mask; and finally obtaining the image change detection result based on the binary change mask. This application's embodiments utilize a stem-layer fusion module to extract and fuse features from image data, effectively replacing traditional channel stitching and improving the quality of fused features. A multi-scale calibration fusion module calibrates and fuses features, effectively addressing issues of poor feature fusion performance, multi-scale redundancy, and noise interference. By completely abandoning the decoder structure of traditional detection models, eliminating all computational overhead and parameter redundancy introduced by the decoder, the model achieves lightweighting at the architectural level, resolving the problems of redundant structure and high computational complexity in current models. In summary, this application's embodiments can achieve both high detection accuracy and a lightweight model with significantly improved computational efficiency.
[0027] The embodiments of this application will be further described below with reference to the accompanying drawings.
[0028] Reference Figure 1 , Figure 1 This is a flowchart of the remote sensing image change detection method provided in the embodiments of this application. The process may specifically include, but is not limited to, steps 110 to 170.
[0029] Step 110: Acquire the dual-temporal remote sensing image data to be detected; Step 120: Input the dual-temporal remote sensing image data into the pre-trained target image change detection model, which includes a stem-layer fusion module, an encoder module, a multi-scale calibration fusion module, and a change detection head connected in sequence; Step 130: Extract and fuse features from dual-temporal remote sensing image data using the stem-layer fusion module to obtain temporal fusion features; Step 140: Perform multi-scale feature extraction on the temporal fusion features through the encoder module to obtain features at multiple different scales; Step 150: Perform calibration and fusion processing on features of multiple different scales using the multi-scale calibration and fusion module to obtain multi-channel fused features; Step 160: Perform change detection using a change detection head based on multi-channel fusion features to obtain a binary change mask; Step 170: Obtain the image change detection result based on the binary change mask.
[0030] In a feasible embodiment, dual-temporal remote sensing imagery refers to spatially registered remote sensing imagery data of the same geographic area acquired at two different time points, denoted as pre-temporal imagery. and post-phase images This is the basic input data for change detection. For example, in urban expansion monitoring scenarios, This is a satellite image of a city acquired in 2018. By comparing satellite imagery of the same city acquired in 2023, the extent of newly added construction land can be detected; in flood disaster assessment scenarios, This is a normal surface image before the flood. Images of inundated areas after a flood can be used to quickly extract the inundated extent; in vegetation change monitoring scenarios... Images acquired in spring, Images of the same region acquired in autumn can be used to analyze seasonal changes in vegetation and differences in growth.
[0031] In one feasible embodiment, the target image change detection model refers to a lightweight remote sensing image change detection model based on an early fusion strategy. It abandons the decoder structure of traditional change detection models and is composed of a stem-layer fusion module, an encoder module, a multi-scale calibration fusion module, and a change detection head connected sequentially. This model receives dual-temporal remote sensing images and outputs pixel-level change detection results. Figure 2 As shown, the training process of the target image change detection model may include, but is not limited to, steps 210 to 290.
[0032] Step 210: Obtain the raw image detection data, which includes dual-temporal remote sensing image data samples and corresponding binary change mask annotations; divide the raw image detection data into a training set and a validation set; Step 220: Construct an initial image change detection model, which includes an initial stem-layer fusion module, an initial encoder module, an initial multi-scale calibration fusion module, and an initial change detection head; Step 230: Obtain the late-stage fusion change detection network; Step 240: Input the training set into the late fusion change detection network for forward propagation to obtain the first prediction probability map; Step 250: Input the training set into the initial image change detection model for forward propagation to obtain intermediate fusion features and a second predicted probability map; Step 260: Based on the first prediction probability map, intermediate fusion features, the second prediction probability map, and the binary change mask labels corresponding to the dual-temporal remote sensing image data samples in the training set, determine the total loss function according to the preset loss function. Step 270: Determine the parameter gradient of the initial image change detection model based on the total loss function, update the parameters of each module in the initial image change detection model based on the parameter gradient, and obtain the updated image change detection model. Step 280: Evaluate the detection performance of the updated image change detection model based on the validation set, and determine the target weight of the model based on the performance evaluation results; Step 290: When the model training reaches the preset number of training rounds or the performance of the validation set detection no longer improves for several consecutive rounds, the target image change detection model is obtained according to the model target weights.
[0033] In a feasible embodiment, the raw image detection data specifically refers to raw remote sensing image change detection data, including dual-temporal remote sensing image data samples and corresponding binary change mask annotations. The binary change mask annotations refer to binary images that correspond one-to-one with the spatial dimensions of the dual-temporal remote sensing images. Pixels in changed areas are marked as 1, and pixels in unchanged areas are marked as 0, used to characterize the extent of land cover change between the two temporal images pixel-by-pixel. The binary change masks can be obtained through manual annotation, thus providing a realistic reference for change areas for model training.
[0034] In a feasible embodiment, before dividing the raw image detection data into training and validation sets, spatial registration, resolution unification, and pixel value normalization can be performed on the dual-temporal remote sensing image data samples to obtain the processed raw image detection data. Specifically, spatial registration is first performed on the dual-temporal remote sensing images to ensure the spatial consistency of the dual-temporal images. Then, resolution unification and pixel value normalization are performed on the images. Pixel value normalization involves scaling the pixel values to the [0,1] range to eliminate the influence of scale and dimensions on subsequent model training.
[0035] In one feasible embodiment, the preprocessed image detection dataset is divided into a training set, a validation set, and a test set according to a preset ratio, such as 7:1:2 or 6:2:2. The training set is used for parameter learning and iterative training of the model, the validation set is used for hyperparameter tuning and model selection during training, and the test set is used for independent performance evaluation of the trained model.
[0036] In one feasible embodiment, before inputting the training set into the late-stage fusion change detection network or the initial image change detection model, data augmentation operations can be performed on the training set data. Data augmentation methods include random horizontal / vertical flipping, random scale cropping, color jittering, Gaussian blurring, etc. Through the above diverse data augmentation strategies, the diversity of training samples can be effectively expanded, thereby improving the generalization ability of the model and avoiding overfitting during training.
[0037] In a feasible embodiment, when constructing the initial image change detection model, a suitable encoder backbone network can be selected according to the actual application requirements, and an initial stem-layer fusion module, an initial multi-scale calibration fusion module, and an initial change detection head can be constructed to form a complete student network. Simultaneously, all parameters of the student network are randomly initialized. The student network can employ a refined early fusion strategy, achieving early deep interaction of dual-temporal features through the stem-layer fusion module, extracting multi-scale deep semantic features through the encoder module, and then replacing the decoder with the multi-scale calibration fusion module to achieve multi-scale feature fusion, semantic enhancement, and adaptive recalibration. Finally, a binary change mask is output through the change detection head. The entire student network has a simple structure, few parameters, and high computational efficiency, achieving extreme lightweighting at the architectural level.
[0038] In a feasible embodiment, the late-fusion change detection network refers to a dual-branch change detection network structure that can independently extract features from two-temporal remote sensing images and then fuse them at the deep feature level or decision level to obtain change detection results. During the training of the image change detection model, the late-fusion change detection network can be used as the teacher network, with its pre-trained weights loaded and all parameters frozen. This prevents the teacher network from participating in gradient updates during the training of the student network (i.e., the initial image change detection model), serving only as a supervisory signal provider. In this embodiment, the teacher network can adopt a high-performance late-fusion twin encoder-decoder architecture currently available, such as ChangeFormer, BIT, RSMamba, HyRet-Change, or any other high-precision change detection network. The core role of the teacher network is to extract high-quality two-temporal features and generate high-precision change prediction results. During model training, its parameters are frozen, it does not participate in gradient updates, and it only provides supervisory information for knowledge distillation to the student network, guiding its training.
[0039] In a feasible embodiment, when training the initial image change detection model, the training hyperparameters can be set first. Specifically, the training hyperparameters may include at least: batch size, preferably 32; learning rate, preferably 3e-4; number of training epochs, preferably 600; the optimizer is AdamW, the weight decay is preferably 0.01, and the beta value is preferably (0.9, 0.999). Furthermore, a linear learning rate decay strategy is adopted, causing the learning rate to linearly decay from the initial value to 0, to improve the stability of the training process. It should be noted that the values of the training hyperparameters in this embodiment are only illustrative examples, and the values of different hyperparameters can be flexibly adjusted according to the actual application scenario. For example, the number of training epochs can also be set to 400. This embodiment does not limit the values of the training hyperparameters.
[0040] In one feasible embodiment, such as Figure 3 As shown, the overall framework of the image change detection model adopts a teacher-student distillation framework, consisting of a teacher network and a student network. The student network is the image change detection model in this embodiment, while the teacher network is a high-performance traditional late-stage fusion change detection network used to provide knowledge distillation supervision for the student network. The network input consists of pairs of bi-temporal remote sensing image samples that have undergone spatial registration, resolution unification, pixel value normalization, and data augmentation. ,in ∈R3×H×W、 ∈R3×H×W are three-channel remote sensing images of the preceding and following time phases, with a resolution of H×W; the network output is a binary change mask M∈RH×W with the same resolution as the input image, where 1 represents the changed area and 0 represents the unchanged area, thus achieving accurate identification and positioning of the changed area.
[0041] In one feasible embodiment, the preprocessed dual-temporal remote sensing image samples are paired By performing a forward propagation on the frozen late-fusion change detection network, the first predicted probability map can be obtained. As a distillation monitoring signal, the same dual-temporal remote sensing image samples are used as a monitoring signal. By performing a forward propagation on the initial image change detection model, intermediate fusion features can be obtained. Second prediction probability map .in, This is the output of the initial multi-scale calibration fusion module.
[0042] In a feasible embodiment, to achieve multi-dimensional supervision of the student network and improve training effectiveness, the preset loss function used in this embodiment includes at least one of the following: cross-entropy loss function, binary cross-entropy loss function, and mean absolute error loss function. Each loss function implements a different supervision function, and its weight coefficients can be optimized according to the actual application scenario.
[0043] In a feasible embodiment, when the loss function includes a cross-entropy loss function, a binary cross-entropy loss function, and a mean absolute error loss function, the process for determining the total loss function is as follows: Figure 4 As shown, steps 410 to 440 may be included, but are not limited to.
[0044] Step 410: Calculate the first loss value using the cross-entropy loss function based on the binary change mask labels corresponding to the second predicted probability map and the dual-temporal remote sensing image data samples; Step 420: Calculate the second loss value based on the intermediate fusion features using the binary cross-entropy loss function; Step 430: Calculate the third loss value using the mean absolute error loss function based on the first and second prediction probability maps; Step 440: Perform weighted fusion based on the first loss value, the second loss value, and the third loss value to obtain the total loss function.
[0045] In one feasible embodiment, the cross-entropy loss can be applied to the predicted probability map of the student network, i.e., the second predicted probability map output by the initial image change detection model. Supervision is performed to guide the student network's predictions to align with the true change mask (i.e., binary change mask annotations), thereby ensuring the model's basic detection accuracy. Binary cross-entropy loss can be applied to the intermediate fusion features output by the initial multi-scale calibration fusion module. Supervision is conducted to enhance the student network's ability to perceive changes in intermediate fusion features, thereby strengthening the effectiveness of multi-scale feature fusion. The mean absolute error loss can be used to predict the probability map of the student network's output. And the predicted probability graph output by the teacher network. Supervision is conducted to achieve knowledge distillation, enabling student networks to learn high-precision predictive information from teacher networks, thereby improving the detection performance of student networks.
[0046] In one feasible embodiment, the first loss value is calculated using the cross-entropy loss function. The calculation can be performed using the second predicted probability map output by the student network. The predicted probability and the corresponding binary transformation mask are input into the cross-entropy loss function. The function compares the predicted probability with the true label pixel by pixel, and finally obtains the first loss value. Second loss value Calculated using the binary cross-entropy loss function. Intermediate fused features can be considered during the calculation. By inputting a binary cross-entropy loss function, which supervises and constrains the changing responses of the intermediate fused features, the second loss value can be obtained. Third loss value Calculated using the mean absolute error loss function. The calculation uses the first predicted probability graph output by the teacher network. The second predicted probability graph output by the student network The two probabilities are input into the mean absolute error loss function, which calculates the absolute difference between the two probability maps pixel by pixel and takes the average value, thus obtaining the third loss value. Then, the weighting coefficients can be used. , , By weighting and fusing the three types of losses separately, the total loss function can be obtained. The mathematical expression is: The weighting coefficients can be set according to the actual application scenario; in this embodiment, they are preferably all set to 1 / 3.
[0047] In one feasible embodiment, after calculation Then, the backpropagation algorithm can be used, based on... The gradients of each parameter in the student network are calculated, and the parameters of the student network are updated using the AdamW optimizer, while the parameters of the teacher network remain unchanged. Furthermore, after each preset number of training epochs, validation set data can be input into the student network to evaluate the model's detection performance, and the weights of the student network with the best performance on the validation set are saved as candidate final training models. Training stops when the model reaches the preset number of training epochs, or when the validation set performance no longer improves for several consecutive epochs, resulting in a fully trained target image change detection model.
[0048] In a feasible embodiment, the stem-layer fusion module of the target image change detection model can be regarded as a primary extraction and fusion module for dual-temporal features. It is mainly used for primary feature extraction of dual-temporal remote sensing images and to achieve refined early fusion, providing high-quality fused features for subsequent encoder feature extraction. The encoder module is mainly responsible for hierarchical multi-scale feature extraction of the temporal fusion features output by the stem-layer fusion module. Through progressive downsampling and feature abstraction, it obtains semantic features at multiple different scales from shallow to deep and from fine to coarse, providing a multi-level feature representation basis for subsequent multi-scale calibration fusion. The multi-scale calibration fusion module integrates the core ideas of Effective Multi-Scale Feature Fusion (EMFF) and Spatial-Spectral Feature Collaboration (SSFC). It is a module without any learnable parameters and can replace the feature recovery and fusion functions of traditional decoders. This module is mainly responsible for achieving semantic enhancement, hierarchical fusion, and adaptive recalibration of encoder multi-scale features through parameterless operations, explicitly modeling spatial-spectral dependencies, suppressing noise response, and solving the problems of poor feature fusion effect, multi-scale redundancy, and noise interference in early fusion schemes. The change detection head is mainly responsible for receiving the multi-channel fusion features output by the multi-scale calibration fusion module, and performing pixel-level change determination based on the fusion features to generate a binary change mask.
[0049] In a feasible embodiment, before inputting the dual-temporal remote sensing image data to be detected into the target image change detection model, preprocessing operations such as spatial registration, resolution unification, and pixel value normalization can be performed on the dual-temporal remote sensing images to ensure that the input data is consistent with the data distribution during the model training phase. If the image size is too large, exceeding the model input requirements or the graphics memory capacity, block cropping can be performed, for example, cropping it into 256×256 image blocks, and completing the change detection of the entire image in a block-by-block manner.
[0050] In a feasible embodiment, the stem-layer fusion module is the core carrier of the refined early fusion strategy. It can be used to replace the simple channel stitching method of existing early fusion schemes, realizing refined extraction and deep interactive fusion of primary features of dual-temporal remote sensing images. This provides high-quality fusion features for subsequent encoder modules and improves the effectiveness of feature extraction. The stem-layer fusion module can perform independent primary feature extraction on dual-temporal images, retaining the primary feature information of each dual-temporal image. Then, through stitching, convolution, and depthwise separable convolution, it achieves deep fusion of the dual-temporal primary features, reducing computational complexity while enhancing the interactivity of dual-temporal features. The specific process of obtaining temporal fusion features by performing feature extraction and fusion processing on dual-temporal remote sensing image data through the stem-layer fusion module is as follows: Primary feature extraction: for previous temporal images Input a dedicated stem-layer convolutional layer (Convpre) to extract primary features from the preceding temporal phase. ; For later phase images Input a dedicated stem-layer convolutional layer (Convpost) to extract post-temporal primary features. .
[0051] The mathematical expression is: Convpre and Convpost are both 3×3 convolutional layers used to convert the original image pixel features into primary semantic features, preserving details such as image edges and textures.
[0052] Feature stitching: combining primary features from two temporal phases and By concatenating along the channel dimension, we obtain the concatenated feature X, mathematically expressed as: X = Concat( , By splicing, the initial fusion of the two-phase features is achieved, preserving the complete information of the two-phase features.
[0053] Deep fusion and dimensionality adjustment: The concatenated feature X is sequentially subjected to 1×1 convolution, depthwise separable convolution (DWConv), and then another 1×1 convolution to achieve deep interactive fusion of features and channel dimension adjustment, generating the final temporal fusion feature F, mathematically expressed as: Among them, 1×1 convolution is used to adjust the number of feature channels to avoid the increase in computational complexity caused by too many channels; depthwise separable convolution reduces computational complexity while fully capturing the spatial local information of features, enhancing the deep interaction of dual-temporal features, and improving the quality of fused features.
[0054] In one feasible embodiment, the encoder module is the core of the feature extraction in the target image change detection model. It is mainly used to extract multi-scale deep semantic features from the temporal fusion features F output by the stem-layer fusion module, providing a rich multi-scale feature foundation for the subsequent multi-scale calibration fusion module. The encoder module in this embodiment is designed according to the principles of universality, flexibility, and compatibility. It supports direct access to various mainstream deep learning backbone networks currently in use without any modification to the backbone network. It can be flexibly selected according to the performance / efficiency requirements of the actual application, specifically including: A. CNN-type backbone networks: such as ResNet-34, ResNet-50, EfficientNet, UNet backbone, etc., with the advantages of high computational efficiency and low deployment difficulty.
[0055] B. Transformer-type backbone networks: such as FocalNet, SegFormer (MIT series), Swin-T, DebiFormer, etc., have the advantage of strong global feature modeling capabilities and are suitable for change detection in complex scenes.
[0056] C. Mamba-type backbone networks: such as RSMamba and ChangeMamba, which have the advantage of global modeling with linear complexity, balancing efficiency and performance.
[0057] In one feasible embodiment, the encoder module performs multi-scale feature extraction on the temporal fusion feature F to obtain features at multiple different scales. For example, the encoder module performs multi-stage downsampling and feature extraction on the fusion feature F to generate four multi-scale features at different scales. (j∈[1,2,3,4]), where These are shallow features, containing low-level information such as the image's edges, textures, and details; These are deep features, containing high-level information such as the semantics of ground features, global structure, and spatial relationships within the image. Multi-scale features. From shallow to deep layers, semantic information is gradually enhanced while detailed information is gradually reduced, providing a rich feature base for hierarchical fusion of the multi-scale calibration fusion module.
[0058] In one feasible embodiment, the multi-scale calibration fusion module can be used to process the multi-scale features extracted by the encoder. Semantic enhancement, hierarchical fusion, resolution restoration, and adaptive recalibration are performed. Spatial-spectral dependencies are explicitly modeled to highlight semantic information in changing regions of multi-scale features, suppress noise response, and generate high-quality multi-channel fused features. The core logic of the multi-scale calibration and fusion module for calibrating and fusing features at multiple different scales includes multi-scale feature resolution unification and hierarchical fusion, as well as spatial-spectral feature collaborative calibration. The overall calibration and fusion process can be summarized as follows: First, multi-scale feature resolution unification is achieved through interpolation; then, semantic enhancement and detail fusion are achieved through parameterless hierarchical fusion. Subsequently, the SSFC parameterless attention mechanism is introduced to generate three-dimensional attention weights for the fused features and perform element-wise adaptive recalibration to suppress redundant responses and noise interference, ultimately generating high-quality multi-channel fused features. The entire module has no learnable parameters and does not increase the model's parameter scale or computational complexity.
[0059] In one feasible embodiment, combined with Figure 5 and Figure 6 As shown, the specific process of step 150, which uses a multi-scale calibration and fusion module to calibrate and fuse features at multiple different scales to obtain multi-channel fused features, may include, but is not limited to, steps 510 to 540.
[0060] Step 510: Perform bilinear interpolation on features at multiple different scales, and then perform feature fusion enhancement on the interpolated features at each scale to obtain preliminary multi-scale fused features; Step 520: Perform feature splitting on the preliminary multi-scale fusion features to obtain the first feature branch, the second feature branch, and the third feature branch; Step 530: Determine the feature relationship weights based on the first feature branch and the second feature branch; Step 540: Obtain multi-channel fusion features based on feature relationship weights and the third feature branch.
[0061] In one feasible embodiment, multiple features at different scales include a first feature. Second feature Third feature and the fourth feature .in, S1 is a shallow feature, containing rich low-level detailed information; S4 is a deep feature, containing high-level semantic information. From arrive The semantic information of the features gradually increases, while the spatial detail information gradually decreases.
[0062] In a feasible embodiment, the process of performing feature fusion enhancement processing on the interpolated features at each scale to obtain preliminary multi-scale fused features can be performed according to the following steps: First, for the interpolated fourth feature... Performing average pooling along the channel dimension yields the fourth pooling feature. Its channel number and third feature Same; secondly, regarding the third feature The mean and hyperbolic tangent are applied to obtain the first attention weight w; then, the pooling fourth feature is processed according to the first attention weight w. Perform weighting, and then pool the fourth feature of the weighted pool. With the third feature By fusing the data, we can obtain enhanced deep features. Subsequently, the deep features of the fusion enhancement were analyzed. Perform average pooling along the channel dimension to obtain pooled features. Pooling features With the second feature After addition, max pooling is performed to obtain the fused enhanced features. Finally, based on this fusion enhancement feature... First feature and the fourth feature Multi-scale feature aggregation is performed to obtain preliminary multi-scale fused features. Specifically, in this process, the multi-scale features output by the encoder are first interpolated to the shallow features using bilinear interpolation. Maintaining the same spatial resolution ensures spatial consistency for subsequent feature fusion; then, the interpolated deep features... Performing average pooling along the channel dimension yields the fourth pooling feature. ,make Channel number and mid-layer features Consistent; then on The mean value is taken along the channel dimension, and an adaptive attention weight (i.e., the first attention weight w) is generated by the tanh activation function; then the fourth feature is pooled using the attention weight w. Weighting is applied to enhance the semantics of deep features, and then combined with... Addition achieves fusion, which can be expressed mathematically as: Where meanc is the channel-level mean operation, AvgPoolc is the channel-level average pooling operation, and ⊙ is the element-wise multiplication operation. This refers to the enhanced mid-to-deep features after fusion. (The text then repeats the description of enhanced mid-to-deep features.) Perform channel-dimensional average pooling to reduce the channel dimension and obtain ,Will Features of the middle and shallow layers After addition, max pooling is performed to highlight significant changes in the features, resulting in fused and enhanced features. The mathematical expression is: MaxPool is a spatial max-pooling operation used to highlight salient features in varying regions and suppress background noise. Furthermore, it integrates and enhances features. shallow features The images are then combined to fully integrate shallow details and mid-to-deep semantic information, capturing fine-grained changes in the image; the fusion result is then combined with the original deep features. By stitching along the channel dimension, preliminary multi-scale fusion features are obtained. The mathematical expression is: .
[0063] In one feasible embodiment, after obtaining preliminary multi-scale fusion features After that, you can... A parameterless attention mechanism, SSFC, is introduced to achieve adaptive recalibration of features and suppress redundant responses and noise interference. Specifically, firstly, the feature is... It is broken down into three feature branches: query Q, key K, and value V. Query Q is the first feature branch, key K is the second feature branch, and value V is the third feature branch. Q is... K and V are obtained through channel average pooling, and are directly derived from... The decomposition yields Q, which serves as a benchmark reference for global prior features; K, which can be used to characterize the feature distribution of each spatial location and spectral channel; and V, which can be used to carry the original multi-scale fused feature information. The three work together through relationship modeling and attention weighting to achieve feature calibration in the spatial-spectral dimension.
[0064] In a feasible embodiment, when determining the feature relationship weights based on the first feature branch Q and the second feature branch K, Gaussian kernel relationship modeling can be performed on K and Q to calculate the feature relationship weight R, the calculation formula of which is: Where σ is a preset hyperparameter used to control the bandwidth of the Gaussian kernel; ε is a local minimum value used to prevent the denominator from being zero. This operation explicitly models the space-spectral dependency through the Gaussian kernel function.
[0065] In a feasible embodiment, when obtaining multi-channel fused features based on feature relation weights and the third feature branch, the relation weights R can be converted into three-dimensional attention weights using a Sigmoid activation function. , and then Element-wise multiplication with V yields the adaptively recalibrated features. Its mathematical expression is: =Sigmoid(R) = ⊙V. (This likely refers to a specific type of V or similar format.) The final output multi-channel fusion feature of the multi-scale calibration fusion module ,Right now: = .
[0066] In one feasible embodiment, the change detection head is the last functional module of the image change detection model, which can be used to process the fusion features output by the multi-scale calibration fusion module. The changes are converted into pixel-level binary change masks to achieve final identification and localization of the changed regions. The change detection head consists of 1×1 convolutional layers and a sigmoid activation function, without a complex layer structure. The change detection head relies on fused features. The specific implementation process for performing change detection and obtaining a binary change mask is as follows: Channel dimension adjustment: merging features Input a 1×1 convolutional layer to convert multi-channel fused features into single-channel features, thereby reducing the dimensionality of the channels. Change probability generation: Input single-channel features into the Sigmoid activation function to map feature values to the [0,1] interval, generating a pixel-level change probability map. The closer the value is to 1, the higher the probability that the pixel is a changing region; the closer the value is to 0, the higher the probability that it is a non-changing region. Binarization: For the probability map of change A fixed threshold (preferably 0.5) is set for binarization to generate a binary transformation mask. The mathematical expression is: =Sigmoid(Conv1×1( )), =Binary( ,0.5).
[0067] In a feasible embodiment, when obtaining the image change detection result based on the binary change mask, if the dual-temporal remote sensing image data is segmented image data, the binary change masks corresponding to all segmented image data can be stitched together according to their original spatial positions to restore the complete change mask of the original image resolution; then the change mask of the original image is visualized, including the marking of change areas and the overlay display with the original image, and finally the intuitive image change detection result is output.
[0068] The method for detecting changes in remote sensing images according to this application will be described below with a specific embodiment.
[0069] like Figure 7 As shown, the detection process of the target image change detection model is as follows: First, acquire the dual-temporal remote sensing image data to be detected, including temporal phase 1 and temporal phase 2. Before inputting the data into the model, preprocess the dual-temporal images. Depending on the actual application scenario, operations such as spatial registration, resolution unification, pixel value normalization, and block cropping can be performed to ensure that the quality of the input data is consistent with the data distribution during the model training phase. The preprocessed dual-temporal images are input into the stem-layer fusion module, which performs primary feature extraction and early fusion of the dual-temporal images to obtain temporal fusion features, providing high-quality fusion input for subsequent multi-scale feature extraction. The temporal fusion features are then input into the encoder module, which performs multi-scale feature extraction processing to obtain features at multiple different scales. (For example to This encompasses multi-level feature representations, ranging from shallow details to deep semantics. Multi-scale features to The input is processed by the multi-scale calibration fusion module, which performs parameterless multi-scale collaborative calibration fusion to obtain multi-channel fusion features. This achieves semantic enhancement and noise suppression of multi-scale features. The multi-channel fused features are then input into the change detection head, and after 1×1 convolution and sigmoid activation, they are converted into pixel-level change probability maps. Then, a binary transformation mask is generated through binarization. This method achieves pixel-level identification of changed areas. If the input dual-temporal remote sensing image data is segmented, the binary change masks corresponding to all segmented image data are stitched together according to their original spatial locations to restore the complete change mask at the original image resolution. Subsequently, the change mask of the original image is visualized, including methods such as marking changed areas and overlaying it with the original image, ultimately outputting intuitive image change detection results.
[0070] Reference Figure 8 , Figure 8 This is an architecture diagram of a remote sensing image change detection system provided in an embodiment of this application. This system can be applied to the remote sensing image change detection method in any of the foregoing embodiments, and can be deployed on various hardware platforms such as servers, edge computing devices, embedded devices, and terminal computers to achieve automated and intelligent processing of remote sensing image change detection. The system 800 consists of multiple core functional modules, and the functions and technical implementations of each module are as follows: The data input module 810 is used to acquire the dual-temporal remote sensing image data to be detected. It supports common remote sensing image formats such as TIFF, PNG, and JPG, and supports multiple input methods such as local file import, network transmission, remote sensing satellite data interface acquisition, and batch folder import.
[0071] The change detection module 820 is used to input dual-temporal remote sensing image data into a pre-trained target image change detection model. This model includes a stem-and-layer fusion module, an encoder module, a multi-scale calibration fusion module, and a change detection head, connected sequentially. The model's detection process includes: feature extraction and fusion processing of the dual-temporal remote sensing image data through the stem-and-layer fusion module to obtain temporal fusion features; multi-scale feature extraction processing of the temporal fusion features through the encoder module to obtain features at multiple different scales; calibration and fusion processing of the features at multiple different scales through the multi-scale calibration fusion module to obtain multi-channel fusion features; and change detection by the change detection head based on the multi-channel fusion features to obtain a binary change mask. The result output module 830 is used to obtain image change detection results based on the binary change mask. Specifically, it can convert the change detection results (such as binary change mask, change probability map) into common image formats for output, and provides a variety of visualization functions, including color rendering of changed areas, overlay display of changed areas and original images, change area statistics, and generation of detection result reports, which are convenient for users to view, analyze and export.
[0072] In one feasible embodiment, the remote sensing image change detection system further includes a data preprocessing module 840, used to convert dual-temporal remote sensing image data into standardized input data. Specifically, this module can perform automated preprocessing of the input dual-temporal images, including at least one of the following functions: spatial registration, resolution unification, pixel value normalization, image segmentation / stitching, and format conversion, providing standardized input data for subsequent feature extraction.
[0073] In one feasible embodiment, the change detection module 820 includes a feature fusion extraction unit 821, a multi-scale calibration fusion unit 822, and a change detection unit 823. The feature fusion extraction unit 821 corresponds to the stem-layer fusion module and the encoder module, realizing early refined fusion of dual-temporal features and multi-scale deep semantic feature extraction, generating multi-scale features for the encoder. The multi-scale calibration fusion unit 822 corresponds to the multi-scale calibration fusion module. It has no learnable parameters and achieves multi-scale feature processing. Semantic enhancement, hierarchical fusion, resolution restoration and adaptive recalibration, explicit modeling of space-spectral dependencies, suppression of noise response, and generation of high-quality fused features. The change detection unit 823 corresponds to the change detection head, which will fuse features. The data is converted into a pixel-level change probability map and then binarized to generate a binary change mask, thus achieving the final identification of the change region.
[0074] It should be noted that the functions of each module of this remote sensing image change detection system 800 correspond to those of the previous embodiments. For details on the working principles of each module of system 800, please refer to the relevant embodiments mentioned above, which will not be repeated here.
[0075] It should be noted that, while ensuring that the core concepts of encoder-only architecture, refined early fusion, parameter-free multi-scale collaborative calibration fusion, and teacher-student distillation training remain unchanged, there are various alternative solutions for its specific technical implementation. These alternative solutions can also achieve the purpose of the embodiments of this application, as follows: Alternatives to the stem-layer fusion module: The stem-layer fusion module in this embodiment uses a fusion method of stem-layer convolution, concatenation, convolution, and depthwise separable convolution. However, it can also be replaced by a fusion method of stem-layer convolution, element-wise addition, and dilated convolution, or by a fusion method of stem-layer convolution, attention-weighted convolution, and pointwise convolution. Alternatively, a fusion method of stem-layer convolution and concatenated group convolution can be used. As long as it can achieve early and refined extraction and deep fusion of dual-temporal features without significantly increasing computational complexity, any fusion method that achieves the same technical effect can be adaptively adopted. This embodiment does not impose any restrictions on this.
[0076] The tanh activation function used in the multi-scale calibration fusion module can be replaced with parameterless activation functions such as Sigmoid and ReLU; the average pooling and max pooling used can be replaced with parameterless pooling operations such as adaptive average pooling and adaptive max pooling; and the Gaussian kernel relation modeling can be replaced with other parameterless distance metrics (such as Euclidean distance and Manhattan distance). As long as no learnable parameters are introduced, and the semantic enhancement, hierarchical fusion, resolution restoration, and adaptive recalibration of multi-scale features can be achieved, the original operations can be replaced. This embodiment does not impose any restrictions on this.
[0077] Alternatives to the teacher network: This application uses a high-precision network with late fusion as the teacher network. It can be replaced by any high-performance change detection network in the current technology, including high-precision networks with early fusion and various change detection networks based on CNN / Transformer / Mamba. As long as the teacher network can generate high-precision prediction results and provide effective knowledge distillation supervision information for the student network, it is applicable.
[0078] Alternative loss functions: This application's embodiments employ a fusion strategy of cross-entropy loss, binary cross-entropy loss, and mean absolute error loss. The mean absolute error loss can be replaced with distilled losses such as mean squared error loss and KL divergence loss. The weight coefficients of each loss function can also be adjusted according to actual needs. As long as a multi-loss fusion strategy is used to achieve multi-dimensional supervision of the student network, the same training effect can be achieved.
[0079] Alternatives to the encoder backbone network: The encoder module in this application supports CNN, Transformer, and Mamba type backbone networks, and can be replaced with any deep learning backbone network in the current technology, such as DenseNet, VisionTransformer (ViT), SWIN-Mamba, etc., without modifying the backbone network, and can be directly connected to the encoder-only architecture.
[0080] Alternative data augmentation methods: The augmentation methods used in this application embodiment, such as random flipping, color jittering, scale cropping, and Gaussian blur, can be replaced by data augmentation methods such as random rotation, brightness / contrast adjustment, salt and pepper noise addition, and random erasure. As long as they can improve the generalization ability of the model and avoid overfitting, they are applicable.
[0081] The above alternative solutions do not change the core concept of the embodiments of this application, namely, the encoder architecture alone, refined early fusion, parameterless multi-scale collaborative calibration fusion, and teacher-student distillation training can all achieve lightweighting of the model and improvement of computational efficiency while ensuring detection performance, and are within the protection scope of the embodiments of this application.
[0082] This application proposes a method and system for detecting changes in remote sensing images that is simple in structure, lightweight, computationally efficient, and highly accurate. It effectively resolves the inherent contradiction between performance and efficiency, meeting the high-precision, high-efficiency, and lightweight requirements for remote sensing image change detection in practical application scenarios such as land resource monitoring, urban planning, and disaster emergency response. The beneficial effects of this application include at least the following: 1. Extremely lightweight design, significantly reducing computational complexity: The decoder structure is completely abandoned, and its function is replaced by a parameterless multi-scale calibration fusion module. At the same time, the computational overhead of the twin encoder is eliminated through an early fusion strategy. It is expected that the number of model parameters, floating-point operations (FLOPs), and inference latency will be significantly reduced, far below existing traditional solutions. This facilitates deployment on edge devices and embedded devices and meets real-time detection requirements.
[0083] 2. High-precision detection, balancing performance and efficiency: The quality of fused features is improved through a refined stem-layer fusion module, and efficient fusion and adaptive recalibration of multi-scale features are achieved through a parameter-free multi-scale calibration fusion module. Spatial-spectral dependencies are explicitly modeled to suppress noise response. The high-precision knowledge of the teacher network is learned by combining the teacher-student distillation framework. It is expected that the student network with only encoder architecture can achieve detection performance comparable to or even better than existing late-stage fusion high-precision schemes, thus solving the inherent contradiction between performance and efficiency in current technologies.
[0084] 3. Simple structure, strong versatility and flexibility: The student network architecture is simple, and the encoder module supports direct access to various CNN, Transformer and Mamba backbone networks without modification. It can be flexibly selected according to the performance / efficiency requirements of actual applications to adapt to different application scenarios.
[0085] 4. Highly efficient training and inference with strong practicality: The training process adopts an end-to-end teacher-student distillation framework, which eliminates the need for complex pre-training and fine-tuning; the inference process requires only one forward propagation of the student network, without the participation of the teacher network, resulting in fast inference speed and suitability for batch automated processing of large-scale remote sensing images.
[0086] 5. The embodiments of this application break the inherent dependence of the current technology on the encoder-decoder architecture, verify that the decoder is a redundant component, and introduce a parameterless collaborative calibration and fusion mechanism to explicitly model the spatial-spectral dependence relationship. This provides a brand-new research idea and technical direction for the lightweight model design of remote sensing image change detection, and has important theoretical significance and technical leading role.
[0087] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for detecting changes in remote sensing images, characterized in that, include: Acquire the dual-temporal remote sensing image data to be detected; The dual-temporal remote sensing image data is input into a pre-trained target image change detection model, which includes a stem-layer fusion module, an encoder module, a multi-scale calibration fusion module, and a change detection head connected in sequence. The stem-layer fusion module performs feature extraction and fusion processing on the dual-temporal remote sensing image data to obtain temporal fusion features. The encoder module performs multi-scale feature extraction on the temporal fusion features to obtain features at multiple different scales. The multi-scale calibration and fusion module performs calibration and fusion processing on the features at multiple different scales to obtain multi-channel fused features. The change detection head performs change detection based on the multi-channel fusion features to obtain a binary change mask. The image change detection result is obtained based on the binary change mask.
2. The remote sensing image change detection method according to claim 1, characterized in that, The process of calibrating and fusing features at multiple different scales through the multi-scale calibration and fusion module to obtain multi-channel fused features includes: Bilinear interpolation is performed on the features at multiple different scales, and feature fusion enhancement is performed on the interpolated features at each scale to obtain preliminary multi-scale fused features. The preliminary multi-scale fusion features are subjected to feature splitting to obtain a first feature branch, a second feature branch, and a third feature branch. The feature relationship weights are determined based on the first feature branch and the second feature branch; Based on the feature relationship weights and the third feature branch, multi-channel fusion features are obtained.
3. The remote sensing image change detection method according to claim 2, characterized in that, The multiple features at different scales include a first feature, a second feature, a third feature, and a fourth feature; The feature fusion and enhancement process performed on the interpolated features at each scale yields preliminary multi-scale fused features, including: The interpolated fourth feature is subjected to average pooling along the channel dimension to obtain a pooled fourth feature, wherein the number of channels of the pooled fourth feature is the same as that of the third feature. The third feature is processed by mean and hyperbolic tangent to obtain the first attention weight; The pooling fourth feature is weighted according to the first attention weight, and the weighted pooling fourth feature is fused with the third feature to obtain the fused enhanced deep feature. The deep features in the fusion enhancement are subjected to channel-dimensional average pooling to obtain pooled features; The pooling feature and the second feature are added together and then max pooling is performed to obtain the fused enhanced feature; Based on the fusion enhancement features, the first feature, and the fourth feature, preliminary multi-scale fusion features are obtained.
4. The remote sensing image change detection method according to claim 1, characterized in that, The training steps of the target image change detection model include: Acquire raw image detection data, which includes dual-temporal remote sensing image data samples and corresponding binary change mask annotations; divide the raw image detection data into a training set and a validation set. An initial image change detection model is constructed, which includes an initial stem-layer fusion module, an initial encoder module, an initial multi-scale calibration fusion module, and an initial change detection head; Obtain a late-stage fusion change detection network; The training set is input into the late fusion change detection network for forward propagation to obtain a first prediction probability map; The training set is input into the initial image change detection model for forward propagation to obtain intermediate fusion features and a second prediction probability map. Based on the first prediction probability map, the intermediate fusion feature, the second prediction probability map, and the binary change mask annotations corresponding to the dual-temporal remote sensing image data samples in the training set, the total loss function is determined according to the preset loss function. The parameter gradient of the initial image change detection model is determined based on the total loss function, and the parameters of each module in the initial image change detection model are updated based on the parameter gradient to obtain the updated image change detection model. The detection performance of the updated image change detection model is evaluated based on the validation set, and the target weight of the model is determined based on the performance test results. When the model training reaches the preset number of training rounds or the performance of the validation set detection no longer improves for several consecutive rounds, the target image change detection model is obtained according to the model target weights.
5. The remote sensing image change detection method according to claim 4, characterized in that, Before dividing the raw image detection data into a training set and a validation set, the method further includes: Spatial registration, resolution unification, and pixel value normalization are performed on the dual-temporal remote sensing image data samples to obtain standard dual-temporal remote sensing image data samples. The standard dual-temporal remote sensing image data samples are subjected to data augmentation processing to obtain the processed image detection raw data.
6. The remote sensing image change detection method according to claim 4, characterized in that, The preset loss function includes at least one of the following: cross-entropy loss function, binary cross-entropy loss function, and mean absolute error loss function.
7. The remote sensing image change detection method according to claim 6, characterized in that, The loss function includes the cross-entropy loss function, the binary cross-entropy loss function, and the mean absolute error loss function; The step of determining the total loss function according to a preset loss function based on the first predicted probability map, the intermediate fusion feature, the second predicted probability map, and the binary change mask annotations corresponding to the dual-temporal remote sensing image data samples in the training set includes: The first loss value is calculated using the cross-entropy loss function based on the binary change mask labeling corresponding to the second predicted probability map and the dual-temporal remote sensing image data sample; The second loss value is calculated based on the intermediate fusion features using the binary cross-entropy loss function. The third loss value is calculated using the mean absolute error loss function based on the first prediction probability map and the second prediction probability map; The total loss function is obtained by weighting and fusing the first loss value, the second loss value, and the third loss value.
8. The remote sensing image change detection method according to claim 1, characterized in that, The step of obtaining the image change detection result based on the binary change mask includes: If the dual-temporal remote sensing image data is segmented image data, the binary change mask of all segmented image data is stitched together to obtain the change mask of the original image. The change mask of the original image is visualized to obtain the image change detection results.
9. A remote sensing image change detection system, characterized in that, The system, applied to the remote sensing image change detection method as described in any one of claims 1-8, comprises: The data input module is used to acquire the dual-temporal remote sensing image data to be detected; The change detection module is used to input the dual-temporal remote sensing image data into a pre-trained target image change detection model. The target image change detection model includes a stem-layer fusion module, an encoder module, a multi-scale calibration fusion module, and a change detection head connected in sequence. The stem-layer fusion module performs feature extraction and fusion processing on the dual-temporal remote sensing image data to obtain temporal fusion features. The encoder module performs multi-scale feature extraction processing on the temporal fusion features to obtain features at multiple different scales. The multi-scale calibration fusion module performs calibration fusion processing on the features at multiple different scales to obtain multi-channel fusion features. The change detection head performs change detection based on the multi-channel fusion features to obtain a binary change mask. The result output module is used to obtain the image change detection result based on the binary change mask.
10. The remote sensing image change detection system according to claim 9, characterized in that, The system also includes a data preprocessing module for converting the dual-temporal remote sensing image data into standardized input data.