A time-series urban land surface change detection method based on SAR and optical image fusion
Patent Information
- Application Number
- CN202611132267.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-29
- Publication Date
- 2026-08-28
AI Technical Summary
在云雨频繁地区,光学影像的可用性大幅降低,导致方法在实际应用中受限;同时,在复杂地表场景下,单一模态特征对建筑物施工变化的判别能力不足
(1)多源遥感数据协同融合,突破单一光学数据在建筑结构变化感知上的能力瓶颈。本发明融合 Sentinel-1 SAR与Sentinel-2光学两类遥感数据源,借助SAR云填充与时序克里金插值完成光学影像云污染修复,并通过DTW算法完成多源影像时空配准,构建完整连续的长时序观测序列。本发明避免了单一遥感数据源易受天气和气候变化影响的问题。相比基于单一数据源的方法,能够充分利用SAR的全天候观测能力和光学影像的丰富光谱纹理信息,在云雨频繁、地表复杂的实际场景中显著提升了变化检测的精度和鲁棒性;
Smart Images

Figure CN122657747A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of multi-source remote sensing temporal change detection technology, and in particular to a method for accurate detection of long-term construction changes in buildings based on SAR and optical image fusion, spatiotemporal alignment and temporal Transformer network. Background Technology
[0002] Multispectral remote sensing imagery, as an important means of acquiring regional surface information, is widely used in the observation and research of urban changes. Optical imagery, with its rich spatial and spectral information, clear texture features, wide swath, and short revisit period, can better reflect differences in ground features and is an important data source for Earth observation. However, optical imagery is limited by its passive imaging characteristics and is susceptible to weather and climate conditions, restricting its application in all-weather monitoring. In contrast, Synthetic Aperture Radar (SAR), through active imaging, is not limited by weather and lighting conditions and can achieve all-weather, all-time surface monitoring.
[0003] Currently, in SAR time series analysis, Multi-Temporal Synthetic Aperture Radar (MT-InSAR) is an advanced method for monitoring surface deformation. This technique, by analyzing multiple SAR images acquired at different times, can monitor minute surface deformations along the radar line of sight with high precision. It is widely used in studies of surface deformation caused by natural processes such as mining, landslides, and volcanic activity, as well as various human activities. Traditional MT-InSAR analysis typically selects long-term coherent and stable permanent scatterers (PS) as the basis for research. However, in rapidly developing urban areas, frequent new construction, demolition, and road reconstruction activities cause abrupt changes in the surface environment, posing challenges to the processing and analysis of long-term data series.
[0004] In recent years, deep learning-based remote sensing change detection methods have made significant progress. Specifically, deep learning-based satellite image time series (SITS) analysis methods have achieved remarkable advancements. Unlike bi-temporal change detection, SITS methods use multi-temporal observation sequences of single pixels as input, capturing the dynamic evolution of the land surface through temporal modeling, demonstrating excellent performance in tasks such as crop classification and land cover mapping. The Temporal Convolutional Neural Network (TCNN) proposed by Pelletier et al. (2019) is a foundational work in SITS pixel-level classification. This method uses three stacked one-dimensional convolutions to extract temporal features, capturing multi-scale temporal dependencies by progressively expanding the temporal receptive field of the convolutional kernels, and finally obtaining the classification result of the entire sequence through global pooling. This method has few parameters, high training efficiency, and has been widely validated in various SITS classification tasks. Garnot and Landrieu's (ICCV 2021) L-TAE (Lightweight Temporal Attention Encoder) further introduces a multi-head temporal self-attention mechanism. It uses a set of learnable lightweight query vectors to perform attention-weighted encoding on the entire temporal sequence, achieving performance comparable to complex models while maintaining extremely low parameter counts. It also supports temporal inputs of unequal length and irregular sampling. However, the above method has the following design limitations, making it difficult to directly apply to urban land surface change detection tasks: (1) The data source is singular and does not utilize multi-source remote sensing data. The above methods use optical remote sensing images as the main data source and do not make full use of the all-weather and all-time observation advantages of SAR images. In areas with frequent cloud and rain, the availability of optical images is greatly reduced, which limits the practical application of the methods; at the same time, in complex surface scenes, the ability of single modal features to distinguish changes in building construction is insufficient.
[0005] (2) Limited temporal receptive field or compressed coding limits the full capture of long temporal change patterns. TempCNN relies on the limited receptive field of three convolutional layers, which makes it difficult to cover the entire construction cycle across multiple observation periods. Although LTAE has a global attention field, it compresses the attention of the entire time series into a dot product between a small number of learnable query vectors and the sequence. The limited capacity of the query vectors constitutes an information bottleneck in the face of long sequences, making it difficult to retain the fine temporal structure of changes throughout the entire construction cycle. Summary of the Invention
[0006] To address the shortcomings of existing technologies, this invention proposes a temporal urban surface change detection method based on the fusion of SAR and optical imagery. By leveraging the synergistic complementarity of SAR and optical multi-source data, heterogeneous branch temporal deep modeling, and time-step classification, it achieves high-precision spatial change detection while accurately pinpointing the specific time points where changes occur. This method is suitable for large-scale, long-term monitoring of urban building construction dynamics. This invention focuses on urban areas, fully utilizing the complementary advantages of optical and SAR imagery. Based on the spatiotemporal distribution characteristics identified through the collaborative identification of multi-source remote sensing data, it constructs a deep temporal change detection model to reveal the evolution patterns of urban built-up areas, providing data support and methodological reference for urban planning.
[0007] Technical solution A method for detecting temporal urban surface changes based on SAR and optical image fusion includes the following steps: Step 1: Using Sentinel-1 SAR and Sentinel-2 optical images as data sources, cloud removal processing of optical images is achieved through SAR-assisted filling and spatiotemporal kriging interpolation. Multi-source image spatiotemporal alignment is achieved by combining DEM terrain correction, orthorectification and dynamic time warping (DTW), and a dataset is constructed based on visual interpretation. Step 2: Construct a heterogeneous branch Transformer deep learning model, extract optical and SAR heterogeneous features through multi-scale temporal convolution, achieve adaptive fusion by using channel attention mechanism, and output change detection results step by step after capturing long temporal dependencies by the Transformer encoder. Step 3: Obtain the final change mask map and change time map through cascaded spatial denoising.
[0008] Beneficial effects Compared with the prior art, the present invention has the following significant advantages: (1) Multi-source remote sensing data fusion overcomes the bottleneck of single optical data in sensing changes in building structures. This invention integrates two types of remote sensing data sources: Sentinel-1 SAR and Sentinel-2 optical data. It uses SAR cloud filling and temporal kriging interpolation to complete cloud pollution repair of optical images, and uses the DTW algorithm to complete spatiotemporal registration of multi-source images to construct a complete and continuous long-term observation sequence. This invention avoids the problem of single remote sensing data sources being susceptible to weather and climate change. Compared with methods based on a single data source, it can make full use of the all-weather observation capability of SAR and the rich spectral texture information of optical images, significantly improving the accuracy and robustness of change detection in real-world scenarios with frequent clouds and rain and complex surface conditions. (2) The combination of heterogeneous branch design and all-pair, all-self attention encoding overcomes the dual limitations of homogeneous channel processing and limited temporal receptive field. This invention constructs a heterogeneous branch temporal change detection network, relying on multi-scale temporal convolution to capture geomorphic features across different time spans, adaptively fusing multi-source heterogeneous features through a channel attention mechanism, and utilizing a temporal Transformer encoder to capture long-term global dependencies. The heterogeneous branch design of this invention enables each modality to independently extract features at the optimal receptive field and network depth. The temporal Transformer encoder breaks through the information bottleneck of fixed receptive field and compressed encoding, and can effectively learn the temporal evolution law of the entire building construction process.
[0009] (3) Dense prediction at each time step enables synchronous and accurate positioning of the spatial distribution of changes and the time nodes of abrupt changes. This invention is designed for the task of detecting changes in the construction time series. It can synchronously and accurately identify the spatial distribution range and abrupt time nodes of changes in surface buildings. It has a significant advantage in the positioning accuracy of time series changes, greatly improving the efficiency and timeliness of urban surface change monitoring. It is suitable for large-scale, long-term dynamic monitoring scenarios of surface construction, and provides reliable technical support for urban planning and management. Attached Figure Description
[0010] Figure 1 This is a diagram illustrating the overall structure of the model processed by the method of this invention. Figure 2a This is a simulation result of the training loss curve of the deep temporal change detection training model according to an embodiment of the present invention. Figure 2b This is a simulation result diagram of the F1 score of the training set and validation set of the deep temporal change detection training model according to an embodiment of the present invention. Figure 3a This is a time series curve of the NIR band reflectance of a certain pixel in a typical building according to an embodiment of the present invention, and a schematic diagram of the time points of change predicted by the model. Figure 3b This is a time series curve of the SAR VV polarization backscattering coefficient of a typical building in an embodiment of the present invention and a schematic diagram of the time points predicted by the model for the change. Detailed Implementation
[0011] The technical solution provided in this application will be further described below with reference to specific embodiments and accompanying drawings. The advantages and features of this application will become clearer from the following description.
[0012] A method for detecting temporal urban surface changes based on SAR and optical image fusion includes the following steps: Step 1: Using Sentinel-1 SAR and Sentinel-2 optical images as data sources, cloud removal processing of optical images is achieved through SAR-assisted filling and spatiotemporal kriging interpolation. Multi-source image spatiotemporal alignment is achieved by combining DEM terrain correction, orthorectification and dynamic time warping (DTW), and a dataset is constructed based on visual interpretation.
[0013] Step 2: Construct a heterogeneous branch Transformer deep learning model, extract optical and SAR heterogeneous features through multi-scale temporal convolution, achieve adaptive fusion using channel attention mechanism, and output change detection results step by step after capturing long temporal dependencies through the Transformer encoder.
[0014] Step 3: Obtain the final change mask map and change time map through cascaded spatial denoising.
[0015] Step 1 specifically involves: The input data includes three types: Sentinel-1 SAR images, Sentinel-2 optical surface reflectance images, and external DEM elevation data, with the time series covering 60 consecutive monthly observation phases. Among them, the VV (Vertical-Vertical) polarization backscattering coefficients are extracted from the SAR images, the optical images are uniformly resampled to 10m resolution, and the DEM is used for image topography and orthorectification.
[0016] Preprocessing steps: (1) Perform terrain correction and orthorectification on SAR and optical images based on external DEM data; (2) Taking advantage of the fact that SAR images are not obscured by clouds, the SAR-assisted filling method is used to repair the pixels in the optical images that are obscured by clouds. (3) Spatio-Temporal Kriging Interpolation is used to further fill in the residual missing values in the optical image; (4) The Dynamic Time Warping (DTW) algorithm is used to perform spatiotemporal alignment of SAR and optical image sequences.
[0017] (5) The remote sensing index time series obtained by band operation of spectral bands and the texture features extracted by the gray level co-occurrence matrix (GLCM) method of spectral bands are stacked and merged with optical spectral features and SAR backscattering coefficient features, and the output is a multi-source time series image stack.
[0018] Step 2 specifically involves: To address the task of detecting building construction processes in multi-source remote sensing time-series data, this invention proposes a depth-based temporal change detection model based on a heterogeneous branch Transformer. This model takes multi-source temporal features constructed from preprocessed Sentinel-1 SAR and Sentinel-2 optical images as input. It employs a multi-branch heterogeneous network structure to extract the temporal evolution patterns of different feature modes, and uses a channel attention mechanism to achieve adaptive feature fusion. Finally, it utilizes a Transformer encoder to capture long-term temporal dependencies, enabling time-step-by-time discrimination of building change states.
[0019] The overall structure of the depth temporal change detection model based on heterogeneous branch Transformer is as follows: Figure 1 As shown, it includes: heterogeneous feature extraction and multi-scale temporal convolution, channel attention fusion, temporal Transformer encoder, and time-step transformation classification head.
[0020] Heterogeneous feature extraction branch: To fully leverage the ability of different types of features in multi-source remote sensing data to represent surface changes, the model is designed with three heterogeneous feature extraction branches to handle optical spectral and exponential features, texture features, and SAR backscattering features, respectively. Each branch employs a differentiated network structure tailored to the characteristics of the input data to achieve optimal feature extraction results.
[0021] Specifically as follows: (1) Optical Spectrum and Remote Sensing Index Branch. Optical Spectrum Branch Input Remote sensing index branch input Where B is the batch sample size. This represents the total number of time phases in the monthly time series. The total number of input channels is used for each sample in this branch. For optical spectroscopy, red, green, blue, near-infrared, and short-wave infrared channels are selected as input channels; for remote sensing indices, Normalized Difference Vegetation Index (NDVI), Normalized Difference Building Index (NDBI), and Normalized Difference Water Index (NDWI) are selected as input channels. Both the optical spectroscopy and index branches share the same cascaded structure. First, the number of input channels is mapped to the hidden layer dimension through a one-dimensional convolutional layer. Then, after batch normalization and ReLU activation, the data is fed into a multi-scale temporal convolution module to further extract spectral variation features across multiple time scales. This module utilizes dilated convolutions with different dilation rates to capture the spectral evolution patterns across multiple time spans in parallel.
[0022] (2) Texture Branch. The input to this branch is an 8-dimensional GLCM temporal texture tensor. It includes high-dimensional spatial statistics such as mean, contrast, and entropy. Due to the high dimensionality and complex spatial structure of texture features, this branch adopts a hierarchical residual network design. First, the input is mapped to the hidden layer dimension through a one-dimensional convolutional layer. After batch normalization (BN) and ReLU processing, it is fed into the texture feature extraction module. The texture feature extraction module includes a skip connection of a 1×1 convolution and a residual path composed of two layers of dilated convolutions. This enhances the network depth while avoiding the gradient vanishing problem and expands the temporal receptive field using dilated convolutions.
[0023] (3) SAR branch. This branch uses the time-series tensor of the single-channel VV polarization backscattering coefficients of Sentinel-1 SAR. This is the input. Considering that SAR data is severely affected by speckle noise, this branch introduces Dropout regularization after processing by the multi-scale temporal convolution module to suppress the propagation of speckle noise and enhance the robustness of features.
[0024] The multi-scale temporal convolution modules used in the optical spectroscopy and remote sensing index branches and the SAR branch are as follows: The multi-scale temporal convolution module is designed to simultaneously capture short-term abrupt changes and long-term gradual trends in the land surface temporal sequence while preserving the original temporal resolution throughout the entire process. Optical spectroscopy, exponential remote sensing, and SAR branches each output features after processing by the multi-scale temporal convolution module. The optical and exponential branches are output directly, while the SAR branch undergoes Dropout regularization and is finally concatenated with the texture branch features along the channel dimension to form the final fused feature. .
[0025] Specifically, the multi-scale temporal convolution module employs a four-way parallel dilated convolution branch structure, using different dilation rates to capture surface change signals at different time scales. The calculation expression for the output features of each dilated convolution branch is as follows: In the formula, These are intermediate temporal feature tensors for the optical, exponential, and SAR branches, respectively, after being mapped by a first-layer one-dimensional convolution and processed by batch normalization and ReLU activation. All four branches use the same convolution kernel size. The coefficient of thermal expansion is configured as follows: ,Regulation Represents no expansion Convolution. The four parallel convolution branches contain three sets of temporal convolutions with different dilation rates and one set of 1×1 basic convolutions; where the dilation coefficients... The branches sequentially obtain temporal receptive fields at time steps 3, 5, and 9. The 1×1 convolutional branch is used to preserve the original resolution features, enabling simultaneous extraction of multi-scale temporal information and basic channel features. The four branches output features... ( After concatenation along the channel dimension, batch-normalized BN and ReLU activation are introduced. Convolution completes cross-branch feature fusion: in, , The first Convolution kernel size and dilation coefficient, This indicates that the tensor is spliced along the channel dimension, and the output is a preliminary fused feature tensor F.
[0026] Channel attention fusion module: Due to the influence of the observation environment, the effective contribution of different remote sensing modal features to change detection tasks varies significantly: cloud cover greatly reduces the reliability of optical image features, and the inherent speckle noise of SAR images weakens the radar feature discrimination capability. Therefore, this invention introduces a channel attention module to achieve adaptive weighted fusion of multi-source heterogeneous temporal features.
[0027] The channel attention module first performs global average pooling (GAP) and global max pooling (GMP) in parallel on the fused features to aggregate global contextual information along the temporal dimension. Then, it generates the attention weight vectors for each channel through a two-layer shared fully connected network. in, , These are the global average pooling and global max pooling operators, respectively. This is a vector channel concatenation operation; , These are the learnable weight matrices for two fully connected layers; Represents the ReLU nonlinear activation function; The Sigmoid gate function is used to constrain the weights to... Interval. Obtain the channel weight vector. Then, the Hadamard product pair splicing feature was used. Perform adaptive recalibration: The adaptive weighting mechanism can dynamically adjust the contribution ratio of each feature branch according to the data quality of the input image: when the SAR data is affected by strong speckle noise, the radar feature weight is automatically reduced; when the optical image has pixel blur after cloud removal, the feature weights of the texture branch and SAR branch are automatically increased, effectively improving the robustness and generalization performance of the model in complex field observation scenarios.
[0028] Timing Transformer Encoder: Temporal features obtained by channel attention weighted fusion The input temporal Transformer encoder leverages the long-range modeling advantages of self-attention mechanisms to capture the global contextual dependencies of temporal features, overcoming the limitations of traditional convolutional and recurrent networks in modeling long-term temporal correlations and adapting to the learning of the global evolutionary patterns of temporal features. Here, T represents the total number of temporal phases in the long-term observation sequence; D is the channel dimension of the fused features at a single time step.
[0029] Furthermore, to enhance the characteristics of temporal dynamic changes and highlight the temporal abrupt changes brought about by construction activities, a temporal differential enhancement strategy is introduced before the feature input encoder. This strategy amplifies the temporal jump signal by using the feature difference between adjacent time steps. The core enhancement method is as follows: In the formula This is the differential enhancement coefficient, used to adjust the contribution weight of the temporal difference features. This embodiment sets... While amplifying abrupt changes in the construction timeline, it suppresses pseudo-changes caused by noise in remote sensing observations.
[0030] This architecture effectively highlights discontinuous change nodes in the time sequence, guiding the self-attention mechanism to focus on key construction time sequence features and avoiding interference from stationary time sequence information. Based on this, a learnable time sequence location encoding is introduced. This injects prior information about temporal location into the model, compensating for the limitation of the self-attention mechanism in not being able to perceive temporal order, and fuses them to obtain the initial input features of the encoder. : Timing Transformer encoder with For input, a multi-layer stacked standard encoding architecture is adopted. Each encoder layer consists of two core sub-modules: Multi-Head Self-Attention (MHA) and Feedforward Network (FFN). Each sub-module is preceded by a LayerNorm layer, and the overall architecture employs a Pre-LayerNorm normalization architecture. Pre-level normalization optimizes the gradient propagation characteristics of deep networks, improving model training stability and providing architectural support for deep temporal feature mining. The encoder's single-layer forward propagation follows the classic architectural logic of "residual connection + normalization": first, the input features are normalized, and then temporal global dependencies are mined through the MHA module, with features updated via residual connections; then, the output features are normalized a second time, and non-linear feature mapping is completed through the feedforward network; finally, the encoded features of the current layer are output through residual connections. , This represents the number of encoder layers. After layer stacking, the fully connected classification head maps independently to each time step, obtaining the time-by-time phase detection results. ,in For the number of categories, .
[0031] Multi-head self-attention is the core modeling module of the encoder. It achieves temporal dependency modeling based on scaled dot product attention. It captures temporal correlation features of different dimensions and scales in parallel through multiple attention heads, taking into account both global dependencies and local detail feature mining. The basic attention calculation mechanism is as follows: The model generates the query matrix through linear projection. Key matrix AND-value matrix : in This is the learnable projective weight matrix.
[0032] Multiple attention heads independently calculate attention weights and then concatenate and fuse them to achieve integrated output of multi-scale temporal features. The feedforward network is a two-layer fully connected nonlinear mapping architecture, relying on activation functions to achieve feature dimension transformation and nonlinear feature purification. At the same time, a Dropout regularization structure is introduced after each sub-module to suppress model overfitting and improve the model's generalization ability from the architectural level.
[0033] Time-step changes in classification header: For the task of temporal change detection, a lightweight temporal classification network is built on the backend of the Transformer encoder to achieve binary classification at each time step, distinguishing between changed and unchanged states in temporal features. The classification head adopts a deep classification structure of "normalization-nonlinear activation-double-layer projection" to perform dimension mapping and category prediction on the single-time-step features output by the encoder. The overall classification inference architecture is simple and adaptable to the characteristics of temporal sequence features. The posterior probability of the predicted category of surface building change output at the t-th time step (temporal phase) is: In the formula For the encoder's final output exist Feature vectors at each time step , The classification network can learn projection parameters. The model ultimately outputs a global temporal probability feature matrix, and outputs the predicted posterior probabilities of the two classes at each time step. After taking the class index corresponding to the largest probability, a discrete change time series is obtained, which is used for joint loss optimization in the training phase, accuracy verification in the evaluation phase, and cascaded post-processing and final detection map generation in the inference phase, forming a complete change detection inference chain.
[0034] Optimization goal: Considering the class imbalance problem (sparse changed samples and a very high proportion of unchanged samples) in construction time sequence change detection tasks, this invention abandons a single loss function and designs a joint loss optimization architecture that combines cross-entropy loss and Dice loss, balancing classification stability and sparse target recognition accuracy. The joint loss function is defined as follows: Among them, cross-entropy loss As a fundamental classification constraint, the deviation between the predicted distribution and the true label at each time step is supervised, providing a stable and uniform gradient signal for model training and ensuring basic classification capability. In the formula, Let be the true label of the t-th time phase, where This indicates that changes in surface structures occurred during this time period. This indicates that no changes in surface structures occurred during this time phase; T represents the total number of time phases in the long-term observation sequence, and t represents the t-th time step in the time series. The model's posterior probability of predicting the type of surface building change at time phase t. The model outputs the predicted posterior probability of the category of no surface building changes at time phase t.
[0035] Dice loss By optimizing the overlap between the predicted and ground truth regions, without relying on sample size balance, this approach effectively focuses on sparse, changing samples, compensates for the shortcomings of cross-entropy loss in class imbalance scenarios, and improves the model's ability to identify key changing regions. In the formula This is a general numerical smoothing term, used solely to avoid computational anomalies and ensure training stability. The two loss functions provide complementary constraints, jointly optimizing model performance from two dimensions: global classification accuracy and matching accuracy in locally varying regions.
[0036] The model employs the AdamW optimizer for parameter updates, improving regularization and generalization capabilities by decoupling weight decay and adaptive learning rate, making it more suitable for training deep temporal networks. Simultaneously, the model incorporates dynamic learning rate scheduling, early stopping mechanisms, and temporal data augmentation strategies to effectively suppress overfitting and enhance its robustness and adaptability in complex temporal scenarios.
[0037] To suppress occasional false detections and edge spikes in pixel-by-pixel temporal classification results, this invention introduces a cascaded spatiotemporal denoising strategy to generate the final detection result. First, temporal consistency is assessed on the temporal-phase classification sequence output by the model, initially filtering out unstable abrupt transitions at the temporal level. Second, area filtering is applied to the binary transformation mask obtained in the previous step to remove spatially isolated small patches. Finally, morphological opening operations are performed using 3×3 rectangular structuring elements to smooth the boundaries of transformation regions and eliminate edge spikes and fine protrusions.
[0038] After the above cascaded denoising, the final output results are: a change area map and a change time series map, which respectively reflect the range and time of occurrence of surface changes.
[0039] Experimental verification This invention selects a city's urban area as the research area. This area has a low and flat terrain, a dense river network, and diverse surface morphology, including plains, waterfront areas, and suburban-rural transition zones. In recent years, the city has experienced rapid urban expansion and urban renewal, leading to complex changes in surface structure and morphology, which places higher demands on the accuracy and timeliness of surface monitoring.
[0040] The experiment used multi-source remote sensing data, including: (1) Sentinel-1 SAR imagery (Level-1C GRDH product) for extracting VV polarization backscattering intensity; (2) Sentinel-2 optical imagery (Level-2A surface reflectance product) after geometric and atmospheric correction; and (3) external DEM data for topographic and orthorectification correction. The time span covered from January 2019 to December 2023, a total of 60 months, with a time resolution of months.
[0041] The experimental dataset was configured with a training set:validation set ratio of 8:2, with approximately 10,000 pixel samples in the training set. The model parameters were batch_size=64, with a maximum training epoch of 100, and training was performed using an NVIDIA GPU within the PyTorch framework. The following metrics were used to evaluate model performance: (1) Spatial F1 score, evaluating the accuracy of spatial detection of changing regions; (2) Temporal F1 score, evaluating the accuracy of temporal detection of changing regions. The training results are as follows: Figure 2a and Figure 2b Experimental results on the validation set show that the model achieves good performance on both spatial and temporal F1 scores. The training and validation losses show a steady decreasing trend, the F1 score gradually increases and tends to converge, and no obvious overfitting is observed.
[0042] Table 1 shows the performance comparison of the proposed model, TempCNN model, and LTAE model on multi-source data containing SAR. Experimental results show that the proposed method achieves the best performance on all validation set evaluation metrics, with TemporalF1 reaching 0.8314 on the validation set, verifying the effectiveness of the method.
[0043] Table 1. Performance comparison of different methods on the validation set. Meanwhile, cascaded spatial denoising processing can remove isolated salt-and-pepper false detection pixels at the edges of buildings while maintaining the integrity of the real change area, significantly reducing the false alarm rate at the edges.
[0044] To visually verify the model's temporal detection accuracy, a typical building change pixel within the study area was selected for visualization analysis. In the building change detection task, the NIR band is highly sensitive to changes in land cover, effectively responding to site clearing in the early stages of construction and subsequent soil restoration. The SAR VV polarization backscattering coefficient significantly responds to the dihedral scattering effect caused by the vertical structure of buildings, capturing the appearance and demolition of the main building structure. Both bands complement each other in characterizing the building change process from spectral and structural dimensions, respectively; therefore, they were selected as typical features for temporal analysis. Construction of this pixel began in March 2020; prior to that, the surface was a mixture of bare soil and low vegetation.
[0045] Change time series diagram as follows Figure 3a and Figure 3b As shown, where, Figure 3a This is a 60-month time-series curve of NIR band reflectance (blue solid line). The red solid line marks the time points predicted by the model of this invention, the blue solid line marks the actual time points, the green solid line marks the time points predicted by the TempCNN model, and the purple solid line marks the time points predicted by the LTAE model. Before construction, NIR reflectance fluctuated with the seasons, with reflectance peaking in summer and decreasing in winter. After construction, NIR reflectance showed a long-term decreasing trend. Figure 3b This is the 60-month time series curve of the SAR VV backscattering coefficient (orange dashed line). The red solid line marks the predicted change time points, the blue solid line marks the actual change time points, the green solid line marks the change time points predicted by the TempCNN model, and the purple solid line marks the change time points predicted by the LTAE model. Before construction, the scattering coefficient was generally low, remaining in the negative range for many years with gentle fluctuations. During construction, the curve showed a sharp jump, exhibiting a very strong positive jump. Subsequently, for a period of time, the scattering coefficient after building formation was generally higher than before construction, but fluctuated to some extent due to vegetation, observation geometry, and speckle noise. Combined with... Figure 3a and Figure 3bThe results shown in the figure, along with the predictions from the proposed model, TempCNN model, and LTAE model, and the actual change time points, reveal that the red and blue solid lines in the optical near-infrared time series and SAR backscattering coefficient time series both coincide in May 2022. The change time points predicted by the method proposed in this invention perfectly match the actual construction start time, making the prediction results more accurate. This result demonstrates that the proposed method can accurately pinpoint the actual abrupt change time points in building construction and possesses reliable time series detection and location capabilities for changes in the temporal sequence of surface buildings.
[0046] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them. After reading this specification, those skilled in the art can make various modifications or variations to the technical solutions without departing from the concept of the present invention, and all such modifications or variations fall within the protection scope of the present invention.
Claims
1. A method for detecting temporal urban surface changes based on SAR and optical image fusion, characterized in that, Includes the following steps: Step 1: Using Sentinel-1 SAR and Sentinel-2 optical images as data sources, cloud removal processing of optical images is achieved through SAR-assisted filling and spatiotemporal kriging interpolation. Multi-source image spatiotemporal alignment is achieved by combining DEM terrain correction, orthorectification and dynamic time warping, and a dataset is constructed based on visual interpretation. Step 2: Construct a heterogeneous branch Transformer deep learning model, extract optical and SAR heterogeneous features through multi-scale temporal convolution, achieve adaptive fusion by using channel attention mechanism, and output change detection results step by step after capturing long temporal dependencies by the Transformer encoder. Step 3: Obtain the final change mask map and change time map through cascaded spatial denoising.
2. The method according to claim 1, characterized in that, Step 1 specifically involves: The input data includes three types: Sentinel-1 SAR imagery, Sentinel-2 optical surface reflectance imagery, and external DEM elevation data. Data preprocessing is as follows: (1) Perform terrain correction and orthorectification on SAR and optical images based on external DEM data; (2) Use SAR-assisted filling method to repair pixels in optical images that are obscured by clouds; (3) Spatiotemporal kriging interpolation is used to fill in the residual missing values in optical images; (4) The dynamic time warping algorithm is used to perform spatiotemporal alignment of SAR and optical image sequences; (5) The remote sensing index time series obtained by band operation of spectral bands and the texture features extracted by gray co-occurrence matrix method of spectral bands are stacked and merged with optical spectral features and SAR backscattering coefficient features, and the output is a multi-source time series image stack.
3. The method according to claim 1, characterized in that, Step 2 specifically involves: Construct a deep temporal change detection model based on heterogeneous branch Transformer, namely a heterogeneous branch Transformer deep learning model; The deep temporal change detection model based on heterogeneous branch Transformer includes: heterogeneous feature extraction and multi-scale temporal convolution, channel attention fusion, temporal Transformer encoder and time-step change classification head; The heterogeneous branch Transformer-based depth temporal change detection model takes multi-source temporal features constructed from preprocessed Sentinel-1 SAR and Sentinel-2 optical images as input. It uses a multi-branch heterogeneous network structure to extract the temporal evolution patterns of different feature modes, and achieves adaptive feature fusion through a channel attention mechanism. Finally, it uses a Transformer encoder to capture long-term temporal dependencies and realizes the discrimination of building change states step by step.
4. The method according to claim 3, characterized in that, The depth temporal change detection model based on heterogeneous branch Transformer includes three heterogeneous feature extraction branches, which respectively process optical spectral and remote sensing index features, texture features, and SAR backscattering features; each branch adopts a differentiated network structure according to the characteristics of the input data. Specifically as follows: (1) Optical Spectrum and Remote Sensing Exponential Branch: Optical spectrum and exponential branch share the same cascade structure; firstly, the number of input channels is mapped to the hidden layer dimension through a one-dimensional convolutional layer, and then after batch normalization and ReLU activation function processing, it is sent to the multi-scale temporal convolution module to extract the spectral change features of multiple time scales, and the dilated convolution with different dilation rates is used to capture the spectral evolution law of multiple time spans in parallel. (2) Texture branch; The input is an 8-dimensional GLCM temporal texture tensor, including high-dimensional spatial statistics such as mean, contrast, and entropy; A hierarchical residual network design is adopted; First, the input is mapped to the hidden layer dimension through a one-dimensional convolutional layer, and after batch normalization (BN) and ReLU processing, it is sent to the texture feature extraction module. The texture feature extraction module includes a skip connection of a 1×1 convolution and a residual path consisting of two layers of dilated convolutions; (3) SAR branch; The temporal tensor of the single-channel VV polarization backscattering coefficients of Sentinel-1 SAR is used as input. After processing by a multi-scale temporal convolution module, Dropout regularization is introduced to suppress the propagation of speckle noise and enhance the robustness of the features.
5. The method according to claim 4, characterized in that, The multi-scale temporal convolution modules used in the optical spectroscopy and remote sensing index branches and the SAR branch are as follows: The multi-scale temporal convolution module employs a 4-way parallel dilated convolution branch structure, using different dilation rates to capture surface change signals at different time scales; the calculation expressions for the output features of each branch are as follows: In the formula, These are intermediate temporal feature tensors from the optical, exponential, and SAR branches, respectively, after being mapped by the first-layer one-dimensional convolutional mapping and processed by batch normalization and ReLU activation; the four dilated convolutional branches uniformly adopt the same kernel size. ; The expansion coefficient is configured as follows ,Regulation Represents no expansion Convolution; four parallel convolution branches contain three sets of temporal convolutions with different dilation rates and one set of 1×1 basic convolutions; where the dilation coefficients... The branches obtain temporal receptive fields of 3, 5, and 9 time steps, respectively. The 1×1 convolutional branch is used to preserve the original resolution features; Output characteristics of the four branches After concatenation along the channel dimension, batch-normalized BN and ReLU activation are introduced. Convolution completes cross-branch feature fusion: in, , The first Convolution kernel size and dilation coefficient, This indicates that the tensor is spliced along the channel dimension, and the output is a preliminary fused feature tensor F.
6. The method according to claim 3, characterized in that, The channel attention fusion module is used to achieve adaptive weighted fusion of multi-source heterogeneous temporal features; The channel attention module first performs global average pooling (GAP) and global max pooling (GMP) in parallel on the fused features to aggregate global contextual information in the temporal dimension. Then, it generates the attention weight vector for each channel through a two-layer shared fully connected network. in, , These are the global average pooling and global max pooling operators, respectively. This is a vector channel concatenation operation; , These are the learnable weight matrices for two fully connected layers; Represents the ReLU nonlinear activation function; The Sigmoid gate function is used to constrain the weights to... interval; Obtain the channel weight vector Then, the Hadamard product pair splicing feature was used. Perform adaptive recalibration: 。 7. The method according to claim 3, characterized in that, The timing Transformer encoder is specifically: Temporal features obtained by channel attention weighted fusion Input a temporal Transformer encoder to capture global contextual dependencies of temporal features; A time-series differential enhancement strategy is introduced to amplify the time-series jump signal by using the feature difference between adjacent time steps. The enhancement method is as follows: In the formula This is the difference enhancement coefficient, used to adjust the contribution weight of the temporal difference feature; Introducing learnable temporal position coding This injects prior information about temporal location into the model, compensating for the limitation of the self-attention mechanism in not being able to perceive temporal order, and fuses them to obtain the initial input features of the encoder. : The timing Transformer encoder is Layer-stacked coding architecture, the input of the first layer is For input.
8. The method according to claim 3, characterized in that, The time-step change classification head is used for time-series change detection tasks. A lightweight time-series classification network is built after the Transformer encoder to achieve time-step binary classification and distinguish between changed and unchanged states in time-series features. The posterior probability of the predicted surface building change category output at time step t is: In the formula For the final output of the Transformer encoder exist Feature vectors at each time step , For classification networks, projection parameters can be learned; The model ultimately outputs a global temporal probability feature matrix, and outputs the predicted posterior probabilities of the two classes at each time step. After taking the class index corresponding to the one with the highest probability, a discrete change time series is obtained. These are used for joint loss optimization in the training phase, accuracy verification in the evaluation phase, and cascaded post-processing and final detection map generation in the inference phase, forming a complete change detection inference chain.
9. The method according to claim 3, characterized in that, The joint loss function of the heterogeneous branch Transformer deep learning model is defined as follows: Among them, cross-entropy loss As a basic classification constraint: In the formula, Let be the true label of the t-th time phase, where This indicates that changes in surface structures occurred during this time period. This indicates that no changes in surface structures occurred during this time phase; T represents the total number of time phases in the long-term observation sequence, and t represents the t-th time step in the time series. The model's posterior probability of predicting the type of surface building change at time phase t. The model's predicted posterior probability for the category of no surface building changes at time phase t; Dice loss The optimization objective is the overlap between the predicted region and the ground truth region, without relying on sample size balance. in This is a general numerical smoothing term.
10. The method according to claim 1, characterized in that, In step 3, A cascaded spatiotemporal denoising strategy is introduced to generate the final detection results: First, temporal consistency is judged on the phase-by-phase classification sequence output by the model to filter out unstable abnormal jumps at the temporal level; second, area filtering is performed on the binary change mask obtained in the previous step to remove spatially isolated small patches; finally, morphological opening operation is performed using 3×3 rectangular structuring elements to smooth the boundaries of the change region and eliminate edge spikes and fine protrusions. After the above cascaded denoising, the final output results are: a change area map and a change time series map, which respectively reflect the range and time of occurrence of surface changes.