Remote sensing image semantic change detection method and device
By introducing a feature extraction module, a semantic association module, and a multi-view change feature fusion module into the semantic change detection of remote sensing images, the problem of ignoring the correlation of semantic features between two time phases in the existing technology is solved, and higher accuracy and more complete change detection are achieved.
Patent Information
- Application Number
- CN202511481569.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-16
- Publication Date
- 2026-01-06
AI Technical Summary
Existing multi-task SCD methods for detecting semantic changes in remote sensing images suffer from problems such as ignoring the correlation of semantic features between two phases, insufficient spectral invariance, and insufficient temporal consistency. This leads to false changes and change noise in the detection results, making it particularly difficult to accurately identify changes in small-sized and complex-shaped objects in complex scenes.
A novel semantic change detection model is adopted, including a feature extraction module, a semantic association module (SCM), and a multi-view change feature fusion module (MCFF). Fine-grained and coarse-grained features are extracted through a weight-shared twin feature extractor and an independent feature extractor. The semantic change features are enhanced by a Bayesian theory fusion pattern, and a change prior information is provided by an auxiliary change detection branch.
It improves the accuracy and boundary consistency of semantic change detection in remote sensing images, reduces false changes, and enhances the ability to identify changed areas, especially improving detection accuracy and completeness in complex scenes.
Smart Images

Figure CN121280903A_ABST
Abstract
Description
Technical Field This disclosure belongs to the field of remote sensing image processing technology, and in particular relates to a method and apparatus for detecting semantic changes in remote sensing images. Background Technology
[0001] Semantic change detection in remote sensing imagery refers to the technique of simultaneously identifying changed areas and change categories based on remote sensing images of the same location at different times, through image processing and mathematical models. This technology has been widely applied in surveying and mapping, urban planning, resource and environmental monitoring, and various dynamic monitoring fields, and can be used to obtain information on land use / land cover type changes. Change detection is generally divided into BCD (Binary Change Detection) and SCD (Semantic Change Detection).
[0002] In recent years, SCD methods based on multi-task learning architectures have generally outperformed other methods and have been widely applied in various change detection scenarios. These methods treat the SCD task as two bi-temporal semantic segmentation tasks and a binary change detection task to distinguish different categories of changed regions while locating the changed regions. By sharing parameters and features to jointly learn multiple related semantic representations, this approach can significantly improve detection results. However, some challenges still exist in this field. Standard three-branch SCD models based on multi-task architectures, such as HRSCD-3 and HRSCD-4, extract bi-temporal semantic information and change information separately, which suffers from the problem of ignoring the correlation between bi-temporal semantic features.
[0003] To address this issue, researchers have utilized spatial cross-attention mechanisms such as Bi-SRNet and HGINet to model semantic association information (SCs). However, this approach lacks explicit modeling of spectral invariance and temporal consistency, which may prevent it from effectively capturing complex spectral-temporal correlations in real-world scenes with varying illumination and seasonal variations. Secondly, popular multi-task SCD methods, such as SCanNet, ChangeMask, and MTSCDNet, generate change features by distinguishing bi-temporal semantic features generated by a shared-parameter Siamese network extractor. However, the bi-temporal semantic information may be over-extracted, making it difficult to capture complex change information and introducing considerable change noise. Furthermore, due to the lack of effective prior information constraints, change information cannot be effectively enhanced, leading to missed detections and incompleteness, especially for objects with complex shapes and small sizes. Summary of the Invention
[0004] To address the aforementioned issues, this disclosure provides a method and apparatus for detecting semantic changes in remote sensing images, and a novel semantic change detection model that can accurately identify potential semantic changes in VHR (Very High Resolution) remote sensing images under different scenarios.
[0005] This invention provides a method for detecting semantic changes in remote sensing images, comprising: Input the dual-temporal images to be detected into the trained semantic change detection model; The trained semantic change detection model is used to detect remote sensing semantic changes, resulting in pre-temporal semantic segmentation results. ST 1. Post-temporal semantic segmentation results ST 2 and fine semantic variation features ; Fine semantic variation features Binarization is performed to obtain a binary change detection map; the semantic segmentation results of the previous time phase are then processed. ST 1 and post-temporal semantic segmentation results ST 2. Masking operations are performed using the binary change detection map to obtain the front-phase semantic detection map and the back-phase semantic detection map; The trained semantic change detection model includes: a feature extraction module, a semantic association module (SCM), and a multi-view change feature fusion module (MCFF), wherein: The feature extraction module is used to extract fine-grained semantic features from bi-temporal images using a weight-shared twin feature extractor. and post-temporal fine-grained semantic features ; and a separate feature extractor was used to extract coarse-grained semantic variation features from the dual-temporal images. ; The Semantic Association Module (SCM) is used for fine-grained semantic features from previous time phases. and post-temporal fine-grained semantic features By sequentially performing similarity semantic information enhancement operations and channel attention operations, enhanced pre-temporal fine-grained semantic features are obtained. and enhanced post-temporal fine-grained semantic features ; The Multi-View Change Feature Fusion (MCFF) module is used to fuse coarse-grained semantic change features. As prior information, based on Bayesian theory fusion patterns and enhanced pre-temporal fine-grained semantic features and enhanced post-temporal fine-grained semantic features Generate refined semantic variation features ; and enhanced pre-temporal fine-grained semantic features and enhanced post-temporal fine-grained semantic features Perform convolution operations separately to obtain the semantic segmentation results of the previous time phase. ST 1 and post-temporal semantic segmentation results ST 2.
[0006] Furthermore, the Semantic Association Module (SCM) is specifically used to refine the semantic features of previous time phases. and post-temporal fine-grained semantic features Perform element-wise multiplication to enhance similar semantic information.
[0007] Furthermore, the Multi-View Change Feature Fusion (MCFF) module is used to fuse coarse-grained semantic change features. As prior information, Bayesian theory is applied to the fusion model to enhance the pre-temporal fine-grained semantic features. and enhanced post-temporal fine-grained semantic features Fine-grained semantic variation features between Enhancement is performed based on coarse-grained semantic change features. and enhanced fine-grained semantic variation features This generates sophisticated semantic variation features.
[0008] Furthermore, the Multi-View Change Feature Fusion (MCFF) module is specifically used to compute enhanced pre-temporal fine-grained semantic features. and enhanced post-temporal fine-grained semantic features Fine-grained semantic variation features between ; to incorporate coarse-grained semantic variation features As prior information, Bayesian theory is applied to fuse patterns to calculate fine-grained semantic change features. posterior probability Enhance fine-grained semantic change features using posterior probability. This results in enhanced fine-grained semantic change features. .
[0009] Furthermore, the Multi-View Change Feature Fusion (MCFF) module is specifically used to fuse coarse-grained semantic change features. Enhanced fine-grained semantic variation features The features are concatenated and convolutional blocks are used to reduce feature dimensionality, generating refined semantic variation features. .
[0010] Furthermore, the semantic change detection model integrates two semantic segmentation tasks for images from different time phases and one change detection task. The semantic change detection model is trained using the following multi-task loss function:
[0011] in, Indicates the total mission loss. This represents the semantic segmentation task loss of previous temporal images. This represents the semantic segmentation task loss for later-phase images. This indicates the loss of the change detection task. This indicates the loss due to semantic change.
[0012] Furthermore, The feature vectors of semantic segmentation results from the previous phase image and the subsequent phase image are calculated using cosine loss.
[0013] The present invention also provides a remote sensing image semantic change detection device for implementing the above method, comprising: an input module, a trained semantic change detection model, and an output module, wherein: The input module is used to input the dual-temporal images to be detected into the trained semantic change detection model; The trained semantic change detection model is used for remote sensing semantic change detection to obtain pre-temporal semantic segmentation results. ST 1. Post-temporal semantic segmentation results ST 2 and fine semantic variation features ; The output module is used for fine-grained semantic variation features. Binarization is performed to obtain a binary change detection map; the semantic segmentation results of the previous time phase are then processed. ST 1 and post-temporal semantic segmentation results ST 2. Masking operations are performed using the binary change detection map to obtain the front-phase semantic detection map and the back-phase semantic detection map.
[0014] The present invention also provides an electronic device, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; Memory, used to store computer programs; The processor implements the above method when executing programs stored in memory.
[0015] The present invention also provides a computer storage medium storing a computer program, which, when executed by a processor, implements the above-described method.
[0016] Compared with the prior art, this disclosure has the following advantages: (1) A robust multi-task network is proposed, which integrates change priors and bitemporal SCs (Semantic Correlations) to identify different semantic changes in VHR remote sensing images. Unlike existing multi-task SCD methods, it incorporates an additional auxiliary CD (Change Detection) branch to provide prior information on changes, thereby enhancing the identification of changed regions and their categories; (2) The SCM module effectively simulates SCs between two time features, reduces spurious changes caused by spectral, illumination and time effects, and is superior to existing cross-time methods (i.e. Cot-SR and TC).
[0017] (3) The MCFF module integrates multi-view change information from different feature extraction branches into a single Bayesian network, enhancing the perception of changed regions and improving regional integrity. Compared with global attention or concatenation operations, MCFF optimizes the semantic features of change information and improves the detection accuracy of different landscapes.
[0018] (4) The above method was applied to the Xiamen (XM) and Jilin No. 1 (JL-1) SCD datasets and compared with 10 typical SCD methods. The SeK values of the classification consistency and accuracy used to evaluate semantic change detection on the XM and JL-1 datasets were 9.28% ~ 65.98% and 5.97% ~ 40.54% higher than the existing methods, respectively. In addition, the detection results showed good boundary consistency and internal integrity.
[0019] Other features and advantages of this disclosure will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the disclosure. The objects and other advantages of this disclosure may be realized and obtained by means of the structures pointed out in the description, claims and drawings. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 A flowchart illustrating a method for detecting semantic changes in remote sensing images according to an embodiment of this disclosure is shown. Figure 2 An architecture diagram of an SCM according to an embodiment of the present disclosure is shown; Figure 3An architecture diagram of the MCFF according to an embodiment of the present disclosure is shown; Figure 4 Visualization of SCD results from JL-1 dataset ablation experiments according to embodiments of this disclosure; Figure 5 An example of farmland change area detected using different time-crossing modules on the JL-1 farmland SCD dataset according to an embodiment of the present disclosure is shown; Figure 6 Examples of farmland SCD results using different methods on the JL-1 dataset according to embodiments of this disclosure are shown; Figure 7 A visualization of farmland SCD results across different patterns on an XM dataset according to an embodiment of the present disclosure is shown; Figure 8 The results of farmland de-agriculturalization on Google Imagery in Xiamen using CPGNet according to embodiments of this disclosure are shown. Figure 9 A visualization of feature maps generated on the JL-1 dataset according to embodiments of the present disclosure is shown. Detailed Implementation
[0022] This invention proposes a novel semantic change detection model, CPGNet, based on a multi-task architecture. It first establishes semantic association information between fine-grained features in both temporal phases to enhance spectral invariance and temporal consistency features, such as invariant information, thereby mitigating the problem of spurious changes. Subsequently, it focuses on capturing multi-view change information generated by different feature extraction branches. Finally, it utilizes effective change information enhancement patterns to improve the detection of inadequate change regions.
[0023] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.
[0024] Figure 1 The diagram illustrates a flowchart of a method for detecting semantic changes in remote sensing images according to an embodiment of this disclosure, including the following steps: Step 1: Input the dual-temporal images to be detected into the trained semantic change detection model; Step 2: The trained semantic change detection model is used for remote sensing semantic change detection to obtain the previous temporal semantic segmentation results. ST 1. Post-temporal semantic segmentation results ST 2 and fine semantic variation features ; Step 3: Fine-grained semantic variation features Binarization is performed to obtain a binary change detection map; the semantic segmentation results of the previous time phase are then processed. ST The semantic segmentation results of 1 and the post-temporal semantic segmentation are respectively ST 2. A masking operation is performed using the binary change detection map to obtain the previous phase semantic detection map T1 and the next phase semantic detection map T2; Among them, semantic change detection models such as Figure 1 As shown, it includes: a feature extraction module, a semantic association module (SCM), and a multi-view change feature fusion module (MCFF), wherein: The feature extraction module is used to extract fine-grained semantic features from bi-temporal images using a weight-shared twin feature extractor. and post-temporal fine-grained semantic features ; and a separate feature extractor was used to extract coarse-grained semantic variation features from the dual-temporal images. Here, the independent feature extractor and the twin feature extractor do not share weights.
[0025] The Semantic Association Module (SCM) is used for fine-grained semantic features from previous time phases. and post-temporal fine-grained semantic features By sequentially performing similarity semantic information enhancement operations and channel attention operations, enhanced pre-temporal fine-grained semantic features are obtained. and enhanced post-temporal fine-grained semantic features ; The Multi-View Change Feature Fusion (MCFF) module is used to fuse coarse-grained semantic change features. As prior information, based on Bayesian theory fusion patterns and enhanced pre-temporal fine-grained semantic features and enhanced post-temporal fine-grained semantic features Generate refined semantic variation features ; and enhanced pre-temporal fine-grained semantic features and enhanced post-temporal fine-grained semantic features Perform convolution operations separately to obtain the semantic segmentation results of the previous time phase. ST 1 and post-temporal semantic segmentation results ST 2.
[0026] The following is a detailed explanation of each of the above modules: 1) The feature extraction module comprises three feature extractors: a weight-shared Siamese feature extractor and an independent feature extractor. All three feature extractors are based on a Multilevel Perceptual Aggregation Network (MPANet). The weight-shared Siamese feature extractor extracts features from the dual-temporal images to capture temporal differences (Long, J., Li, M., Wang, X., & Stein, A. (2024). Semantic change detection using a hierarchical semantic graph interaction network from high-resolution remote sensing images. ISPRS Journal of photogrammetry and remote sensing, 211, 318-335). To avoid feature misleading, an independent MPANet is introduced as a feature extractor to directly capture dynamic changes from the dual-temporal images. The MPANet used extracts and fuses multi-level semantic features through a Transformer network, enhancing the feature representation of different ground features, especially for features with complex shapes and large sizes.
[0027] 2) Semantic Association Module (SCM), its architecture diagram is as follows: Figure 2 As shown.
[0028] Existing methods focus on leveraging spatial attention mechanisms to enhance semantic consistency (SCs) between high-level bi-temporal fine-grained features, but lack explicit modeling of spectral invariance and temporal consistency. Furthermore, high-level temporal features typically exhibit low spatial resolution and high dimensionality, making channel-level semantic consistency modeling crucial. The correlation between bi-temporal semantic features is also significant for improving the identification of invariant components in images from different time phases, thereby reducing spurious changes caused by temporal and seasonal variations.
[0029] To effectively model the semantic consistency between bi-temporal fine-grained features, embodiments of this disclosure employ a Semantic Consensus (SCM) module to learn this consistency. This SCM module models the correlation between bi-temporal fine-grained semantic features to mitigate spurious changes caused by phenological, temporal, and spectral information. Initial change features are then obtained by calculating the differences between the bi-temporal fine-grained semantic features from the Siamese feature extractor. This SCM module enhances the extraction capability for invariant objects by fully integrating similar semantic cues.
[0030] Specifically, assuming pre-temporal fine-grained semantic features and post-temporal fine-grained semantic features These are the pre- and post-temporal fine-grained semantic features obtained from the Siamese Feature Extractor (MPANet). Similarity semantic information is enhanced through the following element-wise multiplication operation: (1) Next, a channel attention operation that combines average pooling (AP) and max pooling (MP) is designed to highlight global and local statistical features for information that remains unchanged in the target object, thereby obtaining an attention weight. : (2) in, , and These are the similarity matrix, the sigmoid function, and the 1×1 convolutional block, respectively.
[0031] Then, the enhanced bi-temporal fine-grained semantic features, which take into account the consistency of bi-temporal fine-grained semantics, are obtained, namely the enhanced pre-temporal fine-grained semantic features. and enhanced pre-temporal fine-grained semantic features : (3) (4) in, This represents a 3×3 convolutional block. This approach further enhances the ability to perceive invariant information, helping to identify changes relevant to the object of interest.
[0032] 3) Multi-view change feature fusion module MCFF, its partial architecture diagram is as follows: Figure 3 As shown.
[0033] Multi-view information of the target of interest helps improve the completeness of detection and accurately distinguish different targets. However, existing semantic change detection research has failed to fully consider change information from different viewpoints. These methods typically extract change information only by calculating the differences between bi-temporal fine-grained features from a Siamese feature extractor. Since bi-temporal fine-grained features may be inaccurate, this approach easily introduces a large amount of noise, leading to a decrease in detection accuracy. Furthermore, this limitation may result in incomplete or missed detection of change regions, especially for small-sized and complex-shaped targets.
[0034] To address these shortcomings, this disclosure presents a Multi-View Variation Feature Fusion (MCFF) module, which fully integrates variation features from different extraction perspectives (multiple views generated by different feature extraction branches), i.e., enhanced pre-temporal fine-grained semantic features. and enhanced post-temporal fine-grained semantic features And coarse-grained semantic change features generated from the auxiliary change detection (ChangeDetection, CD) branch. This enhances the ability to perceive areas of change. Inspired by Bayesian theory, it incorporates coarse-grained semantic change features. As prior information, it further enhances the information on fine-grained variation characteristics.
[0035] Specifically, the enhanced pre-temporal fine-grained semantic features generated by SCM are first calculated. and enhanced pre-temporal fine-grained semantic features From the difference information, fine-grained semantic change features are obtained. Coarse-grained semantic variation features As prior information, the posterior probability of the changed pixel is calculated, thereby determining the initial change position and constraining the change range.
[0036] First, the Sigmoid function is used. To obtain features about coarse-grained semantic changes and fine-grained semantic variation features The probability of change is then used to derive the posterior probability using Bayesian fusion patterns. : (5) Next, use the posterior probability Optimize fine-grained semantic variation features This results in enhanced fine-grained semantic change features. : (6) Finally, the coarse-grained semantic change features Enhanced fine-grained semantic variation features Connect the features and use a 1×1 convolutional block to reduce the feature dimensionality, thereby obtaining fine-grained semantic variation features. : (7) Here, "Cat" represents a splicing operation. This method further enhances the ability to perceive changes in information about objects of interest.
[0037] The semantic change detection model described above integrates three key tasks: two semantic segmentation tasks for images from different time phases, and one change detection task. Figure 1 In this context, the two semantic segmentation tasks are reflected in the transition from the first branch of feature extraction to... ST 1 and from the second branch of feature extraction to ST 2; The change detection task is reflected in the third branch of feature extraction (i.e., the auxiliary CD branch) to... This architecture enables collaborative learning of multiple related tasks, effectively guiding the acquisition of categorical features by identifying changing regions.
[0038] To train the semantic change detection model and obtain a trained model, this disclosure presents a multi-task loss function, which is explained below: In bitemporal semantic segmentation tasks, cross-entropy loss is used as its loss function.
[0039] make and They represent the first The class's label value and predicted probability, where (Class 0 represents the unchanged class). Semantic segmentation loss for dual-temporal images. ( The definition is as follows: (8) in, and These represent the number of pixels in the image and the number of recognized categories, respectively.
[0040] Furthermore, for change detection tasks, a weighted cross-entropy loss is used to alleviate the imbalance between the number of changed pixels and the number of unchanged pixels. Let... This represents the label value of the changed region obtained from the actual label. This represents the predicted probability of the changed region. (Loss of the change detection task) The definition is as follows: (9) in, This indicates the ratio of negative samples to the total sample.
[0041] Furthermore, to improve the consistency between segmentation results and change detection results, this disclosure introduces a semantic change loss. Specifically, this loss is calculated based on the segmentation results of the two-temporal images using a cosine loss formula, as follows: (10) in, and These represent the feature vectors of the semantic segmentation results of the preceding and following time-phase images, respectively.
[0042] Finally, by weighting the losses of all tasks, the total task loss of the multi-task loss function is obtained. The calculation formula is as follows: (11) in, and denoted as the semantic segmentation task loss for the previous and subsequent time phase images, respectively.
[0043] Based on the above method, this disclosure also provides a remote sensing image semantic change detection device corresponding to the above method, for implementing the above method, including: an input module, a trained semantic change detection model, and an output module, wherein: The input module is used to input the dual-temporal images to be detected into the trained semantic change detection model; The trained semantic change detection model is used for remote sensing semantic change detection, yielding pre-temporal semantic segmentation result ST1, post-temporal semantic segmentation result ST2, and refined semantic change features. ; The output module is used for fine-grained semantic variation features. Binarization is performed to obtain a binary change detection map. The binary change detection map is then used to perform masking operations on the previous phase semantic segmentation result ST1 and the subsequent phase semantic segmentation result ST2 to obtain the previous phase semantic detection map and the subsequent phase semantic detection map, respectively.
[0044] The following describes the experimental results for performance evaluation of the disclosed solution: 1. Models included in the comparison: HRSCD-3, HRSCD-4, Bi-SRNet, ChangeMask, SCanNet, MTSCD-Net, CdSC, HGINet, SAM-CD, SCD-SAM 2. Evaluation Metrics: Four pixel-based metrics were used: average crosslinking (mIoU), overall accuracy (OA), separated kappa (SeK), and F1-score. The accuracy of the SCD test results is evaluated using [the following method]. Let TN, TP, FN, and FP represent true negative, true positive, false negative, and false positive, respectively. Next, the IoU and OA metrics are generated using the following method: (12) (13) Next, calculate mIoU based on the IoU values of the changed area and the unchanged area:
[0045] (14) Sek is a measure to mitigate class imbalance by excluding unchanged regions. Specifically, let... Let be the confusion matrix, where, Indicates that it is marked as a category The pixels are classified into categories Quantity, , This represents the total number of all categories. The Sek metric is calculated using the following formula: (15) exist (16) (17) in, and It is along the confusion matrix The sum calculated in the j-th row and j-th column, excluding unchanged elements. .
[0046] The measurement It is the harmonic mean of precision and recall, providing an overall assessment of the variation class.
[0047] (18) in and Recall and recall are defined as follows: (19) (20) 3. Datasets used for testing: Xiamen (XM) farmland non-agricultural SCD dataset and the publicly available JinLin-1 (JL-1) farmland SCD dataset.
[0048] 4. Implementation details and results: The proposed CPGNet was trained for 100 epochs on an NVIDIA GeForce RTX 4090 with 24GB of memory using a batch size of 16. The optimizer was AdamW, and the model was trained with an initial learning rate of 0.0001. It was updated at each iteration according to 0.0001 × ((1-iteration) / total iterations)^1.5. Furthermore, geometric augmentations such as random flipping were applied to the input images to improve training performance.
[0049] Experimental results: 4.1 Ablation Experiment 4.1.1 Module Ablation Ablation tests were conducted by adding the proposed modules (SCM and MCFF) to the JL-1 dataset. Table 1 lists the accuracy values of the SCD results obtained by different methods, demonstrating that the accuracy of all evaluation metrics was significantly improved after including the proposed modules. Notably, the detection accuracy was significantly improved after adding SCM compared to the baseline method (i.e., removing all enhancement modules), especially with a 1.42% improvement in the Sek value. Similarly, it was found that adding the MCFF module improved the Sek value by 1.54%. Reasonably, combining the proposed SCM and MCFF achieved the highest accuracy across all evaluation metrics. These results further demonstrate that SCM and MCFF can effectively capture interesting variation information, reduce spurious variation interference, and alleviate underdetection. Further evaluation of the computational requirements and parameters of the proposed modules revealed that they introduce only a very small amount of computational overhead. See Table 1 for details, where 1M=10. 6 And 1 GFLOPs = 10 9 FLOPs (i.e., floating-point operations).
[0050] Figure 4 The ablation experiments of the proposed module on the JL-1 dataset are shown, where the pre-phase images are represented by (a), (c), and (e), and the post-phase images are represented by (b), (d), and (f). The proposed MCFF improves the integrity of the detection region, such as... Figure 4 The red circles in (a) and (c) illustrate this. Furthermore, incorporating the proposed SCM enhances the identification of unchanged regions and effectively reduces the impact of time differences (…). Figure 4 (b) Blue circle) and misclassification ( Figure 4 (d) The spurious changes caused by the blue circle. By integrating all modules, this method can accurately identify various semantic changes and maintain good detection boundaries even in blurry images. Figure 4 (e) and (f)).
[0051] Table 1 Ablation experiments of SCM and MCFF on the JL-1 farmland SCD dataset.
[0052] 4.1.2 Structural Ablation To evaluate the effectiveness of the auxiliary CD branch, three model variants were designed: 1) CPGNet NA (Remove auxiliary branches), 2) CPGNet Res (MPANet, based on Transformer, was replaced by ResNet34, based on CNN), 3) CPGNet WS (Implementing weight sharing in MPANet). Table 2 lists the implementations using CPGNet, CPGNet, and CPGNet.NA and CPGNet WSRes The evaluation results show that, compared to CPGNet, CPGNet with the addition of an extra auxiliary CD branch is superior. NA The detection accuracy was significantly improved, with mIoU and SeK metrics increasing by 1.37% and 2.66%, respectively. The results indicate that introducing an additional CD branch can effectively enhance change information, thereby improving detection accuracy. Furthermore, the CPGNet embedded transformer network can be observed as an auxiliary branch, compared to the CNN-based CPGNet. Res Compared to CPGNet, it achieved better performance. This success is closely related to the transformer's robust global feature extraction capability, which reduces detection errors. Furthermore, it eliminates the need for an auxiliary CD branch with a weight-sharing structure, which would otherwise lead to lower detection accuracy (i.e., CPGNet). WS This finding suggests that it is beneficial to take into account both time-specific characteristics and the dynamic changes in using different share structures.
[0053] Table 2 CPGNet, CPGNet NA CPGNet and CPGNet WSRes Comparison of SCD results on the JL-1 dataset
[0054] 4.2 Comparative Experiments Using Different Time-Spanning Modules To further evaluate the robustness of the proposed SCM, it was replaced with two recent cross-time modules: the Cot-SR and TC modules, forming the CPG respectively. Cot Net TC Table 3 shows the accuracy results of the three methods on the JL-1 farmland dataset. Clearly, compared to the other two methods, CPGNet achieved the highest detection accuracy across all evaluation metrics. Its SeK values are respectively... Cot Compared to CPGNet TC The results show improvements of 2.62% and 2.97% over CPGNet. Experimental results further confirm the robustness of the proposed SCM in modeling the temporal relationship between bitemporal features, improving the detection accuracy for changes in objects of interest with different appearances.
[0055] Figure 5 The detection error of farmland change on the JL-1 dataset is shown, demonstrating that the proposed CPGNet effectively reduces false negatives (i.e., undetected) and false positives, even in regions with blurred boundaries. Figure 5 (a) and (b)). The method accurately identifies areas of farmland change with high internal compaction in complex scenarios, reducing the time-dependent ( Figure 5 (c) and light differences ( Figure 5 (d) spurious changes caused by [the changes]. However, other methods cannot accurately identify farmland with these changes. The results show that the proposed SCM algorithm fully learns spectral invariance and temporal consistency information, effectively enhancing its ability to perceive unchanged areas.
[0056] Table 3 CPGNet, CPGNet and CPGNet TCCot Comparison of SCD results on the JL-1 dataset
[0057] 4.3 SCD of the JL-1 dataset Four pixel-based measurement methods were employed on the JL-1 SCD dataset, and the proposed method was compared with ten existing techniques (Table 4). Clearly, the proposed CPGNet outperforms Bi-SRNet and MTSCD-Net, achieving improvements of 8.64% and 20.12% in mIoU and 3.03% and 6.04% in SeK, respectively. Furthermore, the proposed method significantly outperforms these recent methods, such as HGINet and SCanNet, in terms of class consistency. F1 scd The values were improved by 2.03% and 7.19%, respectively. It also demonstrated better detection performance compared to SAM-based SCD methods (i.e., SAM-CD and SCD-SAM). Furthermore, SCD-SAM has a higher computational cost but higher detection accuracy compared to SAM-CD. Clearly, methods based on the standard triple branch, such as HRSCD-3 and HRSCD-4, are insufficient to capture semantics and variational relationships, resulting in poor detection accuracy on the JL-1 dataset. Additionally, when using 256×256 pixels as input, the number of parameters and computations were evaluated for all methods (Table 4). The results demonstrate that the method (i.e., CPGNet) achieves superior accuracy while increasing the computational load. Figure 6 The image shows the SCD results detected by various methods when applied to the JL-1 farmland SCD dataset, where the earlier temporal images are represented by (a), (c), and (e), and the later temporal images by (b), (d), and (f). This demonstrates that CPGNet accurately identifies various farmland semantic variations with relatively regular boundaries. Figure 6 (a) and (b)). Furthermore, it effectively mitigates the effects of spurious changes caused by spectral variations. Figure 6 (c) and (d)). Furthermore, our method effectively identifies farmland-to-building transitions of varying scales and maintains high intra-class consistency, while other methods fall short (see [link to relevant documentation]). Figure 6 (e) and (f)).
[0058] Table 4 compares the SCD results on the JL-1 dataset, along with the number of parameters and GFLOPs of floating-point operations.
[0059]
[0060] 4.4 SCD of XM non-agricultural dataset We also compared CPGNet with ten other methods on the XM dataset (Table 5). The results show that CPGNet outperforms all other evaluation methods, achieving the highest scores across all pixel-accuracy-based evaluation measures. Specifically, it outperforms existing methods by 2.24%–21.24% and 9.28%–65.98% in mIoU and SeK measurements, respectively. Furthermore, its method outperforms other methods in class consistency, resulting in a 30.01%–2.39% increase in value. It is worth noting that the CdSC and SCD-SAM methods failed to converge, which may be due to the presence of severely imbalanced pixels in the XM dataset. Furthermore, SCD-SAM exhibited the lowest detection accuracy, but the reasons for this require further investigation.
[0061] Then, the SCD results were displayed on the XM dataset using different methods, such as... Figure 7 As shown in the figure, it was observed that the proposed CPGNet accurately identified "farmland-to-farmland" information, thus effectively detecting shape and boundary information. Figure 7 (a) and (b)). Compared to recent methods like SCanNet and HGINet, this method also demonstrates superior performance in handling complex and changing scenarios. Figure 7 (c) and (d)). Furthermore, it effectively distinguishes different farmland changes and avoids interference from other objects, a challenge faced by other methods (besides HGINet) (see...). Figure 7 (e) and (f)). It is worth noting that in Figure 7 In the scenarios shown, these SAM-based methods (i.e., SAM-CD and SCD-SAM) failed to effectively locate the areas of change and identify the categories of change.
[0062] Finally, the proposed CPGNet is used to display the SCD results across the entire XM image, such as... Figure 8 As shown in the figure, XA and TA represent the Xiang'an and Tong'an regions, respectively. (A)-(D) represent magnified examples of detected changes. This figure demonstrates that CPGNet accurately detects potential semantic changes with different sizes and shapes in the XM region. It also effectively mitigates the effects of time, illumination, and phenological differences, producing satisfactory detection results.
[0063] Table 5 shows the SCD evaluation results of different methods on the XM dataset.
[0064] 5. Technical Effects This invention proposes a novel deep learning network (CPGNet) that considers semantic relationships and change priors in VHR remote sensing image SCDs. The goal is to capture robust change information representations of interesting objects from bi-temporal images using a multi-branch architecture within a multi-task learning paradigm. Secondly, SCs (Sequential Changes) between bi-temporal features are considered to enhance similar semantic representations (e.g., unchanged regions), thereby reducing considerable spurious changes caused by temporal, illumination, and seasonal effects. Finally, unlike existing methods that use attention mechanisms to enhance change information extraction from bi-temporal features, multi-view change features are explicitly introduced as input to the Bayesian network to enhance the perception of changed regions. This approach adds a limited computational cost but significantly improves the detection accuracy of change results (Table 1). Compared to 10 existing methods, the method of this invention demonstrates state-of-the-art (SOTA) performance, resulting in substantial improvements across all evaluation measures on the XM and public JL-1 SCD datasets. 5.1 The effectiveness of the proposed SCM in enhancing the correlation of bi-temporal features Unlike existing methods that focus on modeling SCs in high-level bitemporal features using spatial attention mechanisms, one of the main contributions of this study is to leverage the relationships between these spectral-temporal features to highlight consistent semantic cues. This approach enhances the perception of invariant objects, thereby reducing spurious changes caused by seasonal and illumination differences. Figure 4 The proposed SCM also surpasses two excellent SC modules (see Table 3), namely Cot-SR and TC, because it can fully simulate spectral invariance and time consistency.
[0065] 5.2 The MCFF module enhances the robustness of the representation of change information features. To fully integrate change information from different perspectives, this invention designs a Bayesian MCFF module. This approach alleviates the problems of insufficient and missed detection of change regions, such as... Figure 4As shown in Table 6, to further validate the advantages of the designed MCFF module, it was compared with four popular feature fusion methods: concatenation, multiplication, addition, and global attention. Table 6 presents the evaluation results of different methods on the JL-1 dataset. Clearly, CPGNet using Bayesian fusion achieves the highest accuracy compared to other fusion modalities. More specifically, it outperforms other methods by 0.93% to 3.13% and 0.36% to 1.56% in SeK and mIoU, respectively. This is likely due to the Bayesian fusion modality optimizing the feature distribution, thereby enhancing information about significant changes. Notably, feature fusion via concatenation and element-wise intelligent multiplication outperforms global attention methods in integrating multi-view variation features, demonstrating the effectiveness of the simple fusion strategy. The Bayesian-optimized MCFF module further leverages the synergistic advantages of these operations, achieving the highest SCD accuracy.
[0066] Visualizing the changes using different fusion modes, such as Figure 9 As shown, it is clear that the method integrating the MCFF module significantly enhances the changing features of the object of interest while suppressing irrelevant features. In contrast, other fusion methods, such as feature addition and multiplication, fail to effectively capture changing features, especially for... Figure 9 The scenario shown in (a). Furthermore, using join and global attention operations to fuse multi-view change features can over-extract change features (see [link]). Figure 9 (b)). Conversely, the proposed method effectively alleviates these problems. These results demonstrate the robustness of the proposed MCFF module in fusing and enhancing information about significant changes by optimizing feature distribution. Table 6 shows a comparison of SCD results (%) using different feature fusion modes on the JL-1 dataset.
[0067] Based on the same inventive concept as the above disclosure, this disclosure also provides an electronic device. The electronic device of this disclosure includes at least one processor and at least one memory electrically connected to the processor. The memory is electrically connected to the processor, wherein the memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the method described above.
[0068] It should be noted that the electrical connection between the above-mentioned units does not necessarily mean the connection between lines. The indirect connection method can be applied to the embodiments of this disclosure as long as it achieves the purpose of this disclosure.
[0069] Based on the same inventive concept, this disclosure also provides a computer storage medium storing a computer program, which, when executed by a processor, implements the above-described method.
[0070] Although the present disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure.
Claims
1. A method for detecting semantic change in remote sensing images, characterized in that, Comprise: inputting the double-time-phase images to be detected into the trained semantic change detection model; The trained semantic change detection model performs remote sensing semantic change detection to obtain a previous time phase semantic segmentation result ST 1. a later time phase semantic segmentation result ST 2 and fine semantic change features ; To fine semantic change features A binary change detection map is obtained by performing a binarization process on the binary change detection map; the semantic segmentation result of the previous time phase ST 1and the semantic segmentation result of the subsequent time phase ST 2respectively utilize the binary change detection map to perform a mask operation to obtain a semantic detection map of the previous time phase and a semantic detection map of the subsequent time phase; Wherein, the trained semantic change detection model comprises: a feature extraction module, a semantic correlation module SCM and a multi-view change feature fusion module MCFF, wherein: a feature extraction module configured to extract, using a weight-shared twin feature extractor, pre-time-point fine-grained semantic features from the dual-time-point images and post-time-point fine-grained semantic features ; and extract, using an independent feature extractor, coarse-grained semantic change features from the dual-time-point images ; a semantic correlation module SCM, configured to correlate the preceding-phase fine-grained semantic features and the succeeding-phase fine-grained semantic features in sequence to obtain enhanced preceding-phase fine-grained semantic features and enhanced succeeding-phase fine-grained semantic features ; a multi-view change feature fusion module MCFF configured to fuse the coarse-grained semantic change features as prior information according to Bayesian theory with the mode, the enhanced pre-phase fine-grained semantic features and the enhanced post-phase fine-grained semantic features to generate fine semantic change features ; and perform convolution operations on the enhanced pre-phase fine-grained semantic features and the enhanced post-phase fine-grained semantic features respectively to obtain a pre-phase semantic segmentation result ST 1 and a post-phase semantic segmentation result ST 2.
2. The method of claim 1, wherein, The Semantic Association Module (SCM) is specifically used to analyze fine-grained semantic features from previous time periods. and post-temporal fine-grained semantic features Perform element-wise multiplication to enhance similar semantic information.
3. The method of claim 1, wherein, a multi-view change feature fusion module MCFF for fusing coarse-grained semantic change features as prior information, the enhanced pre-phase fine-grained semantic features and the enhanced post-phase fine-grained semantic features are fused by applying Bayesian theory to enhance the fine-grained semantic change features between them, and based on the coarse-grained semantic change features and the enhanced fine-grained semantic change features to generate refined semantic change features.
4. The method of claim 3, wherein, The multi-view change feature fusion module MCFF is specifically configured to calculate fine-grained semantic change features between enhanced fine-grained semantic features of a previous phase and fine-grained semantic features of a later phase . 5. The method according to claim 3 or 4, characterized in that, The multi-view change feature fusion module MCFF is specifically configured to splice the coarse-grained semantic change features with the enhanced fine-grained semantic change features , and use a convolution block to reduce the feature dimension to generate fine semantic change features .
6. The method of claim 1, wherein, The semantic change detection model integrates two semantic segmentation tasks and a change detection task for different time-phase images, and uses the following multi-task loss function to train the semantic change detection model: wherein, denotes the total task loss, denotes the semantic segmentation task loss for the pre-phase imagery, denotes the semantic segmentation task loss for the post-phase imagery, denotes the change detection task loss, denotes the semantic change loss.
7. The method of claim 6, wherein, The The feature vector based on the semantic segmentation result of the previous phase image and the feature vector based on the semantic segmentation result of the later phase image are calculated by a cosine loss.
8. A device for detecting semantic changes in remote sensing images, characterized in that, For implementing the method of any one of claims 1-7, comprising: an input module, a trained semantic change detection model and an output module, wherein: The input module is used for inputting the double-time-phase images to be detected into the trained semantic change detection model; The trained semantic change detection model is used for remote sensing semantic change detection to obtain a previous time semantic segmentation result ST 1. a later time semantic segmentation result ST 2 and a fine semantic change feature ; The output module is used for fine-grained semantic variation features. Binarization is performed to obtain a binary change detection map; the semantic segmentation results of the previous time phase are then processed. ST 1 and post-temporal semantic segmentation results ST 2. Masking operations are performed using the binary change detection map to obtain the front-phase semantic detection map and the back-phase semantic detection map.
9. An electronic device, comprising: Comprise a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory complete the mutual communication through the communication bus; The memory is used for storing a computer program; The processor is used for executing the program stored on the memory, and realizes the method of any one of claims 1-7.
10. A computer storage medium, characterized in that, The computer storage medium stores a computer program, and the computer program is executed by the processor to realize the method of any one of claims 1-7.