Remote sensing image semantic change detection method and system based on time-space relationship modeling and edge information enhancement

Through the method of spatiotemporal relationship modeling and edge information enhancement, combined with convolutional neural network and visual state space model, the edge blur problem in remote sensing image change detection is solved, and the detection accuracy and consistency are improved, especially in the recognition of changing areas in complex geographic scenes.

CN120388279APending Publication Date: 2025-07-29Chinese People's Liberation Army Cyberspace Force Information Engineering University
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510364935.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-12-27
Filing Date
2025-03-26
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The existing semantic change detection methods for remote sensing images have bottlenecks in the problem of blurring edges of changing objects, and the consistency between the detection results and the actual changing areas needs to be improved, especially in complex geographic scenarios, which is low in object segmentation accuracy, which affects the semantic change detection results.

Method used

Using a method based on spatiotemporal relationship modeling and edge information strengthening, a twin feature encoder is constructed by fusing convolutional neural networks and visual state space models, combining edge perception reinforcement networks, extracting semantic features of remote sensing images and identifying the edges of changing objects, and using a bidirectional spatiotemporal relationship modeler to mine the logical relationship of spatiotemporal change, and constructing a semantic change detection model of remote sensing images.

Benefits of technology

The accuracy of remote sensing image change area detection and semantic category recognition is improved, and the ability to explore potential land object change logic is enhanced. The consistency between the detected change areas and the actual change areas is more accurate, especially in changing areas with irregular shapes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388279A_ABST
    Figure CN120388279A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of remote sensing image processing, in particular to a remote sensing image semantic change detection method and system based on time-space relation modeling and edge information enhancement, and the method comprises the steps: obtaining a front time-phase remote sensing image and a rear time-phase remote sensing image of a target region, and carrying out the preprocessing of the front and rear time-phase remote sensing images; inputting the front-time-phase remote sensing image and the rear-time-phase remote sensing image of the target area into a pre-trained remote sensing image semantic change detection model to obtain a change detection result; wherein the remote sensing image semantic change detection model is fused with a convolutional neural network and a visual state space model VSSM to construct a twinborn feature encoder to extract remote sensing image semantic features, and a bidirectional time-space relationship modeling device based on the visual state space model VSSM excavates a time-space change logic relationship between dual-temporal remote sensing images. And identifying the edge of the dual-time-phase remote sensing image change object by using the edge perception enhancement network. The remote sensing image change detection precision can be improved, and the potential ground feature change logic mining capability is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of remote sensing image processing, and particularly relates to a remote sensing image semantic change detection method and system based on spatio-temporal relationship modeling and edge information enhancement. Background Art

[0002] Remote sensing semantic change detection (SCD) is a technology for identifying the location of surface changes and semantic change information on the Earth's surface by comparing and analyzing remote sensing images at the same geographical location at different times. With the continuous development of China's space remote sensing technology, very-high-resolution (VHR) remote sensing images are becoming increasingly easy to obtain, which provides rich data resources for global surface semantic change detection. Obtaining fine semantic change information from VHR images can continuously and dynamically observe surface changes, and has important application values for land planning, urban management, and sustainable development.

[0003] In recent years, deep learning technology has developed rapidly and has been successfully applied to the intelligent interpretation of remote sensing images. The SCD technology based on deep learning has also become a current research hotspot. Compared with binary change detection, SCD not only locates the changed areas but also provides detailed semantic information before and after the change. Currently, SCD methods can be divided into three categories according to the network structure: namely, feed-forward SCD, siamese SCD, and multi-task SCD. The feed-forward method converts the SCD task into a semantic segmentation task. First, multi-temporal VHR images are mosaicked in the band dimension, and then the mosaicked data is input into a semantic segmentation network to extract change information. This method based on early fusion is vulnerable to the mutual influence of pixels between different phases when extracting semantic change information, resulting in difficult convergence during the training process and insufficient model generalization performance. To solve this problem, siamese neural networks are used for the SCD task. The siamese SCD method uses a dual-branch structure to extract multi-scale semantic features of dual-temporal images respectively, and then fuses the dual-temporal features to obtain the SCD result. This type of method effectively improves the model robustness and the accuracy of semantic change detection. However, SCD is a composite task that requires both obtaining the location of the changed area and extracting the semantic change categories. It is still difficult to extract deep semantic change information using only one semantic segmentation branch. Moreover, this method is still difficult to apply to tasks with a large number of change categories, and there is still the problem of sample class imbalance, and its model applicability needs further study. Since then, multi-task SCD methods have been continuously proposed. According to the backbone network, multi-task SCD methods can be divided into two categories: methods based on Convolutional Neural Network (CNN) and methods based on Transformer. The multi-task SCD method based on CNN has insufficient representation ability for global context information and is still difficult to capture deep and subtle semantic change features. Although the SCD model based on Transformer improves the context information representation ability, this method has a large number of parameters and limited local feature extraction ability.

[0004] Currently, the problem of blurred edges of changed ground objects still restricts the improvement of SCD accuracy. Due to small spectral differences between foreign objects or shadow occlusion, it is difficult to accurately locate the edge contours of complex ground objects. Existing SCD methods usually generate blurred change probability values in the edge area, resulting in a large difference between the detected change edge and the true edge. To improve the accuracy of change detection results, some scholars have proposed introducing object-oriented image analysis technology into the change detection network. However, due to the insufficient performance of object segmentation methods, the object segmentation accuracy is low in complex scenes, which in turn affects the semantic change detection results. Summary of the Invention

[0005] To this end, the present invention provides a remote sensing image semantic change detection method and system based on spatio-temporal relationship modeling and edge information enhancement, which solves the problems of fuzzy edges of changed objects and the need to improve the consistency between the detection results and the actual changed areas in existing remote sensing image semantic detection.

[0006] According to the design scheme provided by the present invention, on the one hand, a remote sensing image semantic change detection method based on spatio-temporal relationship modeling and edge information enhancement is provided, including:

[0007] Obtain the pre-temporal remote sensing image and the post-temporal remote sensing image of the target area, and preprocess the dual-temporal remote sensing images.

[0008] Input the pre-temporal remote sensing image and the post-temporal remote sensing image of the target area into a pre-trained remote sensing image semantic change detection model, and use the remote sensing image semantic change detection model to obtain the semantic change detection result of the dual-temporal remote sensing image of the target area.

[0009] Among them, the remote sensing image semantic change detection model fuses a convolutional neural network and a visual state space model VSSM to construct a twin feature encoder to extract the semantic features of the remote sensing image, a bidirectional spatio-temporal relationship modeler based on the visual state space model VSSM mines the spatio-temporal change logical relationship between the dual-temporal remote sensing images, and an edge-aware enhancement network is used to identify the edges of the changed objects in the dual-temporal remote sensing images.

[0010] As the remote sensing image semantic change detection method based on spatio-temporal relationship modeling and edge information enhancement of the present invention, further, the remote sensing image semantic change detection model includes a twin feature encoder for extracting the semantic features of the dual-temporal remote sensing images, an edge-aware enhancement network for perceiving the edge information of the dual-temporal remote sensing images, a spatio-temporal relationship modeler for reshaping the semantic features of the remote sensing images, a classification decoder for land cover classification of the reshaped semantic features of the remote sensing images, a change area decoder for obtaining the changed area of the dual-temporal remote sensing images by using the semantic features and edge information of the dual-temporal remote sensing images, and an output end for obtaining a semantic change detection map of the dual-temporal remote sensing images by using the land cover classification result and the changed area of the dual-temporal remote sensing images. The spatio-temporal relationship modeler is a dual-temporal spatio-temporal relationship modeler based on the visual state space model VSSM for mining the spatio-temporal connection between the dual-temporal features and performing feature reshaping.

[0011] As a remote sensing image semantic change detection method based on spatiotemporal relationship modeling and edge information enhancement of the present invention, the twin feature encoder is further composed of a ResNet34 network and a visual state space model VSSM, wherein ResNet34 is used to use several layers of convolution kernels to extract the multi-scale semantic features of the input remote sensing image, and embed the visual state space model VSSM in the convolution kernels of the last two layers to obtain the global contextual semantic information of the remote sensing image.

[0012] As a remote sensing image semantic change detection method based on spatiotemporal relationship modeling and edge information enhancement of the present invention, further, the visual state space model VSSM includes a selective scanning branch and an MLP branch, and uses a normalization layer to normalize the output of the selective scanning branch, and performs a point multiplication operation on the normalized semantic features and the semantic features output by the MLP branch, and splices them with the input remote sensing image semantic features in the channel dimension to obtain the semantic features finally output by the twin feature encoder; the selective scanning branch is used to perform point-by-point convolution and nonlinear activation processing on the input remote sensing image semantic features using a depth-wise separable convolution layer and a SiLU activation function, and expand the remote sensing image semantic features through a two-dimensional selective scanning mechanism SS2D to establish a global receptive field, and the MLP branch is used to perform linear operations on the input remote sensing image semantic features.

[0013] As a remote sensing image semantic change detection method based on spatiotemporal relationship modeling and edge information enhancement of the present invention, further, the bidirectional spatiotemporal relationship modeler adopts two scanning methods based on dual-phase cross scanning and space-first then time scanning to perform feature scanning on the semantic features of the remote sensing image, and cross-feeds the scanned features to the visual state space model VSSM, and uses the visual state space model VSSM to selectively memorize the historical information of the semantic features and perform feature reshaping to explore the spatiotemporal relationship between the dual-phase features, wherein each scanning method adopts two scanning orders, forward from before the change to after the change and reverse from after the change to before the change, to perform feature scanning.

[0014] As a remote sensing image semantic change detection method based on spatiotemporal relationship modeling and edge information enhancement of the present invention, further, the edge perception enhancement network utilizes edge detection and edge feature enhancement mechanisms to perceive the edges of changing objects in dual-phase remote sensing images, wherein the edge detection mechanism uses a Gaussian smoothing filter to smooth the remote sensing image and uses the Laplace operator to extract the initial edge features of the remote sensing image; the edge feature enhancement mechanism utilizes the initial edge features to obtain the edges of changing objects in the shallow features extracted by the twin feature encoder and enhances the edges of changing objects by fusing RGB and edge features through a gated fusion unit.

[0015] As a method for remote sensing image semantic change detection based on spatio-temporal relationship modeling and edge information enhancement, further, the training process of the remote sensing image semantic change detection model includes:

[0016] Collect remote sensing image semantic change detection sample data, and perform enhancement processing on the sample data to obtain a remote sensing image semantic change detection sample data set;

[0017] Based on the cross-entropy loss of semantic segmentation of remote sensing images in each time phase, the remote sensing image change region detection loss, and the remote sensing image change semantic consistency loss, construct a joint loss function for model training;

[0018] Based on the joint loss function and using the remote sensing image semantic change detection sample data set, train and optimize the remote sensing image semantic change detection model.

[0019] On the other hand, the present invention also provides a remote sensing image semantic change detection system based on spatio-temporal relationship modeling and edge information enhancement, including: an image acquisition module and a change detection module, where

[0020] The image acquisition module is used to acquire the pre-time-phase remote sensing image and the post-time-phase remote sensing image of the target area, and perform preprocessing on the front and back double-time-phase remote sensing images;

[0021] The change detection module is used to input the pre-time-phase remote sensing image and the post-time-phase remote sensing image of the target area into the pre-trained remote sensing image semantic change detection model, and use the remote sensing image semantic change detection model to obtain the remote sensing image semantic change detection result of the target area double-time-phase remote sensing image;

[0022] Among them, the remote sensing image semantic change detection model fuses a convolutional neural network and a visual state space model VSSM to construct a twin feature encoder to extract remote sensing image semantic features, mines the spatio-temporal change logical relationship between double-time-phase remote sensing images through a two-way spatio-temporal relationship modeler based on the visual state space model VSSM, and uses an edge-aware enhancement network to identify the edges of the change objects in the double-time-phase remote sensing images.

[0023] The beneficial effects of the present invention:

[0024] The present invention combines CNN and VSSM to construct a semantic change detection network model CVS-Net, enabling it to not only have the local feature extraction ability of CNN but also possess the context information modeling ability of the spatial state model. In addition, VSSM is used to construct a dual-temporal semantic feature spatio-temporal modeler to enhance the semantic change detection accuracy, and an edge-aware reinforcement network is utilized to improve the edge detection accuracy of changed objects and reduce the uncertainty of object edges. Further, experimental data is used to verify the semantic change detection performance of the solution in this case. On the SECOND dataset, a separation kappa coefficient (Sek) of 23.95% and an average intersection over union (mIoU) of 72.89% are achieved. On the FZ-SCD dataset, SeK reaches 23.02% and mIoU reaches 72.60%. The experimental data shows that the solution in this case can improve the detection accuracy of the changed area and semantic class recognition in remote sensing images, enhance the ability to mine the logic of potential object changes, and the consistency between the detected changed area and the actual changed area is more accurate, especially in the changed area with irregular shapes. Description of the Drawings

[0025] Figure 1 Schematic diagram of the remote sensing image semantic change detection process based on spatio-temporal relationship modeling and edge information enhancement in the embodiment;

[0026] Figure 2 Schematic diagram of the feature propagation process of the state space module SSM in the embodiment;

[0027] Figure 3 Schematic diagram of the architecture of the remote sensing image semantic encoding detection model in the embodiment;

[0028] Figure 4 Schematic diagram of the network structure of the twin feature encoder in the embodiment;

[0029] Figure 5 Schematic diagram of the working principle of the dual-temporal spatio-temporal feature modeler in the embodiment;

[0030] Figure 6 Schematic diagram of the working principle of the edge-aware reinforcement network in the embodiment;

[0031] Figure 7 Schematic diagram of the FZ-SCD dataset image and semantic change label in the embodiment;

[0032] Figure 8 Schematic diagram of the SCD results of different methods on the SECOND dataset in the embodiment;

[0033] Figure 9 Schematic diagram of the SCD results of different methods on the FZ-SCD dataset in the embodiment;

[0034] Figure 10Schematic of ablation experiment results on the SECOND dataset in the embodiment;

[0035] Figure 11 Schematic of the comparison of the coincidence degree between the changed regions extracted by different methods on the SECOND dataset in the embodiment and the actual changed regions. Detailed implementation manners

[0036] To make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and technical solutions.

[0037] Semantic change detection can simultaneously obtain the ground object change regions and semantic change information, providing information support for application fields such as land use monitoring, urban planning, and sustainable development. Ultra-high-resolution remote sensing images have a more refined representation of ground object information and can capture subtle changes in ground object features on the earth's surface. However, existing semantic change detection methods have insufficient utilization of local and global features of ultra-high-resolution images and ignore the spatio-temporal dependence between different time phases, resulting in inaccurate land cover semantic classification results. In addition, there is a problem of blurred edges in the detected change objects, and the consistency between the detection results and the actual change regions needs to be improved. To address these problems, inspired by the Vision State Space Model (VSSM) with the ability to process long sequences, the embodiment of the present invention provides a semantic change detection method for remote sensing images based on spatio-temporal relationship modeling and edge information enhancement, as Figure 1 shown, including:

[0038] S101. Obtain the pre-phase remote sensing image and the post-phase remote sensing image of the target area, and preprocess the dual-phase remote sensing images before and after;

[0039] S102. Input the pre-phase remote sensing image and the post-phase remote sensing image of the target area into a pre-trained semantic change detection model for remote sensing images, and use the semantic change detection model for remote sensing images to obtain the semantic change detection results of the dual-phase remote sensing images of the target area;

[0040] Wherein, the semantic change detection model for remote sensing images fuses a convolutional neural network and the visual state space model VSSM to construct a twin feature encoder to extract semantic features of remote sensing images, a bidirectional spatio-temporal relationship modeler based on the visual state space model VSSM mines the spatio-temporal change logical relationship between the dual-phase remote sensing images, and an edge-aware enhancement network is used to identify the edges of the change objects in the dual-phase remote sensing images.

[0041] SSM is a mathematical model that describes the behavior of a dynamic system, and it uses a set of first-order differential equations (continuous-time system) or difference equations (discrete-time system) to represent the internal state evolution of the system:

[0042] h′(t) = Ah(t) + Bx(t) (1)

[0043] Meanwhile, another set of equations is used to describe the relationship between the system state and the output:

[0044] y(t) = Ch(t) + Dx(t) (2)

[0045] The forward propagation process of the features in the SSM is as Figure 2 shown, and its hidden state h at each moment t is calculated based on the current input x t and the hidden state h at the previous moment t-1 The coefficient matrix A stores the essence of all previous historical information, and the SSM updates the spatial hidden state h(t) at the next moment based on the coefficient matrix A. Assuming there is an input signal x(t), this signal is first multiplied by the matrix B, and then the result of multiplying the hidden state h(t) by the matrix A is added to obtain the updated hidden state h′(t). Then, the matrix C is used to convert the state into an output, and it is concatenated with the direct signal from the input to the output provided by the matrix D to obtain the final output.

[0046] In the embodiment of this case, the architecture of the remote sensing image semantic change detection model CVS-Net is as Figure 3 shown. CVS-Net follows the multi-task SCD paradigm and simultaneously performs land cover classification and regional change detection tasks. In the encoding stage, the extraction of semantic features and edge information is decoupled into two paths. The multi-level semantic features of the dual-temporal images are extracted by the twin CNN-VSS feature encoder, while the edge-aware strengthening branch (BAS) is used to extract edge information. To mine the spatio-temporal correlation between the deep dual-temporal semantic features, the spatio-temporal relationship modeler (ST-VSSM) is used to reshape the dual-temporal semantic features. Then, the fused multi-level semantic features and edge features are input into the change region decoder to obtain the change region results. The dual-temporal land cover classification results are obtained by the land cover classification decoder. Finally, the classification results are mapped to the change region to generate the dual-temporal remote sensing semantic change detection SCD results.

[0047] Specifically, the remote sensing image semantic change detection model can be designed to include a twin feature encoder for extracting semantic features of dual-temporal remote sensing images, an edge perception enhancement network for perceiving edge information of dual-temporal remote sensing images, a spatio-temporal relationship modeler for reshaping the semantic features of remote sensing images, a classification decoder for classifying the reshaped semantic features of remote sensing images into land cover types, a change region decoder for obtaining the change regions of dual-temporal remote sensing images by using the semantic features and edge information of dual-temporal remote sensing images, and an output end for obtaining a semantic change detection map of dual-temporal remote sensing images by using the land cover classification results and the change regions of dual-temporal remote sensing images. The spatio-temporal relationship modeler is a dual-temporal spatio-temporal relationship modeler based on the visual state space model VSSM for mining the spatio-temporal connection between dual-temporal features and performing feature reshaping.

[0048] Among them, the twin feature encoder is composed of a ResNet34 network and a visual state space model VSSM. Among them, ResNet34 is used to extract multi-scale semantic features of the input remote sensing image by using several layers of convolutional kernels, and the visual state space model VSSM is embedded in the convolutional kernels of the last two layers respectively to obtain the global context semantic information of the remote sensing image. The visual state space model VSSM includes a selective scanning branch and an MLP branch, and uses a normalization layer to normalize the output of the selective scanning branch. The normalized semantic features are multiplied point by point with the semantic features output by the MLP branch, and then concatenated with the semantic features of the input remote sensing image in the channel dimension to obtain the final output semantic features of the twin feature encoder. The selective scanning branch is used to perform pointwise convolution and non-linear activation processing on the semantic features of the input remote sensing image by using a depthwise separable convolutional layer and a SiLU activation function, and expand the semantic features of the remote sensing image through a two-dimensional selective scanning mechanism SS2D to establish a global receptive field. The MLP branch is used to perform a linear operation on the semantic features of the input remote sensing image.

[0049] Combining CNN and VSSM to construct a twin feature encoder CNN-VSS aims to simultaneously improve the local and global feature extraction capabilities of the change detection network. As Figure 4As shown in the figure, CNN-VSS consists of an improved ResNet34 network and a visual state space module. The dual-temporal images I1 and I2 are first input into the ResNet34 network, and through the repeated convolutional operations in the conv1_x to conv5_x stages, multi-scale semantic features of the images are extracted. To prevent local information loss, ResNet34 is improved by referring to the network design of HRNet, and the downsampling operations in the residual structures from conv3_x to conv5_ are removed, so that the feature resolution is always maintained at a relatively high size (1 / 8 of the input image size). To model the global context semantic information, a VSSM is embedded after each of the conv4_x to conv5_x stages. The VSSM contains two feature processing branches, namely the selective scanning branch and the MLP branch. The dual-temporal semantic features are input into the selective scanning branch and the MLP branch simultaneously after passing through the initial linear embedding layer. In the selective scanning branch, the features pass through a 3×3 depthwise separable convolutional layer and the SiLU activation function, and then are input into the two-dimensional selective scanning module (SS2D). The MLP branch performs simple linear operations on the input features. Finally, the output of SS2D passes through a normalization layer LN, and then performs a dot product operation with the output of the MLP branch, and is concatenated with the original input features in the channel dimension to obtain the output features. Four scales of dual-temporal features are obtained through CNN-VSS For subsequent operations.

[0050] Among them, the bidirectional spatio-temporal relationship modeler uses two scanning methods, namely dual-temporal cross-scanning and spatial-temporal scanning (first spatial then temporal), to scan the semantic features of remote sensing images, and cross-feeds the scanned features into the visual state space model VSSM. The visual state space model VSSM selectively memorizes the historical information of semantic features and performs feature reshaping to explore the spatio-temporal relationship between dual-temporal features. Among them, each scanning method uses two scanning orders, namely the forward order from before change to after change and the reverse order from after change to before change, to scan features.

[0051] To fully explore the spatio-temporal connection of dual-temporal images and guide the network to learn the spatio-temporal change logic between images, a dual-temporal spatio-temporal relationship modeler (ST-VSSM) based on VSSM, as Figure 5As shown. Inspired by the scanning method of Visual Mamba, the dual-temporal scanning strategy in the embodiments of this case includes two scanning methods, namely dual-temporal cross-scanning and spatial-temporal scanning. In addition, the spatio-temporal connection in image change detection is related to both the positive and reverse temporal order analyses of the image changing over time: the positive temporal order analysis helps to understand the trends and laws of changes, and the reverse temporal order analysis helps to reveal hidden laws or anomalies. Therefore, each scanning method of ST-VSSM adopts a bidirectional scanning strategy, that is, the scanning order from before change to after change (forward) and from after change to before change (backward). The dual-temporal features are cross-fed into the VSSM, and by virtue of the VSSM's ability to selectively remember historical information and solve long-distance dependencies, the spatio-temporal relationship mining between the dual-temporal features is realized.

[0052] As Figure 5 shown, the dual-temporal features output by the last stage of the encoder are sliced into N×N sub-feature blocks through a patch embedding operation, denoted as Taking the forward scanning order as an example, in dual-temporal cross-scanning, first is input into the VSSM, and the VSSM dynamically adjusts matrix A according to this input, and is passed to the next input Repeat the above operation, and the VSSM will also receive the previous input information, thereby reshaping the dual-temporal time relationship features. In spatial-temporal scanning, the T1 feature sequence is first input into the VSSM in sequence, and then the T2 feature sequence is input into the VSSM. Finally, the reshaped features obtained by each scanning method through forward and reverse sequence modeling are fused by matrix addition. Then, the bidirectional fusion features obtained by the two scanning methods are concatenated in the channel dimension and smoothed by a 1×1 convolution operation to obtain the final dual-temporal features.

[0053] Among them, the edge-aware enhancement network uses an edge detection and edge feature enhancement mechanism to perceive the edges of change objects in dual-temporal remote sensing images. Among them, the edge detection mechanism uses a Gaussian smoothing filter to smooth the remote sensing image and uses a Laplacian operator to extract the initial edge features of the remote sensing image; the edge feature enhancement mechanism uses the initial edge features to obtain the edges of change objects in the shallow features extracted by the Siamese feature encoder and enhances the edges of change objects by fusing RGB and edge features through a gated fusion unit.

[0054] As Figure 6 shown, a Gaussian smoothing filter is used to reduce the sensitivity of the image to noise; subsequently, a 3×3 Laplacian operator is used to generate the initial edge feature B LO . However, directly using the single-channel BLO Input into the deep model results in insufficient learning of high-level semantic information. Therefore, in the embodiments of this case, an edge feature enhancement mechanism is utilized to obtain high-level semantic features through convolutional operations. The edge perception reinforcement mechanism only uses the conv1 and conv2_x of ResNet34 to enhance edge features, as Figure 6 shown. This shallow network structure can ensure that the model learns more refined spatial details and rich context information. Two levels of edge features of the dual-temporal image are obtained through the edge enhancement module and Then, a gated fusion module (GFM) is used to fuse the RGB and edge feature maps. The GFM first concatenates the image features and edge features in the channel dimension, then reduces the dimension of the concatenated features through a convolutional layer, and then uses a Sigmoid layer to calculate the modal weights. Multiply this tensor element-wise with the features of the two modalities respectively, and finally concatenate the weighted features of the two modalities in the channel dimension. Finally, the fused features and the multi-scale semantic features of the image are input into the change decoder. The decoder first upsamples the high-level semantic features, then concatenates the features in the channel dimension and performs convolutional dimension reduction, and obtains the pixel-level change probability through the Softmax function.

[0055] Specifically, the training process of the remote sensing image semantic change detection model can be designed to include:

[0056] Collect remote sensing image semantic change detection sample data, and perform enhancement processing on the sample data to obtain a remote sensing image semantic change detection sample data set;

[0057] Construct a joint loss function for model training based on the cross-entropy loss of semantic segmentation of each temporal remote sensing image, the remote sensing image change region detection loss, and the remote sensing image change semantic consistency loss;

[0058] Based on the joint loss function and using the remote sensing image semantic change detection sample data set, train and optimize the remote sensing image semantic change detection model.

[0059] The joint loss function is jointly trained by four loss functions, where the multi-class cross-entropy loss function L s is used to train the semantic classification task; the DiceLoss function L R is used to train the change region detection task; the change semantic consistency loss function L sc is used to ensure the consistency of semantic information and change information. The calculation formula of the joint loss function L scd can be expressed as:

[0060] L scd = w1 * L s1 + w2 * L s2 + w3 * L R + w4 * Lsc (3)

[0061] Among them, L s1 and L s2 are the semantic segmentation losses of each temporal phase image, and w1 to w4 are weight coefficients with a sum of 1. In this study, all weight coefficients are set to 0.25.

[0062] Furthermore, based on the above method, the present invention also provides a remote sensing image semantic change detection system based on spatio-temporal relationship modeling and edge information enhancement, including: an image acquisition module and a change detection module. Among them,

[0063] The image acquisition module is used to acquire the pre-temporal remote sensing image and the post-temporal remote sensing image of the target area, and preprocess the dual-temporal remote sensing images.

[0064] The change detection module is used to input the pre-temporal remote sensing image and the post-temporal remote sensing image of the target area into the pre-trained remote sensing image semantic change detection model, and use the remote sensing image semantic change detection model to obtain the semantic change detection result of the dual-temporal remote sensing image of the target area.

[0065] Among them, the remote sensing image semantic change detection model fuses a convolutional neural network and a visual state space model VSSM to construct a siamese feature encoder to extract the semantic features of the remote sensing image, uses a bidirectional spatio-temporal relationship modeler based on the visual state space model VSSM to mine the spatio-temporal change logical relationship between the dual-temporal remote sensing images, and uses an edge-aware enhancement network to identify the edges of the changed objects in the dual-temporal remote sensing images.

[0066] To verify the effectiveness of the solution of this case, the following further explanation is made in combination with experimental data:

[0067] Four accuracy metrics commonly used in the SCD task are adopted, including overall accuracy (OA), mean intersection over union (mIoU), separated Kappa coefficient (SeK), and F scd score. OA is the ratio of the number of correctly classified pixels to the total number of pixels in the image. Assume Q = {q ij} is the confusion matrix, where q ij represents the number of pixels classified into class i, and j represents its true label class (where 0 represents no change). The calculation formula of OA is:

[0068]

[0069] To comprehensively evaluate the SCD results with uneven classes, two metrics, mIoU and SeK, are introduced to evaluate the detection ability of the changed area and the discrimination ability of the changed semantics respectively. The calculation formula of mIoU is as follows:

[0070] mIoU = (IoUnc +IoU c ) / 2 (5)

[0071]

[0072] Among them, nc represents unchanged and c represents changed. SeK is calculated according to the confusion matrix Q = {q ij}, where However Its calculation formula is:

[0073]

[0074] F scd Focus on evaluating the detection accuracy of the changed area, through precision (P scd ) and recall (R scd ) calculated as follows:

[0075]

[0076] Validation experiments were carried out on the SECOND dataset and the Area A dataset (FZ-SCD). SECOND is the SCD benchmark dataset, containing 2968 pairs of image samples with an image size of 512×512. The training set, validation set, and test set can be divided in a ratio of 7:2:1.

[0077] The FZ-SCD dataset is a self-made dataset for the experiment. The research area is as Figure 7 shown, and the red box is the test area. The FZ-SCD dataset contains two GF-2 remote sensing images, taken on December 7, 2016 and February 18, 2020 respectively. For each image, the panchromatic band and the multispectral band are fused using the Gram-Schmidt method, and the fused GF-2 image is resampled to 0.8m, and then the two-temporal images are georegistered. The FZ-SCD dataset includes 6 main land cover classes, namely bare land, buildings, vegetation, water, roads, and others. The dataset is uniformly cropped to a size of 512×512 pixels. After sample enhancement operations, a total of 2200 pairs of samples are obtained, and the division ratio of the training set, validation set, and test set is also 7:2:1.

[0078] The proposed solution in this case is compared and analyzed with seven other mainstream SCD methods, which are: HRSCD.str4, Bi-SRNet, ChangeMamba, SCanNet, TED, SMNet, and MTSCD-Net. All comparative experiments and ablation experiments are carried out in a hardware environment equipped with an RTX 4090 24G GPU. To ensure the fairness of experimental comparison, CVS-Net and the comparison models use the same dataset, hyperparameter configuration, and code running environment. All methods are based on the Ubuntu 20.04 operating system and are implemented using the Python programming language. When training the CVS-Net model, Adam is selected as the optimizer, and the number of training epochs is set to 100, the batch size is 4, and the initial learning rate is 10 -4 .

[0079] 1. Comparative analysis of experimental results on the SECOND dataset

[0080] On the SECOND dataset, the experimental results of the CVS-Net method and the other seven mainstream SCD methods are shown in Table 1. CVS-Net achieved the best SCD results, outperforming other methods in all four key metrics. OA, mIoU, and SeK reached 88.26%, 73.07%, and 23.95% respectively. Among the comparison methods, SCanNet performed the best, followed by ChangeMamba. SCanNet utilizes the spatio-temporal dependence relationship of dual-temporal images to improve the accuracy of SCD and uses spatio-temporal constraints to guide semantic information learning. However, this method has poor recognition accuracy for ground object categories with small inter-class differences, such as the distinction between low vegetation and trees, as shown in Figure 8 the red box area in (a). ChangeMamba first introduced the Mamba architecture into the field of remote sensing change detection and used VSSM as the encoder to fully learn global context information. However, this architecture does not make full use of shallow features such as geometry and texture in high-resolution remote sensing images, resulting in limited ability in extracting details of changed ground objects, as shown in Figure 8 the red box areas in (a) and (c).

[0081] In contrast, CVS-Net combines the advantages of CNN and VSSM. In the early encoding stage, CNN is used to fully capture low-level semantic information; in the later stage, the VSSM module is introduced to enhance the global modeling ability. As Figure 8 shown, CVS-Net achieved the best results in both change semantic category recognition and change area extraction. Especially in the red box areas in Figure 8 (a) and (c), only CVS-Net accurately recognized the ground object categories in the changed areas. In addition, Figure 8(a)-(c) in it further confirm that CVS-Net has significant advantages in maintaining the integrity of changed features, mainly because CVS-Net integrates and strengthens edge information.

[0082] Table 1 Precision evaluation results of different SCD models on the SECOND dataset

[0083]

[0084] 2. Comparative analysis of experimental results on the FZ-SCD dataset

[0085] Table 2 lists the precision evaluation results of each method on the FZ-SCD dataset. Among them, CVS-Net in the solution of this case performs the best, and OA, mIoU, SeK and F SCD are 86.91%, 72.71%, 23.30% and 62.05% respectively. Figure 9 shows the result instances of each method on the FZ-SCD dataset. Through comparative analysis, it is found that CVS-Net shows higher precision in identifying changed semantic information, especially when distinguishing between easily confused feature categories, such as sparse grassland and bare land, as shown by the red box in Figure 9. In addition, CVS-Net can capture the changed area more completely, and the detection results are more consistent with the actual changed area, as Figure 9 shown by the purple box in

[0086] Table 2 Precision evaluation of SCD results of each method on the FZ-SCD dataset

[0087]

[0088] 3. Ablation experiment

[0089] 1). Verification of the effectiveness of each module

[0090] To prove the effectiveness of each module in CVS-Net, the improved ResNet34 is used as the basic SCD network in the experiment. Subsequently, the proposed twin CNN-VSS feature encoder, ST-VSSM and BAS are sequentially added to the ResNet34 basic network. To quantitatively evaluate the contribution of each module to the improvement of the model performance, the contribution index G is defined according to four evaluation indicators. First, calculate the ratio of the increase (or decrease) of each precision index after adding one module to that after adding all modules, and then take the average value of the increase (or decrease) ratio of each index. The calculation formula is:

[0091]

[0092] Among them, and They respectively represent the accuracy evaluation results of CVS-Net, the basic model, and the \(i\)-th one after adding a certain module.

[0093] As can be seen from Table 3, with the addition of each module, the recognition ability of the SCD model for the changed area and the accuracy of the changed semantic classification are gradually improved. SeK after adding all modules is improved by 3.65% compared with the basic model, while GTC is decreased by 3.40%. The contribution index \(G\) of ST-VSSM is the largest, which is 37.73%, indicating that the spatio-temporal relationship between the two-temporal features is effectively captured, and the spatio-temporal relationship is an effective factor for improving SCD. Figure 10 It shows the change detection results of each model in the ablation experiment. There are many missed detections in the detection results of the basic model, and the results are irregular and seriously fragmented. After adding the CNN-VSS feature extractor, thanks to the effective long-distance information modeling ability of the spatial state model, all accuracy indicators have been significantly improved, especially the problem of missed detection of large buildings has been effectively improved. However, the semantic classification accuracy is still poor, and this problem is solved after embedding ST-VSSM. The Sek value is improved by 1.52% compared with the model that only uses the CNN-VSS feature encoder. With the integration of the edge perception enhancement branch BAS, the shape of the changed ground objects is effectively constrained.

[0094] Table 3 Accuracy of SCD ablation experiment on the SECOND dataset

[0095]

[0096] 2) Effectiveness of the CNN-VSS feature encoder

[0097] The CNN-VSS feature encoder in the solution of this case is compared and analyzed with the encoders in the mainstream semantic segmentation networks to verify its effectiveness. Specifically, the encoder in the mainstream SCD network is used to replace the CNN-VSS feature encoder, and other modules of CVS-Net are kept unchanged, so as to construct four CVS-Net variant models, which are respectively:

[0098] (1) V1: Use the CNN-based ResNet34 network as the encoder of CVS-Net;

[0099] (2) V2: Use the encoder of the ChangeFormer model, which is composed of alternating multi-head self-attention layers and multi-layer perceptron modules.

[0100] (3) V3: Use the encoder in BOTNet, which adopts a hybrid CNN-Transformer architecture. In the early stage, CNN (ResNet34) is used for feature encoding, and a multi-head self-attention module based on Transformer is used in the last encoding stage;

[0101] (4) V4: Use VMamba in ChangeMamba as the encoder, which contains four encoding stages. The first encoding stage contains a Stem and a VSS module, and the other three encoding stages each contain a downsampling layer and a VSS module.

[0102] Table 4 shows the SCD results using different encoders on the public dataset SECOND. All metrics of the SCD results using the CNN-VSS encoder exceed those of the mainstream SCD methods based on CNN, Transformer, and Mamba. In addition, V3, which combines CNN and Transformer to construct a feature encoder, obtains the second-best accuracy, surpassing V1, V3, and V4 based on a single backbone network. This shows that when extracting remote sensing image features, combining backbone networks of multiple architectures to construct a feature extractor is an effective mechanism for characterizing the change features of remote sensing ground objects. In addition, the parameter quantity of the Mamba-based method is large, while the CNN-VSS feature encoder in the solution of this case not only improves the local and global feature extraction capabilities of images but also effectively controls the model training cost and computational consumption.

[0103] Table 4 Comparison of SCD result accuracies using different feature encoders

[0104]

[0105] 3) Effectiveness of the spatio-temporal modeler

[0106] To verify the effectiveness of the spatio-temporal modeler, the dual-temporal feature spatio-temporal modeling method in the solution of this case is compared with four other different spatio-temporal modeling methods, which are respectively:

[0107] (1) STM_1: Use the SS2D module for spatio-temporal modeling in all encoding stages;

[0108] (2) STM_2: Use the decoder of ChangeMamba for spatio-temporal relationship modeling;

[0109] (3) STM_3: Perform spatio-temporal relationship modeling in the decoder. After fusing the dual-temporal features in the channel dimension, they are fed into the SS2D for spatio-temporal relationship modeling.

[0110] As can be seen from Table 5, the spatio-temporal relationship modeling method in the solution of this case obtains the best SCD performance. Its SeK value has increased by 1.23%, 0.82% and 1.52% respectively compared with the spatio-temporal relationship modeling methods of STM_1, STM_2 and STM_3. The significant improvement of the Sek value strongly proves that the solution of this case has higher accuracy in identifying the types of ground objects in the changed area. This excellent performance is mainly due to the fact that the solution of this case models the bidirectional spatio-temporal relationship of the changed ground objects, thus greatly enhancing the model's ability to deeply explore the potential logic of ground object changes. Further analysis shows that compared with the accuracy performance of the spatio-temporal modeling technology not adopted, the accuracy of using the spatio-temporal modeling method has shown varying degrees of improvement. This comparison result strongly proves that the spatio-temporal relationship modeling is indeed an effective and key technology to improve the SCD accuracy. Among the comparison methods, STM_3 performs the worst, and the SCD results it obtains are significantly lower than those of other methods. This is mainly because it simply stitches the dual-temporal features in the channel dimension and fails to fully explore and utilize the deep information in the spatio-temporal relationship, resulting in insufficient detection accuracy.

[0111] Table 5 SCD accuracy evaluation results using different spatio-temporal modeling methods

[0112]

[0113] 4) Comparative analysis of change area segmentation errors

[0114] In the SCD method based on the multi-task architecture, accurately identifying the changed area is crucial for improving the overall SCD accuracy. To qualitatively verify the effectiveness of the edge information enhancement module in improving the segmentation accuracy of the changed area, the experiment analyzed the consistency degree between the changed areas detected by each method and the reference changed area from both quantitative and qualitative perspectives. The global segmentation error defines a consistency degree evaluation index. Existing SCD methods have a higher extraction accuracy for large and regular changed areas, but have insufficient extraction accuracy for changed areas with complex shapes. To focus on evaluating the segmentation performance of the SCD method for complex changed areas, a consistency degree evaluation index weighted based on the object shape complexity is used Specifically, on the basis of calculating the original global segmentation error, the calculation of the Boyce-Clark radius shape index for each object is added. The Boyce-Clark radius shape index (hereinafter referred to as SBC) measures the shape regularity of a polygon by calculating the ratio of the radius of the inscribed circle to the radius of the circumscribed circle of the polygon. The larger its value, the more complex the shape, and vice versa. The calculated SBC is normalized, and the reciprocal of the normalized SBC of each object is used as the weight of the global segmentation error. The specific calculation formula is:

[0115]

[0116] Where obj represents the changed object, w k is the GTC weight of the kth object, Normal represents the normalization operation, r i is the radius length from the centroid of an object to the edge of the object, and n is the number of radiation radii with equal angle difference, which is set to 32 in the experiment.

[0117] Table 6: The degree of agreement between the change regions extracted by each method and the reference change regions

[0118]

[0119] Table 6 lists the consistency of each method on the SECOND dataset. Figure 11 The degree of agreement between the change regions extracted by each method in the four test areas and the reference change regions is visualized. Figure 11 It can be seen that compared with other methods, the degree of consistency between the change area obtained by this solution and the actual change area Higher, and the edges of the changing areas are more accurate, especially in the changing areas with irregular shapes, such as Figure 11 In As shown. In contrast, HRSCD.str4, Bi-SRNet, ChangeMamba, ScanNet and TED all have different degrees of edge blur or incomplete area problems when detecting such irregular shaped changing areas. In urban environments, the detection of changing areas is more difficult due to the interference of complex elements such as buildings and roads. Areas B and C are both densely built areas in the city. The solution in this case can still accurately extract the changing areas, and there are fewer false detections and missed detections, which demonstrates its strong anti-interference ability. Other methods have more false detections and missed detections in this area, affecting the accuracy of the detection results. Area D has a narrow and long changing area. The solution in this case also performs well in this area. The changing area is complete and the edges are clear.

[0120] The experimental results on the SECOND and FZ-SCD datasets show that:

[0121] (1) Compared with mainstream SCD methods such as HRSCD.str4, Bi-SRNet, ChangeMamba, ScanNet, and TED, the proposed CVS-Net in this case demonstrated optimal detection performance on both the SECOND and FZ-SCD datasets, effectively improving the overall accuracy of semantic change detection. CVS-Net significantly reduced the missed detection cases, with mIoU reaching 72.89% and 72.60% respectively, and the accuracy of change semantic recognition was significantly improved, with the SeK index increasing by an average of 2.75% and 1.93% on the two datasets. These results strongly verified the effectiveness of the proposed solution in this case.

[0122] (2) Compared with the SCD results of feature extractors based on CNN, Transformer, and Mamba, the CNN-VSS feature extractor effectively improved the accuracy of change region detection and semantic category recognition. The mIoU increased by up to 1.68%, and the F scd increased by up to 1.71%. In addition, while enhancing the feature representation ability, CNN-VSS also took into account the model complexity, and its number of parameters was only 22.05 Mb. ST-SS2D can effectively model the bidirectional spatio-temporal relationship of changing ground objects, thereby enhancing the model's ability to mine the potential logic of ground object changes. Compared with the STM_1, STM_2, and STM_3 methods, its SeK value increased by 1.23%, 0.82%, and 1.52% respectively.

[0123] (3) The proposed solution in this case achieved the highest degree of consistency between the obtained change region and the actual change region. The consistency evaluation index C reached 92.97%, and the edge of the change region was more accurate, especially in the change region with irregular shapes.

[0124] Unless otherwise specifically stated, the relative steps, numerical expressions, and values of the components and steps set forth in these embodiments do not limit the scope of the present invention.

[0125] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the system disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0126] The units and method steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those of ordinary skill in the art can use different methods to implement the described functions for each specific application, but such implementation is not considered to exceed the scope of the present invention.

[0127] Those of ordinary skill in the art can understand that all or part of the steps in the above methods can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium, such as a read-only memory, a magnetic disk, or an optical disc, etc. Optionally, all or part of the steps of the above embodiments can also be implemented using one or more integrated circuits. Correspondingly, each module / unit in the above embodiments can be implemented in the form of hardware or in the form of a software functional module. The present invention is not limited to any specific form of the combination of hardware and software.

[0128] Finally, it should be noted that the above-described embodiments are only specific embodiments of the present invention, used to illustrate the technical solutions of the present invention, rather than limiting them. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions described in the foregoing embodiments or easily conceive of changes, or make equivalent replacements for some of the technical features; and these modifications, changes, or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A semantic change detection method for remote sensing images based on spatio-temporal relationship modeling and edge information enhancement, characterized in that, It includes: obtaining the pre-temporal remote sensing image and the post-temporal remote sensing image of the target area, and preprocessing the dual-temporal remote sensing images; Inputting the pre-temporal remote sensing image and the post-temporal remote sensing image of the target area into a pre-trained remote sensing image semantic change detection model, and using the remote sensing image semantic change detection model to obtain the semantic change detection result of the dual-temporal remote sensing image of the target area; Among them, the remote sensing image semantic change detection model fuses a convolutional neural network and a visual state space model VSSM to construct a twin feature encoder to extract the semantic features of the remote sensing image, mines the spatio-temporal change logical relationship between the dual-temporal remote sensing images based on the bidirectional spatio-temporal relationship modeler of the visual state space model VSSM, and uses an edge-aware reinforcement network to identify the edges of the changed objects in the dual-temporal remote sensing images.

2. The semantic change detection method for remote sensing images based on spatio-temporal relationship modeling and edge information enhancement according to claim 1, characterized in that The remote sensing image semantic change detection model includes a twin feature encoder for extracting the semantic features of the dual-temporal remote sensing images, an edge-aware reinforcement network for perceiving the edge information of the dual-temporal remote sensing images, a spatio-temporal relationship modeler for reshaping the semantic features of the remote sensing images, a classification decoder for classifying the land cover of the reshaped semantic features of the remote sensing images, a change area decoder for obtaining the changed area of the dual-temporal remote sensing images by using the semantic features and edge information of the dual-temporal remote sensing images, and an output end for obtaining a semantic change detection map of the dual-temporal remote sensing images by using the land cover classification result and the changed area of the dual-temporal remote sensing images. The spatio-temporal relationship modeler is a dual-temporal spatio-temporal relationship modeler constructed based on the visual state space model VSSM for mining the spatio-temporal connection between the dual-temporal features and performing feature reshaping.

3. The semantic change detection method for remote sensing images based on spatio-temporal relationship modeling and edge information enhancement according to claim 1 or 2, characterized in that, The twin feature encoder is composed of a ResNet34 network and a visual state space model VSSM. Among them, ResNet34 is used to extract the multi-scale semantic features of the input remote sensing image by using several convolutional kernels, and the visual state space model VSSM is embedded in the convolutional kernels of the last two layers respectively to obtain the global context semantic information of the remote sensing image.

4. The semantic change detection method for remote sensing images based on spatio-temporal relationship modeling and edge information enhancement according to claim 3, characterized in that The visual state space model VSSM includes a selective scan branch and an MLP branch, and uses a normalization layer to normalize the output of the selective scan branch. The normalized semantic features are dot-multiplied with the semantic features output by the MLP branch, and then concatenated with the semantic features of the input remote sensing image in the channel dimension to obtain the final output semantic features of the twin feature encoder. The selective scan branch is used to perform pointwise convolution and non-linear activation processing on the semantic features of the input remote sensing image by using a depthwise separable convolutional layer and a SiLU activation function, and expand the semantic features of the remote sensing image through a two-dimensional selective scan mechanism SS2D to establish a global receptive field. The MLP branch is used to perform a linear operation on the semantic features of the input remote sensing image.

5. The semantic change detection method for remote sensing images based on spatio-temporal relationship modeling and edge information enhancement according to claim 1 or 2, characterized in that, The two-way spatio-temporal relationship modeling device performs feature scanning on the semantic features of the remote sensing image by using two scanning methods, namely, dual-temporal cross-scanning and spatial-first and temporal-second scanning, and cross-feeds the scanned features to the visual state space model VSSM. The visual state space model VSSM selectively memorizes the historical information of the semantic features and performs feature reshaping to mine the spatio-temporal relationship between the dual-temporal features. Among them, each scanning method performs feature scanning in two scanning orders, namely, the forward order from before change to after change and the reverse order from after change to before change.

6. The semantic change detection method for remote sensing images based on spatio-temporal relationship modeling and edge information enhancement according to claim 1 or 2, characterized in that The edge-aware enhancement network uses an edge detection and edge feature enhancement mechanism to perceive the edges of the changed objects in the dual-temporal remote sensing image. Among them, the edge detection mechanism uses a Gaussian smoothing filter to smooth the remote sensing image and uses a Laplace operator to extract the initial edge features of the remote sensing image; the edge feature enhancement mechanism uses the initial edge features to obtain the edges of the changed objects in the shallow features extracted by the siamese feature encoder and enhances the edges of the changed objects by fusing the RGB and edge features through a gated fusion unit.

7. The semantic change detection method for remote sensing images based on spatio-temporal relationship modeling and edge information enhancement according to claim 1, characterized in that The training process of the remote sensing image semantic change detection model includes: Collecting remote sensing image semantic change detection sample data and performing enhancement processing on the sample data to obtain a remote sensing image semantic change detection sample data set; Constructing a joint loss function for model training based on the cross-entropy loss of semantic segmentation of each temporal remote sensing image, the detection loss of the changed area of the remote sensing image, and the semantic consistency loss of the remote sensing image change; Based on the joint loss function and using the remote sensing image semantic change detection sample data set to train and optimize the remote sensing image semantic change detection model.

8. A remote sensing image semantic change detection system based on spatio-temporal relationship modeling and edge information enhancement, characterized in that, Including: an image acquisition module and a change detection module, where The image acquisition module is used to acquire the pre-temporal remote sensing image and the post-temporal remote sensing image of the target area and perform preprocessing on the dual-temporal remote sensing images before and after; The change detection module is used to input the pre-temporal remote sensing image and the post-temporal remote sensing image of the target area into the pre-trained remote sensing image semantic change detection model, and use the remote sensing image semantic change detection model to obtain the semantic change detection result of the dual-temporal remote sensing image of the target area; Among them, the remote sensing image semantic change detection model fuses a convolutional neural network and a visual state space model VSSM to construct a siamese feature encoder to extract the semantic features of the remote sensing image, mines the spatio-temporal change logic relationship between the dual-temporal remote sensing images based on the two-way spatio-temporal relationship modeling device of the visual state space model VSSM, and uses the edge-aware enhancement network to identify the edges of the changed objects in the dual-temporal remote sensing image.

9. An electronic device, characterized in that, It includes: At least one processor, and a memory coupled to the at least one processor; Among them, the memory stores a computer program, and the computer program can be executed by the at least one processor to implement the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, and when the computer program is executed, it can implement the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Remote sensing image semantic change detection method based on time sequence remote sensing Mama

    CN120580599A

  • Multi-source remote sensing image real-time seamless splicing method based on adaptive generative adversarial network

    CN121921191A