Multi-task remote sensing semantic change detection method and system based on visual state space model
By introducing the feature sharing mechanism and cross-scan mechanism of visual state space model in remote sensing image processing, the correlation problem of semantic segmentation and change detection in multi-task remote sensing semantic change detection is solved, and the detection accuracy and consistency are improved.
Patent Information
- Application Number
- CN202510340373.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-29
AI Technical Summary
The existing multitasking remote sensing semantic change detection methods lack relevance in semantic segmentation and change detection subtasks, and cannot effectively utilize the time and spatial information of the image, resulting in inconsistent output results or errors.
Using a method based on visual state space model, the feature sharing mechanism is introduced and the cross-scanning mechanism is improved in the feature extraction encoder and the change detection decoder, and the correlation between semantic feature information and semantic auxiliary information is enhanced.
The utilization rate of image time and spatial information is improved, the accuracy and consistency of semantic segmentation and change detection are enhanced, and the integrity and rationality of the boundary contour of the changing area are ensured.
Smart Images

Figure CN120388300A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and further relates to a multi-task remote sensing semantic change detection method, which can be used for urban planning and management, ecosystem detection and disaster assessment. Background Art
[0002] Remote sensing technology is an effective means to perceive the surface environment, obtain the distribution of natural resources, and analyze land use. The surface environment in which humans live is changing all the time. Accurately mastering the land cover type and its changes is of great significance for national territorial space planning, ecological environment monitoring, disaster assessment, etc.
[0003] Remote sensing change detection is a technology that uses two or more remote sensing images acquired at different times to detect surface changes occurring in the same geographical area. Using high-resolution multi-temporal optical remote sensing images to detect surface change elements is a commonly used change detection method. With the continuous development of remote sensing technology, remote sensing image change detection has become a popular research field in remote sensing technology, playing an important role in fields such as urban planning, land cover analysis, disaster assessment, ecosystem monitoring, and resource management.
[0004] At present, the research on remote sensing change detection methods at home and abroad mainly focuses on binary change detection methods. The detection process of traditional binary change detection methods is to input a pair of multi-temporal remote sensing images, and through image preprocessing, feature extraction, change information extraction, and change result output, a binary change prediction map is obtained. This map describes the binary change types at each pixel in the image, and divides specific geographical elements in the image into change regions and non-change regions. The current binary change detection methods mainly have the following problems: ① Binary change detection focuses on detecting the positions of changed pixels between multi-temporal images, without considering the categories of changed pixels, and cannot depict the semantic change information required in subsequent applications; ② Conventional binary change detection methods mainly focus on buildings and production land, while in actual applications, the changes of multiple types of land elements in the image scene are usually concerned. Single-class change detection methods cannot meet the scene requirements of multi-class change detection, and the generality is insufficient.
[0005] Remote sensing semantic change detection is based on conventional binary change detection. It processes the semantic information in the extracted features, identifies and distinguishes various land cover types in multi-temporal images, obtains the semantic segmentation results of the corresponding images, and then combines with the binary change image obtained from change detection for joint processing to get the change situation of specific land types. Currently, the commonly used strategy for various change detection methods is the post-classification detection method, that is, multi-temporal remote sensing images are separately divided into semantic regions, and then the semantic segmentation results are directly compared to identify change types. However, this strategy makes the semantic types in multi-temporal images independent, ignoring the spatio-temporal correlation between multi-temporal images. If there is a deviation in the classification of one temporal phase, it will lead to incorrect change results.
[0006] In the research on remote sensing semantic change detection methods, Zhao S et al. introduced the Mamba architecture based on the state space model into the dense prediction task of high-resolution remote sensing images in their published paper "RS-Mamba for Large Remote Sensing Image Dense Prediction", proposed the RS-Mamba method that can be used for semantic segmentation and change detection, designed an omni-directional selective scanning module to selectively scan the remote sensing image in multiple directions, so as to extract spatial features in multiple directions. This method can efficiently process large-size remote sensing images and has good global modeling ability, but there are still two deficiencies: one is that only by modifying the scanning method of Mamba to enhance the global receptive field, it ignores the role of local information in the dense prediction task and does not make good use of the local information in the image. The other is that the semantic segmentation and change detection tasks are relatively independent and cannot complete the multi-task semantic change detection task simultaneously.
[0007] Niu Y et al. proposed a new symmetric multi-task network SMNet that fuses global and local information in their published paper "SMNet: Symmetric Multi-Task Network for Semantic Change Detection in Remote Sensing Images Based on CNN and Transformer". Based on convolutional neural network and Transformer, it uses a hybrid unit composed of a pre-activated residual block PR and a transformation block TB to construct a pre-activated residual change backbone PRTB, so as to obtain richer semantic features with local and global information from bi-temporal images, and optimizes the model through multi-task prediction branches and a custom loss function to improve the accuracy of semantic change detection. However, due to the lack of correlation between the semantic detection decoder and the change detection decoder, there are differences in the obtained semantic segmentation results and change detection results, and it cannot well reflect semantic changes.
[0008] Cui F et al. proposed a multi-task semantic change detection method MTSCD-Net based on Swin Transformer in their published paper "MTSCD-Net: A network based on multi-task learning for semantic change detection of bitemporal remote sensing images". It uses a Siamese semantic-aware encoder to extract multi-scale features, and an aggregation module to combine the features. Then, a change information extraction module is designed to enhance the feature expression ability by fully fusing two-level difference features. In the decoder stage, the spatial attention weight map is obtained using the features of the change detection sub-task, providing position prior information for the features of the semantic segmentation sub-task, enhancing the correlation between the two sub-tasks, and improving the detection quality. Since this method mainly focuses on processing the spatial change information between images and ignores the impact of temporal change information on change detection, there are some pseudo-changes in the change detection results, and the reduction of the boundary contour of the change results is not good.
[0009] In summary, the existing multi-task remote sensing semantic change detection methods either lack relevance between the semantic segmentation and change detection sub-tasks, the sub-tasks cannot guide each other, and the output results lack consistency; or they do not fully utilize the spatio-temporal change information in the change detection task and ignore the impact of temporal changes between images, resulting in errors in the output results. Summary of the Invention
[0010] The purpose of the present invention is to provide a multi-task remote sensing semantic change detection method based on a visual state space model for the above-mentioned deficiencies of the existing technologies, so as to improve the utilization rate of image time and space information, enhance the relevance between the semantic segmentation and change detection sub-tasks, and improve the detection accuracy.
[0011] The technical idea to achieve the purpose of the present invention is to improve the utilization rate of image time and space information by introducing a feature sharing mechanism and improving the cross-scanning mechanism in the feature extraction encoder and the change detection decoder; and enhance the relevance between the semantic segmentation and change detection tasks by fusing change feature information and semantic auxiliary information.
[0012] According to the above idea, the implementation scheme of the present invention includes the following:
[0013] 1. A multi-task remote sensing semantic change detection method based on a visual state space model, characterized by comprising:
[0014] 1) Design two feature extraction encoders, two semantic decoders, and a change detection decoder according to the visual state space model;
[0015] 2) Obtain dual-temporal high-resolution remote sensing images and use two feature extraction encoders to extract their semantic features, thereby obtaining two-way feature-enhanced semantic feature information;
[0016] 3) The semantic feature information of the two feature enhancements is input into two semantic decoders respectively. The semantic segmentation map of the corresponding time phase is obtained through four semantic information extraction stages, and the semantic information of different stages is obtained from the semantic decoders as semantic auxiliary information.
[0017] 4) The semantic feature information and semantic auxiliary information are imported into the change detection decoder, and a binary change map is obtained through four change information extraction stages;
[0018] 5) Using the binary change map as a mask, the semantic segmentation map is cropped to remove the non-changing area to obtain the final semantic change detection result.
[0019] Furthermore, two feature extraction encoders are designed based on the visual state space model, and their structures are the same, both including an image input layer, an image block layer and four visual state space layers with the same structure connected in sequence, wherein:
[0020] The image input layer is used to adjust the input three-channel RGB image to a resolution of 256×256 and output it to the image blocking layer;
[0021] The image segmentation layer is used to segment the resized image into 64 non-overlapping image blocks of 32×32 pixels, record the original position information of each image block, and output the segmented image group to the visual state space layer;
[0022] The four visual state space layers are used to extract key features from the divided image group. They are connected in series through a downsampling layer. The downsampling layer downsamples the feature image processed by the previous visual state space layer, adjusts the image size to half of the original size, and outputs it to the next visual state space layer for use. Through layer-by-layer extraction and dynamic downsampling, a deep image semantic feature map is obtained as the result output of the feature extraction encoder.
[0023] Furthermore, the change detection decoder designed according to the visual state space model uses the bi-temporal feature maps obtained by processing the two feature extraction encoders as input, and processes the bi-temporal feature maps through four different processing stages, wherein:
[0024] The first processing stage is used to complete spatio-temporal state space processing and the upsampling layer. That is, first divide the dual-temporal feature map into 8×8 image patches, and perform cross-scanning in four different directions: horizontally to the right, horizontally to the left, vertically downwards, and vertically upwards to obtain four groups of one-dimensional sequences to represent the spatio-temporal change information in the image; then perform image resolution expansion processing on the one-dimensional sequences through upsampling to obtain the feature map of the first stage;
[0025] The second processing stage is used to extract spatio-temporal features from the feature map of the first stage. Through re-partitioning and scanning processing, deeper spatio-temporal information is captured. The spatio-temporal features extracted from different processing layers are fused to obtain a fused feature map, and image resolution expansion processing is performed on the fused feature map, and then the second-stage feature map is obtained through upsampling;
[0026] The third processing stage is used to extract deeper spatio-temporal features from the feature map of the second stage, fuse these features, and then obtain the feature map of the third stage through upsampling;
[0027] The fourth processing stage further extracts deeper spatio-temporal features from the feature map of the third stage, further fuses these features, and finally obtains a high-resolution and accurate change feature map through upsampling after multiple processes and fusions. Its output is the result of the change detection decoder.
[0028] 2. A multi-task remote sensing semantic change detection device based on a visual state space model, comprising:
[0029] Feature extraction module 1 is used to extract semantic feature information from dual-temporal high-resolution optical remote sensing images, and input the extracted dual-temporal semantic feature information to the semantic segmentation module and the change detection module for processing;
[0030] Semantic segmentation module 2 is used to extract semantic category information of various ground objects in the remote sensing image from the dual-temporal semantic feature information, and restore it to a semantic segmentation map of the original image size as the result of semantic segmentation for output;
[0031] Change detection module 3 is used to extract change information of dual-temporal remote sensing images from the dual-temporal semantic feature information, restore it to a binary change map of the original image size, and then apply the binary change map as a mask to the semantic segmentation map to obtain a land change type map of the change area as the result of semantic change detection for output.
[0032] Compared with the prior art, the present invention has the following advantages:
[0033] First, in the feature extraction encoder and the change detection encoder designed based on the visual state space model of the present invention, by using the global-local dual-branch design to obtain multi-scale feature information of the image, and through the feature sharing mechanism to fuse the global and local feature information, the model can more effectively combine these two types of features, improving the ability to extract spatio-temporal information.
[0034] Second, in the feature extraction encoder and the change detection encoder designed based on the visual state space model of the present invention, the cross-scanning mechanism is improved, so as to establish a spatio-temporal change relationship between two images at different time phases, fully excavate the changes in the time dimension and the structural changes in the space of the image, thereby improving the detection efficiency and accuracy of change detection.
[0035] Third, in the present invention, since semantic feature information and semantic auxiliary information are introduced into the change detection decoder, the semantic segmentation result and the change detection feature information are fused, and the category information in the semantic segmentation result is used to guide the feature extraction of change detection from the perspective of image semantic categories, ensuring the integrity and rationality of the boundary contours of the changed regions, enhancing the relevance between the semantic segmentation and the change detection subtasks, and improving the accuracy of change detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 is the implementation flowchart of the multi-task remote sensing semantic change detection method provided in Embodiment 1 of the present invention;
[0037] Figure 2 is the structural diagram of the multi-task remote sensing semantic change detection model designed in the method of the present invention;
[0038] Figure 3 is Figure 2 the structural diagram of the feature extraction encoder designed according to the visual state space model in
[0039] Figure 4 is Figure 2 the structural diagram of the semantic segmentation decoder designed according to the visual state space model in
[0040] Figure 5 is Figure 2 the structural diagram of the change detection encoder designed according to the visual state space model in
[0041] Figure 6 is the structural block diagram of the multi-task remote sensing semantic change detection device provided in Embodiment 2 of the present invention;
[0042] Figure 7 is the experimental result diagram of Simulation Experiment 1 in Embodiment 2 of the present invention;
[0043] Figure 8It is the experimental result graph of Simulation Experiment 2 in Embodiment 2 of the present invention. Specific Embodiment
[0044] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0045] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, other embodiments obtained by those of ordinary skill in the art without creative work shall fall within the protection scope of the present invention.
[0046] It should be noted that the step numbers in the specification and claims of the present invention are only for clear description of the implementation solutions of the present invention and for easy understanding, and their sequence numbers are not limited.
[0047] Embodiment 1: Multitask Remote Sensing Semantic Change Detection Method Based on Visual State Space Model
[0048] Refer to Figure 1 , and the implementation steps of the embodiment are as follows.
[0049] Step 1, obtain remote sensing data and divide it into a training set and a test set.
[0050] 1.1) Obtain high-resolution remote sensing optical images of the same area at different time nodes, such as several months or several years apart, from satellite data sources, ensuring that the spectral bands, spatial resolutions, and imaging conditions of the images are as consistent as possible to reduce noise and errors;
[0051] 1.2) Precisely align the two images through georegistration technology to ensure that the same geographical coordinates correspond to the same pixel positions in the two images;
[0052] 1.3) Select the area of interest in the image, and intercept sub-images with a size of 1024×1024 pixels from it to form paired image pairs, and label them with specific land cover types to constitute a remote sensing semantic change detection data set;
[0053] 1.4) Divide the data set into a training set and a test set according to a ratio of 7:3.
[0054] Step 2, construct a semantic change detection model.
[0055] Design two feature extraction encoders, two semantic segmentation decoders, and one change detection decoder according to the visual state space model to construct a multitask remote sensing semantic change detection model.
[0056] Refer toFigure 2 , the implementation of this step includes the following:
[0057] 2.1) Refer to Figure 3 , design two feature extraction encoders with the same structure. Each feature extraction encoder is composed of an image input layer, an image block division layer, and four visual state space layers connected in sequence. Among them:
[0058] The image input layer is mainly used to receive and process the input three-channel RGB image, and adjust the input three-channel RGB image to a resolution of 256×256;
[0059] The image block division layer is used to evenly divide the adjusted image into an 8×8 grid according to the size of 32×32 pixels, obtaining 64 non-overlapping image blocks. During the block division process, record the position information of each image block in the original image so that the source of each image block can be accurately located during subsequent processing;
[0060] The visual state space layer is used to extract semantic feature information in the image, and its implementation is as follows:
[0061] First, scan the grouped images after block division in four different directions: horizontally to the right, horizontally to the left, vertically downwards, and vertically upwards, and flatten them into the following four groups of one-dimensional sequences:
[0062] S1 = [F(0,0), F(0,1), F(0,2),..., F(0,7), F(1,0),..., F(7,7)]
[0063] S2 = [F(0,7), F(0,6), F(0,5),..., F(0,0), F(1,7),..., F(7,0)]
[0064] S3 = [F(0,0), F(1,0), F(2,0),..., F(7,0), F(0,1),..., F(7,7)]
[0065] S4 = [F(7,0), F(6,0), F(5,0),..., F(0,0), F(7,1),..., F(0,7)]
[0066] Among them, F is the input image group, each group has 8×8 image blocks, and the coordinate range of each group of images is from 0 to 7. S1 is the one-dimensional sequence scanned horizontally to the right, S2 is the one-dimensional sequence scanned horizontally to the left, S3 is the one-dimensional sequence scanned vertically downwards, and S4 is the one-dimensional sequence scanned vertically upwards;
[0067] Next, for each one-dimensional sequence obtained by scanning, set the initial state variable h0 to a zero vector, and update the system state h at each position in order using the state space equation for each position in the sequence i , and calculate the output y at the current position i :
[0068]
[0069] where x i ∈R L represents the i-th element in the input sequence, h i ∈R N is the state variable at the current i-th position, h i-1 ∈R N is the state variable at the previous position i - 1, y i ∈R L represents the system output at the current i-th position, i = 1, 2, …, 64. are the state matrix, input matrix, output matrix, and direct transmission matrix of the system respectively. These matrices together define the dynamic behavior of the system. L is the sequence length, and N is the state space size.
[0070] Next, take the system output y at each position in the four groups of sequences i as the weight coefficient, and arrange them in order to obtain four groups of one-dimensional weight sequences. Then arrange these four groups of one-dimensional weight sequences in the original order, restore them to a two-dimensional weight image, and integrate them by superposition to obtain the weight matrix W n×n :
[0071]
[0072] where Q k is the two-dimensional weight image restored from the one-dimensional weight sequence. The range of k is from 1 to 4, representing the first to fourth groups of one-dimensional weight sequences respectively. n×n represents the image size, and the range of n is from 16 to 128;
[0073] Finally, add the weight matrix W n×n to the original image group F n×n input to the current visual state space layer to obtain the output O n×n of the visual state space layer:
[0074] O n×n = f(F n×n + W n×n + b)
[0075] where F n×n is the original image group, W n×nis the input two-dimensional weight matrix, b is the bias term, f is the activation function, and the output is O n×n is a two-dimensional feature map with the same size as the original image group;
[0076] The four visual state space layers are connected in series through a downsampling layer. The downsampling layer is used to adjust the size of the feature image and downsample the image to half of its original size. These four visual state space layers sequentially extract multi-level semantic features from shallow to deep in the image and output the final semantic feature map. At the same time, the first-stage intermediate feature map, the second-stage intermediate feature map, and the third-stage intermediate feature map are sequentially output through the first three visual state space layers for use in the corresponding level feature fusion processing in the subsequent semantic segmentation decoder.
[0077] 2.2) Refer to Figure 4 , design two semantic segmentation decoders with the same structure. Each semantic segmentation decoder can be divided into four stages hierarchically. The first three stages have the same structure and are sequentially composed of a visual state space layer, an upsampling layer, and a feature fusion layer. The last stage is sequentially composed of a visual state space layer and an upsampling layer, where:
[0078] The visual state space layer has the same structure as the visual state space layer in the above feature extraction decoder and is used to model the global spatial context information of the input feature map and reconstruct high-resolution semantic category information from deep semantic information;
[0079] The upsampling layer upsamples the input image by using bilinear interpolation, calculates the weighted average of adjacent four pixels to estimate the value of the new pixel, and expands the image size to 2 times its original size to reconstruct high-resolution output;
[0080] The feature fusion layer is used to integrate feature maps of different levels at multiple scales, and its implementation is as follows:
[0081] For the feature map F1 after upsampling in the first stage of the semantic segmentation encoder, add it to the third-stage intermediate feature map F mid3 in the feature extraction encoder, and then perform fusion processing to obtain the fusion result F out1 :
[0082] F out1 = F merge1 + N(f(W (3) * F merge1 + b2))
[0083] where, F merge1 = F mid3 + f(W (1) * F1 + b1) is the feature map after adding and combining F1 and F mid3 , W(1) is a 1×1 convolution kernel, W (3) is a 3×3 convolution kernel, b3 and b4 are two different bias terms for convolution, f is an activation function, N is a normalization process, F out1 is the feature map that fuses the key information in F merge1 in
[0084] For the feature map F2 after upsampling in the second stage of the semantic segmentation encoder, add it to the intermediate feature map F mid2 in the second stage of the feature extraction encoder, and then perform a fusion process to obtain the fusion result F out2 :
[0085] F out2 = F merge2 + N(f(W (3) * F merge2 + b4))
[0086] where F merge2 = F mid2 + f(W (1) * F2 + b3) is the feature map after adding and combining F2 and F mid2 in which W (1) is a 1×1 convolution kernel, W (3) is a 3×3 convolution kernel, b3 and b4 are two different bias terms for convolution, f is an activation function, N is a normalization process, F out2 is the feature map that fuses the key information in F merge2 in
[0087] For the feature map F3 after upsampling in the third stage of the semantic segmentation encoder, add it to the intermediate feature map F mid1 in the first stage of the feature extraction encoder to get F merge3 , and then perform a fusion process to obtain the fusion result F out3 :
[0088] F out3 = F merge3 + N(f(W (3) * F merge3 + b6))
[0089] where F merge3 = F mid1 + f(W (1) * F3 + b5) is the feature map after adding and combining F3 and F mid3 in which W (1) is a 1×1 convolution kernel, W (3) is a 3×3 convolution kernel, b5 and b6 are two different bias terms for convolution, f is an activation function, N is a normalization process, F out3 is the feature map that fuses the key information in F merge3Feature maps of key information
[0090] Take the fused feature maps F out1 、F out2 、F out3 as inputs and pass them to the second, third, and fourth semantic segmentation stages for processing respectively. At the same time, take these three fused feature maps as semantic auxiliary information and record them as the first, second, and third layer auxiliary semantic feature maps in sequence for the subsequent change detection decoder to process and use.
[0091] 2.3) Refer to Figure 5 and design a change detection decoder that can be divided into four stages hierarchically. The first stage of it is composed of a visual state space layer and an upsampling layer connected in sequence. The structures of the latter three stages are the same, and each is composed of a visual state space layer, a feature fusion layer, and an upsampling layer connected in sequence, where:
[0092] The visual state space layer is used to model the spatio-temporal change relationship of the input dual-temporal feature maps and reconstruct high-resolution change information from deep semantic information. First, it integrates the input dual-temporal feature maps by channel juxtaposition, and then processes the integrated feature maps with the same structure as the visual state space layer in the feature extraction encoder:
[0093] For the visual state space layer in the first stage of the change detection decoder, use the semantic feature maps output by two feature extraction encoders as inputs to process and obtain change feature maps;
[0094] For the visual state space layers in the latter three stages of the change detection decoder, use the first, second, and third layer auxiliary semantic feature maps provided by two semantic segmentation decoders as inputs respectively to process and obtain the intermediate change feature maps corresponding to the stages;
[0095] The upsampling layer has the same structure as the upsampling layer in the semantic segmentation encoder and is used to double the image size to reconstruct high-resolution outputs;
[0096] The feature fusion layer is used to integrate multi-level image change information and fuse the change feature maps processed in the previous stage and the intermediate change feature maps in the current stage. Its structure is the same as the feature fusion layer in the semantic segmentation decoder.
[0097] 2.4) After connecting two feature extraction encoders with two semantic segmentation decoders in one-to-one correspondence, then connect these two feature extraction encoders together with a change detection encoder to jointly form a complete remote sensing semantic change detection model.
[0098] Step 3, train the remote sensing semantic change detection model.
[0099] 3.1) Design the loss function of the remote sensing semantic change detection model according to the characteristics of remote sensing semantic change data:
[0100] For the multi-class change regions in the image, based on the original cross-entropy loss function, according to the proportion of samples of different land cover types, different weight coefficients are assigned to different classes. The weight coefficient used is inversely proportional to the proportion of the samples, and the weights are normalized. In this way, the weighted multi-class cross-entropy loss function L is set. wce :
[0101]
[0102] where N is the number of land cover classes, y s is the true label of this class, p s is the predicted probability of the model for this class, and w s is the weight coefficient for this class;
[0103] For the unchanged regions in the image, the binary cross-entropy loss function L bce is used to measure the difference between the true label and the predicted change:
[0104] L bce = -[y b log p b + (1 - y b ) log(1 - p b )]
[0105] where y s is the true label of the unchanged class, and p s is the predicted probability of the model for the unchanged class;
[0106] To balance the classification accuracy and segmentation performance of the model, the Dice loss function L Dice is used to evaluate the regional similarity of the segmentation results:
[0107]
[0108] where P is the predicted set of the segmentation class, Y is the true label set of the segmentation class, and |P ∩ Y| represents the intersection of P and Y;
[0109] The above loss functions are jointly used to form the loss function Loss of the designed remote sensing semantic change detection model:
[0110] Loss = L wce + L bce + L Dice ;
[0111] 3.2) Input the training set into the model for training, set the initial learning rate η0 = 0.1, and the maximum number of iterations T = 100;
[0112] 3.3) Dynamically adjust the learning rate using the exponential decay method:
[0113] Each time the model is trained, it will process 8 training samples and obtain the corresponding predicted output. Calculate the loss function based on the predicted output and the true labels of the training set, then use the backpropagation algorithm to calculate the gradient of the loss function with respect to the model parameters, and update the model parameters according to the calculated gradient using the stochastic gradient descent optimizer to obtain the updated model parameters;
[0114] After the model has made a complete pass through the entire training set, that is, after one T, input the validation set into the model with updated parameters to obtain the predicted output of the model. Calculate the loss function, overall prediction accuracy, F1 score, intersection over union, Kappa coefficient and other performance metrics based on the predicted output and the true labels as the performance of the model at the current T stage;
[0115] According to the performance, if the model performance no longer improves, adjust the learning rate η at the current stage through exponential decay t :
[0116]
[0117] where η0 is the initial learning rate, t represents the current training stage, γ is the decay rate used to control the speed of learning rate decay, and k is the decay step used to control the frequency of learning rate decay;
[0118] 3.4) Repeat step 3.3) until the loss function converges or the maximum number of iterations is reached to obtain the trained remote sensing semantic change detection model.
[0119] Step 4, use the trained remote sensing semantic change detection model to perform multi-task remote sensing semantic change detection and output the semantic change results.
[0120] 4.1) Semantic feature extraction:
[0121] Use the pair of double-temporal raw images with 1024×1024 pixels obtained in step 1 as the model input, and input them into the two feature extraction encoders of the trained remote sensing semantic change detection model respectively. Sequentially extract the semantic features of the images from shallow to deep through the four feature extraction layers in the feature extraction encoder, and output the final deep semantic feature map through the fourth module for input to the subsequent semantic segmentation decoder and change detection decoder. Output the first intermediate feature map, the second intermediate feature map, and the third intermediate feature map level by level from the first three feature extraction modules for participating in the feature fusion processing of the corresponding levels in the subsequent semantic segmentation decoder;
[0122] 4.2) Input the semantic feature maps of the two time phases extracted in step 4.1) into the semantic segmentation decoder corresponding to the trained remote sensing semantic change detection model respectively. Through the processing of the four semantic extraction stages of the semantic segmentation decoder, gradually restore the corresponding semantic segmentation map S n×n , and its implementation includes:
[0123] In the first semantic information extraction stage, input the input feature map through the visual state space layer to learn the image semantic information, and then through upsampling, restore the feature map size to the same as the third intermediate feature map in step 4.1), and then input them together into the feature fusion layer of the encoder to restore the first semantic information map of the image and send it to the second stage;
[0124] In the second semantic information extraction stage, after processing the first semantic information map through the visual state space layer and the upsampling layer, restore it to the same size as the second intermediate feature map in step 4.1), and input them together into the feature fusion layer of the encoder for processing, restore the second semantic information map of the image, and send it to the third stage;
[0125] In the third semantic information extraction stage, after processing the second semantic information map through the visual state space layer and the upsampling layer, restore it to the same size as the first intermediate feature map in step 4.1), and input them together into the feature fusion layer of the encoder for processing, restore the third semantic information map of the image, and send it to the fourth stage;
[0126] In the fourth semantic information extraction stage, process the third semantic information map through the visual state space layer to restore it to land cover type information, and then restore it to a semantic segmentation map with the same size as the original image through the upsampling layer, which is the output result of the semantic decoder. At the same time, input the semantic information maps obtained in the first three stages as semantic auxiliary information into the corresponding change information extraction stage.
[0127] 4.3) Input the semantic feature maps of the dual-time phases extracted in step 4.1) into the change detection decoder. At the same time, with the help of the semantic auxiliary information processed by the semantic decoder in step 4.2), gradually restore the binary change map M in the dual-time phase images through four change information extraction stages n×n , and the implementation is as follows:
[0128] In the first change information extraction stage, take the dual-time phase semantic feature maps as the input, extract the change information of the image through the visual state space layer, and then obtain the first-stage change feature map through upsampling, and restore its size to the same as the first semantic auxiliary feature information map obtained in step 4.2), and send it to the second stage;
[0129] In the second change information extraction stage, the change feature map of the first stage is used as the input, and the two-phase first semantic auxiliary information maps obtained in step 4.2) learn the spatio-temporal change relationship through the visual state space layer, and then Figure 1 integrate the deep semantic information and the shallow spatial information through the image fusion layer, and finally obtain the change feature map of the second stage through upsampling, and restore its size to the same as the second semantic auxiliary feature information map obtained in step 4.2), and send it to the third stage;
[0130] In the third change information extraction stage, the change feature map of the second stage is used as the input, and the two-phase second semantic auxiliary information maps obtained in step 4.2) learn the spatio-temporal change relationship through the visual state space layer, and then Figure 1 integrate the deep semantic information and the shallow spatial information through the image fusion layer, and finally obtain the change feature map of the third stage through upsampling, and restore its size to the same as the third semantic auxiliary feature information map obtained in step 4.2), and send it to the fourth stage;
[0131] In the fourth change information extraction stage, the change feature map of the third stage is used as the input, and the two-phase third semantic auxiliary information maps obtained in step 4.2) learn the spatio-temporal change relationship through the visual state space layer, and then Figure 1 integrate the deep semantic information and the shallow spatial information through the image fusion layer, and finally obtain a binary change map with the same size as the original image through upsampling, which is the output of the change detection decoder.
[0132] 4.4) Combine semantic segmentation and change detection results to output semantic change results:
[0133] Use the semantic segmentation map S n×n obtained in step 4.2) n×n and the binary change map M obtained in step 4.3) n×n to perform mask screening by pixel-by-pixel multiplication, that is, retain the semantic classification results of the changed areas corresponding to the value of 1 in the binary change map, remove the results of the non-changed areas corresponding to the value of 0 in the binary change map, and output the land change type map of the changed areas, which is the final result R
[0134] R n×n = M n×n × S n×n
[0135] where M n×n is a binary change map with a size of n×n, and S n×n is an image segmentation map with a size of n×n.
[0136] Example 2: Multitask Remote Sensing Semantic Change Detection Device
[0137] Referring to Figure 6 , this example includes a feature extraction module, two semantic segmentation modules, and a change detection module. The outputs of the two single-temporal phases of the feature extraction module are respectively connected to the inputs of the two semantic segmentation modules, and the outputs of these two temporal phases are jointly connected to the input of the change detection module.
[0138] The feature extraction module receives high-resolution optical remote sensing images of two temporal phases, extracts semantic feature information therefrom, and outputs the semantic feature maps of the two temporal phases to the semantic segmentation module and the change detection module for processing;
[0139] The semantic segmentation module receives the semantic feature maps extracted by the feature extraction module, extracts the category information of various ground objects in the remote sensing image from the semantic feature information of the image, and restores it to a semantic segmentation map of the original image size, and outputs it as the result of semantic segmentation to the change detection module;
[0140] The change detection module receives the two-temporal-phase semantic feature maps extracted by the feature extraction module, extracts the change information in the remote sensing image from the two-temporal-phase semantic feature information, restores it to a binary change map of the original image size, and then applies the binary change map as a mask to the received semantic segmentation map to obtain a specific land change type map of the change area, and outputs it as the result of semantic change detection.
[0141] The effect of the present invention can be further illustrated by the following simulation experiment results.
[0142] I. Simulation Conditions
[0143] The simulation experiments were all carried out on a server equipped with an Nvidia GeForce RTX3090 graphics processor. The computer operating system was Ubuntu 22.04.4 LTS, and the remote sensing semantic change detection device used in the experiment was built based on the PyTorch framework.
[0144] The publicly available SECOND dataset and Landsat-SCD dataset were used as experimental test data, and both of these datasets were divided into a training set and a test set at a ratio of 7:3, without separately setting a validation set.
[0145] In the experiment, the same experimental parameters were used for the present invention and the comparative method. The batchsize of the input image data was set to 8, the initial learning rate was set to 0.1, the learning rate was dynamically adjusted using the exponential decay method, the stochastic gradient descent optimizer was used to update the model parameters, and the maximum number of iterations was set to 50 epochs.
[0146] II. Simulation Content
[0147] Simulation 1: The proposed method and the existing three remote sensing semantic change detection methods, namely TED, Bi-SRNet, and SCanNet, are respectively used for remote sensing semantic change detection on the SECOND dataset. Three representative groups of comparison results are selected from the detection results, as Figure 7 shown. Among them, Group 7(a) consists of two remote sensing images before and after the construction of a factory building in the same area, depicting the common transformation process from natural landform to artificial landform in the change detection task; Group 7(b) consists of two remote sensing images before and after the construction of a residential area in the same area, including a change area with a complex geometric structure; Group 7(c) consists of two remote sensing images before and after the demolition of rural houses in the same area, including an irregular change area and a vegetation pseudo-change area caused by different seasons.
[0148] It can be seen from the comparison results that in high-resolution remote sensing scenarios such as the SECOND dataset, compared with the other three commonly used remote sensing semantic change detection methods, the semantic change detection results of the proposed method show better classification and segmentation accuracy. For the water body and building areas in Group 7(a), the proposed method can achieve more refined and complete surface object segmentation; for the building areas in Group 7(b), the proposed method can more clearly detect the geometric connection relationship between houses, indicating that the proposed method has better perception ability for the global context space features in the image; for the farmland area on the left in Group 7(c), the proposed method can more effectively master the spatio-temporal change relationship of the same type of land cover, better distinguish the pseudo-changes existing in the image, and has a lower false detection probability.
[0149] Simulation 2: The proposed method and the existing three remote sensing semantic change detection methods, namely TED, Bi-SRNet, and SCanNet, are respectively used for remote sensing semantic change detection on the Landsat-SCD dataset. Three relatively representative groups of comparison results are selected from the detection results, as Figure 8 shown. Among them, Group 8(a) consists of two remote sensing images before and after the construction of cultivated land and water conservancy facilities in the same area, depicting the common transformation process from natural landform to artificial landform in the change detection task; Group 8(b) consists of two remote sensing images before and after the shrinkage of a lake in the same area, including a complex continuous irregular water body change scene; Group 8(c) consists of two remote sensing images before and after the construction of a town and cultivated land in the same area, including change area information of various sizes and scales.
[0150] It can be seen from the comparison results that in remote sensing scenarios with medium resolution and large scale such as the Landsat-SCD dataset, the present invention has better detection effects. Among them, for the changed areas in group 8(a), the present invention can more accurately classify the specific land cover types in the changed areas, achieving the classification effect closest to the true labels; for the water body areas in group 8(b), the present invention has achieved the best effect in restoring the edges of the changed areas; for the multiple building areas in group 8(c), the present invention can effectively identify and segment buildings with different sizes.
[0151] The above experimental comparison results show that the present invention has better detection accuracy and classification consistency in the remote sensing semantic change detection task, and has good versatility in the semantic change detection tasks of different remote sensing scenarios.
Claims
1. A multi-task remote sensing semantic change detection method based on a visual state space model, characterized in that Including: 1) Design two feature extraction encoders, two semantic decoders, and a change detection decoder according to the visual state space model; 2) Obtain dual-temporal high-resolution remote sensing images, and use the two feature extraction encoders to extract their semantic features to obtain two-way feature-enhanced semantic feature information; 3) Input the two-way feature-enhanced semantic feature information into the two semantic decoders respectively, obtain the semantic segmentation maps corresponding to the respective time phases through four semantic information extraction stages, and obtain the semantic information at different stages from the semantic decoders as semantic auxiliary information; 4) Import the semantic feature information and semantic auxiliary information into the change detection decoder, and obtain a binary change map through four change information extraction stages; 5) Use the binary change map as a mask to perform a cropping operation on the semantic segmentation map, and remove the non-changing regions to obtain the final semantic change detection result.
2. The method according to claim 1, wherein In step 1), two feature extraction encoders are designed according to the visual state space model, and their structures are the same. Each includes an image input layer, an image block layer, and four identical visual state space layers connected in sequence, where: The image input layer is used to adjust the input three-channel RGB image to a resolution of 256×256 and output it to the image block layer; The image block layer is used to divide the resized image into 64 non-overlapping image blocks according to the size of 32×32 pixels, record the original position information of each image block, and output the grouped images to the visual state space layer; The four visual state space layers are used to extract the key features in the grouped images. They are connected in series through a downsampling layer. The downsampling layer downsamples the feature image processed by the previous visual state space layer, adjusts the image size to half of the original, and outputs it for use by the next visual state space layer. Deep image semantic feature maps are obtained through layer-by-layer extraction and dynamic downsampling and output as the results of the feature extraction encoder.
3. The method according to claim 2, characterized in that, The realization of obtaining the deep image semantic feature maps through layer-by-layer extraction and downsampling includes the following: First, cross-scan the grouped images in four different directions: horizontally to the right, horizontally to the left, vertically downwards, and vertically upwards, and flatten them into four groups of one-dimensional sequences, which are expressed as follows: S1 = [F(0,0), F(0,1), F(0,2),..., F(0,7), F(1,0),..., F(7,7)] S2 = [F(0,7), F(0,6), F(0,5),..., F(0,0), F(1,7),..., F(7,0)] S3 = [F(0,0), F(1,0), F(2,0),..., F(7,0), F(0,1),..., F(7,7)] S4 = [F(7,0), F(6,0), F(5,0),..., F(0,0), F(7,1),..., F(0,7)] Among them, F is the input image group, each group has 8×8 image patches, and the dimension of each group of images ranges from 0 to 7. S1 is a one-dimensional sequence scanned horizontally to the right, S2 is a one-dimensional sequence scanned horizontally to the left, S3 is a one-dimensional sequence scanned vertically downward, and S4 is a one-dimensional sequence scanned vertically upward; Next, for each one-dimensional sequence, use the state-space equation to calculate the state variable weight h of each element in the sequence i and the output weight y i : where x i ∈R L represents the i-th element in the input sequence, y i ∈R L represents the system output at the current i-th position, h i ∈R N is the state variable at the current i-th position, h i-1 ∈R N is the state variable at the previous position i - 1, i = 1, 2, …, 64, are respectively the state matrix, input matrix, output matrix and direct transfer matrix of the system, L is the sequence length, N is the state space size, and a one-dimensional weight sequence is obtained after calculation. Secondly, arrange the obtained one-dimensional weight sequence in the original order to restore it to a two-dimensional weight image, and integrate the four groups of weight images by superimposing them to obtain the weight matrix W n×n : Among them, Q k is a two-dimensional weight image restored from a group of one-dimensional weight sequences. The range of k is from 1 to 4, representing the first to fourth groups of one-dimensional weight sequences respectively. n×n represents the image size, and the range of n is from 16 to 128; Finally, add the weight matrix to the original image group input to the current visual state space layer to obtain the feature map processed by the state space model, which is the processing result O of the visual state space layer n×n : O n×n = f(F n×n + W n×n + b) Among them, F n×n is the original image group, W n×n is the input two-dimensional weight matrix, both of which have a size of n×n. b is the bias term, and the three are added together. f is the activation function.
4. The method according to claim 1, characterized in that In step 1), two semantic decoders are designed according to the visual state space model. They have the same structure and both use the feature extraction encoder to process the obtained feature map with a size of 16×16 as the input, and are divided into four different processing stages. The first three stages are sequentially connected with a visual state space layer, an upsampling layer, and a feature fusion layer, and the last stage is sequentially connected with a visual state space layer and an upsampling layer, where: The visual state space layer is used to apply the visual state space model to model the global spatial context information of the input feature map and reconstruct the high-resolution semantic category information from the deep semantic information; The upsampling layer uses bilinear interpolation to upsample the input image, calculates the weighted average of four adjacent pixels to estimate the value of the new pixel, expands the image size to twice the original, and reconstructs the high-resolution output. The feature fusion layer is used to integrate multi-source, multi-scale or different-level feature information, and its processing process is as follows: First, receive the feature map F processed by the upsampling layer curr , and use it together with the intermediate feature map F obtained at the corresponding stage in the feature extraction encoder mid as inputs to perform a feature map addition operation to obtain the merged feature map F merge : F merge = F mid + f(W (1) * F curr + b1) Among them, W (1) is a 1×1 convolutional kernel used to perform channel dimensionality reduction on the input image F curr b1 is the bias term of the convolution, and f is the activation function; Secondly, perform feature fusion on the added image to obtain the fused feature map F out : F out = F merge + N(f(W (3) * F merge + b2)) Among them, W (3) is a 3×3 convolutional kernel, b2 is the bias term of the convolution, f is the activation function, and N is the normalization process; Finally, output the obtained fused feature map F out to the next stage for processing. At the same time, take the feature maps obtained in the first three processing stages as semantic auxiliary information, and divide them into the first, second, and third layers of semantic auxiliary information in order for use by the subsequent change detection decoder.
5. The method according to claim 1, wherein In step 1), a change detection decoder is designed according to the visual state space model, which uses the dual-temporal feature map processed by two feature extraction encoders as the input, and processes the dual-temporal feature map through four different processing stages, where: The first processing stage is used to complete the spatio-temporal state space processing and the upsampling layer, that is, first divide the dual-temporal feature map into 8×8 image patches, and perform cross-scanning in four different directions: horizontally to the right, horizontally to the left, vertically downward, and vertically upward to obtain four groups of one-dimensional sequences to represent the spatio-temporal change information in the image; then perform image resolution expansion processing on the one-dimensional sequence through upsampling to obtain the feature map of the first stage; The second processing stage is used to extract spatio-temporal features from the feature map of the first stage. Through re-partitioning and scanning processing, deeper spatio-temporal information is captured. The spatio-temporal features extracted from different processing layers are fused to obtain the fused feature map, and the fused feature map is subjected to image resolution expansion processing, and then the feature map of the second stage is obtained through upsampling; The third processing stage is used to extract deeper spatio-temporal features from the feature map of the second stage, fuse these features, and then obtain the feature map of the third stage through upsampling; The fourth processing stage further extracts deeper spatio-temporal features from the feature map of the third stage, further fuses these features, and finally obtains a high-resolution and accurate change feature map through upsampling after multiple processing and fusions. Its output is the result of the change detection decoder.
6. The method according to claim 1, wherein In step 2), two feature extraction encoders are used to extract the semantic features of the dual-temporal high-resolution remote sensing images. The dual-temporal remote sensing images are respectively input into the two feature extraction encoders, and the semantic features of the images are extracted from shallow to deep through four feature extraction modules in the feature extraction encoder in sequence. The first intermediate feature map, the second intermediate feature map, and the third intermediate feature map are output level by level through the first three feature extraction modules for the feature fusion processing of the corresponding levels in the subsequent semantic decoder; the final deep semantic feature map is output through the fourth module.
7. The method according to claim 1, wherein In step 3), the semantic segmentation maps corresponding to the respective time phases are obtained through four semantic information extraction stages, and the semantic information at different stages is obtained from the semantic decoder as semantic auxiliary information. Its implementation includes the following: In the first semantic information extraction stage, the input feature map is used to learn the image semantic information through the visual state space layer, and then the first semantic feature map is obtained through upsampling. After restoring the size of the first semantic feature map to be the same as that of the third intermediate feature map in step 2), they are input into the feature fusion layer in the encoder together to restore the first semantic information of the image and send it to the second stage; In the second semantic information extraction stage, after the first semantic information is processed through the visual state space layer and the upsampling layer, the second semantic feature map is obtained. Then it is input into the feature fusion layer in the encoder together with the second intermediate feature map in step 2) for processing to restore the second semantic information of the image and send it to the third stage; In the third semantic information extraction stage, after the second semantic information is processed through the visual state space layer and the upsampling layer, the third semantic feature map is obtained. Then it is input into the feature fusion layer in the encoder together with the first intermediate feature map in step 2) for processing to restore the third semantic information of the image and send it to the fourth stage; In the fourth semantic information extraction stage, the third semantic information is processed through the visual state space layer to be restored to the feature information of the land cover type, and then through the upsampling layer, the feature information is restored to the semantic segmentation map with the same size as the original image, which is the output result of the semantic decoder. The semantic information obtained in the first three stages is used as semantic auxiliary information and input into the corresponding change information extraction stage.
8. The method according to claim 1, wherein In step 4), a binary change map is obtained through four change information extraction stages. Its implementation includes the following: In the first change information extraction stage, the dual-temporal semantic feature maps extracted in step 2) are used as the input, and the change information of the image is extracted through the visual state space layer. Then the first-stage change feature map is obtained through upsampling, and its size is restored to be the same as that of the third semantic feature map obtained in step 3) and sent to the second stage; In the second change information extraction stage, the change feature map of the first stage is used as the input, and the first semantic auxiliary information obtained in step 3) is used to learn the spatio-temporal change relationship through the visual state space layer to obtain the spatio-temporal change feature map. Then, the spatio-temporal change feature map and the change feature map of the first stage are integrated with deep semantic information and shallow spatial information through the image fusion layer. Finally, the second stage change feature map is obtained through upsampling, and its size is restored to be the same as that of the second semantic feature map obtained in step 3) and sent to the third stage; In the third change information extraction stage, the change feature map of the second stage is used as the input, and the second semantic auxiliary information obtained in step 3) is used to learn the spatio-temporal change relationship through the visual state space layer to obtain the spatio-temporal change feature map. Then, the spatio-temporal change feature map and the change feature map of the second stage are integrated with deep semantic information and shallow spatial information through the image fusion layer. Finally, the third stage change feature map is obtained through upsampling, and its size is restored to be the same as that of the first semantic feature map obtained in step 3) and sent to the fourth stage; In the fourth change information extraction stage, the change feature map of the third stage is used as the input, and the third semantic auxiliary information obtained in step 3) is used to learn the spatio-temporal change relationship through the visual state space layer to obtain the spatio-temporal change feature map. Then, the spatio-temporal change feature map and the change feature map of the third stage are integrated with deep semantic information and shallow spatial information through the image fusion layer. Finally, a binary change map with the same size as the original image is obtained through upsampling, which is the output of the change detection decoder.
9. The method according to claim 1, wherein In step 5), a cropping operation is performed on the semantic segmentation map to remove the non-change regions to obtain the final semantic change detection result, and its implementation includes: The binary change map obtained in step 4) and the semantic segmentation map obtained in step 3) are subjected to mask screening by element-wise pixel multiplication, that is, the semantic classification results of the changed regions corresponding to the value of 1 in the binary change map are retained, and the results of the non-changed regions corresponding to the value of 0 in the binary change map are removed. Finally, the land change type map of the changed region is output, which is the result R of semantic change detection m×n : R m×n = M m×n × S m×n Among them, M m×n is a binary change diagram with dimensions of m×n, and S m×n is an image segmentation diagram with dimensions of m×n.
10. A multi-task remote sensing semantic change detection device based on a visual state space model, comprising: A feature extraction module, configured to extract semantic feature information from high-resolution optical remote sensing images of two temporal phases, and input the extracted semantic feature information of the two temporal phases to a semantic segmentation module and a change detection module for processing; A semantic segmentation module, configured to extract semantic category information of various ground objects in the remote sensing image from the semantic feature information of the two temporal phases, and restore it to a semantic segmentation map with the size of the original image, and output it as the result of semantic segmentation; A change detection module, configured to extract change information of the two-temporal-phase remote sensing image from the semantic feature information of the two temporal phases, and restore it to a binary change map with the size of the original image, and then use the binary change map as a mask to be applied to the semantic segmentation map to obtain a land change type map of the change region, and output it as the result of semantic change detection.
Citation Information
Cited By
Casting surface defect automatic detection method
CN121169855A
Remote sensing image segmentation method and system based on lightweight UMFormer
CN122434963A
A method and system for detecting semantic changes in remote sensing images
CN122574684A