Remote sensing image semantic change detection method and device based on Mamba model
By introducing the state space model of the Mamba model, the problem of insufficient receptive field in semantic change detection of remote sensing images is solved, and accurate change area and category detection of remote sensing images is achieved, thereby improving the accuracy and robustness of detection.
Patent Information
- Application Number
- CN202411635675.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-15
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2044-11-15
AI Technical Summary
Existing deep learning models cannot effectively capture long-range dependencies between pixels in semantic change detection in remote sensing images, resulting in inaccurate detection.
A remote sensing image semantic change detection method based on the Mamba model is adopted. By introducing a state space model with time-varying parameters, global modeling is achieved to capture the complex relationships and long-distance semantic associations between images. Feature processing is performed using the Mamba encoder, difference feature extraction module, feature fusion enhancement module and decoder.
It significantly improves the ability to recognize subtle changes, improves the accuracy and robustness of semantic change detection, and can accurately detect changed areas and categories.
Smart Images

Figure CN119580258B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing and computer vision, and in particular relates to a remote sensing image semantic change detection method and device based on a Mamba model. Background Art
[0002] Semantic change detection is an important research area in computer vision and image processing. Its primary goal is to identify changes in ground object categories in multi-temporal imagery. With advances in remote sensing data acquisition technology, particularly the development of high-resolution satellite imagery and drone technology, ground observations have become increasingly sophisticated. This makes semantic change detection promising in areas such as land cover monitoring, urban expansion tracking, disaster response, and resource management. Traditional change detection methods focus on binary detection of changed areas and fail to provide detailed categorization information before and after the change. Semantic change detection, on the other hand, not only identifies which areas have changed but also further identifies the specific category of the change.
[0003] In existing technologies, the changed areas are detected by analyzing change vectors, followed by classification. With the rapid development of deep learning technology, end-to-end deep learning models have been widely used in semantic change detection tasks. By processing multi-temporal images simultaneously, such models avoid the error accumulation problem that may arise from step-by-step processing in traditional methods. Typical deep learning architectures, such as fully convolutional networks (FCNs), construct a three-branch FCN architecture, in which two branches are used to extract semantic information and the other branch extracts change information. However, the main problem with convolutional neural networks such as FCNs is that their limited receptive field cannot capture long-range dependencies between pixels, resulting in inaccurate detection. Summary of the Invention
[0004] To address the problems of the above-mentioned prior art, the present invention uses a Mamba-based remote sensing semantic change detection model to process remote sensing images. This introduces time-varying parameters into the state-space model, achieving global modeling with linear complexity. This effectively models the relationship between any region and the entire image, solving the problem of insufficient receptive field in convolutional neural networks. For semantic change detection tasks, in complex areas containing different objects, the interaction between different feature layers is modeled from a global perspective, which can capture the complex relationship and long-range semantic associations between any region and the entire image. Through more comprehensive contextual information, the ability to recognize subtle changes is significantly improved, thereby accurately detecting changed areas and categories, and improving the accuracy and robustness of semantic change detection in complex environments.
[0005] Specifically, the present invention provides a remote sensing image semantic change detection method based on the Mamba model, the method comprising: constructing a remote sensing semantic change detection model, the model comprising a Mamba encoder branch, a difference feature extraction module, a feature fusion enhancement module, a Mamba decoder branch, and a classifier; using the remote sensing semantic change detection model to detect remote sensing images comprises:
[0006] Acquire paired dual-time remote sensing images and preprocess the dual-time remote sensing images;
[0007] The preprocessed dual-time remote sensing image is input into the Mamba encoder branch for feature extraction to obtain multi-scale features;
[0008] The multi-scale features are processed using a difference feature extraction module to obtain multi-scale difference features;
[0009] The multi-scale features and multi-scale difference features are input into the feature fusion enhancement module for fusion to obtain semantic features;
[0010] The Mamba decoder branch is used to decode semantic features and multi-scale difference features respectively;
[0011] The decoded semantic features and multi-scale difference features are input into the classifier and subjected to masking operation to obtain the semantic change map of the dual-temporal remote sensing image.
[0012] To achieve the above objectives, the present invention further provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements any of the above-mentioned remote sensing image semantic change detection methods based on the Mamba model.
[0013] To achieve the above-mentioned objectives, the present invention also provides a remote sensing image semantic change detection device based on the Mamba model, comprising a processor and a memory; the memory is used to store a computer program; the processor is connected to the memory and is used to execute the computer program stored in the memory, so that the remote sensing image semantic change detection device based on the Mamba model performs any of the above-mentioned remote sensing image semantic change detection methods based on the Mamba model.
[0014] Beneficial effects of the present invention:
[0015] The present invention proposes a remote sensing image semantic change detection method based on the Mamba model. The method can accurately detect the changed areas and specific change categories of the changed areas in dual-temporal remote sensing images. The present invention utilizes Mamba encoder branches to extract image features. A difference feature extraction module expands and splices image features in multiple directions, and then generates a feature sequence after processing using a state-space model, which can more accurately identify the differences between dual-temporal images. The feature fusion enhancement module strengthens the category representation of the changed areas, suppresses irrelevant background noise and subtle unchanged areas, and ensures the accuracy of the change features in complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 This is a flow chart of the remote sensing image semantic change detection method based on the Mamba model of the present invention;
[0017] Figure 2 This is a structural diagram of the remote sensing semantic change detection model of the present invention;
[0018] Figure 3 Schematic diagram of the VSS block structure of the present invention;
[0019] Figure 4 Schematic diagram of the SS2D structure of the present invention;
[0020] Figure 5 Schematic diagram of the structure of the difference feature extraction module of the present invention;
[0021] Figure 6 It is a structural diagram of the feature fusion enhancement module of the present invention. DETAILED DESCRIPTION
[0022] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0023] The invention provides a remote sensing image semantic change detection method based on a Mamba model. The method comprises: constructing a remote sensing semantic change detection model, wherein the model comprises a Mamba encoder branch, a difference feature extraction module, a feature fusion enhancement module, a Mamba decoder branch, and a classifier; detecting remote sensing images using the remote sensing semantic change detection model, comprising: acquiring paired dual-temporal remote sensing images and preprocessing the dual-temporal remote sensing images; inputting the preprocessed dual-temporal remote sensing images into a Mamba encoder branch for feature extraction to obtain multi-scale features; using the difference feature extraction module to process the multi-scale features to obtain multi-scale difference features; inputting the multi-scale features and the multi-scale difference features into a feature fusion enhancement module for fusion to obtain semantic features; using the Mamba decoder branch to respectively decode the semantic features and the multi-scale difference features; inputting the decoded semantic features and the multi-scale difference features into a classifier and performing a masking operation to obtain a semantic change map of the dual-temporal remote sensing image.
[0024] The above remote sensing semantic change detection model is a trained remote sensing semantic change detection model.
[0025] The SECOND dataset was used to train a remote sensing semantic change detection model. The dataset contains 4,662 pairs of 512×512 aerial images with a resolution of 0.5-3 meters per pixel. The dataset focuses on six major land cover categories: bare ground, trees, low vegetation, water bodies, buildings, and playgrounds. This dataset uses land cover map pairs and non-change masks to represent change, distinguishing between changed and unchanged pixels within the same category.
[0026] A specific implementation method of a remote sensing image semantic change detection method based on the Mamba model is as follows: Figure 1 As shown, the method includes:
[0027] S1: Data preprocessing, performing preprocessing operations on the paired dual-temporal remote sensing images I1 and I2, including normalization, random inversion, etc.
[0028] S2: Construct the Mamba encoder branch, differential feature extraction module, feature fusion enhancement module, Mamba decoder branch and classifier;
[0029] S3: The preprocessed dual-temporal remote sensing image is used to extract multi-scale image features of the dual-temporal remote sensing image using the Mamba encoder.
[0030] S4: The extracted multi-scale image features are input into the difference feature extraction module to extract the multi-scale difference features. The multi-scale image features and the multi-scale difference features are input into the feature fusion enhancement module, and the difference features are used to enhance the image features to obtain the semantic features.
[0031] S5: The difference features are input into the Mamba decoder for decoding, and then a binary classifier is used to generate a binary change map. The semantic features are input into the Mamba decoder for decoding, and then a semantic classifier is used to generate a semantic category map.
[0032] S6: Perform a masking operation on the binary change map and the semantic map, and finally obtain the semantic change map of the dual-temporal remote sensing image.
[0033] In this embodiment, the image preprocessing is to preprocess the paired dual-temporal remote sensing images I1 and I2 simultaneously; specifically, the images of the dual-temporal high-resolution remote sensing image training set are subjected to data enhancement by random flipping, random blurring, random color adjustment, and normalization.
[0034] In this embodiment, the Mamba encoder branch adopts the VMamba backbone network. The Mamba encoder branch consists of five stages. The first is the Patch Embedding operation to increase the image dimension to 96. The number of VSS blocks in the remaining four stages is [2, 2, 9, 2]. The difference feature extraction module expands and splices the feature maps of the two images in multiple directions, and then uses the state space model to generate a change feature sequence. The sequence is then divided into two sequences and a weight matrix is obtained by Sigmoid. The weight matrix is multiplied with the features of the two images and then subtracted to obtain the difference feature. The feature fusion enhancement module generates semantic features by fusing the difference features with the image features and combining the attention mechanism to perform weighted processing on the image features. Each block of the Mamba decoder is composed of VSS blocks, and then an upsampling module is used to double the size of the input features.
[0035] Specifically, the two Mamba encoder branches are weight-shared, where the Mamba encoder branch uses the VMamba backbone network, which consists of five stages. The first stage is Patch Embedding, and the second to fifth stages are Figure 2 The VSS block is composed of the following: Figure 3 As shown in Figure 2, it includes a linear layer, a 3×3 depth-separable convolution, a SS2D layer, and a linear layer, where the SS2D layer that completes the feature extraction is shown in Figure 2. Figure 4 As shown in the figure, SS2D includes selective scanning and state space model. The second to fourth stages also include downsampling operations, which increase the channel dimension by two times and reduce the size by half. The image features of each stage will be used as a difference extraction module and a feature fusion enhancement module.
[0036] The difference feature extraction module is as follows Figure 5As shown in the figure, it consists of a selective scanning operation and an SSM module. First, the difference feature extraction module performs a scanning operation on the image features, and then the sequence splicing is input into the SSM module, and the corresponding addition and separation are performed to obtain the feature map. The corresponding attention weight is generated by the Sigmoid function, and then the attention weight is fused with the image features. Finally, the results are subtracted to obtain the difference feature.
[0037] Feature fusion enhancement module such as Figure 6 As shown in the figure, it consists of a 3×3 convolutional layer and a 1×1 convolutional layer to form a convolution operation. Each convolutional layer is followed by Batchnorm and ReLU, which are used to generate the fully connected layer 1 for scaling parameters and the fully connected layer 2 for offset parameters. First, the difference features are concatenated with the image features, and the features are extracted through the convolution operation. Then, the scaling parameters and offset parameters are obtained through the fully connected layer. Finally, the scaling parameters and offset parameters are fused with the image features to generate semantic features.
[0038] The Mamba decoder branches are two semantic Mamba decoders, which share weights, and a binary Mamba decoder. The semantic Mamba decoder consists of four stages, each of which is composed of VSS blocks. The binary Mamba decoder consists of four stages, each of which is composed of VSS blocks. All decoders use bilinear interpolation for upsampling in each stage.
[0039] The classifiers are semantic classifier and binary classifier, both of which are composed of convolution layers with a convolution kernel of 1×1.
[0040] In this embodiment, the use of the Mamba encoder to extract multi-scale features includes: extracting features of the dual-phase images I1 and I2. The dual-phase image of size 512×512 is directly passed into the VMamba backbone network loaded with a pre-trained model. After Patch Emebedding in the Vmamba backbone network, the image dimension becomes 96 and the size becomes 1 / 4 of the original size. It then passes through 2, 2, 9, and 2 VSS blocks in sequence; four different directional feature sequences are obtained through selective scanning, and then the state space model processes the feature sequence by a method of globally processing context information. Finally, the four different directional sequences are added together and the image features are obtained after passing through a linear layer. In the Mamba encoder branch, a total of four different scale feature maps are extracted, and their channel dimensions are 96, 192, 384, and 768, respectively.
[0041] In this embodiment, each VSS block includes a linear layer, a 3×3 depthwise separable convolution, an SS2D module, a linear layer, where the SS2D includes a 2D selective scan operation and a state space model (SSM) that relies on the intermediate implicit state h(t)∈R NMapping one-dimensional input function x(t)∈R to output y(t)∈R, the specific formula is:
[0042] h′(t)=Ah(t)+Bx(t)
[0043] y(t)=Ch(t)
[0044] where A∈R N×N Represents the state matrix, B∈R N×1 , C∈R N×1 Represents the projection parameter, SS2D discretizes this continuous system, introduces the time scale parameter Δ, and uses a fixed discretization rule to convert A and B into discrete parameters and Typically, zero-order hold is used as the discretization rule, which can be defined as follows:
[0045]
[0046] After discretization, SSM can be linearly recursively expressed as:
[0047]
[0048] y(t)=Ch(t)
[0049] It can also be expressed as a global convolution:
[0050]
[0051] in Represents a structured convolution kernel, L represents the length of the input sequence x, and the first three units except the first unit also include a downsampling with a coefficient of 2, so that the output feature map is 1 / 2 of the input.
[0052] In this embodiment, four image features at different scales are input into the difference feature extraction module to obtain four difference features at different scales. The difference features at four different scales are used to perform feature fusion enhancement with the image features to obtain semantic features, which specifically includes the following steps:
[0053] Step 1: First, scan the image features to obtain a pair of feature sequences containing four directions. Then, concatenate the feature sequences with the same scanning direction and input them into the SSM module. Add them up to obtain the change features. The change features are divided into two feature maps. Then, Sigmoid is used to obtain the weight matrix, which is multiplied by the image features. Then, the results are subtracted to obtain the difference features.
[0054] Step 2: Concatenate the difference features with the image features, and then perform a convolution operation. The number of channels is halved through a 3×3 convolution layer, and features are further extracted through a 1×1 convolution layer. The difference features are input into the fully connected layer 1 and the fully connected layer 2 to obtain the scaling parameter and offset parameter. The scaling parameter is multiplied with the feature after passing through the Sigmoid activation function, and then added to the offset parameter to generate the semantic feature.
[0055] The difference features of the four scales are input into the binary Mamba decoder for decoding, and the semantic features of the four scales are input into the semantic Mamba decoder for decoding. The binary Mamba decoder and the semantic Mamba decoder are Unet architectures, which are both composed of four blocks. The difference feature map and the semantic feature map are obtained through the VSS block and the upsampling module.
[0056] The difference feature map is subjected to a classifier to generate a binary change map, and the semantic feature map is subjected to a classifier to generate a final semantic category map. Specifically, for the generated difference feature map, a classifier is constructed to perform binary classification in the channel dimension to generate a final binary change map; for the semantic feature map, a classifier is constructed to perform multi-classification in the channel dimension to generate a semantic category map. The classifier is specifically implemented as follows: a 2D convolution is used to transform the channel dimension to generate the final binary change map and semantic category map. The obtained binary change map and I1 and I2 semantic category maps are upsampled to their original size. The binary change map is masked with the I1 and I2 semantic category maps to finally obtain the semantic change map of the dual-temporal remote sensing image.
[0057] A remote sensing image semantic change detection system based on the Mamba model, the system comprising: a preprocessing module, a Mamba encoder branch, a difference feature extraction module, a feature fusion enhancement module, a Mamba decoder branch, and a classifier;
[0058] The preprocessing module is used to perform image enhancement on the input dual-temporal remote sensing image;
[0059] The Mamba encoder branch is used to extract features of the dual-temporal remote sensing image;
[0060] The difference feature extraction module is used to extract difference features from image features;
[0061] The feature fusion enhancement module is used to perform feature fusion using difference features and image features;
[0062] The Mamba decoder branch is used to decode difference features and semantic features;
[0063] The classifier is used to obtain a binary change map from the difference feature map, obtain a semantic category map from the semantic feature map, judge whether there is a change in the dual-temporal remote sensing image based on the binary change map, and judge the semantic category of the dual-temporal remote sensing image based on the semantic category map.
[0064] The system implementation of the present invention is the same as the specific implementation of the method.
[0065] In one embodiment of the present invention, the present invention further includes a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-mentioned remote sensing image semantic change detection methods based on the Mamba model.
[0066] Those skilled in the art will appreciate that all or part of the steps in the above-described method embodiments can be implemented using hardware associated with a computer program. The aforementioned computer program can be stored in a computer-readable storage medium. When executed, the program performs the steps in the above-described method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.
[0067] A remote sensing image semantic change detection device based on the Mamba model comprises a processor and a memory; the memory is used to store a computer program; the processor is connected to the memory and is used to execute the computer program stored in the memory, so that the remote sensing image semantic change detection device based on the Mamba model performs any of the above-mentioned remote sensing image semantic change detection methods based on the Mamba model.
[0068] Specifically, the memory includes various media that can store program codes, such as ROM, RAM, magnetic disk, USB flash drive, memory card or optical disk.
[0069] Preferably, the processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0070] The above embodiments further illustrate the purpose, technical solutions and advantages of the present invention in detail. It should be understood that the above embodiments are only preferred implementation plans of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made to the present invention within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A remote sensing image semantic change detection method based on the Mamba model, characterized in that: include: Build a remote sensing semantic change detection model, which includes a Mamba encoder branch, a difference feature extraction module, a feature fusion enhancement module, a Mamba decoder branch, and a classifier; The remote sensing semantic change detection model is used to detect remote sensing images, including: Acquire paired dual-time remote sensing images and preprocess the dual-time remote sensing images; The preprocessed dual-time remote sensing image is input into the Mamba encoder branch for feature extraction to obtain multi-scale features; The multi-scale features are processed using a difference feature extraction module to obtain multi-scale difference features. Specifically, the following steps are performed: the image features of the dual-temporal remote sensing image are subjected to a linear layer and a 3×3 depthwise separable convolution, and the convolution results are expanded along four different directions; the two sequences expanded along the same direction are concatenated and input into the SSM module to generate a feature sequence; the feature sequences are summed and output through a linear layer; the output features are separated into two feature maps, and two attention weights are obtained after a Sigmoid activation function; the two attention weights are used to perform dot product with the features of the dual-temporal remote sensing image; the elements of the corresponding dot product results are subtracted and the absolute value is taken to obtain the difference features; The multi-scale features and multi-scale difference features are input into the feature fusion enhancement module for fusion to obtain semantic features; The Mamba decoder branch is used to decode semantic features and multi-scale difference features respectively; The decoded semantic features and multi-scale difference features are input into the classifier and subjected to masking operation to obtain the semantic change map of the dual-temporal remote sensing image.
2. The remote sensing image semantic change detection method based on the Mamba model according to claim 1, characterized in that: The preprocessing of the dual-time remote sensing image includes randomly flipping, randomly blurring and randomly adjusting the color of the remote sensing image to obtain an enhanced dual-time remote sensing image.
3. The remote sensing image semantic change detection method based on the Mamba model according to claim 1, characterized in that: The Mamba encoder branch consists of five cascade units, of which the first cascade unit is Patch Embedding, which consists of a 3×3 convolutional layer; the second to fifth cascade units include 2, 2, 9, and 2 VSS blocks respectively; each VSS block consists of a linear layer, a 3×3 depth-separable convolution, an SS2D module and a linear layer; the SS2D module includes a 2D selective scan operation and a state-space model.
4. The remote sensing image semantic change detection method based on the Mamba model according to claim 1, characterized in that: The feature fusion enhancement module consists of a 3×3 convolutional layer, a 1×1 convolutional layer, two fully connected layers, and the activation function ReLU.
5. The remote sensing image semantic change detection method based on the Mamba model according to claim 4, characterized in that: The feature fusion enhancement module fuses multi-scale features and multi-scale difference features, including: splicing multi-scale features and multi-scale difference features; halving the number of channels of the spliced features through a 3×3 convolution layer, and further extracting features through a 1×1 convolution layer; inputting the further feature-extracted difference features into the fully connected layer 1 to obtain the scaling parameters, and inputting the difference features into the fully connected layer 2 to obtain the offset parameters; the scaling parameters are multiplied by the image features after passing through the Sigmoid activation function, and then added to the offset parameters to generate semantic features.
6. The remote sensing image semantic change detection method based on the Mamba model according to claim 1, characterized in that: The Mamba decoder branch consists of four cascaded units, each unit includes 1 VSS block, and each VSS block includes a linear layer, a 3×3 depth-wise separable convolution, an SS2D module, and a linear layer.
7. The remote sensing image semantic change detection method based on the Mamba model according to claim 1, characterized in that: The classifier consists of a binary classifier and a semantic classifier, both of which are composed of a convolutional layer with a 1×1 convolution kernel. The binary classifier performs binary classification on the decoded difference features to obtain a binary change map. The semantic classifier classifies the decoded semantic features to obtain a semantic category map of the bi-temporal remote sensing image. A masking operation is performed on the binary change map and the semantic category map of the bi-temporal remote sensing image to obtain the semantic change map of the bi-temporal remote sensing image.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: The computer program is executed by a processor to implement the remote sensing image semantic change detection method based on the Mamba model according to any one of claims 1 to 7.
9. A remote sensing image semantic change detection device based on the Mamba model, characterized in that: The invention comprises a processor and a memory; the memory is used to store a computer program; the processor is connected to the memory and is used to execute the computer program stored in the memory, so that the remote sensing image semantic change detection device based on the Mamba model executes the remote sensing image semantic change detection method based on the Mamba model according to any one of claims 1 to 7.