Remote sensing image change detection method and device based on iterative Mamba architecture
By adopting the iterative Mamba architecture in remote sensing image change detection, combining the Mamba feature extractor, state space change detection module and iterative diffusion model, the existing methods are solved in the problem of insufficient accuracy and noise interference in high-resolution image processing, and efficient and accurate change detection is achieved.
Patent Information
- Application Number
- CN202411198502.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-29
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2044-08-29
AI Technical Summary
Existing remote sensing image change detection methods have problems of insufficient accuracy and noise interference when processing high-resolution images, and deep learning methods such as CNN and transformers have challenges in computing resource and data dependencies.
Using an iterative Mamba architecture method, high-dimensional and low-dimensional feature maps are extracted through the Mamba feature extractor, combined with the state space change detection module to capture long-frequency change features, and reduce the noise impact through feature fusion and global hybrid attention modules, and finally noise correction is performed through the iterative diffusion model.
It significantly improves the accuracy and efficiency of remote sensing image change detection, reduces the consumption of computing resources, is suitable for environments with limited resources, and ensures the fidelity of the change detection results.
Smart Images

Figure CN119205638B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a remote sensing image change detection method and device based on iterative Mamba architecture. Background Art
[0002] Remote sensing image change detection (CD) is the process of analyzing remote sensing images taken at different times to identify changes in ground objects. It has a wide range of applications in land use management, environmental monitoring, resource assessment, and disaster assessment. Existing change detection methods mainly include methods based on traditional image processing technology and methods based on deep learning.
[0003] Traditional change detection methods usually rely on differential calculations at the pixel level, feature level, and decision level, but these methods often face problems such as insufficient accuracy and noise interference when processing high-resolution remote sensing images. Pixel-level methods detect changes by comparing the grayscale or color values of images. Although simple and intuitive, they are sensitive to noise and difficult to handle complex scenes. Feature-level methods perform change detection by extracting features such as texture and shape of images, which can reduce the impact of noise to a certain extent, but are prone to losing detail information in high-resolution images. Decision-level methods combine multiple detection results and generate the final change detection results through voting or other strategies, but their computational complexity is high and they are susceptible to individual false detections.
[0004] In recent years, deep learning techniques, especially Convolutional Neural Networks (CNNs) and Transformers models, have performed well in the field of image processing, but these methods also have some limitations in change detection tasks.
[0005] Convolutional neural networks perform well in capturing local features, but they are inadequate in capturing long-distance dependencies and global information, which limits their performance in complex change detection tasks. Especially when processing large-scale images, CNNs have a limited receptive field and it is difficult to effectively utilize the global information of the image, resulting in inaccurate change detection results. In addition, CNNs require a large amount of labeled data during training, which is a big challenge in the field of remote sensing images.
[0006] The transformer model can capture global features well, but its computational complexity is high and it is difficult to operate efficiently in resource-limited environments. The transformer models each pixel of the image through the self-attention mechanism, which can well capture the long-distance dependencies in the image, but the computational complexity of this method increases quadratically with the increase of the image size, resulting in huge consumption of computing resources when processing high-resolution images. In addition, the transformer model is highly dependent on training data and also requires a large amount of labeled data.
[0007] Commonly used image information fusion methods are prone to information loss or redundancy when processing multi-temporal images, affecting the accuracy of change detection. Information fusion of multi-temporal images is a key step in change detection, but existing methods often introduce redundant information or lose important change information during information extraction and fusion, thus affecting the accuracy of detection results. For example, simple image difference methods are prone to introduce noise, while feature extraction-based methods may ignore subtle changes.
[0008] Therefore, it is urgent to develop a new remote sensing image change detection method. Summary of the invention
[0009] The main purpose of the present invention is to provide a remote sensing image change detection method and device based on iterative Mamba architecture to solve the problem of low accuracy and efficiency of remote sensing image change detection in the prior art.
[0010] To achieve the above object, the first aspect of the present invention provides a remote sensing image change detection method based on iterative Mamba architecture, the method comprising:
[0011] Inputting the first remote sensing image and the second remote sensing image into the Mamba feature extractor respectively to extract a multi-dimensional feature map set of the first remote sensing image and a multi-dimensional feature map set of the second remote sensing image respectively, wherein the multi-dimensional feature map set includes at least a high-dimensional feature map and a low-dimensional feature map;
[0012] Inputting the multi-dimensional feature map set of the first remote sensing image and the multi-dimensional feature map set of the second remote sensing image into the state space change detection module, and obtaining the multi-dimensional long-frequency change feature map set of the first remote sensing image and the second remote sensing image through state space modeling, wherein the multi-dimensional long-frequency change map feature set includes at least a high-dimensional change feature map and a low-dimensional change feature map;
[0013] Performing feature fusion on the multi-dimensional long-frequency change map feature set to obtain a low-dimensional fused change feature map;
[0014] Inputting the high-dimensional change feature map and the low-dimensional fused change feature map into a global hybrid attention module for feature fusion to generate a fused output feature map;
[0015] The fusion output feature map is subjected to noise correction through an iterative diffusion model to generate a remote sensing image change detection map.
[0016] Further, the Mamba feature extractor includes a linear embedding layer and N encoder layers, the encoder layer includes a VSS block and a patch merging layer, and the step of inputting the first remote sensing image into the Mamba feature extractor to extract a multi-dimensional feature map set of the first remote sensing image includes:
[0017] The first remote sensing image is processed into non-overlapping patches by the linear embedding layer, each patch is linearly embedded into a preset feature space, and the initial feature vectors of all patches are set to form an initial feature map;
[0018] The initial feature map is input into the first encoder layer for encoder layer processing; wherein the encoder layer processing step includes scanning the image features of the feature map through the VSS block, integrating the contextual relationship, and generating the feature map processed by the VSS block; the feature map processed by the VSS block is input into the patch merging layer, a preset number of adjacent patches in the feature map processed by the VSS block are linearly transformed, and spliced to form a super patch, and the feature vectors of all super patches are set to form a new feature map, and the feature dimension of the new feature map is higher than the feature dimension of the initial feature map;
[0019] The new feature map is input into the next encoder layer, and the steps of the encoder layer processing are repeated until the encoder processing of the N encoder layers is completed, and the set of feature maps output by each encoder layer is used as the multi-dimensional feature map set of the first remote sensing image, and the multi-dimensional feature map set includes at least a high-dimensional feature map and a low-dimensional feature map.
[0020] Furthermore, the step of inputting the multidimensional feature map set of the first remote sensing image and the multidimensional feature map set of the second remote sensing image into the state space change detection module, and obtaining the multidimensional long-frequency change feature map of the first remote sensing image and the second remote sensing image through state space modeling includes:
[0021] Inputting the feature map of the first remote sensing image output by the i-th encoder layer and the feature map of the second remote sensing image output by the i-th encoder layer into the state space change detection module, processing them through the state space model, and outputting the captured i-th layer long-frequency change feature map; wherein 0<i≤N;
[0022] The set of long-frequency variation feature maps of each layer is used as the multi-dimensional long-frequency variation feature set map, the long-frequency variation feature output of the high-dimensional feature map of the first remote sensing image and the high-dimensional feature map of the second remote sensing image is used as the high-dimensional long-frequency variation feature map, and the long-frequency variation feature output of the low-dimensional feature map of the first remote sensing image and the low-dimensional feature map of the second remote sensing image is used as the low-dimensional long-frequency variation feature map.
[0023] Furthermore, the step of fusing the multi-dimensional long-frequency change feature set to obtain a low-dimensional fused change feature map includes:
[0024] Upsampling the high-dimensional change feature map so that the spatial resolution of the high-dimensional change feature map matches the resolution of the low-dimensional change feature map;
[0025] The upsampled high-dimensional change feature map is fused with the low-dimensional change feature map by element-wise addition to obtain a low-dimensional fused change feature map.
[0026] Furthermore, the step of inputting the high-dimensional change feature map and the low-dimensional fused change feature map into the global hybrid attention module for feature fusion to generate a fused output feature map includes:
[0027] Perform channel expansion on the high-dimensional change feature map and the low-dimensional fused change feature map, respectively, to obtain an expanded high-dimensional change feature map and an expanded low-dimensional fused change feature map, respectively;
[0028] Performing feature fusion on the expanded high-dimensional change feature map and the expanded low-dimensional fused change feature map to generate a fused feature map;
[0029] Using 1×1 convolution and Sigmoid activation function on the fused feature map to generate a global attention feature map;
[0030] Splitting the global attention feature map into a high-dimensional attention feature map and a low-dimensional attention feature map;
[0031] Multiplying the high-dimensional attention feature map and the low-dimensional attention feature map by the initial feature map element by element, respectively, to generate a weighted high-dimensional feature map and a weighted low-dimensional feature map, respectively;
[0032] The weighted high-dimensional feature map and the weighted low-dimensional feature map are reconstructed to generate the final fused output feature map.
[0033] Furthermore, the step of performing noise correction on the fusion output feature map by an iterative diffusion model to generate a remote sensing image change detection map includes:
[0034] Obtain the initialized input noise feature as the starting point of iteration;
[0035] The fusion output feature map is iteratively optimized multiple times through an iterative diffusion model to finally generate the remote sensing image change detection map, wherein each iterative optimization includes a forward process and a reverse process.
[0036] In the forward process, the input noise feature is gradually added to the change feature map of the input iterative forward process until a pure noise map is accumulated;
[0037] In the reverse process, the noise estimation model is used to perform noise estimation on the image input to the reverse process to generate a noise estimation value, and the denoised image is calculated according to the reverse denoising formula.
[0038] Furthermore, the noise estimation value is:
[0039]
[0040] Where t is the number of iterations, and is the changing feature, D represents the decoder of the noise estimation model, GHAT represents the global hybrid attention mechanism, and Represent low-dimensional and high-dimensional change characteristics respectively, is the input noise characteristic.
[0041] Furthermore, the reverse denoising formula is:
[0042]
[0043] Where t is the number of iterations, represents noise, and are diffusion model parameters, is the noise estimate, is the mean square error of the noise estimate, is the denoised image after the t-1th iteration.
[0044] Furthermore, the noise estimation model adopts a Unet structure, including a second encoder and a second decoder.
[0045] The second encoder receives input features, the input features including input noise features and input change feature maps, extracts layer by layer and generates output feature maps of each layer, wherein the output feature maps of each layer include different dimensional information of the input features;
[0046] The second decoder is jump-connected to the second encoder, and the output feature map of the corresponding layer in the second encoder is passed to the corresponding layer in the second decoder, and the resolution of the output feature map is restored layer by layer until the difference between the resolution of the output feature map and the input change feature map is less than a preset threshold, thereby obtaining a preliminary change detection map;
[0047] The preliminary change detection map is taken as input through an iterative optimization mechanism, and the preliminary change detection map is iteratively processed multiple times through multi-dimensional features and a global hybrid attention mechanism, and the change detection results are gradually corrected until a preset iteration stop condition is met, and the final remote sensing image change detection map is output.
[0048] A second aspect of the present invention provides a remote sensing image change detection device based on an iterative Mamba architecture, the device comprising:
[0049] A feature extraction module, used for inputting the first remote sensing image and the second remote sensing image into a Mamba feature extractor respectively, so as to extract a multi-dimensional feature map set of the first remote sensing image and a multi-dimensional feature map set of the second remote sensing image respectively, wherein the multi-dimensional feature map set includes at least a high-dimensional feature map and a low-dimensional feature map;
[0050] A state space change detection module, used to obtain a multi-dimensional long-frequency change feature map set of the first remote sensing image and the second remote sensing image through state space modeling, wherein the multi-dimensional long-frequency change map feature set includes at least a high-dimensional change feature map and a low-dimensional change feature map;
[0051] A low-dimensional feature generation module, used to perform feature fusion on the multi-dimensional long-frequency change map feature set to obtain a low-dimensional fused change feature map;
[0052] A global hybrid attention module, used for performing feature fusion on the high-dimensional change feature map and the low-dimensional fused change feature map to generate a fused output feature map;
[0053] The iterative diffusion module is used to perform noise correction on the fusion output feature map through an iterative diffusion model to generate a remote sensing image change detection map.
[0054] The remote sensing image change detection method and device based on the iterative Mamba architecture proposed in the present invention extracts high-dimensional and low-dimensional feature map sets of remote sensing images by introducing the Mamba feature extractor, making full use of the multi-dimensional information in the image, thereby significantly improving the accuracy of change detection. The state space change detection module is used for modeling, the multi-dimensional long-frequency change features between images are captured and extracted, and the multi-dimensional long-frequency change feature map set is fused through the feature fusion technology to obtain a low-dimensional fused change feature map. This process effectively integrates the complementary information of features of different dimensions and reduces the influence of redundancy and noise. The fusion of high-dimensional change feature maps and low-dimensional fused change feature maps is achieved through the introduction of the global hybrid attention module, ensuring the integrity of important change information and improving the fidelity of change detection. The design of the iterative Mamba architecture gradually optimizes the feature extraction and change detection process in an iterative manner while maintaining high performance, effectively reducing the computational complexity, and can greatly reduce the consumption of computing resources while ensuring detection accuracy, thereby improving the efficiency of change detection, and is suitable for environments with limited resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 It is a flow chart of a remote sensing image change detection method based on iterative Mamba architecture in one embodiment of the present invention;
[0056] Figure 2 It is a schematic block diagram of the structure of a remote sensing image change detection device based on iterative Mamba architecture in one embodiment of the present invention;
[0057] Figure 3 It is a schematic block diagram of the structure of a computer device according to an embodiment of the present invention.
[0058] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0059] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0060] Reference Figure 1 The embodiment of the present invention discloses a remote sensing image change detection method based on iterative Mamba architecture, the method comprising:
[0061] S1, inputting the first remote sensing image and the second remote sensing image into the Mamba feature extractor respectively, so as to extract a multi-dimensional feature map set of the first remote sensing image and a multi-dimensional feature map set of the second remote sensing image respectively, wherein the multi-dimensional feature map set includes at least a high-dimensional feature map and a low-dimensional feature map;
[0062] S2, inputting the multi-dimensional feature map set of the first remote sensing image and the multi-dimensional feature map set of the second remote sensing image into the state space change detection module, and obtaining the multi-dimensional long-frequency change feature map set of the first remote sensing image and the second remote sensing image through state space modeling, wherein the multi-dimensional long-frequency change map feature set includes at least a high-dimensional change feature map and a low-dimensional change feature map;
[0063] S3, performing feature fusion on the multi-dimensional long-frequency change map feature set to obtain a low-dimensional fused change feature map;
[0064] S4, inputting the high-dimensional change feature map and the low-dimensional fused change feature map into a global hybrid attention module for feature fusion to generate a fused output feature map;
[0065] S5. Perform noise correction on the fusion output feature map through an iterative diffusion model to generate a remote sensing image change detection map.
[0066] In this embodiment, in the above step S1, the above first remote sensing image and the second remote sensing image are remote sensing images of the same area at different times or under different conditions. The two images are respectively input into a pre-trained Mamba feature extractor. The above Mamba feature extractor is a model based on a convolutional neural network (CNN) or other deep learning architecture, which can automatically learn and extract low-dimensional and high-dimensional features in an image. The input remote sensing image is subjected to multi-layer convolution, pooling and other operations of the Mamba feature extractor to gradually abstract the multi-dimensional features of the image.
[0067] In the above step S2, the difference between the two images in the feature space is analyzed by the state space modeling technology, and the significant changes in the long time scale are identified, so as to obtain a set of multi-dimensional long-frequency change feature maps. The high-dimensional change feature map reveals the changes at the macro level such as the type of landform; the low-dimensional change feature map captures the changes at the micro level such as the edge and texture.
[0068] In the above step S3, feature fusion is performed on the multi-dimensional long-frequency change feature map set obtained in the S2 stage, and the long-frequency change feature maps of different dimensions are merged to form a low-dimensional fused change feature map containing richer information.
[0069] In the above step S4, the global hybrid attention module not only considers the importance of local features, but also incorporates global context information, making the fused output feature map more comprehensive and accurate.
[0070] In the above step S5, the iterative diffusion model gradually smoothes the noise area in the image by simulating the natural diffusion process of pixel values in the image, while maintaining the edge and detail information of the image.
[0071] The fused output feature map after noise correction is clearer and more accurate, and can better reflect the actual changes in the remote sensing image. The resulting remote sensing image change detection map can be used in many fields such as environmental monitoring, urban planning, and disaster assessment.
[0072] This embodiment introduces the Mamba feature extractor to extract the high-dimensional and low-dimensional feature map sets of remote sensing images, making full use of the multi-dimensional information in the image, thereby significantly improving the accuracy of change detection. The state space change detection module is used for modeling, capturing and extracting the multi-dimensional long-frequency change features between images, and the multi-dimensional long-frequency change feature map set is fused through feature fusion technology to obtain a low-dimensional fused change feature map. This process effectively integrates the complementary information of features of different dimensions and reduces the influence of redundancy and noise. The fusion of high-dimensional change feature maps and low-dimensional fused change feature maps is achieved through the introduction of a global hybrid attention module, ensuring the integrity of important change information and improving the fidelity of change detection. The design of the iterative Mamba architecture gradually optimizes the feature extraction and change detection process through an iterative manner while maintaining high performance, effectively reducing the computational complexity, and can significantly reduce the consumption of computing resources while ensuring detection accuracy, thereby improving the efficiency of change detection, and is suitable for environments with limited resources.
[0073] In a specific embodiment, the Mamba feature extractor includes a linear embedding layer and N encoder layers, the encoder layer includes a VSS block and a patch merging layer, and the step S1 of inputting the first remote sensing image into the Mamba feature extractor to extract a multi-dimensional feature map set of the first remote sensing image includes:
[0074] S101, performing block processing on the first remote sensing image through the linear embedding layer, dividing the first remote sensing image into non-overlapping patches, linearly embedding each patch into a preset feature space, and forming an initial feature map with a set of initial feature vectors of all patches;
[0075] Specifically, the first remote sensing image input is divided into non-overlapping patches, such that each patch is of size Then, the linear embedding layer embeds each patch into a preset feature space, such as a high-dimensional feature space, to form an initial feature map .
[0076]
[0077] S102, inputting the initial feature map into the first encoder layer for encoder layer processing; wherein the encoder layer processing step includes scanning the image features of the feature map through the VSS block, integrating the contextual relationship, and generating the feature map processed by the VSS block; inputting the feature map processed by the VSS block into the patch merging layer, linearly transforming a preset number of adjacent patches in the feature map processed by the VSS block, and splicing them to form a super patch, and the feature vectors of all super patches are set to form a new feature map, and the feature dimension of the new feature map is higher than the feature dimension of the initial feature map;
[0078] The VSS block includes depthwise convolution and SiLU activation function.
[0079]
[0080] In order to reduce the spatial resolution of the feature map layer by layer and increase the number of channels, a patch merging layer is set at the end of each encoder layer. This layer downsamples the features by concatenating adjacent patches and performing linear transformations.
[0081]
[0082] Among them, the patch merging operation can be expressed as:
[0083] = ,
[0084] &
[0085] Through the above steps, the Mamba feature extractor is able to extract multi-dimensional features from the input image and maintain efficient computing performance while capturing long-range dependencies.
[0086] S103, input the new feature map to the next encoder layer, repeat the encoder layer processing steps, until the encoder processing of the N encoder layers is completed, and the set of feature maps output by each encoder layer is used as the multi-dimensional feature map set of the first remote sensing image, and the multi-dimensional feature map set includes at least high-dimensional feature maps and low-dimensional feature maps. After being processed by each encoder layer, the dimension of the feature map of the first remote sensing image gradually increases, and the spatial resolution gradually decreases. Among them, the high-dimensional feature map is the feature map output by the last encoder layer, and relatively, the feature maps with lower dimensions output by other layers are used as low-dimensional feature maps.
[0087] In a specific embodiment, the step of inputting the second remote sensing image into the Mamba feature extractor to extract a multi-dimensional feature map set of the second remote sensing image is the same as the aforementioned steps S101-S103, including:
[0088] S111, performing block processing on the second remote sensing image through the linear embedding layer, dividing the second remote sensing image into non-overlapping patches, linearly embedding each patch into a preset feature space, and forming an initial feature map with a set of initial feature vectors of all patches;
[0089] S112, inputting the initial feature map into the first encoder layer for encoder layer processing; wherein the encoder layer processing step includes scanning the image features of the feature map through the VSS block, integrating the contextual relationship, and generating the feature map processed by the VSS block; inputting the feature map processed by the VSS block into the patch merging layer, linearly transforming a preset number of adjacent patches in the feature map processed by the VSS block, and splicing them to form a super patch, and the feature vectors of all super patches are set to form a new feature map, and the feature dimension of the new feature map is higher than the feature dimension of the initial feature map;
[0090] S113. Input the new feature map to the next encoder layer, repeat the encoder layer processing steps until the encoder processing of the N encoder layers is completed, and use the set of feature maps output by each encoder layer as the multi-dimensional feature map set of the first remote sensing image, wherein the multi-dimensional feature map set includes at least a high-dimensional feature map and a low-dimensional feature map.
[0091] In a specific embodiment, the step S2 of inputting the multi-dimensional feature map set of the first remote sensing image and the multi-dimensional feature map set of the second remote sensing image into a state space change detection module (State Space Change Detection, VSS-CD), and obtaining the multi-dimensional long-frequency change feature map of the first remote sensing image and the second remote sensing image through state space modeling, includes:
[0092] S201, inputting the feature map of the first remote sensing image output by the i-th encoder layer and the feature map of the second remote sensing image output by the i-th encoder layer into the state space change detection module, processing them through the state space model, and outputting the captured i-th layer long-frequency change feature map; wherein 0<i≤N;
[0093] S202, taking the set of long-frequency variation feature maps of each layer as the multi-dimensional long-frequency variation feature set map, outputting the long-frequency variation features of the high-dimensional feature map of the first remote sensing image and the high-dimensional feature map of the second remote sensing image as the high-dimensional long-frequency variation feature map, and outputting the long-frequency variation features of the low-dimensional feature map of the first remote sensing image and the low-dimensional feature map of the second remote sensing image as the low-dimensional long-frequency variation feature map.
[0094] In this embodiment, the VSS-CD module performs state space modeling through a linear time-invariant system (LTI), which specifically includes the following equations:
[0095]
[0096]
[0097] in, Indicates the hidden state, and Respectively represent the first remote sensing image and the second remote sensing image in The characteristics of the layer, , , and are the parameters of the state-space model.
[0098] In order to realize the model in discrete time, it is discretized by zero-order hold (ZOH), and it is obtained:
[0099]
[0100]
[0101] in, , , and is the discretized system matrix, which is calculated as follows:
[0102]
[0103] In actual implementation, the first-order Taylor expansion approximation :
[0104]
[0105] In the VSS-CD module, enter the characteristics and First it passes through a linear embedding layer, then a deep convolution and SiLU activation function, and finally enters the state space model for processing. The final output is Represents the captured changing characteristics.
[0106]
[0107]
[0108]
[0109] Through the above steps, the VSS-CD module can effectively capture the long-frequency change features in the before and after change images and improve the accuracy and fidelity of change detection.
[0110] In a specific embodiment, the step S3 of performing feature fusion on the multi-dimensional long-frequency change feature set to obtain a low-dimensional fused change feature map includes:
[0111] S301, upsampling the high-dimensional change feature map so that the spatial resolution of the high-dimensional change feature map matches the resolution of the low-dimensional change feature map;
[0112] S302: Fusing the upsampled high-dimensional change feature map with the low-dimensional change feature map by element-wise addition to obtain a low-dimensional fused change feature map.
[0113] In this embodiment, in the above step S301, the high-dimensional change feature map is upsampled to increase the number of pixels of the image, thereby improving the spatial resolution of the image to match the spatial resolution of the low-dimensional change feature map for subsequent feature fusion.
[0114] In the above step S302, the upsampled high-dimensional change feature map is fused with the low-dimensional change feature map, and a low-dimensional fused change feature map is generated by element-wise addition, which can directly combine the information of the two feature maps at the same position to generate a low-dimensional fused change feature map containing richer features.
[0115] The low-dimensional fused change feature map not only retains the detail information of the low-dimensional change feature map, but also incorporates the semantic information of the high-dimensional change feature map, which helps to improve the accuracy of subsequent change detection.
[0116] In a specific embodiment, the step S4 of inputting the high-dimensional change feature map and the low-dimensional fused change feature map into the global hybrid attention module for feature fusion to generate a fused output feature map includes:
[0117] S401, performing channel expansion on the high-dimensional change feature map and the low-dimensional fused change feature map, respectively, to obtain an expanded high-dimensional change feature map and an expanded low-dimensional fused change feature map, respectively;
[0118] Assume that the high-dimensional change feature of the input is , the low-dimensional change feature is , first through Convolution expansion channel number:
[0119]
[0120] The expanded high-dimensional change features and low-dimensional change features are and .
[0121] S402: Fusing the expanded high-dimensional change feature map and the expanded low-dimensional fused change feature map to generate a fused feature map.
[0122] Concatenate high-dimensional features and low-dimensional features in the channel dimension , the concatenated feature map .
[0123] Concatenate the global feature map with the concatenated feature map to generate a fused feature map
[0124]
[0125] The final fusion feature map .
[0126] S403, using 1×1 convolution and Sigmoid activation function on the fused feature map to generate a global attention feature map: , global attention feature map .
[0127] S404, dividing the global attention feature map into a high-dimensional attention feature map and a low-dimensional attention feature map;
[0128]
[0129]
[0130] Among them, the high-dimensional attention feature map , low-dimensional attention feature map .
[0131] S405, multiplying the high-dimensional attention feature map and the low-dimensional attention feature map by the initial feature map element by element, respectively, to generate a weighted high-dimensional feature map and a weighted low-dimensional feature map respectively;
[0132]
[0133]
[0134] Weighted high-dimensional feature map , weighted low-dimensional feature map .
[0135] S406: reconstruct the weighted high-dimensional feature map and the weighted low-dimensional feature map to generate the final fused output feature map. .
[0136] In this embodiment, the GHAT module realizes the effective fusion of high-dimensional and low-dimensional features through the above steps, and improves the accuracy and meticulousness of change detection. Through the cross-attention mechanism, GHAT can capture the interactive relationship between features in a global scope and generate a more accurate and reliable change feature map.
[0137] The step S5 of performing noise correction on the fusion output feature map by using an iterative diffusion model to generate a remote sensing image change detection map comprises:
[0138] S501, obtaining an initialized input noise feature as an iteration starting point;
[0139] The input noise feature is usually a randomly generated matrix with the same size as the fused output feature map, and the element values follow a certain distribution (such as Gaussian distribution or uniform distribution) to simulate the initial noise in the image. These noise features serve as the starting point of the iterative diffusion model and are used in the subsequent forward and reverse iterations.
[0140] S502, performing multiple iterations of optimization on the fusion output feature map through an iterative diffusion model, and finally generating the remote sensing image change detection map, wherein each iteration of the optimization includes a forward process and a reverse process,
[0141] In the forward process, the input noise feature is gradually added to the change feature map of the input iterative forward process until a pure noise map is accumulated;
[0142] In the reverse process, the noise estimation model is used to perform noise estimation on the image input to the reverse process to generate a noise estimation value, and the denoised image is calculated according to the reverse denoising formula.
[0143] Through the above forward and reverse iterative processes, the iterative diffusion model decomposes the complex image generation problem into a series of simple denoising sub-problems by gradually adding and removing noise, effectively correcting the noise in the fusion output feature map. In each iteration, the model generates a denoised image based on the current image state and noise characteristics. The forward process simulates the accumulation process of noise and provides the necessary information for reverse denoising; the reverse process gradually removes the noise in the image through noise estimation and reverse denoising formulas, and restores a clear remote sensing image change detection map, thereby improving the signal-to-noise ratio of the remote sensing image change detection map and retaining the detail information in the remote sensing image change detection map.
[0144] In a specific embodiment, the input noise characteristic of the initialization is assumed to be , after T iterations, the final change detection graph is obtained The process of each iteration can be expressed as:
[0145]
[0146] in, represents noise, and are diffusion model parameters, is the noise estimate.
[0147] In the forward process, noise is gradually added to the image, so that it gradually changes from the original image to pure noise. Suppose the original image is , after T iterations, the final noisy image is obtained . The forward process can be expressed as:
[0148]
[0149] in, is the dimension parameter of the noise. By accumulating the noise addition process, the noise image at any time can be obtained:
[0150]
[0151] in, .
[0152] In a specific embodiment, in the reverse process, the noise is gradually removed and the noisy image is restored to the original image. The reverse process is represented by the following formula:
[0153]
[0154] in, is the estimated mean, is the estimated variance.
[0155] In the reverse process, the key is to accurately estimate the noise. Suppose the image at the current moment is , the model is estimated through the noise network Generate a noise estimate:
[0156]
[0157] Among them, E and D represent encoder and decoder respectively, GHAT represents the global hybrid attention mechanism, and Represent low-dimensional and high-dimensional change characteristics respectively, Represents the noise characteristics.
[0158] Based on the noise estimate, calculate the denoised image in the reverse process:
[0159]
[0160] The final inverse denoising formula is:
[0161]
[0162]
[0163] Where t is the number of iterations, represents noise, and are diffusion model parameters, is the noise estimate, is the mean square error of the noise estimate, is the denoised image after the t-1th iteration.
[0164] In a specific embodiment, the noise estimation model adopts a Unet structure, including a second encoder and a second decoder.
[0165] The second encoder receives input features, the input features including input noise features and input change feature maps, extracts layer by layer and generates output feature maps of each layer, wherein the output feature maps of each layer include different dimensional information of the input features;
[0166] The second decoder is jump-connected to the second encoder, and the output feature map of the corresponding layer in the second encoder is passed to the corresponding layer in the second decoder, and the resolution of the output feature map is restored layer by layer until the difference between the resolution of the output feature map and the input change feature map is less than a preset threshold, thereby obtaining a preliminary change detection map;
[0167] The preliminary change detection map is taken as input through an iterative optimization mechanism, and the preliminary change detection map is iteratively processed multiple times through multi-dimensional features and a global hybrid attention mechanism, and the change detection results are gradually corrected until a preset iteration stop condition is met, and the final remote sensing image change detection map is output.
[0168] Specifically, in the noise estimation model, the key to noise estimation is to encode and decode the input noise features and change features to generate a high-precision change detection map. Suppose the input noise feature is , the change characteristics are First, the input feature map is processed by the second encoder to generate a multi-dimensional feature representation:
[0169]
[0170]
[0171] in, Represents the output feature map of the i-th layer of the second encoder.
[0172] In the second decoder part, the multi-dimensional features of the second encoder are gradually restored to the original resolution through upsampling and skip connections:
[0173]
[0174]
[0175]
[0176]
[0177] Finally, the second decoder outputs a high-precision remote sensing image change detection map:
[0178]
[0179] In order to further improve the accuracy of change detection, NEUNet gradually corrects and refines the change detection results through an iterative optimization mechanism. In each iteration, the change detection results are gradually corrected through multi-dimensional features and a global hybrid attention mechanism, and finally a high-precision remote sensing image change detection map is generated.
[0180] Reference Figure 2 An embodiment of the present invention further provides a remote sensing image change detection device based on iterative Mamba architecture, the device comprising:
[0181] The feature extraction module 10 is used to input the first remote sensing image and the second remote sensing image into the Mamba feature extractor respectively, so as to extract a multi-dimensional feature map set of the first remote sensing image and a multi-dimensional feature map set of the second remote sensing image respectively, wherein the multi-dimensional feature map set includes at least a high-dimensional feature map and a low-dimensional feature map;
[0182] The state space change detection module 20 is used to obtain a multi-dimensional long-frequency change feature map set of the first remote sensing image and the second remote sensing image through state space modeling, wherein the multi-dimensional long-frequency change map feature set includes at least a high-dimensional change feature map and a low-dimensional change feature map;
[0183] A low-dimensional feature generation module 30 is used to perform feature fusion on the multi-dimensional long-frequency change map feature set to obtain a low-dimensional fused change feature map;
[0184] A global hybrid attention module 40, used for performing feature fusion on the high-dimensional change feature map and the low-dimensional fused change feature map to generate a fused output feature map;
[0185] The iterative diffusion module 50 is used to perform noise correction on the fusion output feature map through an iterative diffusion model to generate a remote sensing image change detection map.
[0186] In this embodiment, for the specific implementation of each module in the above-mentioned device embodiment, please refer to the above-mentioned method embodiment, which will not be described in detail here.
[0187] Reference Figure 3 In an embodiment of the present invention, a computer device is also provided. The computer device may be a server, and its internal structure may be as follows: Figure 3 As shown. The computer device includes a processor, a memory, a display screen, an input device, a network interface and a database connected through a system bus. Among them, the processor designed by the computer is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store the corresponding data in this embodiment. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, the above-mentioned method for remote sensing image change detection based on the iterative Mamba architecture is implemented.
[0188] Those skilled in the art will understand that Figure 3 The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present invention, and does not constitute a limitation on the computer device to which the solution of the present invention is applied.
[0189] An embodiment of the present invention further provides a computer-readable storage medium on which a computer program is stored, and when the computer program is executed by a processor, the above method is implemented. It can be understood that the computer-readable storage medium in this embodiment can be a volatile readable storage medium or a non-volatile readable storage medium.
[0190] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media provided by the present invention and used in the embodiments may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double-speed data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM.
[0191] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, device, article or method including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, device, article or method. In the absence of further restrictions, an element defined by the sentence "includes a ..." does not exclude the presence of other identical elements in the process, device, article or method including the element.
[0192] The above description is only a preferred embodiment of the present invention, and does not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A remote sensing image change detection method based on iterative Mamba architecture, characterized in that: The method comprises: Inputting the first remote sensing image and the second remote sensing image into the Mamba feature extractor respectively to extract a multi-dimensional feature map set of the first remote sensing image and a multi-dimensional feature map set of the second remote sensing image respectively, wherein the multi-dimensional feature map set includes at least a high-dimensional feature map and a low-dimensional feature map; Inputting the multi-dimensional feature map set of the first remote sensing image and the multi-dimensional feature map set of the second remote sensing image into the state space change detection module, and obtaining a multi-dimensional long-frequency change feature map set composed of the first remote sensing image and the second remote sensing image through state space modeling, wherein the multi-dimensional long-frequency change feature map set includes at least a high-dimensional change feature map and a low-dimensional change feature map; Performing feature fusion on the multi-dimensional long-frequency change feature map set to obtain a low-dimensional fused change feature map; Inputting the high-dimensional change feature map and the low-dimensional fused change feature map into a global hybrid attention module for feature fusion to generate a fused output feature map; Performing noise correction on the fusion output feature map through an iterative diffusion model to generate a remote sensing image change detection map; The Mamba feature extractor includes a linear embedding layer and N encoder layers, the encoder layer includes a VSS block and a patch merging layer, and the step of inputting the first remote sensing image into the Mamba feature extractor to extract a multi-dimensional feature map set of the first remote sensing image includes: The first remote sensing image is processed into non-overlapping patches by the linear embedding layer, each patch is linearly embedded into a preset feature space, and the initial feature vectors of all patches are set to form an initial feature map; The initial feature map is input into the first encoder layer for encoder layer processing; wherein the encoder layer processing step includes scanning the image features of the feature map through the VSS block, integrating the contextual relationship, and generating the feature map processed by the VSS block; the feature map processed by the VSS block is input into the patch merging layer, a preset number of adjacent patches in the feature map processed by the VSS block are linearly transformed, and spliced to form a super patch, and the feature vectors of all super patches are set to form a new feature map, and the feature dimension of the new feature map is higher than the feature dimension of the initial feature map; The new feature map is input into the next encoder layer, and the steps of the encoder layer processing are repeated until the encoder processing of N encoder layers is completed, and the set of feature maps output by each encoder layer is used as the multi-dimensional feature map set of the first remote sensing image, and the multi-dimensional feature map set includes at least a high-dimensional feature map and a low-dimensional feature map.
2. The remote sensing image change detection method based on iterative Mamba architecture according to claim 1 is characterized in that: The step of inputting the multidimensional feature map set of the first remote sensing image and the multidimensional feature map set of the second remote sensing image into the state space change detection module, and obtaining the multidimensional long-frequency change feature map set composed of the first remote sensing image and the second remote sensing image through state space modeling, comprises: Inputting the feature map of the first remote sensing image output by the i-th encoder layer and the feature map of the second remote sensing image output by the i-th encoder layer into the state space change detection module, processing them through the state space model, and outputting the captured i-th layer long-frequency change feature map; wherein 0<i≤N; The set of long-frequency variation feature maps of each layer is used as the multi-dimensional long-frequency variation feature map set, the long-frequency variation feature output of the high-dimensional feature map of the first remote sensing image and the high-dimensional feature map of the second remote sensing image is used as the high-dimensional long-frequency variation feature map, and the long-frequency variation feature output of the low-dimensional feature map of the first remote sensing image and the low-dimensional feature map of the second remote sensing image is used as the low-dimensional long-frequency variation feature map.
3. The remote sensing image change detection method based on iterative Mamba architecture according to claim 2 is characterized in that: The step of performing feature fusion on the multi-dimensional long-frequency change feature map set to obtain a low-dimensional fused change feature map comprises: Upsampling the high-dimensional change feature map so that the spatial resolution of the high-dimensional change feature map matches the resolution of the low-dimensional change feature map; The upsampled high-dimensional change feature map is fused with the low-dimensional change feature map by element-wise addition to obtain a low-dimensional fused change feature map.
4. The remote sensing image change detection method based on iterative Mamba architecture according to claim 1, characterized in that: The step of inputting the high-dimensional change feature map and the low-dimensional fused change feature map into the global hybrid attention module for feature fusion to generate a fused output feature map comprises: Perform channel expansion on the high-dimensional change feature map and the low-dimensional fused change feature map, respectively, to obtain an expanded high-dimensional change feature map and an expanded low-dimensional fused change feature map, respectively; Performing feature fusion on the expanded high-dimensional change feature map and the expanded low-dimensional fused change feature map to generate a fused feature map; Using 1×1 convolution and Sigmoid activation function on the fused feature map to generate a global attention feature map; Splitting the global attention feature map into a high-dimensional attention feature map and a low-dimensional attention feature map; Multiplying the high-dimensional attention feature map and the low-dimensional attention feature map by the initial feature map element by element, respectively, to generate a weighted high-dimensional feature map and a weighted low-dimensional feature map, respectively; The weighted high-dimensional feature map and the weighted low-dimensional feature map are reconstructed to generate the final fused output feature map.
5. The remote sensing image change detection method based on iterative Mamba architecture according to claim 1, characterized in that: The step of performing noise correction on the fused output feature map by an iterative diffusion model to generate a remote sensing image change detection map comprises: Obtain the initialized input noise feature as the starting point of iteration; The fusion output feature map is iteratively optimized multiple times through an iterative diffusion model to finally generate the remote sensing image change detection map, wherein each iterative optimization includes a forward process and a reverse process. In the forward process, the input noise feature is gradually added to the change feature map of the input iterative forward process until a pure noise map is accumulated; In the reverse process, the noise estimation model is used to perform noise estimation on the image input to the reverse process to generate a noise estimation value, and the denoised image is calculated according to the reverse denoising formula.
6. The remote sensing image change detection method based on iterative Mamba architecture according to claim 5, characterized in that: The noise estimate is: Where t is the number of iterations, and is the changing feature, D represents the decoder of the noise estimation model, GHAT represents the global hybrid attention mechanism, and Represent low-dimensional and high-dimensional change characteristics respectively, is the input noise characteristic.
7. The remote sensing image change detection method based on iterative Mamba architecture according to claim 5, characterized in that: The reverse denoising formula is: Where t is the number of iterations, represents noise, and are diffusion model parameters, is the noise estimate, is the mean square error of the noise estimate, is the denoised image after the t-1th iteration.
8. The remote sensing image change detection method based on iterative Mamba architecture according to claim 5, characterized in that: The noise estimation model adopts a Unet structure, including a second encoder and a second decoder. The second encoder receives input features, the input features including input noise features and input change feature maps, extracts layer by layer and generates output feature maps of each layer, wherein the output feature maps of each layer include different dimensional information of the input features; The second decoder is jump-connected to the second encoder, and the output feature map of the corresponding layer in the second encoder is passed to the corresponding layer in the second decoder, and the resolution of the output feature map is restored layer by layer until the difference between the resolution of the output feature map and the input change feature map is less than a preset threshold, thereby obtaining a preliminary change detection map; The preliminary change detection map is taken as input through an iterative optimization mechanism, and the preliminary change detection map is iteratively processed multiple times through multi-dimensional features and a global hybrid attention mechanism, and the change detection results are gradually corrected until a preset iteration stop condition is met, and the final remote sensing image change detection map is output.
9. A remote sensing image change detection device based on iterative Mamba architecture, characterized in that: The device comprises: A feature extraction module, used for inputting the first remote sensing image and the second remote sensing image into a Mamba feature extractor respectively, so as to extract a multi-dimensional feature map set of the first remote sensing image and a multi-dimensional feature map set of the second remote sensing image respectively, wherein the multi-dimensional feature map set includes at least a high-dimensional feature map and a low-dimensional feature map; A state space change detection module, used to obtain a multi-dimensional long-frequency change feature map set consisting of the first remote sensing image and the second remote sensing image through state space modeling, wherein the multi-dimensional long-frequency change feature map set includes at least a high-dimensional change feature map and a low-dimensional change feature map; A low-dimensional feature generation module is used to perform feature fusion on the multi-dimensional long-frequency change feature map set to obtain a low-dimensional fused change feature map; A global hybrid attention module, used for performing feature fusion on the high-dimensional change feature map and the low-dimensional fused change feature map to generate a fused output feature map; An iterative diffusion module, used for performing noise correction on the fusion output feature map through an iterative diffusion model to generate a remote sensing image change detection map; The Mamba feature extractor includes a linear embedding layer and N encoder layers, the encoder layer includes a VSS block and a patch merging layer, and the feature extraction module performs the step of inputting the first remote sensing image into the Mamba feature extractor to extract a multi-dimensional feature map set of the first remote sensing image, including: The first remote sensing image is processed into non-overlapping patches by the linear embedding layer, each patch is linearly embedded into a preset feature space, and the initial feature vectors of all patches are set to form an initial feature map; The initial feature map is input into the first encoder layer for encoder layer processing; wherein the encoder layer processing step includes scanning the image features of the feature map through the VSS block, integrating the contextual relationship, and generating the feature map processed by the VSS block; the feature map processed by the VSS block is input into the patch merging layer, a preset number of adjacent patches in the feature map processed by the VSS block are linearly transformed, and spliced to form a super patch, and the feature vectors of all super patches are set to form a new feature map, and the feature dimension of the new feature map is higher than the feature dimension of the initial feature map; The new feature map is input into the next encoder layer, and the steps of the encoder layer processing are repeated until the encoder processing of N encoder layers is completed, and the set of feature maps output by each encoder layer is used as the multi-dimensional feature map set of the first remote sensing image, and the multi-dimensional feature map set includes at least a high-dimensional feature map and a low-dimensional feature map.
Citation Information
Patent Citations
Remote sensing image change detection method based on twinborn multi-scale difference feature fusion
CN113420662A
Remote sensing image change detection method and device, electronic equipment and storage medium
CN117789046A