Image change detection method, apparatus, equipment, and medium based on state space
By using a state-space-based Siamese network architecture, and leveraging lightweight state-space modules and pixel-level differential fusion processing, the shortcomings of spatiotemporal information processing in remote sensing image change detection are addressed, achieving efficient change detection.
Patent Information
- Application Number
- CN202511114614.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-11
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-08-11
Smart Images

Figure CN120612496B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular to an image change detection method, apparatus, device and medium based on state space. Background Technology
[0002] Remote sensing image change detection identifies changes in ground features by comparing remote sensing images from different times and generates change mask images. With the development of deep learning, convolutional neural networks (CNNs) and transformer architectures have become mainstream methods, but these methods have limitations in capturing large-scale spatial information and computational efficiency.
[0003] Current state-space model-based visual models (such as VMamba and ChangeMamba) have improved computational efficiency, but still face challenges such as image heterogeneity differences, background semantic variations, and non-target semantic interference.
[0004] Therefore, how to accurately process the spatiotemporal information in remote sensing images and improve the accuracy of change detection is a technical problem that urgently needs to be solved. Summary of the Invention
[0005] This application provides a state-space-based image change detection method, apparatus, device, and medium, which achieves the technical effect of accurately processing spatiotemporal information in remote sensing images and improving the accuracy of change detection.
[0006] To achieve the above objectives, the main technical solutions adopted in this application include:
[0007] In a first aspect, embodiments of this application provide an image change detection method based on state space, the method comprising:
[0008] Acquire dual-temporal remote sensing images; the dual-temporal remote sensing images include a first-temporal remote sensing image before the change and a second-temporal remote sensing image after the change; the dual-temporal remote sensing images are obtained by taking pictures of the same geographic area at at least two different time points;
[0009] The dual-temporal remote sensing images are respectively input into two weight-shared Siamese network branches, and encoded through a lightweight state space module with multiple encoding stages to obtain multiple first-temporal feature maps at different scales corresponding to the first-temporal remote sensing image, and multiple second-temporal feature maps at different scales corresponding to the second-temporal remote sensing image; the lightweight state space module is used to extract global features and remove redundant features.
[0010] The first temporal feature map and the second temporal feature map are input into the temporal state space feature fusion module. Through pixel-level difference and fusion processing, the temporal state space fused feature is obtained. The temporal state space feature fusion module is used to fuse spatial features from different coding stages.
[0011] The spatiotemporal state space feature refinement module in multiple decoding stages decodes the spatiotemporal state space fusion features, the first spatiotemporal feature map, and the second spatiotemporal feature map to obtain the target change mask map corresponding to the dual-temporal remote sensing image; the spatiotemporal state space feature refinement module is used to refine and restore spatial details.
[0012] This embodiment provides a state-space-based image change detection method, apparatus, device, and medium. By acquiring dual-temporal remote sensing images, including both before and after a change, it reflects the geographical features of different time points, providing foundational data for subsequent change detection. During processing, a lightweight state-space module across multiple coding stages of a Siamese network extracts feature maps at different scales, effectively extracting global features and eliminating redundant information, avoiding unnecessary noise interference. Next, a temporal state-space feature fusion module, through pixel-level difference and fusion processing, combines the feature maps from both temporal stages to achieve deep fusion of spatial features from different coding stages, thereby enhancing the expressive power of spatiotemporal information. Finally, a spatiotemporal state-space feature refinement module across multiple decoding stages refines and restores the fused features, further extracting spatiotemporal change features from the image and generating a high-precision change mask. This refinement process ensures that each stage can more accurately capture spatiotemporal changes in the image, and the final output change mask more precisely reflects the changed areas in the image, effectively addressing the shortcomings of traditional remote sensing image change detection methods in spatiotemporal information processing and significantly improving the accuracy of change detection.
[0013] In one embodiment, the lightweight state space module includes a first feature transformation component, a lightweight network layer, a second feature transformation component, and a third feature transformation component connected in sequence; the input position of the first feature transformation component is connected to the output position of the second feature transformation component via a jump, and the input position of the second feature transformation component is connected to the output position of the third feature transformation component via a jump.
[0014] In one implementation, the encoding process for the lightweight state space module at any encoding stage is performed as follows:
[0015] The first feature transformation component performs a first feature transformation on the received input image to obtain a first feature transformation image; the first feature transformation component includes a layer normalization layer, a fully connected layer, and a depthwise convolution layer connected in sequence.
[0016] The first feature transformation image is divided into two parts proportionally by the lightweight network layer. The first part of the first feature transformation image is cross-scanned a specified number of times to obtain a global feature map containing global features. The second part of the first feature transformation image is subjected to identity mapping to obtain a local feature map with redundant features removed. The global feature map and the local feature map are then concatenated by channels to obtain a lightweight feature map.
[0017] The lightweight feature map is subjected to a second feature transformation by the second feature transformation group to obtain a second feature transformation image; the second feature transformation component includes a layer normalization layer and a fully connected layer connected in sequence.
[0018] The response is based on the first stitched image obtained by stitching the input image received by the lightweight state space module and the second feature transformation image. The third feature transformation component performs a third feature transformation on the first stitched image to obtain a third feature transformation image. The third feature transformation component includes a depthwise convolution, a layer normalization layer and a fully connected layer connected in sequence.
[0019] The response is based on the second stitched image obtained by stitching the first stitched image and the third feature transformation image, the second stitched image is determined as the temporal feature map of the encoding stage corresponding to the lightweight state space module, and the temporal feature map is used as the output image.
[0020] This embodiment extracts preliminary features from the input image through a first feature transformation component and uses a lightweight network layer to segment the feature map, thereby obtaining a global feature map and a local feature map. This helps to capture global and local changes in the image. Next, a second feature transformation component further optimizes the feature representation. Combined with a stitching strategy based on a lightweight network layer, it can fuse the spatiotemporal information of the image, thereby improving the accuracy of change detection. Finally, after processing by a third feature transformation component, a temporal feature map is generated as the output image of the lightweight state space module. This process effectively utilizes feature transformation and stitching strategies at different levels to extract more accurate spatiotemporal features from remote sensing images, providing more precise information for change detection and solving the problem that traditional methods struggle to handle complex spatiotemporal relationships.
[0021] In one embodiment, the step of inputting the first temporal feature map and the second temporal feature map into a temporal state space feature fusion module, and obtaining temporal state space fused features through pixel-level difference and fusion processing, includes:
[0022] The first temporal feature map of the penultimate coding stage and the second temporal feature map of the penultimate coding stage are processed by pixel-level difference to obtain the first difference feature map corresponding to the penultimate coding stage.
[0023] The first temporal feature map and the first difference feature map of the penultimate encoding stage are fused to obtain a first temporal fused feature map; and the second temporal feature map and the first difference feature map of the penultimate encoding stage are fused to obtain a second temporal fused feature map.
[0024] The first phase feature map and the second phase feature map of the last coding stage of the plurality of coding stages are processed by pixel-level difference to obtain the second difference feature map corresponding to the last coding stage.
[0025] The first phase feature map and the second difference feature map of the last encoding stage are fused to obtain a third phase fused feature map; and the second phase feature map and the second difference feature map of the last encoding stage are fused to obtain a fourth phase fused feature map.
[0026] Based on the first, second, third, and fourth temporal fusion feature maps, a fusion convolution process is performed to obtain the temporal state space fusion feature.
[0027] This embodiment uses pixel-level differential processing to obtain difference feature maps at different encoding stages, which helps to capture subtle temporal changes in the image. Subsequently, by fusing the temporal feature maps with the difference feature maps, the spatiotemporal feature representation of the image is further enhanced, making the features of each temporal phase more comprehensive and reducing the spatiotemporal information that a single feature map might miss. This layer-by-layer feature fusion and temporal differential strategy can more accurately handle complex spatiotemporal relationships in remote sensing images, and has significant application value, especially in change detection tasks. By performing differential and fusion operations separately at different encoding stages, more refined spatiotemporal features can be effectively extracted. Finally, through fusion convolution processing, a comprehensive temporal state-space fusion feature is generated, improving the accuracy and precision of change detection, thus solving the problem that traditional methods struggle to accurately capture complex spatiotemporal changes when processing remote sensing images. This technique can significantly improve the effect of change detection in remote sensing images and is suitable for practical application scenarios requiring extremely precise spatiotemporal information.
[0028] In one embodiment, the spatiotemporal state space feature refinement module includes a lightweight state space module, a first convolutional component, and a second convolutional component connected in sequence; the input position of the first convolutional component is connected to the output position of the second convolutional component via a jump.
[0029] In one implementation, the decoding process for the spatiotemporal state space feature refinement module at any decoding stage is performed as follows:
[0030] The lightweight state space module extracts features from the received input features to obtain a lightweight graph.
[0031] A first spatial feature map is obtained by performing a first convolution operation on the lightweight map using the first convolution component; the first convolution component includes a 1×1 convolutional layer.
[0032] The second convolutional component performs a second convolutional operation on the first spatial feature map to obtain a second spatial feature map; the second convolutional component includes a 1×1 convolutional layer, a 3×3 convolutional layer, a normalization layer and a 1×1 convolutional layer connected in sequence.
[0033] The response is based on a new spatial feature map obtained by stitching together the first spatial feature map and the second spatial feature map. The new spatial feature map is determined as the change mask map corresponding to the decoding stage of the spatiotemporal state spatial feature refinement module, and the change mask map is used as the output image.
[0034] This embodiment refines the input features using a lightweight state space module and multi-layer convolutional operations (1×1 convolutional layers, 3×3 convolutional layers, etc.), enabling the effective extraction and fusion of different spatial features. In this process, the combined design of the first and second convolutional components not only captures image details but also enhances the expression of spatiotemporal features. Finally, a new spatial feature map generated by stitching together the first and second spatial feature maps further improves the accuracy of change detection results. The change mask map in this process, as the output image, more accurately reflects the spatiotemporal changes in remote sensing images, achieving efficient processing and accurate detection of complex spatiotemporal information. This solves the problem of capturing spatiotemporal changes in traditional methods when processing remote sensing images, significantly improving the overall performance of the change detection task.
[0035] In one embodiment, the method further includes:
[0036] The output image of the lightweight state space module in the encoding stage is determined as the corresponding image of the input image in the decoding stage;
[0037] The stitched image obtained by stitching the first spatial feature map and the corresponding image is used as the input image of the second convolutional component.
[0038] In one implementation, the temporal state space fusion features, the first temporal feature map, and the second temporal feature map are decoded by a spatiotemporal state space feature refinement module in multiple decoding stages to obtain a target change mask map corresponding to the dual-temporal remote sensing image, including:
[0039] The temporal state space fusion features are decoded through the spatiotemporal state space feature refinement module of the first decoding stage in the decoding stage to obtain the initial change mask map;
[0040] The initial change mask image, the first temporal feature image and the second temporal feature image corresponding to the encoding stage are decoded by the spatiotemporal state space feature refinement module of the next decoding stage in the decoding stage until all the decoding stages are completed, and the target change mask image is obtained.
[0041] This embodiment decodes the spatiotemporal state space fusion features using the spatiotemporal state space feature refinement module in the first decoding stage to initially obtain a change mask map. Then, the initial change mask map, along with the first and second temporal feature maps output from the corresponding encoding stage, are input into the spatiotemporal state space feature refinement module in the next decoding stage for processing. This process is repeated until all decoding stages are completed, thereby obtaining the target change mask map. This progressive refinement method allows each decoding stage to more accurately extract spatiotemporal change features from the image. The resulting change mask map more precisely reflects the changed areas in the image, improving the accuracy of change detection and meeting the needs of high-precision remote sensing image processing.
[0042] In one embodiment, each of the twin network branches includes a first-stage coding unit, at least one second-stage coding unit, and a third-stage coding unit connected in sequence; both the second-stage coding unit and the third-stage coding unit include the lightweight state-space module; the input data of the first-stage coding unit is the dual-temporal remote sensing image, the input data of the initial second-stage coding unit includes the initial feature map output by the first-stage coding unit; the input data of subsequent second-stage coding units includes the temporal feature map output by the preceding second-stage coding unit; and the input data of the third-stage coding unit is the temporal feature map output by the last second-stage coding unit.
[0043] Specifically, for any twin network branch, the first-stage coding unit extracts features from the dual-temporal remote sensing image to obtain the corresponding initial feature map;
[0044] The initial feature map is sequentially input into at least one second-stage coding unit and the third-stage coding unit for feature extraction, resulting in multiple corresponding temporal feature maps at different scales.
[0045] Secondly, embodiments of this application provide a state-space-based image change detection device, the device comprising:
[0046] An image acquisition unit is used to acquire dual-temporal remote sensing images; the dual-temporal remote sensing images include a first-temporal remote sensing image before the change and a second-temporal remote sensing image after the change; the dual-temporal remote sensing images are obtained by taking pictures of the same geographic area at at least two different time points;
[0047] The encoding processing unit is used to input the dual-temporal remote sensing images into two weight-shared Siamese network branches, and perform encoding processing through a lightweight state space module with multiple encoding stages to obtain multiple first-temporal feature maps of different scales corresponding to the first-temporal remote sensing image and multiple second-temporal feature maps of different scales corresponding to the second-temporal remote sensing image; the lightweight state space module is used to extract global features and remove redundant features.
[0048] The temporal fusion unit is used to input the first temporal feature map and the second temporal feature map into the temporal state space feature fusion module, and obtain the temporal state space fusion feature through pixel-level difference and fusion processing; the temporal state space feature fusion module is used to fuse spatial features of different coding stages;
[0049] The decoding processing unit is used to decode the temporal state space fusion features, the first temporal feature map, and the second temporal feature map through a spatiotemporal state space feature refinement module in multiple decoding stages to obtain the target change mask map corresponding to the dual-temporal remote sensing image; the spatiotemporal state space feature refinement module is used to refine and restore spatial details.
[0050] Thirdly, embodiments of this application provide a computer device, including:
[0051] The system includes a memory and a processor, which are interconnected. The memory stores computer instructions, and the processor executes these computer instructions to perform the state-space-based image change detection method described above.
[0052] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer instructions, which are used to cause a computer to execute the state-space-based image change detection method described above. Attached Figure Description
[0053] To more clearly illustrate the technical solutions in the specific embodiments of this application or the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0054] Figure 1 A flowchart illustrating a state-space-based image change detection method, apparatus, device, and medium provided in this application embodiment;
[0055] Figure 2 A schematic diagram of two weight-sharing twin network branches provided in an embodiment of this application;
[0056] Figure 3 A flowchart illustrating the encoding process of a lightweight state space module for any encoding stage, as provided in this application embodiment;
[0057] Figure 4 A schematic diagram of cross-scan provided for an embodiment of this application;
[0058] Figure 5 A schematic diagram of a lightweight state space module provided in an embodiment of this application;
[0059] Figure 6 A flowchart of step S5 provided in an embodiment of this application;
[0060] Figure 7 A schematic diagram of the temporal state-space feature fusion module provided in an embodiment of this application;
[0061] Figure 8 A flowchart illustrating the decoding process of the spatiotemporal state space feature refinement module for any decoding stage, provided in an embodiment of this application.
[0062] Figure 9 A schematic diagram of the spatiotemporal state space feature refinement module provided in an embodiment of this application;
[0063] Figure 10 A flowchart of step S7 provided in an embodiment of this application;
[0064] Figure 11 The figure shown is a comparison of the experimental results of binary change detection on the WHU-CD change detection dataset;
[0065] Figure 12 The figure shown is a comparison of the experimental results of binary change detection on the LEVIR-CD change detection dataset;
[0066] Figure 13A block diagram of a state-space-based image change detection device provided in this application embodiment;
[0067] Figure 14 A schematic diagram of the structure of a computer device provided in an embodiment of the present application. Detailed Implementation
[0068] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0069] Remote sensing image change detection is a key technology widely used in fields such as natural disaster assessment, urban planning, environmental surveys, and land monitoring. Its main objective is to identify changes in ground features and generate change mask images by comparing remote sensing images taken at different times. With the development of computer technology, change detection methods have evolved from traditional machine learning to deep learning. Deep neural networks have gradually replaced traditional manual feature extraction methods and further improved the accuracy of change detection by learning labeled change features.
[0070] Current deep learning methods are mainly based on Convolutional Neural Networks (CNNs) and Transformer architectures, employing an encoder-decoder structure for image change recognition. In the encoder stage, the network extracts deep semantic information by compressing features; in the decoder stage, operations such as upsampling restore the dimensionality of the feature map, ultimately outputting a mask image for change detection. However, these methods have certain limitations. For example, CNNs are limited by a finite receptive field, making it difficult to capture a wide range of spatial contextual information; while Transformers, although capable of capturing longer-range dependencies, have high computational costs, leading to high training and deployment costs.
[0071] To overcome these shortcomings, state-space model-based visual models (such as VMamba and ChangeMamba) have recently been proposed, demonstrating excellent performance in natural language processing and visual tasks by effectively reducing computational cost and improving model performance. The ChangeMamba method combines spatiotemporal modeling mechanisms to achieve accurate change detection through three frameworks. Although these methods improve computational efficiency to some extent, they still face challenges such as image heterogeneity differences, background semantic variations, and non-target semantic interference, affecting the accuracy of change detection.
[0072] Therefore, how to accurately process the spatiotemporal information in remote sensing images and improve the accuracy of change detection results is a technical problem that urgently needs to be solved.
[0073] To address the aforementioned technical problems, according to embodiments of this application, an image change detection method, apparatus, device, and medium based on state space are provided. It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowcharts, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0074] This embodiment provides a state-space-based image change detection method, apparatus, device, and medium. Figure 1 A flowchart illustrating a state-space-based image change detection method, apparatus, device, and medium provided in this application embodiment is shown below. Figure 1 As shown, the process includes the following steps:
[0075] Step S1: Acquire dual-temporal remote sensing images; dual-temporal remote sensing images include a first-temporal remote sensing image before the change and a second-temporal remote sensing image after the change; dual-temporal remote sensing images are obtained by taking pictures of the same geographic area at at least two different time points.
[0076] Specifically, dual-temporal remote sensing images refer to two remote sensing images of the same geographic area taken at different times, including a first-temporal remote sensing image before the change and a second-temporal remote sensing image after the change. These two images represent the state of the same geographic area at two different points in time.
[0077] Step S3: Input the dual-temporal remote sensing images into two weight-shared Siamese network branches respectively, and encode them through a lightweight state space module with multiple encoding stages to obtain multiple first-temporal feature maps of different scales corresponding to the first-temporal remote sensing image and multiple second-temporal feature maps of different scales corresponding to the second-temporal remote sensing image; the lightweight state space module is used to extract global features and remove redundant features.
[0078] Specifically, dual-temporal remote sensing images (including a first-temporal image before the change and a second-temporal image after the change) are input into two weight-sharing Siamese network branches. Each branch is encoded using lightweight state-space modules with multiple encoding stages. These modules efficiently extract global features from the images while removing redundant information, thereby reducing computational cost and improving processing speed. Through this process, the network generates multiple temporal feature maps at different scales corresponding to the first and second-temporal images. These feature maps not only contain rich semantic information but also ensure the compactness and effectiveness of the features, providing high-quality input for subsequent change detection tasks.
[0079] In one implementation, each Siamese network branch includes a first-stage coding unit, at least one second-stage coding unit, and a third-stage coding unit connected in sequence. Both the second-stage and third-stage coding units include a lightweight state-space module. The input data for the first-stage coding unit is a dual-temporal remote sensing image. The input data for the initial second-stage coding unit includes the initial feature map output by the first-stage coding unit. The input data for subsequent second-stage coding units includes the temporal feature map output by the preceding second-stage coding unit. The input data for the third-stage coding unit is the temporal feature map output by the last second-stage coding unit. Specifically, for any Siamese network branch, feature extraction is performed on the dual-temporal remote sensing image using the first-stage coding unit to obtain a corresponding initial feature map. The initial feature map is then sequentially input into at least one second-stage coding unit and one third-stage coding unit for feature extraction, resulting in multiple temporal feature maps at different scales.
[0080] Specifically, please refer to Figure 2 This diagram illustrates two weight-sharing Siamese network branches provided in an embodiment of this application. Each Siamese network branch includes a first-stage coding unit, at least one second-stage coding unit, and a third-stage coding unit connected sequentially. The coding stage includes at least one second-stage coding unit and a third-stage coding unit, here consisting of three second-stage coding units and one third-stage coding unit.
[0081] The first-phase remote sensing image before the change and the second-phase remote sensing image after the change are respectively input into two Siamese network branches with shared weights. Image patch embedding is a technique that divides the input image into multiple patches, each of which can be regarded as an independent input unit. This technique helps the model better capture local features while reducing computational cost. The Stem backbone is the initial part of the Siamese network branches, used for preliminary feature extraction and downsampling of the input image. It typically includes multiple convolutional and pooling layers to extract low-level features and reduce the resolution of the feature map. The first-phase remote sensing image has dimensions (h, w, 3). After passing through the image patch embedding in the first-stage coding unit and the Stem backbone to obtain the feature map, the dimensions become (h / 4, w / 4, 32). Then, it enters the second-stage coding unit and the third-stage coding unit. Each second-stage coding unit includes a lightweight state space module and downsampling. The third-stage coding unit also includes a lightweight state space module. After passing through the first-stage coding unit, the image undergoes sequential downsampling as it passes through the lightweight state space module in the second-stage coding unit. This gradually reduces the resolution of the first-phase feature map extracted in each coding stage, from 1 / 4, 1 / 8, 1 / 16, and finally to 1 / 32, while the channel dimension gradually increases. The coding stage of the second-phase remote sensing image follows the same process. Finally, the first and second-phase feature maps output from the third coding stage, as well as the first and second-phase feature maps output from the fourth coding stage, are input into the temporal state space feature fusion module.
[0082] Step S5: Input the first phase feature map and the second phase feature map into the phase state space feature fusion module. Through pixel-level difference and fusion processing, the phase state space fusion feature is obtained. The phase state space feature fusion module is used to fuse spatial features of different coding stages.
[0083] Specifically, the temporal state space feature fusion module integrates spatial features from different encoding stages through pixel-level differencing and fusion processing to generate temporal state space fusion features. This process not only captures the differences between two temporal feature maps but also fuses their common features, thereby generating a comprehensive feature representation that provides richer information for subsequent change detection.
[0084] Step S7: The spatiotemporal state space feature refinement module in multiple decoding stages decodes the spatiotemporal state space fusion features, the first spatiotemporal feature map, and the second spatiotemporal feature map to obtain the target change mask map corresponding to the dual-temporal remote sensing image; the spatiotemporal state space feature refinement module is used to refine and restore spatial details.
[0085] Specifically, the temporal state-space fusion features, as well as the first and second temporal feature maps, are input into a multi-stage temporal state-space feature refinement module. These modules progressively decode the feature maps, refining and restoring spatial details to ultimately generate a target change mask map corresponding to the dual-temporal remote sensing image. The generated target change mask map is a two-dimensional image, where the value of each pixel indicates whether that pixel has changed. Typically, the pixel values of the change mask map are 0 or 1, where 1 indicates a change and 0 indicates no change. This embodiment not only enhances the spatial information of the feature maps but also progressively restores the image details through multi-stage processing, enabling the final generated change mask map to more accurately reflect the changed areas in the dual-temporal remote sensing image.
[0086] Please continue reading Figure 2 The temporal state-space fusion features are input to the spatiotemporal state-space feature refinement module in the first decoding stage. The output feature image is upsampled to increase its resolution, resulting in a corresponding change mask. This change mask is then input to the spatiotemporal state-space feature refinement module in the next stage, the second decoding stage. After processing, it is similarly upsampled to increase its resolution, yielding the corresponding change mask for the second decoding stage. This process continues until the fourth stage encoding outputs the final target change mask. Before generating the final target change mask, classification and thresholding are performed to convert pixel values into binary images. Typically, a fixed threshold (e.g., 0.5) is used to distinguish between changed and unchanged pixels. Therefore, the final target change mask is a two-dimensional image, where 1 represents change and 0 represents no change.
[0087] Starting from the second decoding stage, the input also includes the output of the corresponding encoding stage. Specifically, the input of the second decoding stage includes the first and second temporal feature maps output from the first encoding stage; the input of the third decoding stage includes the first and second temporal feature maps output from the second encoding stage; and the input of the fourth decoding stage includes the first and second temporal feature maps output from the third encoding stage. That is, skip connections are used to link to the shallow features corresponding to the encoding stage to supplement the spatial structural details in the feature maps.
[0088] This embodiment provides a state-space-based image change detection method, apparatus, device, and medium. By acquiring dual-temporal remote sensing images, including both before and after a change, it reflects the geographical features of different time points, providing foundational data for subsequent change detection. During processing, a lightweight state-space module across multiple coding stages of a Siamese network extracts feature maps at different scales, effectively extracting global features and eliminating redundant information, avoiding unnecessary noise interference. Next, a temporal state-space feature fusion module, through pixel-level difference and fusion processing, combines the feature maps from both temporal stages to achieve deep fusion of spatial features from different coding stages, thereby enhancing the expressive power of spatiotemporal information. Finally, a spatiotemporal state-space feature refinement module across multiple decoding stages refines and restores the fused features, further extracting spatiotemporal change features from the image and generating a high-precision change mask. This refinement process ensures that each stage can more accurately capture spatiotemporal changes in the image, and the final output change mask more precisely reflects the changed areas in the image, effectively addressing the shortcomings of traditional remote sensing image change detection methods in spatiotemporal information processing and significantly improving the accuracy of change detection.
[0089] Figure 3 The flowchart provided in this application describes the encoding process of a lightweight state space module for any encoding stage. The process may include the following steps:
[0090] Step S301: Perform a first feature transformation on the received input image through the first feature transformation component to obtain a first feature transformation image; the first feature transformation component includes a layer normalization layer, a fully connected layer, and a depthwise convolution layer connected in sequence.
[0091] Specifically, the main purpose of layer normalization is to adjust the distribution of the input image to a standard range, typically a distribution with a mean of 0 and a standard deviation of 1. Standardizing the distribution of the input image reduces scale differences during gradient descent, thereby accelerating the training process and improving model stability. Fully connected layers learn the mapping relationship between input and output features, transforming the input features into higher-level feature representations. Depthwise convolutions extract local features by sliding the convolution kernel across the input image. These local features can capture information such as edges and textures in the image. Specifically, the input image is first normalized to ensure a stable distribution, providing a good foundation for subsequent feature extraction. Then, the normalized feature vector is transformed into a higher-level feature representation through a fully connected layer, integrating global information. Finally, deep convolutions extract local features, capturing information such as edges and textures in the image, resulting in the first feature-transformed image. Preferably, the first feature-transformed image can also be processed using the SiLU (Sigmoid Linear Unit) activation function to further enhance the expressive power of the features before being input into the lightweight network layer.
[0092] Step S303: The first feature transformation image is divided into two parts proportionally by a lightweight network layer. The first feature transformation image of the first part is cross-scanned a specified number of times to obtain a global feature map containing global features. The first feature transformation image of the second part is subjected to identity mapping to obtain a local feature map with redundant features removed. The global feature map and the local feature map are then concatenated by channels to obtain a lightweight feature map.
[0093] Specifically, the lightweight network layer employs a lightweight SS2D (State Space 2D) layer, whose main function is to efficiently process image feature maps while reducing computational cost and feature redundancy. By dividing the feature map into two parts and processing them differently, the lightweight SS2D layer can simultaneously capture global and local features, thereby improving the model's performance and efficiency. The lightweight SS2D layer proportionally divides the input first feature transformation image into two parts. The first part of the first feature transformation image has dimensions (h, w, λc), where λ is a proportional parameter with a value range of 0 ≤ λ ≤ 1. The second part of the first feature transformation image has dimensions (h, w, (1-λ)c). The first part of the first feature transformation image enters the S6 block (the sixth-generation evolution of the Structured State Space Model (SSM)) and performs a specified number of cross-scanning operations to obtain a global feature map containing global features. Cross-scanning is an efficient feature extraction method that captures global contextual information by scanning the feature map from different directions (such as from top left to bottom right and from bottom right to top left). This mechanism enables global information exchange, enhancing the global feature representation of the feature map. Please refer to [link / reference]. Figure 4This is a schematic diagram of cross-scanning provided in an embodiment of this application. This embodiment performs cross-scanning from top left to bottom right and from bottom right to top left. The first feature transformation image in the second part undergoes identity mapping, which directly passes the input feature map to the output without any transformation. This operation reduces feature redundancy, preserves local feature information, and reduces computational cost. Finally, the global feature map and the local feature map are concatenated along the channel dimension to obtain a lightweight feature map of dimension (h, w, c). This concatenation operation preserves information from both global and local features while reducing redundancy.
[0094] It should be noted that the first feature transformation image in the first part enters the S6 block (the sixth-generation evolution of the Structured State-Space Model (SSM)). Since the input data is an image, it is essentially discrete data, not continuous data. Therefore, the S6 block needs to be discretized to correspond to discrete image data. This not only eliminates the need for additional continuous-to-discrete conversion but also leverages the optimizations of GPUs and other hardware for discrete tensor operations, enabling efficient parallel computation of the S6 block. Specifically, the S6 block can be represented by the following linear ordinary differential equation:
[0095]
[0096]
[0097] Where u(t) is the input image patch; h(t) is the intermediate representation of the image patch; A∈R N×N Let B be the state matrix, describing the dynamic changes of h(t); B∈R N×1 Given the input projection, describe how u(t) affects h(t); C∈R 1×N For the output projection, describe how h(t) affects the output image y(t); D∈R 1 These are the weight parameters.
[0098] When discretizing, for the time interval [t] a ,t b ], latent state variables The continuous-time state change process can be represented as:
[0099]
[0100] Where τ is the integral over the time variable.
[0101] By sampling the time scale parameter Δ, the above continuous time state h(t) can be obtained. b The updated discretized formula is as follows:
[0102]
[0103] Step S305: Perform a second feature transformation on the lightweight feature map through the second feature transformation group to obtain the second feature transformation image; the second feature transformation component includes a layer normalization layer and a fully connected layer connected in sequence.
[0104] Specifically, the mean and standard deviation of each feature channel in the lightweight feature map are calculated by normalizing the map. These statistics are used to normalize the feature channels to a mean of 0 and a standard deviation of 1. The feature vector after layer normalization is then transformed into a higher-level feature representation through a fully connected layer. The fully connected layer learns the mapping relationship between input and output features, integrates global information, and enhances the expressive power of the features, resulting in a second feature transformation image. This feature map contains a higher-level feature representation, providing richer information for subsequent feature fusion and change detection.
[0105] Step S307: In response to the first stitched image obtained by stitching the input image and the second feature transformation image received by the lightweight state space module, the first stitched image is subjected to a third feature transformation by the third feature transformation component to obtain the third feature transformation image; the third feature transformation component includes a depthwise convolution, a layer normalization and a fully connected layer connected in sequence.
[0106] Specifically, skip connections reduce feature loss during propagation and preserve more detailed information. Here, the input image received by the lightweight state space module is concatenated with the second feature-transformed image through skip connections to obtain the first concatenated image. This concatenation operation preserves both the original information of the input image and the information after feature transformation, providing richer context for subsequent feature extraction. The first concatenated image is then convolved using depthwise convolution to extract local features. The convolved feature map is then normalized using layer normalization to ensure that the input of each layer has the same distribution, reducing internal covariate shifts. The normalized feature map is then linearly transformed using a fully connected layer to extract higher-level feature representations, ultimately yielding the third feature-transformed image.
[0107] Step S309: In response to the second stitched image obtained by stitching the first stitched image and the third feature transformation image, the second stitched image is determined as the temporal feature map of the encoding stage corresponding to the lightweight state space module, and the temporal feature map is used as the output image.
[0108] Specifically, the first stitched image is stitched together using skip connections and the third feature transformation image to obtain the second stitched image. This second stitched image is then designated as the temporal feature map of the encoding stage corresponding to the lightweight state space module. This feature map contains the original information of the input image, the information after two feature transformations, and the information after the third feature transformation, forming a comprehensive feature representation.
[0109] In one embodiment, the lightweight state space module includes a first feature transformation component, a lightweight network layer, a second feature transformation component, and a third feature transformation component connected in sequence; the input position of the first feature transformation component is connected to the output position of the second feature transformation component via a jump, and the input position of the second feature transformation component is connected to the output position of the third feature transformation component via a jump.
[0110] Specifically, please refer to Figure 5 This is a schematic diagram of a lightweight state space module provided in an embodiment of this application. Each lightweight state space module includes a first feature transformation component, a lightweight network layer, a second feature transformation component, and a third feature transformation component. The input position of the first feature transformation component is connected to the output position of the second feature transformation component via a jump, and the input position of the second feature transformation component is connected to the output position of the third feature transformation component via a jump. The first feature transformation component includes a layer normalization layer, a fully connected layer, and a depthwise convolution layer connected in sequence. The second feature transformation component includes a layer normalization layer and a fully connected layer connected in sequence. The third feature transformation component includes a depthwise convolution layer, a layer normalization layer, and a fully connected layer connected in sequence. The lightweight state space module first processes the received input image through the layer normalization layer, the fully connected layer, and the depthwise convolution layer in the first feature transformation component to obtain a first feature-transformed image. The lightweight network layer divides the first feature transformation image into two parts proportionally. The first part of the first feature transformation image undergoes a specified number of cross-scans to obtain a global feature map containing global features. The second part of the first feature transformation image undergoes an identity mapping to obtain a local feature map with redundant features removed. The global and local feature maps are then concatenated through channels to obtain the lightweight feature map. The lightweight feature map is then passed sequentially through layer normalization and fully connected layers in the second feature transformation component to obtain the second feature transformation image. Here, the input image received by the lightweight state space module is concatenated with the second feature transformation image through skip connections to obtain the first concatenated image. The first concatenated image is then passed sequentially through depthwise convolution, layer normalization, and fully connected layers in the third feature transformation component to obtain the third feature transformation image. The first concatenated image is then concatenated with the third feature transformation image through skip connections to obtain the second concatenated image. This second concatenated image is determined as the temporal feature map of the encoding stage corresponding to the lightweight state space module, and the temporal feature map is used as the output image corresponding to this lightweight state space module.
[0111] This embodiment extracts preliminary features from the input image through a first feature transformation component and uses a lightweight network layer to segment the feature map, thereby obtaining a global feature map and a local feature map. This helps to capture global and local changes in the image. Next, a second feature transformation component further optimizes the feature representation. Combined with a stitching strategy based on a lightweight network layer, it can fuse the spatiotemporal information of the image, thereby improving the accuracy of change detection. Finally, after processing by a third feature transformation component, a temporal feature map is generated as the output image of the lightweight state space module. This process effectively utilizes feature transformation and stitching strategies at different levels to extract more accurate spatiotemporal features from remote sensing images, providing more precise information for change detection and solving the problem that traditional methods struggle to handle complex spatiotemporal relationships.
[0112] Figure 6 The flowchart for step S5 provided in the embodiments of this application may include the following steps:
[0113] Step S51: The first phase feature map and the second phase feature map of the penultimate coding stage are processed by pixel-level difference to obtain the first difference feature map corresponding to the penultimate coding stage.
[0114] Specifically, the first phase feature map of the penultimate encoding stage S-1 Second phase feature map Pixel-level difference processing is performed to obtain the first difference feature map. :
[0115]
[0116] Step S53: The first phase feature map and the first difference feature map of the penultimate encoding stage are fused to obtain the first phase fused feature map; and the second phase feature map and the first difference feature map of the penultimate encoding stage are fused to obtain the second phase fused feature map.
[0117] Specifically, the fusion process is as follows:
[0118]
[0119]
[0120] in, This is the first difference feature map of the penultimate encoding stage S-1; This is the first phase feature map of the penultimate encoding stage S-1; This is the second phase feature map of the penultimate encoding stage S-1; This is the first phase fusion feature map; This is the second phase fusion feature map.
[0121] Step S55: The first phase feature map of the last coding stage and the second phase feature map of the last coding stage are processed by pixel-level difference to obtain the second difference feature map corresponding to the last coding stage.
[0122] Specifically, the first phase feature map of the last encoding stage S. Second phase feature map Pixel-level difference processing is performed to obtain the second difference feature map. :
[0123]
[0124] Step S57: The first phase feature map and the second difference feature map of the last encoding stage are fused to obtain the third phase fused feature map; and the second phase feature map and the second difference feature map of the last encoding stage are fused to obtain the fourth phase fused feature map.
[0125] Specifically, the fusion process is as follows:
[0126]
[0127]
[0128] in, This is the second difference feature map of the last encoding stage S; This is the first phase feature map of the last encoding stage S; This is the second phase feature map of the last encoding stage S; This is a feature map of the third phase fusion. This is the fourth phase fusion feature map.
[0129] Step S59: Perform fusion convolution processing based on the first phase fusion feature map, the second phase fusion feature map, the third phase fusion feature map and the fourth phase fusion feature map to obtain the phase state space fusion feature.
[0130] Specifically, the first phase fusion feature map Feature map fused with the third phase Perform pixel-level fusion and fuse the feature maps from the second time phase. Feature map fused with the fourth phase Pixel-level fusion is performed, and the two fused results are then convolved and channel-mixed to output a temporal state space fusion feature with double the number of channels.
[0131]
[0132]
[0133]
[0134] Among them, F TM This refers to the temporal state-space fusion characteristics; This is the feature map of the first branch after fusion; This is the feature map of the second branch after fusion.
[0135] Please see Figure 7 This is a schematic diagram of the temporal state space feature fusion module provided in an embodiment of this application. When there are four encoding stages, the penultimate encoding stage is the third encoding stage, and the last encoding stage is the fourth encoding stage. Specifically: the T1 temporal feature map and the T2 temporal feature map of the third encoding stage are processed by pixel-level difference to obtain a first difference feature map, which is then fused with the T1 temporal feature map and the T2 temporal feature map of the third encoding stage respectively to obtain the fused temporal fusion feature map (first difference feature fusion).
[0136]
[0137]
[0138]
[0139] in, This is the first difference feature map of the third encoding stage; This is the T1 phase feature map of the third encoding stage; This is the T2 phase feature map of the third encoding stage; This is a T1 phase fusion feature map; This is a T2 phase fusion feature map.
[0140] The fourth encoding stage repeats the above process, performing a second differential feature fusion:
[0141]
[0142]
[0143]
[0144] in, This is the second difference feature map of the fourth encoding stage; This is the T1 phase feature map of the fourth encoding stage; This is the T2 phase feature map of the fourth encoding stage; This is a feature map of the third phase fusion. This is the fourth phase fusion feature map.
[0145] The result of fusing the two differential features is fused with the output of the first differential feature fusion at the pixel level. After convolution and channel mixing, a temporal state space fusion feature with double the number of channels is finally output.
[0146]
[0147]
[0148]
[0149] Among them, F TM This refers to the temporal state-space fusion characteristics; This is the feature map of the first branch after fusion; This is the feature map of the second branch after fusion.
[0150] This embodiment uses pixel-level differential processing to obtain difference feature maps at different encoding stages, which helps to capture subtle temporal changes in the image. Subsequently, by fusing the temporal feature maps with the difference feature maps, the spatiotemporal feature representation of the image is further enhanced, making the features of each temporal phase more comprehensive and reducing the spatiotemporal information that a single feature map might miss. This layer-by-layer feature fusion and temporal differential strategy can more accurately handle complex spatiotemporal relationships in remote sensing images, and has significant application value, especially in change detection tasks. By performing differential and fusion operations separately at different encoding stages, more refined spatiotemporal features can be effectively extracted. Finally, through fusion convolution processing, a comprehensive temporal state-space fusion feature is generated, improving the accuracy and precision of change detection, thus solving the problem that traditional methods struggle to accurately capture complex spatiotemporal changes when processing remote sensing images. This technique can significantly improve the effect of change detection in remote sensing images and is suitable for practical application scenarios requiring extremely precise spatiotemporal information.
[0151] Figure 8 The flowchart provided in this application embodiment illustrates the decoding process of the spatiotemporal state space feature refinement module for any decoding stage. This process may include the following steps:
[0152] Step S701: The received input features are extracted using the lightweight state space module to obtain a lightweight map.
[0153] Specifically, the spatiotemporal state space refinement module receives features from the output of the previous stage or image features from the corresponding decoding stage. The lightweight state space module extracts features from the received input features. The feature extraction method of the lightweight state space module is the same as that in steps S301 to S309, and will not be repeated here.
[0154] Step S703: Perform a first convolution operation on the lightweight map using the first convolution component to obtain a first spatial feature map; the first convolution component includes a 1×1 convolutional layer.
[0155] Specifically, a first spatial feature map is obtained by performing a convolution operation on the lightweight map using a 1×1 convolutional layer to adjust the number of channels in the feature map.
[0156] Step S705: Perform a second convolution operation on the first spatial feature map using the second convolution component to obtain a second spatial feature map; the second convolution component includes a 1×1 convolutional layer, a 3×3 convolutional layer, a normalization layer, and a 1×1 convolutional layer connected in sequence.
[0157] Specifically, the number of channels in the first spatial feature map is adjusted by a 1×1 convolutional layer, local features are extracted by a 3×3 convolutional layer, features are normalized, the number of channels is adjusted again by a 1×1 convolutional layer, and finally the second spatial feature map is output.
[0158] Step S707: In response to the new spatial feature map obtained by stitching together the first spatial feature map and the second spatial feature map, the new spatial feature map is determined as the change mask map corresponding to the decoding stage of the spatiotemporal state spatial feature refinement module, and the change mask map is used as the output image.
[0159] Specifically, a new spatial feature map is obtained by splicing the first spatial feature map with the second spatial feature map through skip connections. The feature map after skip connection processing enhances feature propagation and fusion, and the new spatial feature map integrates feature information at different levels. The new spatial feature map is determined as the change mask map corresponding to the decoding stage of the spatiotemporal state spatial feature refinement module and used as the output image.
[0160] In one implementation, the spatiotemporal state space feature refinement module includes a lightweight state space module, a first convolutional component, and a second convolutional component connected in sequence; the input position of the first convolutional component is connected to the output position of the second convolutional component via a jump.
[0161] In one implementation, the output image of the lightweight state space module in the encoding stage is determined as the corresponding image of the input image in the decoding stage; the stitched image obtained by stitching the first spatial feature map and the corresponding image is used as the input image of the second convolutional component.
[0162] Specifically, please refer to Figure 9 This is a schematic diagram of the spatiotemporal state space feature refinement module provided in an embodiment of this application. Taking the spatiotemporal state space feature refinement module corresponding to the second decoding stage as an example, the module includes a lightweight state space module, a first convolutional component, and a second convolutional component. The first convolutional component includes a 1×1 convolutional layer. The second convolutional component includes a 1×1 convolutional layer, a 3×3 convolutional layer, a normalized layer, and a 1×1 convolutional layer connected in sequence. The input position of the first convolutional component is connected to the output position of the second convolutional component via a skip connection. The lightweight state space module receives features from the previous encoding stage or the output of the previous decoding stage and processes them to obtain a lightweight map. Figure 9 The first decoding stage receives the temporal state space fusion features output from the previous stage, while the second to fourth encoding stages receive the outputs from the previous decoding stages.
[0163] In the spatiotemporal state space feature refinement module corresponding to the second decoding stage provided in this application embodiment, the features output from the previous encoding stage are first received and feature extracted through a lightweight state space module to obtain a lightweight map. The lightweight map is further feature extracted through a 1×1 convolutional layer in the first convolutional component to obtain a first spatial feature map. The temporal feature map output from the first encoding stage is obtained, including the first temporal feature map output from the first encoding stage and the second temporal feature map output from the second encoding stage. The first temporal feature map output from the first encoding stage, the second temporal feature map output from the second encoding stage, and the first spatial feature map are concatenated to obtain an updated first spatial feature map. The updated first spatial feature map is sequentially passed through a 1×1 convolutional layer, a 3×3 convolutional layer, normalization, and a 1×1 convolutional layer in the second convolutional component to obtain a second spatial feature map. The updated first spatial feature map is concatenated with the second spatial feature map through skip connections to obtain a new spatial feature map. The new spatial feature map is determined as the change mask map corresponding to the decoding stage of the spatiotemporal state space feature refinement module, and the change mask map is used as the output image.
[0164] This embodiment refines the input features using a lightweight state space module and multi-layer convolutional operations (1×1 convolutional layers, 3×3 convolutional layers, etc.), enabling the effective extraction and fusion of different spatial features. In this process, the combined design of the first and second convolutional components not only captures image details but also enhances the expression of spatiotemporal features. Finally, a new spatial feature map generated by stitching together the first and second spatial feature maps further improves the accuracy of change detection results. The change mask map in this process, as the output image, more accurately reflects the spatiotemporal changes in remote sensing images, achieving efficient processing and accurate detection of complex spatiotemporal information. This solves the problem of capturing spatiotemporal changes in traditional methods when processing remote sensing images, significantly improving the overall performance of the change detection task.
[0165] Figure 10 The flowchart for step S7 provided in the embodiments of this application may include the following steps:
[0166] Step S71: The temporal state space fusion features are decoded through the spatiotemporal state space feature refinement module of the first decoding stage in the decoding stage to obtain the initial change mask map.
[0167] Step S73: The initial change mask map, the first temporal feature map and the second temporal feature map corresponding to the encoding stage are decoded by the spatiotemporal state space feature refinement module of the next decoding stage in the decoding stage, until all decoding stages are completed, and the target change mask map is obtained.
[0168] Specifically, the temporal state space fusion feature F TM The initial change mask image M1 is obtained by refining the spatiotemporal state space features of the first decoding stage in the decoding stage. The initial change mask image M1 and the first temporal feature image F corresponding to the encoding stage are then used for decoding. T1 Second phase characteristic diagram F T2 Decoding is performed through the spatiotemporal state space feature refinement module of the next decoding stage until all decoding stages are completed, resulting in the target change mask map M. final .
[0169] This embodiment decodes the spatiotemporal state space fusion features using the spatiotemporal state space feature refinement module in the first decoding stage to initially obtain a change mask map. Then, the initial change mask map, along with the first and second temporal feature maps output from the corresponding encoding stage, are input into the spatiotemporal state space feature refinement module in the next decoding stage for processing. This process is repeated until all decoding stages are completed, thereby obtaining the target change mask map. This progressive refinement method allows each decoding stage to more accurately extract spatiotemporal change features from the image. The resulting change mask map more precisely reflects the changed areas in the image, improving the accuracy of change detection and meeting the needs of high-precision remote sensing image processing.
[0170] It should also be noted that the total loss function in this embodiment consists of two parts:
[0171] The first part is the Binary Cross-Entropy Loss: a statistical optimization used for pixel-level classification. During network parameter optimization, the binary cross-entropy loss is minimized iteratively to obtain a more accurate and complete change mask map.
[0172]
[0173] Among them, Y w,h The true label is 1 in the changing region and 0 in the unchanged region; Y p w,h The predicted value represents the probability of a change at position (w,h); W is the image width; and H is the image height.
[0174] The second part is the Dice loss function: by maximizing the similarity between the segmentation prediction result and the real label, it reflects the overlap between the predicted mask and the real label.
[0175]
[0176] Where Y is the true label matrix, Y p This is the matrix of predicted values.
[0177] The total loss function is expressed as:
[0178]
[0179] Among them, L total L is the total loss function; bce For binary cross-entropy loss; L dice This is the Dice loss function.
[0180] To illustrate the technical effects achieved by the embodiments of this application, qualitative and quantitative analyses were performed on two publicly available platform datasets to verify the performance of the state-space-based image change detection method provided in this application. These include the WHU-CD dataset, a publicly available dataset specifically designed for change detection tasks, containing a pair of remote sensing images from New Zealand. The images have a resolution of 32507 pixels × 15354 pixels and a spatial resolution of 0.2 meters per pixel, and were taken in April 2012 and April 2016, respectively, covering an area of 20.5 square kilometers. The dataset features detailed annotations for changes in buildings within the images. For ease of processing, the images were cropped into non-overlapping 256×256 pixel patches and randomly divided into a training set (6096 images), a validation set (762 images), and a test set (762 images) in an 8 / 1 / 1 ratio. The dataset reference is: S. Ji, S. Wei, and M. Lu, “Fully convolutional networks for multisource building extraction from an open aerial and satellite imagery data set,” IEEE Transactions on Geoscience and Remote Sensing, vol.57, no.1, pp.574–586, 2018. The LEVIR-CD dataset, specifically developed for change detection tasks, contains 637 pairs of Google Earth images. The images have a resolution of 1024×1024 pixels, a spatial resolution of 0.5 meters per pixel, and a time span ranging from 5 to 14 years. The dataset also features detailed annotations of building changes in the images. The images are cropped into non-overlapping 256×256 pixel patches and randomly divided into a 7 / 1 / 2 ratio for the training set (7120 images), validation set (1024 images), and test set (2048 images). The dataset reference is: K. Chen, Z. Zou, and Z. Shi, “Building extraction from remotesensing images with sparse token transformers,” Remote Sensing, vol.13, no.21, p.4441, 2021.
[0181] In the qualitative analysis, the method of this application generates a binary change mask map corresponding to the input image and compares it with the ground truth labels provided in the dataset and the results of other methods to observe whether the mask is clear and complete and whether there are any missegmentations. In the quantitative analysis, objective evaluation indicators commonly used in the field of change detection are used for numerical comparison. These indicators are calculated entirely by the computer based on the input data and the generated change mask map, ensuring the objectivity and accuracy of the calculation. The evaluation indicators used in this embodiment include: precision (P in the table), recall (R in the table), F1 score (F1 in the table), intersection-over-union (IoU) of the change categories, and overall accuracy (OA) of all categories. The calculation formulas for these indicators are as follows:
[0182]
[0183]
[0184]
[0185]
[0186]
[0187] Based on these metrics, the method in this embodiment was compared with five existing multi-temporal remote sensing image change detection neural network methods: FC-EF, FC-Siam-Conc, FC-Siam-Diff, IFN, and BIT. Specific comparison results are shown in Tables 1 and 2.
[0188] This embodiment implements the proposed method using the PyTorch deep learning framework, employing an NVIDIA Tesla A40 GPU (48GB VRAM) for both training and inference. Network parameters are optimized using the Adam optimizer, with a momentum of 0.9, weight decay of 0.0001, an initial learning rate of 0.001, and a batch size of 8. 100 training epochs are performed on the training set, with inference evaluation conducted on the validation set every 10 epochs to assess the training performance. After training, the trained model is used for inference on the test set, and evaluation metrics are calculated to assess the experimental results.
[0189] During training, other common data augmentation techniques were also employed, including random flipping, random scaling (0.8 to 1.2 times), random cropping, Gaussian blur, and random color jitter.
[0190] The comparison results are shown in Table 1, which presents the experimental results of binary change detection on the WHU-CD change detection dataset.
[0191] Table 1 shows the experimental results of binary change detection on the WHU-CD change detection dataset.
[0192]
[0193] Note: Higher values indicate better results.
[0194] As shown in Table 1, the experimental results of this application achieved the best performance in all four metrics: recall, F1 score, intersection-over-union ratio (IoU), and overall accuracy. It ranked third in accuracy, only behind the IFN and BIT methods. Specifically, the recall was improved by 1.26%, the F1 score by 0.53%, and the IoU by 0.88%. In quantitative analysis, this application performed excellently on the WHU-CD dataset, surpassing existing state-of-the-art methods. In qualitative analysis, combined with… Figure 11 The comparison of the experimental results of binary change detection on the WHU-CD change detection dataset shown in the figure demonstrates that the mask generated by this application is clear and complete, with neat edge contours and fewer mis-segmented pixels, exhibiting higher accuracy and reliability.
[0195] The comparison results are shown in Table 2, which presents the experimental results of binary change detection on the LEVIR-CD change detection dataset.
[0196] Table 2 shows the experimental results of binary change detection on the LEVIR-CD change detection dataset.
[0197]
[0198] Note: Higher values indicate better results.
[0199] As shown in Table 2, the experimental results of this application achieved the best performance in three metrics: F1 score, intersection-over-union ratio (IoU), and overall accuracy. It ranked second in precision and recall, only behind the IFN and BIT methods. Specifically, the F1 score was improved by 1.85%, and the IoU by 3.08%. In quantitative analysis, this application performed exceptionally well on the LEVIR-CD dataset, surpassing existing state-of-the-art methods. In qualitative analysis, combined with… Figure 12 The comparison of binary change detection results on the LEVIR-CD change detection dataset shown in the figure demonstrates that the mask generated in this application is clear, complete, and has neat edge contours. Furthermore, it maintains high segmentation accuracy even when processing a large number of buildings, and its overall performance has reached an advanced level.
[0200] Accordingly, please refer to Figure 13 A block diagram of a state-space-based image change detection device provided in this application embodiment, the device comprising:
[0201] The image acquisition unit 101 is used to acquire dual-temporal remote sensing images; the dual-temporal remote sensing images include a first-temporal remote sensing image before the change and a second-temporal remote sensing image after the change; the dual-temporal remote sensing images are obtained by taking pictures of the same geographic area at at least two different time points;
[0202] The encoding processing unit 103 is used to input the dual-temporal remote sensing images into two Siamese network branches with shared weights, and perform encoding processing through a lightweight state space module with multiple encoding stages to obtain multiple first-temporal feature maps of different scales corresponding to the first-temporal remote sensing image and multiple second-temporal feature maps of different scales corresponding to the second-temporal remote sensing image; the lightweight state space module is used to extract global features and remove redundant features.
[0203] The temporal fusion unit 105 is used to input the first temporal feature map and the second temporal feature map into the temporal state space feature fusion module, and obtain the temporal state space fusion feature through pixel-level difference and fusion processing; the temporal state space feature fusion module is used to fuse spatial features of different coding stages;
[0204] The decoding processing unit 107 is used to decode the temporal state space fusion features, the first temporal feature map, and the second temporal feature map through a spatiotemporal state space feature refinement module in multiple decoding stages to obtain the target change mask map corresponding to the dual-temporal remote sensing image; the spatiotemporal state space feature refinement module is used to refine and restore spatial details.
[0205] In some optional implementations, the lightweight state space module includes a first feature transformation component, a lightweight network layer, a second feature transformation component, and a third feature transformation component connected in sequence; the input position of the first feature transformation component is connected to the output position of the second feature transformation component via a jump, and the input position of the second feature transformation component is connected to the output position of the third feature transformation component via a jump.
[0206] In some alternative implementations, the encoding process for the lightweight state space module at any encoding stage is performed as follows:
[0207] The first feature transformation is performed on the received input image by the first feature transformation component to obtain the first feature transformation image; the first feature transformation component includes a layer normalization, a fully connected layer and a depthwise convolution connected in sequence.
[0208] The first feature transformation image is divided into two parts proportionally by a lightweight network layer. The first part of the first feature transformation image is cross-scanned a specified number of times to obtain a global feature map containing global features. The second part of the first feature transformation image is subjected to identity mapping to obtain a local feature map with redundant features removed. The global feature map and the local feature map are then concatenated through channels to obtain a lightweight feature map.
[0209] The lightweight feature map is subjected to a second feature transformation by the second feature transformation group to obtain the second feature transformation image; the second feature transformation component includes a layer normalization layer and a fully connected layer connected in sequence;
[0210] The response is based on the first stitched image obtained by stitching the input image received by the lightweight state space module and the second feature transformation image. The third feature transformation component performs a third feature transformation on the first stitched image to obtain the third feature transformation image. The third feature transformation component includes a depthwise convolution, a layer normalization layer and a fully connected layer connected in sequence.
[0211] The response is based on the second stitched image obtained by stitching the first stitched image and the third feature transformation image. The second stitched image is determined as the temporal feature map of the encoding stage corresponding to the lightweight state space module, and the temporal feature map is used as the output image.
[0212] In some alternative implementations, the time-phase fusion unit 105 includes:
[0213] The first phase feature map of the penultimate coding stage and the second phase feature map of the penultimate coding stage are processed by pixel-level difference to obtain the first difference feature map corresponding to the penultimate coding stage.
[0214] The first phase feature map and the first difference feature map of the penultimate encoding stage are fused to obtain the first phase fused feature map; and the second phase feature map and the first difference feature map of the penultimate encoding stage are fused to obtain the second phase fused feature map.
[0215] The first phase feature map and the second phase feature map of the last coding stage are processed by pixel-level difference to obtain the second difference feature map corresponding to the last coding stage.
[0216] The first phase feature map and the second difference feature map of the last encoding stage are fused to obtain the third phase fused feature map; and the second phase feature map and the second difference feature map of the last encoding stage are fused to obtain the fourth phase fused feature map.
[0217] Fusion convolution is performed based on the first, second, third, and fourth phase fusion feature maps to obtain the phase state space fusion features.
[0218] In some optional implementations, the spatiotemporal state space feature refinement module includes a lightweight state space module, a first convolutional component, and a second convolutional component connected in sequence; the input position of the first convolutional component is connected to the output position of the second convolutional component via a jump.
[0219] In some optional implementations, the decoding process for the spatiotemporal state space feature refinement module at any decoding stage is performed in the following manner:
[0220] The lightweight state space module extracts features from the received input features to obtain a lightweight graph.
[0221] The first convolution operation is performed on the lightweight map by the first convolution component to obtain the first spatial feature map; the first convolution component includes a 1×1 convolutional layer;
[0222] The second spatial feature map is obtained by performing a second convolution operation on the first spatial feature map through the second convolution component; the second convolution component includes a 1×1 convolutional layer, a 3×3 convolutional layer, a normalization layer and a 1×1 convolutional layer connected in sequence.
[0223] The response is based on a new spatial feature map obtained by stitching together the first spatial feature map and the second spatial feature map. The new spatial feature map is determined as the change mask map corresponding to the decoding stage of the spatiotemporal state spatial feature refinement module, and the change mask map is used as the output image.
[0224] In some alternative embodiments, the apparatus further includes:
[0225] The output image of the lightweight state space module in the encoding stage is determined as the corresponding image of the input image in the decoding stage;
[0226] The stitched image obtained by stitching the first spatial feature map and the corresponding image is used as the input image of the second convolutional component.
[0227] In some optional implementations, the decoding processing unit 107 includes:
[0228] The temporal state space fusion features are decoded through the spatiotemporal state space feature refinement module in the first decoding stage to obtain the initial change mask map.
[0229] The initial change mask map, the first phase feature map and the second phase feature map corresponding to the encoding stage are decoded by the spatiotemporal state space feature refinement module of the next decoding stage in the decoding stage, until all decoding stages are completed, and the target change mask map is obtained.
[0230] In some optional implementations, each Siamese network branch includes a first-stage coding unit, at least one second-stage coding unit, and a third-stage coding unit connected in sequence; both the second-stage and third-stage coding units include a lightweight state-space module; the input data for the first-stage coding unit is a dual-temporal remote sensing image, the input data for the initial second-stage coding unit includes the initial feature map output by the first-stage coding unit, the input data for subsequent second-stage coding units includes the temporal feature map output by the preceding second-stage coding unit, and the input data for the third-stage coding unit is the temporal feature map output by the last second-stage coding unit.
[0231] Specifically, for any twin network branch, feature extraction is performed on the dual-temporal remote sensing image through the first-stage coding unit to obtain the corresponding initial feature map;
[0232] The initial feature map is sequentially input into at least one second-stage coding unit and one third-stage coding unit for feature extraction, resulting in multiple temporal feature maps of different scales.
[0233] Further functional descriptions of the above modules and units are the same as those in the corresponding embodiments described above, and will not be repeated here.
[0234] In this embodiment, an image change detection device based on state space is presented in the form of a functional unit. Here, a unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.
[0235] Please see Figure 14 , Figure 14 This application provides a schematic diagram of the structure of a computer device, as shown in the embodiment of the present application. Figure 14As shown, the computer device includes one or more processors 10, memory 20, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components communicate with each other via different buses and can be mounted on a common motherboard or otherwise installed as needed. The processors can process instructions executed within the computer device, including instructions stored in or on memory to display graphical information of a GUI on external input / output devices (such as display devices coupled to the interfaces). In some alternative implementations, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple computer devices can be connected, each providing some of the necessary operations (e.g., as a server array, a group of blade servers, or a multiprocessor system). Figure 14 Take a processor 10 as an example.
[0236] Processor 10 may be a central processing unit, a network processor, or a combination thereof. Processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The programmable logic device may be a complex programmable logic device (CAMP), a field-programmable gate array (FPGA), a general-purpose array logic (GPA), or any combination thereof.
[0237] The memory 20 stores instructions executable by at least one processor 10 to cause the at least one processor 10 to perform the method shown in the above embodiments.
[0238] The memory 20 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the computer device. Furthermore, the memory 20 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some alternative embodiments, the memory 20 may optionally include memory remotely located relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0239] The memory 20 may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as flash memory, hard disk or solid-state drive; the memory 20 may also include a combination of the above types of memory.
[0240] The computer device also includes a communication interface 30 for communicating with other devices or communication networks.
[0241] This application also provides a computer-readable storage medium. The methods described in this application can be implemented in hardware or firmware, or implemented as recordable on a storage medium, or implemented as computer code downloaded over a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and subsequently stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code. When the software or computer code is accessed and executed by the computer, processor, or hardware, the methods shown in the above embodiments are implemented.
[0242] The apparatus, module, or unit described in the above embodiments can be implemented by a computer chip or entity, or by a product having a certain function. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.
[0243] For ease of description, the above devices are described separately by function as various units. Of course, in implementing this application, the functions of each unit can be implemented in one or more software and / or hardware.
[0244] Those skilled in the art will understand that the embodiments of this application can be provided as methods or apparatus. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0245] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, and devices according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0246] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0247] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0248] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0249] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the apparatus embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0250] The foregoing is merely an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.
[0251] Although embodiments of this application have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of this application, and such modifications and variations all fall within the scope defined by the appended claims.
Claims
1. A state-space-based image change detection method, characterized in that, The method includes: Acquire dual-temporal remote sensing images; the dual-temporal remote sensing images include a first-temporal remote sensing image before the change and a second-temporal remote sensing image after the change; the dual-temporal remote sensing images are obtained by taking pictures of the same geographic area at at least two different time points; The dual-temporal remote sensing images are respectively input into two weight-shared Siamese network branches, and encoded through a lightweight state space module with multiple encoding stages to obtain multiple first-temporal feature maps at different scales corresponding to the first-temporal remote sensing image, and multiple second-temporal feature maps at different scales corresponding to the second-temporal remote sensing image; the lightweight state space module is used to extract global features and remove redundant features. The first and second temporal feature maps are input into the temporal state space feature fusion module, and a temporal state space fusion feature is obtained through pixel-level difference and fusion processing. The temporal state space feature fusion module is used to fuse spatial features from different coding stages. Specifically, inputting the first and second temporal feature maps into the temporal state space feature fusion module and obtaining the temporal state space fusion feature through pixel-level difference and fusion processing includes: performing pixel-level difference processing on the first and second temporal feature maps of the penultimate coding stage to obtain a first difference feature map corresponding to the penultimate coding stage; and performing fusion processing on the first temporal feature map and the first difference feature map of the penultimate coding stage to obtain a first temporal fusion feature map. The process involves fusing the second temporal feature map of the penultimate encoding stage with the first difference feature map to obtain a second temporal fusion feature map; performing pixel-level difference processing on the first temporal feature map of the last encoding stage and the second temporal feature map of the last encoding stage to obtain a second difference feature map corresponding to the last encoding stage; fusing the first temporal feature map of the last encoding stage with the second difference feature map to obtain a third temporal fusion feature map; fusing the second temporal feature map of the last encoding stage with the second difference feature map to obtain a fourth temporal fusion feature map; and performing fusion convolution processing based on the first temporal fusion feature map, the second temporal fusion feature map, the third temporal fusion feature map, and the fourth temporal fusion feature map to obtain the temporal state space fusion feature. The spatiotemporal state space feature refinement module in multiple decoding stages decodes the spatiotemporal state space fusion features, the first spatiotemporal feature map, and the second spatiotemporal feature map to obtain the target change mask map corresponding to the dual-temporal remote sensing image; the spatiotemporal state space feature refinement module is used to refine and restore spatial details.
2. The method according to claim 1, characterized in that, The lightweight state space module includes a first feature transformation component, a lightweight network layer, a second feature transformation component, and a third feature transformation component connected in sequence; the input position of the first feature transformation component is connected to the output position of the second feature transformation component via a jump, and the input position of the second feature transformation component is connected to the output position of the third feature transformation component via a jump.
3. The method according to claim 2, characterized in that, The encoding process for the lightweight state space module at any encoding stage is performed as follows: The first feature transformation component performs a first feature transformation on the received input image to obtain a first feature transformation image; the first feature transformation component includes a layer normalization layer, a fully connected layer, and a depthwise convolution layer connected in sequence. The first feature transformation image is divided into two parts proportionally by the lightweight network layer. The first part of the first feature transformation image is cross-scanned a specified number of times to obtain a global feature map containing global features. The second part of the first feature transformation image is subjected to identity mapping to obtain a local feature map with redundant features removed. The global feature map and the local feature map are then concatenated by channels to obtain a lightweight feature map. The lightweight feature map is subjected to a second feature transformation by the second feature transformation group to obtain a second feature transformation image; the second feature transformation component includes a layer normalization layer and a fully connected layer connected in sequence. The response is based on the first stitched image obtained by stitching the input image received by the lightweight state space module and the second feature transformation image. The third feature transformation component performs a third feature transformation on the first stitched image to obtain a third feature transformation image. The third feature transformation component includes a depthwise convolution, a layer normalization layer and a fully connected layer connected in sequence. The response is based on the second stitched image obtained by stitching the first stitched image and the third feature transformation image, the second stitched image is determined as the temporal feature map of the encoding stage corresponding to the lightweight state space module, and the temporal feature map is used as the output image.
4. The method according to claim 1, characterized in that, The spatiotemporal state space feature refinement module includes a lightweight state space module, a first convolutional component, and a second convolutional component connected in sequence; the input position of the first convolutional component is connected to the output position of the second convolutional component via a jump.
5. The method according to claim 4, characterized in that, The decoding process for the spatiotemporal state space feature refinement module at any decoding stage is performed as follows: The lightweight state space module extracts features from the received input features to obtain a lightweight graph. A first spatial feature map is obtained by performing a first convolution operation on the lightweight map using the first convolution component; the first convolution component includes a 1×1 convolutional layer. The second convolutional component performs a second convolutional operation on the first spatial feature map to obtain a second spatial feature map; the second convolutional component includes a 1×1 convolutional layer, a 3×3 convolutional layer, a normalization layer and a 1×1 convolutional layer connected in sequence. The response is based on a new spatial feature map obtained by stitching together the first spatial feature map and the second spatial feature map. The new spatial feature map is determined as the change mask map corresponding to the decoding stage of the spatiotemporal state spatial feature refinement module, and the change mask map is used as the output image.
6. The method according to claim 5, characterized in that, The method further includes: The output image of the lightweight state space module in the encoding stage is determined as the corresponding image of the input image in the decoding stage; The stitched image obtained by stitching the first spatial feature map and the corresponding image is used as the input image of the second convolutional component.
7. The method according to claim 1, characterized in that, The temporal state space fusion features, the first temporal feature map, and the second temporal feature map are decoded by a spatiotemporal state space feature refinement module through multiple decoding stages to obtain the target change mask map corresponding to the dual-temporal remote sensing image, including: The temporal state space fusion features are decoded through the spatiotemporal state space feature refinement module of the first decoding stage in the decoding stage to obtain the initial change mask map; The initial change mask image, the first temporal feature image and the second temporal feature image corresponding to the encoding stage are decoded by the spatiotemporal state space feature refinement module of the next decoding stage in the decoding stage until all the decoding stages are completed, and the target change mask image is obtained.
8. The method according to claim 1, characterized in that, Each of the twin network branches includes a first-stage coding unit, at least one second-stage coding unit, and a third-stage coding unit connected in sequence; both the second-stage coding unit and the third-stage coding unit include the lightweight state-space module; the input data of the first-stage coding unit is the dual-temporal remote sensing image, the input data of the initial second-stage coding unit includes the initial feature map output by the first-stage coding unit; the input data of subsequent second-stage coding units includes the temporal feature map output by the preceding second-stage coding unit; the input data of the third-stage coding unit is the temporal feature map output by the last second-stage coding unit. Specifically, for any twin network branch, the first-stage coding unit extracts features from the dual-temporal remote sensing image to obtain the corresponding initial feature map; The initial feature map is sequentially input into at least one second-stage coding unit and the third-stage coding unit for feature extraction, resulting in multiple corresponding temporal feature maps at different scales.
9. An image change detection device based on state space, characterized in that, The device includes: An image acquisition unit is used to acquire dual-temporal remote sensing images; the dual-temporal remote sensing images include a first-temporal remote sensing image before the change and a second-temporal remote sensing image after the change; the dual-temporal remote sensing images are obtained by taking pictures of the same geographic area at at least two different time points; The encoding processing unit is used to input the dual-temporal remote sensing images into two weight-shared Siamese network branches, and perform encoding processing through a lightweight state space module with multiple encoding stages to obtain multiple first-temporal feature maps of different scales corresponding to the first-temporal remote sensing image and multiple second-temporal feature maps of different scales corresponding to the second-temporal remote sensing image; the lightweight state space module is used to extract global features and remove redundant features. A temporal fusion unit is used to input the first temporal feature map and the second temporal feature map into a temporal state space feature fusion module, and obtain temporal state space fusion features through pixel-level difference and fusion processing; the temporal state space feature fusion module is used to fuse spatial features of different coding stages; wherein, inputting the first temporal feature map and the second temporal feature map into the temporal state space feature fusion module, and obtaining temporal state space fusion features through pixel-level difference and fusion processing includes: processing the first temporal feature map of the penultimate coding stage and the second temporal feature map of the penultimate coding stage through pixel-level difference processing to obtain a first difference feature map corresponding to the penultimate coding stage; and performing fusion processing on the first temporal feature map of the penultimate coding stage and the first difference feature map to obtain a first temporal fusion feature map. The process involves: fusing the second temporal feature map of the penultimate encoding stage with the first difference feature map to obtain a second temporal fused feature map; performing pixel-level difference processing on the first temporal feature map of the last encoding stage and the second temporal feature map of the last encoding stage to obtain a second difference feature map corresponding to the last encoding stage; fusing the first temporal feature map of the last encoding stage with the second difference feature map to obtain a third temporal fused feature map; fusing the second temporal feature map of the last encoding stage with the second difference feature map to obtain a fourth temporal fused feature map; and performing fusion convolution processing on the first temporal fused feature map, the second temporal fused feature map, the third temporal fused feature map, and the fourth temporal fused feature map to obtain the temporal state space fused feature map. The decoding processing unit is used to decode the temporal state space fusion features, the first temporal feature map, and the second temporal feature map through a spatiotemporal state space feature refinement module in multiple decoding stages to obtain the target change mask map corresponding to the dual-temporal remote sensing image; the spatiotemporal state space feature refinement module is used to refine and restore spatial details.
10. A computer device, characterized in that, include: A memory and a processor are communicatively connected, the memory stores computer instructions, and the processor executes the state-space-based image change detection method according to any one of claims 1 to 8 by executing the computer instructions.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing a computer to perform the state-space-based image change detection method according to any one of claims 1 to 8.