Time sequence remote sensing image change detection method fusing spatial features

The U-shaped neural network with LSTM and spatial-token modules addresses the challenges of pseudo-changes and alignment in remote sensing change detection by learning seasonal patterns and enhancing feature extraction, thereby improving the detection of genuine land cover changes.

CN120318667APending Publication Date: 2025-07-15HOHAI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410050431.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-12
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The prior art is difficult to effectively screen out land cover changes caused by human activities in remote sensing image change detection, especially under high resolution conditions, there are pseudo-change information and image proofreading deviations, and the model has insufficient modeling ability to model time information.

Method used

The fusion method of long-term memory structure and spatial feature extraction module is adopted to learn seasonal information of multi-time phase remote sensing images through U-shaped symmetric network structure, and the global spatial features are extracted in combination with the multi-head cross-attention mechanism to enhance the resolution adaptability and pseudo-change recognition ability of the model.

Benefits of technology

The model's ability to recognize pseudo-change images is improved, and the land cover changes caused by human activities can be effectively screened out, which improves the accuracy of change detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318667A_ABST
    Figure CN120318667A_ABST
Patent Text Reader

Abstract

The invention discloses a time sequence remote sensing image change detection method fusing spatial features, and relates to the field of remote sensing change detection. The method comprises an encoder, a decoder and a spatial feature extraction module. The encoder is composed of five layers, and each layer comprises a convolutional neural network and a long-short term memory module. In each layer, the long and short time memory module orderly processes the input image sequence; the decoder is composed of five layers, each layer is composed of an up-sampling layer and a convolutional neural network, and a jump connection mode is used between the encoder and the decoder. The spatial feature extraction module is composed of a cross attention mechanism. The method has the beneficial effects that time sequence information and global spatial position information in an input image sequence can be concerned at the same time; a long-short-term memory structure is used for screening seasonal periodic features in a remote sensing image, and the recognition capability of a model on a pseudo-change image is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of remote sensing image change detection, and specifically relates to a method for remote sensing image change detection of time series integrating spatial features Background Art

[0002] Change detection is the process of determining changes in the state of land cover based on multiple observations at different times, providing important data support for solving the change detection of land cover. With the changes of the times, people have transformed natural resources while exploring the earth's resources. Such changes can be divided into two categories: one is caused by natural environmental factors, such as the changes of trees with seasons, and the changes of lakes during the flood season and the dry season; the other is the changes caused by human activities, such as the changes of residential land and farmland. We mainly study the latter to study the changes of land resources

[0003] In a macroscopic sense, this task is to learn the feature distribution of target category changes through a multi-layer neural network structure, so as to predict the change category. In a microscopic sense, this task is to perform semantic-level classification on each pixel in the input image to determine which pixel has changed. For this task, the feature extraction ability is crucial. Under different seasons, lighting, and vegetation distribution conditions, the model needs to screen and model the target category. Ignoring the periodic change information in the natural environment, focus on the changes of farmland and urban residential land. Change detection based on high resolution is still very challenging, mainly due to the following two points. First, there is a lot of pseudo-change information in the change detection task, and the model needs to screen relevant features according to the semantic information in the image. Second, compared with single-image, there is an image calibration deviation in bi-temporal images

[0004] In recent years, neural networks have been very successful in the field of remote sensing image processing. For example, remote sensing target detection, land cover classification, and change detection. In land cover classification, the focus of the task is to overcome the limitations in convolution and the problem of different scales of targets. Generally, dilated convolution or large convolution kernels are used to extract spatial context information. Although both are pixel tasks, the difference between land cover classification and change detection lies in the temporal information. In change detection, multiple time-dimensional images need to be input, so the model needs to improve the modeling ability for temporal information Summary of the Invention

[0005] Objective of the present invention: For the above-mentioned pseudo-change problem, we adopt the long short-term memory structure, take multi-temporal information as a time series input, and use this module to learn the periodic change rules of seasonal information (vegetation, coastline) in the sequence, and perform targeted feature extraction on the input image sequence to achieve the effect of feature screening. Considering that the long short-term memory structure is mainly composed of convolutional operations, aiming at the limitations of convolutional operations, the present invention finally uses a spatial feature extraction module to strengthen the global spatial features. In addition, the spatial feature extraction module can increase the resolution adaptability of the model.

[0006] The technical solution of the present invention is as follows: A method for detecting changes in time series remote sensing images integrating spatial features, characterized by comprising the following steps:

[0007] Step 1: Data sequence preprocessing: Obtain remote sensing images of the same location at different times, and obtain the final image sequence through cropping and alignment.

[0008] Step 2: Create a network model: Construct a U-shaped symmetric network structure. The first half of the network is called the encoder, and the second half is called the decoder. Make the remote sensing images of the same location at different times into an input sequence, put it into the model for training, and set the target loss function for model training;

[0009] Step 3: Define the spatial feature extraction module: Model the global spatial information of the features, and use the multi-head cross-attention mechanism based on high-dimensional vectors to extract feature information in different subspaces.

[0010] Furthermore, a method for detecting changes in time series remote sensing images integrating spatial features according to claim 1, characterized in that the specific steps of the data sequence preprocessing in step 1 are as follows: Obtain a set of remote sensing images of the same location at different times, set the image cropping size, crop the images in the set into the same size according to the size, classify them according to the location, and integrate the images at the same location into an input sequence.

[0011] Furthermore, a method for detecting changes in time series remote sensing images integrating spatial features according to claim 1, characterized in that the specific steps of creating the network model in step 2 are as follows: The encoder has a total of five layers of structures. Each layer uses a combination of a convolutional module and a long short-term memory module. As the number of layers deepens, the spatial size of the feature map also decreases. The spatial size of each layer is reduced by half compared to the previous layer, and at the same time the number of channels will also increase accordingly. Denote the input time phase as X i (i >= 2) H and W are the length and width of the image, C is the original input channel of the image, and i represents the number of time phases. In each layer, merge Xi into a vector T is the number of time series, which is calculated by putting it into the convolutional module Then put X out into the long short-term memory module in sequence according to the time sequence for calculation.

[0012] The calculation process of the long short-term memory module is as follows: First, merge the input and and calculate the corresponding output features through forget gate convolution, input gate convolution and output gate convolution, where σ(·) is the sigmoid function, Ft: represents 3x3 convolution operation, f t represents the forgetting feature, i t represents the input feature, σ t represents the output feature. As shown in 1), finally, obtain the final hidden layer state through the output gate and the current cell state and pass Hi to the next layer. As Figure 4 shown.

[0013] The decoder has a total of five-layer structure, and each layer is composed of an upsampling layer and a convolutional module. The number of channels of the convolutional module in each layer corresponds to that in the encoder. In each layer, first perform the upsampling operation, then put the sampled features into the convolutional module for calculation, and finally obtain the output where H’, W” and H”W” respectively represent the sizes of the features in the operation.

[0014] Furthermore, according to a method for change detection of time series remote sensing images integrating spatial features described in claim 1, it is characterized in that the specific steps of defining the spatial feature extraction module in step 3 are as follows:

[0015] The spatial feature extraction module consists of two parts: a high-dimensional vector acquisition module and a cross-attention module. The high-dimensional vector acquisition module first reduces the dimension of the feature through convolution operation to obtain generates an attention map using convolution operation, then uses this attention map to perform weighted summation on different parts of the input tensor, and finally obtains a high-dimensional vector representing semantic information through the multi-head self-attention mechanism (MSA) L represents the vector length.

[0016] The cross-attention mechanism is to use the high-dimensional vector to strengthen the features of the input . First, the linear layer performs linear mapping on the high-dimensional vector and the input , maps X1 to maps T to and Then, an attention weight is obtained by multiplying K and Q. Next, the weight after the softmax operation is multiplied by the vector V to obtain the output. Finally, Y0 is upsampled and transformed through the output layer to obtain the final result. Where H0 and W0 represent the sizes of the features in the operation, and H and W represent the sizes of the final output.

[0017] Furthermore, in the change detection network method described in claim 3, it is characterized in that: in the above structure, each layer has a skip connection, which connects the result of each encoding layer with the result of the upsampling of the corresponding decoding layer to restore the detailed information of the features.

[0018] The specific implementation of the system is the same as the method.

[0019] A computer device, characterized in that: the computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the above computer program, it implements the method for detecting changes in time-series remote sensing images by fusing spatial features as described in any one of claims 1-5.

[0020] A computer-readable storage medium, characterized in that: the computer-readable storage medium stores a computer program for executing the method for detecting changes in time-series remote sensing images by fusing spatial features as described in any one of claims 1-5.

[0021] Beneficial effects: Compared with the prior art, the present invention has the following advantages: The model can simultaneously focus on the temporal information and global spatial position information in the input image sequence, and uses the long short-term memory structure to screen the seasonal periodic features in the remote sensing images, improving the model's ability to identify pseudo-change images. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 It is a flowchart of the method according to an embodiment of the present invention;

[0023] Figure 2 It is a framework diagram of the method for fusing spatial features of time-series U-shaped neural network in a specific embodiment of the present invention;

[0024] Figure 3 It is a framework diagram of the spatial feature extraction module in a specific embodiment of the present invention.

[0025] Figure 4 It is a framework diagram of the long short-term memory module in a specific embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0026] The present invention will be further described in conjunction with the accompanying drawings and specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by this application.

[0027] Embodiment 1

[0028] The overall structure of the ST-UNet structure is described in detail in Figure 2 . In the network structure, the extraction capabilities of long short-term memory and Transformer for spatio-temporal feature information are fully utilized to make up for the limitations existing in convolutional operations. Our model is first composed of a time feature extraction module of the UNet structure, and then the time features rich in local spatial position information are sent into the STM to extract global spatial position information. Finally, the output features based on spatio-temporal feature information are obtained. In this network, the input image is a multi-temporal image, and the long short-term memory module is used to learn and screen the seasonal laws of periodic changes in the image sequence. Multiple temporal inputs are used to update the feature states, and the forget gate is used to select the important and unimportant information of the memory, so as to improve the model's recognition ability for pseudo-changing images.

[0029] As Figure 2 shown, our ST-UNet contains two main components: 1) A Siamese neural network based on long short-term memory, which is used to process time series input images and uses multiple temporal inputs to update the feature states. 2) A Spatial-Token-Module (STM) module extracts the global spatial information in the image; while reducing the computational complexity, the head cross-attention mechanism is used to extract the feature information in different sub-spaces.

[0030] In the long short-term memory module, we use an adjusted network structure. Compared with the parameter-free bilinear interpolation downsampling of the previous model, convolution is used for downsampling, and the convolution kernel can better record image sampling. By analyzing the effective receptive field of the network, it is found that the receptive field after the fifth layer of the model still cannot reach the whole image, so the number of convolutions in the improved model is increased to enlarge the effective receptive field of the features, which is convenient for the model to model high-dimensional features.

[0031] Embodiment 2

[0032] To verify the effectiveness of the present invention, Table 1 shows the accuracy results of different CD methods in LEVIR-CD. As shown in the table, ST-UNet performs better in the dual-temporal phase. Compared with past networks, ST-UNet leads by 0.9% in the F1 value. It can be seen that the combination of long-short-term memory and token transformer in the model is good, which can capture different feature information in the dual-temporal images while paying attention to spatial information.

[0033] Training details: Our model uses an NVIDIA 3090 GPU during training and testing. Random flipping, random rotation, and translation transformation are used during training; the training optimizer is AdamW, where the parameter beta is set to 0.99 - 0.999, the learning rate is initialized to 1e-3, the cosine annealing strategy is used, the minimum learning rate is 1e-8, and the batch size is 8. The total number of training epochs is 200. Cross Entropy Loss without weights is mainly used.

[0034] Matters not covered by the present invention are well-known techniques.

[0035] The above embodiments are only for illustrating the technical concept and features of the present invention, and their purpose is to enable those skilled in the art to understand the content of the present invention and implement it accordingly. It should not be used to limit the protection scope of the present invention. Any equivalent changes or modifications made according to the spirit and essence of the present invention should be covered within the protection scope of the present invention.

[0036] Table 1 Accuracy Results Table in LEVIR-CD

[0037]

Claims

1. A method for change detection of time - series remote sensing images integrating spatial features, characterized in that, Including the following steps: Step 1: Data sequence preprocessing: Obtain remote sensing images at the same location at different times, and obtain the final image sequence through cropping and alignment. Step 2: Create a network model: Construct a U-shaped symmetric network structure. The first half of the network is called the encoder, and the second half is called the decoder. Make the remote sensing images at the same location at different times into an input sequence, put it into the model for training, and set up the target loss function for model training. Step 3: Define the spatial feature extraction module: Perform global spatial information modeling on the features, and use the multi-head cross-attention mechanism based on high-dimensional vectors to extract feature information in different subspaces.

2. A method for change detection of time series remote sensing images integrating spatial features according to claim 1, characterized in that, The specific steps of the data sequence preprocessing in Step 1 are as follows: Obtain a set of remote sensing images at the same location at different times, set the image cropping size, crop the images in the set into the same size according to the size, classify them according to the location, and integrate the images at the same location into an input sequence.

3. A method for change detection of time series remote sensing images integrating spatial features according to claim 1, characterized in that, The specific steps for creating the network model in step 2 are as follows: The encoder has a total of five layers. Each layer combines a convolutional module and a long short-term memory module. As the number of layers deepens, the spatial dimension of the feature map decreases, and the spatial dimension of each layer is reduced by half compared to the previous layer. At the same time, the number of channels also increases. Denote the input phase as H and W are the length and width of the image, C is the original input channel of the image, and i represents the number of phases. In each layer, Xi is merged into a vector T is the number of time series, which is calculated by putting it into the convolutional module Then, Xout is sequentially put into the long short-term memory module for calculation in the order of time before and after. The calculation process of the long short-term memory module is as follows: First, the input and are merged, and the corresponding output features are calculated through the forget gate convolution, input gate convolution, and output gate convolution, where σ(·) is the sigmoid function, Ft: represents the 3x3 convolution operation, f t represents the forget feature, represents the input feature, represents the output feature. Finally, the final hidden layer state is obtained through the output gate and the current cell state and Hi is passed to the next layer. The decoder has a total of five layers of structure. Each layer is composed of an upsampling layer and a convolutional module. The number of channels in the convolutional module of each layer corresponds to that in the encoder. In each layer, first perform the upsampling operation, then put the sampled features into the convolutional module for calculation, and finally obtain the output. Where H’, W’ and H”W” respectively represent the sizes of the features in the operation.

4. A method for detecting changes in time-series remote sensing images that fuses spatial features according to claim 1, characterized in that The specific steps of defining the spatial feature extraction module in Step 3 are as follows: The spatial feature extraction module consists of two parts: a high-dimensional vector acquisition module and a cross-attention module. The high-dimensional vector acquisition module first reduces the dimension of the features through convolution operations to obtain generate an attention map using convolution operations, then use the attention map to perform weighted summation on different parts of the input tensor, and finally obtain a high-dimensional vector representing semantic information through the multi-head self-attention mechanism (MSA). L represents the vector length. The cross-attention mechanism uses a high-dimensional vector to strengthen the features of the input . First, the linear layer performs a linear mapping on the high-dimensional vector and the input , mapping X1 to mapping T to and . Then, an attention weight is obtained by multiplying K and Q . Next, the weight after the softmax operation is multiplied by the vector V to obtain the output . Finally, Y0 is upsampled and transformed by the output layer to obtain the final result where H0 and W0 represent the sizes of the features in the operation, and H and W represent the sizes of the final output 5. The change detection network method according to claim 3, wherein: In the above structure, each layer has a skip connection, which connects the result of each encoding layer with the result of the corresponding decoding layer after upsampling to restore the detailed information of the features.

6. A computer device, characterized in that: The computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the above computer program, it implements the method for detecting changes in time-series remote sensing images that fuses spatial features according to any one of claims 1-5.