Remote Sensing Image Building Change Detection Method, Device, Equipment and Storage Medium

Through the combination of twin convolutional neural networks and long-term memory networks, a combined loss function is constructed and parameters are optimized, which solves the sample imbalance in building change detection on rural homesteads, and improves the accuracy of building change detection in remote sensing images.

CN116310776BActive Publication Date: 2025-06-13GUANGZHOU URBAN PLANNING & DESIGN SURVEY RES INST
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211100910.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-09-13
Filing Date
2022-09-09
Publication Date
2025-06-13
Estimated Expiration
2042-09-09

AI Technical Summary

Technical Problem

The existing remote sensing image building change detection methods lack research on the self-built houses of villagers on rural homesteads, and there is a problem of unbalanced training positive and negative samples, and it is impossible to effectively combine texture features and different phase features, resulting in insufficient detection accuracy.

Method used

A twin convolutional neural network and a long and short-term memory network are combined to construct a combined loss function through feature extraction, stacking, training and decoding operations, and the parameters are updated using the Adam optimization algorithm to improve the imbalance problem of training data.

Benefits of technology

The accuracy of remote sensing image building change detection is improved, especially the change detection of buildings on rural homesteads, effectively extracting texture and spectral characteristics, and improving the problem of sample imbalance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116310776B_ABST
    Figure CN116310776B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, apparatus, device and storage medium for detecting building changes in remote sensing images. The method includes: acquiring the front and back two-phase images of training data; encoding operation: processing the front and back two-phase images by using a siamese convolutional neural network to obtain initial change information; training operation: training the initial change information by using a long short-term memory network to obtain fine change information; decoding operation: restoring the fine change information by using deconvolution to obtain an initial change detection result; evaluating the loss of the initial change detection result according to a combined loss function, updating the parameters of the siamese convolutional neural network and the long short-term memory network by using the Adam optimization algorithm, and executing the encoding operation, the training operation, the decoding operation and the evaluation operation again; when it is determined that the loss value does not decrease, updating the initial change detection result to the final change detection result. The present invention can improve the problem of unbalanced positive and negative samples in training and improve the detection accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of remote sensing technology, and particularly to a method, device, equipment and storage medium for detecting building changes in remote sensing images. Background Art

[0002] Change detection of remote sensing images is to extract the changed ground object information from two-phase or multi-phase images. Change detection of buildings in remote sensing images is to extract the change characteristics of buildings, a type of artificial ground objects, so as to determine whether the buildings have changed. Change detection of remote sensing images has been widely applied in fields such as land use change investigation and urban expansion analysis.

[0003] Currently, the method for detecting building changes in remote sensing images has always been a research hotspot in the field of remote sensing applications. Generally, it can be classified into traditional methods and learning-based methods. Traditional methods mainly include numerical operation methods, transformation methods and advanced model methods. Learning-based methods can mainly be divided into methods based on random forests, methods based on support vector machines and deep learning methods represented by convolutional neural networks.

[0004] The existing methods for detecting building changes in remote sensing mainly have the following problems: First, the existing models do not specifically study the self-built houses of villagers on rural homesteads. The self-built houses of villagers on rural homesteads mainly have characteristics such as small house area, irregular house shape, and small spacing between houses. Second, the existing change detection methods lack a deep learning network structure that combines texture features and different-phase features. Third, the existing training methods do not effectively optimize the problem of unbalanced building change detection samples. In summary, a method for detecting building changes in remote sensing images for buildings on rural homesteads is needed. Summary of the Invention

[0005] The present invention provides a method, device, equipment and storage medium for detecting building changes in remote sensing images, which can improve the problem of unbalanced positive and negative samples in training, thereby improving the accuracy of detecting building changes in remote sensing images.

[0006] In a first aspect, the present invention provides a method for detecting building changes in remote sensing images, including the following steps:

[0007] Obtain the front and rear two-phase images of the training data;

[0008] Encoding operation: Use a siamese convolutional neural network to perform feature extraction and feature stacking on the front and rear two-phase images to obtain initial change information;

[0009] Training operation: Use a long short-term memory network to train the initial change information to obtain refined change information;

[0010] Decoding operation: The fine change information is restored by deconvolution to obtain the initial change detection result;

[0011] Construct a combined loss function, with the initial value of i being 1, and perform the i-th evaluation operation: Perform loss evaluation on the initial change detection result according to the combined loss function to obtain the i-th loss value;

[0012] Optimization operation: Use the Adam optimization algorithm to update the parameters of the siamese convolutional neural network and the long short-term memory network, and perform the encoding operation, the training operation, the decoding operation, and the evaluation operation again, updating the i-th loss value to the (i + 1)-th loss value;

[0013] Judge whether i satisfies i less than the preset number of training times N. If so, increment i by one and return to execute the optimization operation; if not, perform the following difference judgment operation:

[0014] Judge whether the difference between the (i + 1)-th loss value and the i-th loss value is less than the preset threshold. If so, update the initial change detection result to the final change detection result of the building; if not, perform the optimization operation to update the i-th loss value and the (i + 1)-th loss value, and return to the difference judgment operation.

[0015] Preferably, the encoding operation includes:

[0016] Feature extraction is performed according to multiple groups of convolutional modules with different scales, and the input features of the previous group of convolutional modules are calculated to obtain features of different scales; wherein, the siamese convolutional neural network is composed of N groups of convolutional modules spliced in sequence;

[0017] Stack the features of different scales along the channel direction to obtain the initial stacked features, and slice the initial stacked features along the channel direction to obtain the initial sliced features;

[0018] The initial sliced features are sequentially input into the long short-term memory network for training to obtain the trained features;

[0019] Arrange the trained features according to the initial positions to obtain the initial features after training;

[0020] Activate the initial features after training to obtain the feature extraction result corresponding to the current convolutional module, and sequentially pass through N groups of convolutional modules to obtain the initial change information.

[0021] Preferably, the training operation includes:

[0022] Construct a spatial attention feature network and a channel attention feature network;

[0023] Based on the long short-term memory network, the spatial attention feature network, and the channel attention feature network, perform feature recalibration on the initial change information to obtain channel attention features and spatial attention features;

[0024] Multiply the spatial attention feature, the channel attention feature, and the initial transformation information to obtain fine-grained change information.

[0025] Preferably, the performing feature recalibration on the initial change information to obtain channel attention features includes:

[0026] Perform global average pooling operation on the initial change information along the channel direction to obtain a compressed one-dimensional feature vector;

[0027] Input the one-dimensional feature vector into the long short-term memory network for training to obtain a first intermediate feature vector;

[0028] Input the first intermediate feature vector into a network symmetric to the long short-term memory network for learning to obtain channel attention features.

[0029] Preferably, the performing feature recalibration on the initial change information to obtain spatial attention features includes:

[0030] Slice the initial change information pixel by pixel along the channel direction to obtain sliced features;

[0031] Input the sliced features into the long short-term memory network for training to obtain a second intermediate feature vector;

[0032] Reassemble the second intermediate feature vector according to the positions between different pixel features to obtain a two-dimensional spatial compression feature;

[0033] Extract features from the two-dimensional spatial compression feature according to the convolutional layer to obtain spatial attention features.

[0034] Preferably, the decoding operation includes:

[0035] Input the fine-grained change information into N groups of transposed convolution modules to obtain an initial output feature map;

[0036] Subtract the initial output feature map from the feature map corresponding to the initial change information to obtain a feature difference map;

[0037] Convert the feature difference map according to N groups of transposed convolution modules to obtain a set of change detection results, and determine the Nth detection result in the set of change detection results as the initial change detection result.

[0038] Preferably, evaluating the loss of the initial change detection result according to the combined loss function to obtain the loss value at the i-th time, includes:

[0039] According to the images of two adjacent time phases, obtaining the ground truth of the training data;

[0040] Resampling the ground truth by using bilinear interpolation method to obtain a set of ground truth images;

[0041] Corresponding the set of ground truth images and the set of change detection results in sequence, and adding them into the combined loss function to calculate and obtain a set of losses;

[0042] Weighting the set of losses to obtain the loss value at the i-th time.

[0043] In a second aspect, the present invention provides a remote sensing image building change detection device, including:

[0044] A data acquisition module, configured to acquire images of two adjacent time phases of training data;

[0045] An encoding module, configured to perform feature extraction and feature stacking on the images of two adjacent time phases by using a siamese convolutional neural network to obtain initial change information;

[0046] A training module, configured to train the initial change information by using a long short-term memory network to obtain fine change information;

[0047] A decoding module, configured to restore the fine change information by using deconvolution to obtain an initial change detection result;

[0048] An evaluation module, configured to construct a combined loss function, with the initial value of i being 1, perform the i-th evaluation operation: evaluating the loss of the initial change detection result according to the combined loss function to obtain the loss value at the i-th time;

[0049] An optimization module, configured to perform an optimization operation: using the Adam optimization algorithm to update the parameters of the siamese convolutional neural network and the long short-term memory network, and performing the encoding operation, the training operation, the decoding operation and the evaluation operation again, updating the loss value at the i-th time to the loss value at the (i + 1)-th time;

[0050] A first judgment module, configured to judge whether i satisfies i less than a preset number of training times N, if so, adding 1 to i and returning to perform the optimization operation; if not, performing a difference judgment operation;

[0051] The second judgment module is used to perform a difference judgment operation: judge whether the difference between the (i + 1)-th loss value and the i-th loss value is less than a preset threshold. If so, update the initial change detection result to the final change detection result of the building. If not, perform the optimization operation to update the i-th loss value and the (i + 1)-th loss value, and return to the difference judgment operation.

[0052] In a third aspect, the present invention further provides a terminal device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, the remote sensing image building change detection method described in any one of the above is implemented.

[0053] In a fourth aspect, the present invention further provides a computer-readable storage medium, which includes a stored computer program. When the computer program runs, it controls the device where the computer-readable storage medium is located to execute the remote sensing image building change detection method described in any one of the above.

[0054] Compared with the prior art, the present invention has the following beneficial effects:

[0055] By combining a siamese convolutional neural network and a long short-term memory network, the present invention can effectively extract the change information in the images of two different time phases and obtain an initial change detection result. By using a combined loss function to evaluate the loss of training data and combining the Adam optimization algorithm to update the parameters of the siamese convolutional neural network and the long short-term memory network, the texture features and spectral features in the images can be effectively extracted, the problem of unbalanced positive and negative samples in training can be improved, and thus the accuracy of remote sensing image building change detection can be improved. At the same time, the present invention can be applied to the buildings on rural homesteads. Description of the Drawings

[0056] Figure 1 is a schematic flowchart of the remote sensing image building change detection method provided by the first embodiment of the present invention;

[0057] Figure 2 is a schematic diagram of the remote sensing image building change detection provided by the embodiment of the present invention;

[0058] Figure 3 is another schematic diagram of the remote sensing image building change detection provided by the embodiment of the present invention;

[0059] Figure 4 is a schematic diagram of the encoding module provided by the embodiment of the present invention;

[0060] Figure 5 is a schematic diagram of the channel attention mechanism structure in the training stage provided by the embodiment of the present invention;

[0061] Figure 6 It is a schematic structural diagram of the spatial attention mechanism in the training stage provided by an embodiment of the present invention;

[0062] Figure 7 It is a schematic structural diagram of a remote sensing image building change detection device provided by the second embodiment of the present invention. Specific implementation manners

[0063] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0064] Referring to Figure 1 、 Figure 2 , the first embodiment of the present invention provides a method for detecting building changes in remote sensing images, including the following steps:

[0065] S11, obtaining the images of two time phases before and after the training data;

[0066] S12, encoding operation: using a siamese convolutional neural network to perform feature extraction and feature stacking on the images of the two time phases before and after, to obtain initial change information;

[0067] S13, training operation: using a long short-term memory network to train the initial change information, to obtain fine-grained change information;

[0068] S14, decoding operation: using deconvolution to restore the fine-grained change information, to obtain an initial change detection result;

[0069] S15, constructing a combined loss function, with the initial value of i being 1, performing the i-th evaluation operation: and performing loss evaluation on the initial change detection result according to the combined loss function, to obtain the i-th loss value;

[0070] S16, optimization operation: using the Adam optimization algorithm to update the parameters of the siamese convolutional neural network and the long short-term memory network, and performing the encoding operation, the training operation, the decoding operation, and the evaluation operation again, updating the i-th loss value to the (i + 1)-th loss value;

[0071] S17, judging whether i satisfies i less than a preset number of training times N. If so, adding 1 to i and returning to perform the optimization operation; if not, performing the following difference judgment operation;

[0072] S18. Determine whether the difference between the (i + 1)-th loss value and the i-th loss value is less than a preset threshold. If so, update the initial change detection result to the final change detection result of the building. If not, perform the optimization operation to update the i-th loss value and the (i + 1)-th loss value, and return to the difference judgment operation.

[0073] It should be noted that this embodiment is designed and implemented using the Python language, and the deep learning framework used is TensorFlow.

[0074] In step S11, acquire the pre- and post-phase images of the training data. Exemplarily, for the drawn ground truth map, assign the pixels of the changed part as 1 and the unchanged part as 0. Further, for the training data, the training data of a single training sample can be obtained in a random batch manner.

[0075] In one implementation, the training data and the test data can also be split in a ratio of 7:3.

[0076] In this embodiment, both the training data and the test data are cropped to a size of 256 * 256 or 512 * 512. Using the above cropping size, when performing deconvolution in step S14, the output result can be restored to the size of the original input image without the need for upsampling. Preferably, the size of the input image in this embodiment is 256 * 256 * 3.

[0077] In step S12, the siamese convolutional neural network includes a feature extraction module and a feature stacking module.

[0078] Specifically, the encoding operation specifically includes:

[0079] Extract features according to the feature extraction module; wherein, the feature extraction module uses two groups of convolutional neural networks with shared weights;

[0080] Stack the two-phase features extracted by the feature extraction module according to the feature stacking module to obtain the stacked features;

[0081] Learn the stacked features using a convolutional module to obtain the initial change information.

[0082] It should be noted that the initial change information can be presented in the form of a feature map.

[0083] Specifically, for the feature extraction module, two groups of convolutional neural networks with shared weights are used to extract features. The two groups of convolutional neural networks with shared weights are composed of twenty convolutional modules. Each convolutional module consists of a convolutional layer, BatchNormolization, and a non-linear activation function ReLU function. Each convolutional layer uses a 3*3 convolution, and the stride of a single convolutional module is 2.

[0084] Among them, the meaning of shared weights is that the convolutional kernels of the convolutional layers at corresponding positions in the two channels for extracting features adopt the same set of weight parameters. After every five convolutional modules, a max pooling layer is used to reduce the feature map to half of its original size, and the number of feature channels is increased to twice the original. After twenty convolutional modules, the H and W of the feature map become one-sixteenth of the original, and the number of feature channels increases to sixteen times the original.

[0085] In this embodiment, the sizes of the feature maps of the four groups of convolutional modules in the feature extraction module are 256*256*4, 128*128*16, 64*64*64, and 32*32*256 in sequence. In addition, after every five convolutional modules, the change information at the current stage is obtained by taking the difference between the feature maps of the previous and current time phases. This difference will be connected to the deconvolutional layer of the same size in the decoding process by means of skip connection.

[0086] Specifically, for the feature stacking module, the two-time-phase features extracted by the shared-weight feature extraction are stacked in the channel direction to obtain the stacked features. Then, two convolutional modules are used to further learn the stacked features. Each convolutional module consists of a convolutional layer, Batch Normolization, and a non-linear activation function ReLU function, and each convolutional layer uses a 3*3 convolution.

[0087] In this embodiment, the sizes of the feature maps in the feature stacking module are 32*32*256, 32*32*128, and 32*32*64 in sequence. The feature stacking module mainly performs a feature compression in the channel direction and does not change the height and width of the feature map.

[0088] In step S13, the training operation specifically includes:

[0089] Slice the initial change information pixel by pixel along the channel direction to obtain the sliced features;

[0090] Input the sliced features into the long short-term memory network for training and output the trained features;

[0091] Re-piece together the trained features according to the positions between different pixel features to obtain the refined change information.

[0092] It should be noted that the fine-grained change information can be presented in the form of a feature map.

[0093] In this example, an improved network GRU of the long short-term memory network is used for feature learning and training. First, the feature map (with a size of 32*32*64) extracted from the Siamese neural network is traversed for the channel direction features of pixels, and the feature size is 1*1*C. Then, the feature is sliced and extracted. After that, through the learning and training of the long short-term memory network, the trained features are concatenated according to the original input, and the size of the concatenated feature map is H*W*C.

[0094] In step S14, a continuous eight-layer decoding network is used to restore the fine-grained change information in step S13 to the size of the original input image, obtaining the initial change detection result.

[0095] Specifically, the eight-layer decoding network consists of four consecutive transposed convolution modules. Each transposed convolution module consists of a transposed convolution layer and a convolution layer. The transposed convolution layer restores the size of the current feature map to twice the original size, and then further processes it using the Batch Normolization and ReLU activation functions. In the convolution layer of the module, first, the initial change information in step S12 and the current feature map are stacked, and then a 1*1 convolution is used to reduce the current feature dimension by half. After passing through four consecutive groups of transposed convolution modules, the feature map is restored to the size of H*W*1. Finally, a softmax function is used to convert the feature map into the probabilities corresponding to pixel changes and unchanged, and the initial change detection result of the building is output.

[0096] In step S15, a combined loss function is constructed. With the initial value of i being 1, the i-th evaluation operation is performed: the loss evaluation of the initial change detection result is carried out according to the combined loss function to obtain the i-th loss value. Specifically, the Focal loss function and the Dice loss function are combined to obtain the combined loss function. The combined loss function can handle the segmentation problem in irregular small-area buildings and can increase the stability during the training process.

[0097] Among them, the calculation formula of the Focal loss is as follows:

[0098]

[0099] Among them, L fl is the focal loss, α is the balance factor, γ is the training weight factor for simple samples, y′ is the sample prediction value, and y is the sample true label.

[0100] The calculation formula of the Dice loss is as follows:

[0101]

[0102] Among them, L Dice is the Dice loss, y' is the sample prediction value, and y is the sample true label.

[0103] The calculation formula of the combined loss function is as follows:

[0104] L combination = L fl + λL Dice

[0105] Among them, L combination is the combined loss, and λ is the balance coefficient. Specifically, in this embodiment, λ is 0.5.

[0106] In order to reduce the possibility of overfitting of the network during the training process, the network will add a regularization method for optimization. Specifically, in this embodiment, the L2 regularization method is used for optimization, and the regularization loss will also be added to the final combined loss function.

[0107] In step S16, perform the optimization operation: use the Adam optimization algorithm to update the parameters of the siamese convolutional neural network and the long short-term memory network, and perform the encoding operation, the training operation, the decoding operation, and the evaluation operation again, and update the i-th loss value to the (i + 1)-th loss value.

[0108] Specifically, the process of optimizing using the Adam optimization algorithm is as follows:

[0109]

[0110]

[0111]

[0112]

[0113]

[0114] Among them, is the first-order estimation moment parameter, is the second-order estimation moment parameter, is the exponential decay rate of the first estimation, β 2 is the exponential decay rate of the second estimation, ε is a smoothing term, w t is the weight of the network, ΔL t is the gradient of backpropagation.

[0115] In this embodiment, performing the encoding operation, the training operation, the decoding operation, and the evaluation operation again, and updating the i-th loss value to the (i + 1)-th loss value is a process of backpropagation of data. During the backpropagation process, a decaying learning rate strategy is adopted for dynamic adjustment of the learning rate. Generally, in the early stage of network training, the learning difficulty is relatively small, and a larger learning rate is required. In the later stage of network training, more refined adjustment of network parameters is needed.

[0116] In step S17, it is judged whether i satisfies i < the preset number of training times N. If so, i is incremented by one and the optimization operation is returned to be executed; if not, the difference judgment operation in step S18 is executed:

[0117] It is judged whether the difference between the (i + 1)-th loss value and the i-th loss value is less than the preset threshold. If so, the initial change detection result is updated to the final change detection result of the building; if not, the optimization operation is executed to update the i-th loss value and the (i + 1)-th loss value, and then it returns to the difference judgment operation.

[0118] It should be noted that the value of the number of training times N can be preset as needed. For example, when N is taken as 500, it means that after 500 optimization updates of the siamese convolutional neural network and the long short-term memory network, whether to output the final change detection result of the building will be considered.

[0119] In addition, the threshold is set according to the requirements of detection accuracy. The higher the detection accuracy requirement, the smaller the set threshold. Exemplarily, when the threshold is taken as 0.1, when the difference between the (i + 1)-th loss value and the i-th loss value is less than 0.1, it can be considered that the (i + 1)-th loss value no longer decreases. At this time, the siamese convolutional neural network and the long short-term memory network are optimal, and the initial change detection result at this time can be updated to the final change detection result of the building and finally output.

[0120] For the convenience of understanding the present invention, in combination with Figures 3 - 6 , the following will further describe some preferred embodiments of the present invention.

[0121] In one implementation, the encoding operation includes:

[0122] Feature extraction is performed according to multiple groups of convolutional modules with different scales, and the input features of the previous group of convolutional modules are calculated to obtain features with different scales; wherein, the siamese convolutional neural network is composed of N groups of convolutional modules spliced in sequence;

[0123] The features with different scales are stacked along the channel direction to obtain an initial stacked feature, and the initial stacked feature is sliced along the channel direction to obtain an initial sliced feature;

[0124] The initial slice features are sequentially input into a long short-term memory network for training to obtain trained features;

[0125] The trained features are arranged according to the initial positions to obtain the initial features after training;

[0126] The initial features after training are activated to obtain a feature extraction result corresponding to the current convolutional module, and the initial change information is obtained by passing through N groups of convolutional modules in sequence.

[0127] It should be noted that during the encoding process, the adopted siamese convolutional neural network is composed of N groups of convolutional modules with different scales spliced in sequence. Each group of convolutional modules is used to execute the above encoding operation steps, and N is an integer greater than or equal to 2.

[0128] In specific implementation, first, convolutional modules with multiple different scales are used to extract features of different scales. The convolutional modules are represented as Conv ∈ {R 1×1 , R 3×3 , R 5×5 , R 7×7}, and then the input features of the previous group of convolutional modules are calculated to obtain features of different scales. Among them, the input features are represented as:

[0129]

[0130] The features of different scales are represented as:

[0131]

[0132] Then, the features of different scales are stacked along the channel direction to obtain the initial stacked features, represented as: Then, the initial stacked features are sliced along the channel direction to obtain the initial slice features, represented as:

[0133] F set ={f concat_slice1 ∈ R 1×1×2NC , f concat_slice2 ∈ R 1×1×2NC …, f concat_slice ∈ R 1×1×2NC}

[0134] Next, the stacked features are subjected to pixel-by-pixel feature extraction through an LSTM. The initial slice features are input into a long short-term memory network LSTM for training, and the output trained features are represented as:

[0135] F set2 ={f initial_slice1 ∈ R 1×1×2NC , f initial_slice2 ∈ R1×1×2NC …, f initial_slice ∈ ℝ 1×1×2NC}

[0136] Then, arrange the trained features according to the initial positions to obtain the initial features after training:

[0137]

[0138] Finally, pass the initial features after training through the BN operation (Batch normalization) and the leaky - relu activation function to obtain the feature extraction result of this group of convolutional modules, denoted as:

[0139] f initial = BatchNormalization(leaky_relu(f initial_raw ))

[0140] Then, pass through N groups of convolutional modules in sequence to obtain a group of results of initial feature extraction, as the initial change information, denoted as:

[0141]

[0142] wherein, as the input of the training operation.

[0143] In one implementation, the training operation includes:

[0144] Construct a spatial attention feature network and a channel attention feature network;

[0145] According to the long - short - term memory network, the spatial attention feature network, and the channel attention feature network, perform feature recalibration on the initial change information to obtain channel attention features and spatial attention features;

[0146] Multiply the spatial attention features, the channel attention features, and the initial transformation information to obtain the refined change information.

[0147] It should be noted that in this embodiment, the long - short - term memory spatio - spectral joint attention mechanism unit is used to train the initial change information. Among them, the spatial attention feature network is denoted as:

[0148]

[0149] The channel attention feature network is denoted as:

[0150] f ChannelAttention ∈ ℝ 1×1×2NC

[0151] The refined change information is denoted as:

[0152]

[0153] Among them, the main calculation formula of the LSTM is as follows:

[0154] σ = Sigmoid(x)

[0155] f t = σ(W f [h t-1 , x t + b f )

[0156] i t = σ(W i [h t-1 , x t + b i )

[0157]

[0158]

[0159] o t = σ(W o [h t-1 , x t + b o )

[0160] h t = o t tanh(C t )

[0161] Among them, x t is the input at the t-th moment, h t-1 is the hidden layer state at the (t - 1)-th moment, h t is the hidden layer state at the t-th moment, where i t is the input gate at the t-th moment, f t is the forget gate at the t-th moment, is the temporary cell state, C t is the current cell state, C t-1 is the cell state at the (t - 1)-th moment, o t is the value of the output gate. W and b are the weight parameters and bias parameters of the network respectively.

[0162] Furthermore, the feature recalibration of the initial change information to obtain the channel attention feature includes:

[0163] Performing global average pooling operation on the initial change information along the channel direction to obtain a compressed one-dimensional feature vector;

[0164] Input the one-dimensional feature vector into the long short-term memory network for training to obtain a first intermediate feature vector;

[0165] Input the first intermediate feature vector into a network symmetric to the long short-term memory network for learning to obtain channel attention features.

[0166] Specifically, perform global average pooling operation on the initial change information along the channel direction, so that the initial feature is compressed from a three-dimensional tensor into a one-dimensional feature vector (with a length of C). The formula used is as follows:

[0167]

[0168] Then, input the one-dimensional feature vector into the long short-term memory network for learning to obtain a first intermediate feature vector. The length of the input feature of this part of the long short-term memory network is C, and the length of the output feature is N*C (N>=2). The formula is expressed as:

[0169] f ChannelProgress (k) = LSTM(f average )

[0170] Input the first intermediate feature vector into the second long short-term memory network for learning to obtain channel attention features. Among them, the second long short-term memory network is a network symmetric to the long short-term memory network. The length of the input feature of this part of the long short-term memory network is N*C, and the length of the output feature is C. The formula is expressed as:

[0171] f ChannelAttention (k) = LSTM(f ChannelProgress (k))

[0172] Furthermore, the feature recalibration of the initial change information to obtain spatial attention features includes:

[0173] Slice the initial change information pixel by pixel along the channel direction to obtain the sliced features;

[0174] Input the sliced features into the long short-term memory network for training to obtain a second intermediate feature vector;

[0175] Reassemble the second intermediate feature vector according to the positions between different pixel features to obtain a two-dimensional spatial compression feature;

[0176] Extract features from the two-dimensional spatial compression feature according to the convolutional layer to obtain spatial attention features.

[0177] Specifically, slice the initial change information pixel by pixel along the channel direction to obtain the sliced features. The set of sliced features is expressed as:

[0178] F spatial_slice ={f spatial_slice1 ∈R 1×1×2NC , f spatial_slice2 ∈R 1×1×2NC ,…};

[0179] Input the sliced feature f slice ∈R 1×1×8C into the long short-term memory network for training, and output the trained feature, i.e., the second intermediate feature vector f reduce ∈R 1×1×1 , which is expressed as:

[0180] F spatial_reduce ={f spatial_reduce1 ∈R 1×1×1 , f spatial_reduce2 ∈R 1×1×1 ,…}

[0181] Then, reassemble according to the positions between different pixel features to obtain the two-dimensional space compression feature, which is expressed as:

[0182]

[0183] Next, use a 3*3 convolutional layer to extract features from the two-dimensional space compression feature to obtain a process vector:

[0184] Then use a 3*3 convolutional layer to continue extracting features from the process vector to obtain the spatial attention feature, which is expressed as:

[0185]

[0186] Finally, multiply the spatial attention feature ChannelAttention ∈R 1×1×2NC and the channel attention feature f and the initial change information

[0187] f rebuild (i, j, k)=f SpatialAttention (i, j)·f ChannelAttention (k)·f initial (i, j, k)

[0188] wherein,

[0189] In one implementation, the decoding operation includes:

[0190] Input the fine change information into N sets of deconvolution modules to obtain an initial output feature map;

[0191] Subtract the feature map corresponding to the initial change information from the initial output feature map to obtain a feature difference map;

[0192] Convert the feature difference map according to N sets of deconvolution modules to obtain a set of change detection results, and determine the Nth detection result in the set of change detection results as the initial change detection result.

[0193] It should be noted that the decoding operation consists of N sets of deconvolution modules, and the initial output feature map is obtained by inputting the fine change information, which is expressed as:

[0194]

[0195] By subtracting the initial output feature map DC set and the feature map with the same size as the initial change information F set a feature difference map Diff set is obtained, which is expressed as:

[0196]

[0197] Diff set is converted into a set of change detection results through N sets of 1*1 deconvolution modules Conv1×1, which is expressed as:

[0198]

[0199] Among them, is used as the final result of the decoding operation, that is, the initial change detection result; the rest is regarded as the intermediate result of the decoding module.

[0200] In one implementation, the loss evaluation of the initial change detection result according to the combined loss function to obtain the ith loss value includes:

[0201] Obtain the ground truth of the training data according to the two-phase images before and after;

[0202] Resample the ground truth using bilinear interpolation to obtain a set of ground truth images;

[0203] Correspond the set of ground truth images and the set of change detection results in sequence and add them to the combined loss function to calculate a set of losses;

[0204] Weight the set of losses to obtain the ith loss value.

[0205] It should be noted that the combined loss strategy in this embodiment consists of multiple parts of losses. The ground truth of the training data is represented as: T ∈ I H×W×1 , the images of the two adjacent time phases are represented as I 1 , I 2 ∈ R H×W×C . The ground truth is 0 or 1, corresponding to the images of the two adjacent time phases being unchanged and changed, respectively.

[0206] T ∈ I is resampled by bilinear interpolation to obtain a set of ground truth images, which is represented as: H×W×1

[0207]

[0208] Then, T set and O set are respectively added to the combined loss function in sequence to calculate a set of losses:

[0209] L set ={L 1 , L 2 ,…L N}

[0210] The set of losses is weighted to finally obtain the loss value at the i-th time:

[0211] L final = λ 1 L 1 + λ 2 L 2 + λ 3 L 3 +…+ λ N L N

[0212] By combining a siamese convolutional neural network and a long short-term memory network, the present invention can effectively extract the change information in the images of the two adjacent time phases and obtain an initial change detection result; the long short-term memory squeeze-and-excitation unit proposed in this patent is used to optimize and learn the initial change information, so as to learn the temporal features. The training data is evaluated for loss through a combined loss function, and the parameters of the siamese convolutional neural network and the long short-term memory network are updated by combining with the Adam optimization algorithm, which can effectively improve the problem of unbalanced positive and negative samples in training, thereby improving the accuracy of building change detection in remote sensing images. At the same time, the present invention can be applied to buildings on rural homesteads.

[0213] Referring to Figure 7 , the second embodiment of the present invention provides a device for building change detection in remote sensing images, including:

[0214] A data acquisition module, configured to acquire the images of the two adjacent time phases of the training data;​

[0215] An encoding module, configured to use a Siamese convolutional neural network to extract and stack features from the front and back two-phase images, so as to obtain initial change information;

[0216] A training module, configured to use a long short-term memory network to train the initial change information to obtain fine-grained change information;

[0217] A decoding module, configured to use deconvolution to restore the fine-grained change information to obtain an initial change detection result;

[0218] An evaluation module, configured to construct a combined loss function, with the initial value of i being 1, and perform the i-th evaluation operation: perform loss evaluation on the initial change detection result according to the combined loss function to obtain the i-th loss value;

[0219] An optimization module, configured to perform an optimization operation: use the Adam optimization algorithm to update the parameters of the Siamese convolutional neural network and the long short-term memory network, and perform the encoding operation, the training operation, the decoding operation, and the evaluation operation again, and update the i-th loss value to the (i + 1)-th loss value;

[0220] A first determination module, configured to determine whether i satisfies i less than a preset number of training times N. If so, increment i by 1 and return to perform the optimization operation; if not, perform a difference determination operation;

[0221] A second determination module, configured to perform a difference determination operation: determine whether the difference between the (i + 1)-th loss value and the i-th loss value is less than a preset threshold. If so, update the initial change detection result to the final change detection result of the building; if not, perform the optimization operation to update the i-th loss value and the (i + 1)-th loss value, and return to the difference determination operation.

[0222] An embodiment of the present invention further provides a terminal device. The terminal device includes: a processor, a memory, and a computer program stored in the memory and executable on the processor, such as a remote sensing image building change detection program. When the processor executes the computer program, the steps in the above-mentioned embodiments of various remote sensing image building change detection methods are implemented, such as Figure 1 The step S11 shown. Alternatively, when the processor executes the computer program, the functions of each module / unit in the above-mentioned device embodiments are implemented, such as the loss evaluation module.

[0223] Exemplarily, the computer program may be divided into one or more modules / units, which are stored in the memory and executed by the processor to implement the present invention. The one or more modules / units may be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the terminal device.

[0224] The terminal device may be a computing device such as a desktop computer, a notebook, a palm computer, and a smart tablet. The terminal device may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that the above components are only examples of the terminal device and do not constitute a limitation on the terminal device. It may include more or fewer components than the above, or combine certain components, or different components. For example, the terminal device may further include input / output devices, network access devices, a bus, etc.

[0225] The so-called processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), graphics processing units (GPUs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor is the control center of the terminal device and connects various parts of the entire terminal device through various interfaces and lines.

[0226] The memory may be used to store the computer program and / or modules. The processor realizes various functions of the terminal device by running or executing the computer program and / or modules stored in the memory, and by calling the data stored in the memory. The memory may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created according to the use of the terminal device (such as audio data, a phone book, etc.), etc. In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart media card (SMC), a flash card, at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0227] Among them, if the modules / units integrated in the terminal device are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-described embodiment methods of the present invention, it can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0228] It should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the device embodiments provided by the present invention, the connection relationship between the modules indicates that there is a communication connection between them, which can be specifically implemented as one or more communication buses or signal lines. Those of ordinary skill in the art can understand and implement it without creative effort.

[0229] The above-described specific embodiments have further elaborated on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. It is particularly pointed out that for those skilled in the art, any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for detecting building changes in remote sensing images, characterized in that, it includes: Obtain the images of two time phases before and after the training data; Encoding operation: Use a siamese convolutional neural network to extract features and stack features from the images of the two time phases before and after, to obtain initial change information; Training operation: Use a long short-term memory network to train the initial change information to obtain fine change information; Decoding operation: Use deconvolution to restore the fine change information to obtain an initial change detection result; Construct a combined loss function, with the initial value of i being 1, and perform the i-th evaluation operation: Perform loss evaluation on the initial change detection result according to the combined loss function to obtain the i-th loss value; Optimization operation: Use the Adam optimization algorithm to update the parameters of the siamese convolutional neural network and the long short-term memory network, and perform the encoding operation, the training operation, the decoding operation and the evaluation operation again, and update the i-th loss value to the (i + 1)-th loss value; Judge whether i satisfies i less than the preset number of training times N. If so, add 1 to i and return to perform the optimization operation; if not, perform the following difference judgment operation: Judge whether the difference between the (i + 1)-th loss value and the i-th loss value is less than the preset threshold. If so, update the initial change detection result to the final change detection result of the building; if not, perform the optimization operation to update the i-th loss value and the (i + 1)-th loss value, and return to the difference judgment operation; Among them, the decoding operation includes: Input the fine change information into N groups of deconvolution modules to obtain an initial output feature map; Subtract the feature map corresponding to the initial change information from the initial output feature map to obtain a feature difference map; According to N groups of deconvolution modules, convert the feature difference map to obtain a set of change detection results, and determine the N-th detection result in the set of change detection results as the initial change detection result; The performing loss evaluation on the initial change detection result according to the combined loss function to obtain the i-th loss value includes: According to the images of the two time phases before and after, obtain the result truth value of the training data; Use bilinear interpolation to resample the result truth value to obtain a set of truth images; Correspond the set of truth images and the set of change detection results in sequence, and add them to the combined loss function to calculate and obtain a set of losses; Weight the set of losses to obtain the i-th loss value.

2. The method for detecting building changes in remote sensing images according to claim 1, characterized in that, the encoding operation includes: Perform feature extraction according to multiple groups of convolution modules with different scales, and calculate the output features of the previous group of convolution modules to obtain features with different scales; among them, the siamese convolutional neural network is composed of N groups of convolution modules spliced in sequence; Stack the features with different scales along the channel direction to obtain an initial stacked feature, and slice the initial stacked feature along the channel direction to obtain an initial sliced feature; Input the initial slice features into a long short-term memory network for training in sequence to obtain the trained features; Arrange the trained features according to the initial positions to obtain the initial features after training; Activate the initial features after training to obtain the feature extraction results corresponding to the current convolutional module, and sequentially pass through N groups of convolutional modules to obtain the initial change information.

3. The remote sensing image building change detection method according to claim 1, characterized in that, the training operation includes: Construct a spatial attention feature network and a channel attention feature network; According to the long short-term memory network, the spatial attention feature network and the channel attention feature network, perform feature recalibration on the initial change information to obtain channel attention features and spatial attention features; Multiply the spatial attention feature, the channel attention feature and the initial change information to obtain the fine change information.

4. The remote sensing image building change detection method according to claim 3, characterized in that, the performing feature recalibration on the initial change information to obtain channel attention features includes: Perform global average pooling operation on the initial change information along the channel direction to obtain a compressed one-dimensional feature vector; Input the one-dimensional feature vector into the long short-term memory network for training to obtain a first intermediate feature vector; Input the first intermediate feature vector into a network symmetric to the long short-term memory network for learning to obtain channel attention features.

5. The remote sensing image building change detection method according to claim 3, characterized in that, the performing feature recalibration on the initial change information to obtain spatial attention features includes: Perform slicing processing on the initial change information pixel by pixel along the channel direction to obtain the sliced features; Input the sliced features into the long short-term memory network for training to obtain a second intermediate feature vector; Reassemble the second intermediate feature vector according to the positions between different pixel features to obtain a two-dimensional spatial compression feature; Extract features from the two-dimensional spatial compression feature according to the convolutional layer to obtain spatial attention features.

6. A remote sensing image building change detection device, characterized in that, comprises: A data acquisition module for acquiring the front and back two-phase images of the training data; An encoding module for performing feature extraction and feature stacking on the front and back two-phase images by using a siamese convolutional neural network to obtain initial change information; A training module for training the initial change information by using a long short-term memory network to obtain fine change information; A decoding module for restoring the fine change information by using deconvolution to obtain the initial change detection result; An evaluation module for constructing a combined loss function, with the initial value of i being 1, perform the i-th evaluation operation: perform loss evaluation on the initial change detection result according to the combined loss function to obtain the i-th loss value; An optimization module for performing optimization operations: using the Adam optimization algorithm to update the parameters of the siamese convolutional neural network and the long short-term memory network, and executing the encoding module, the training module, the decoding module, and the evaluation module again, updating the i-th loss value to the (i + 1)-th loss value; A first judgment module for judging whether i satisfies i less than a preset number of training times N. If so, increment i by one and return to execute the optimization operation; if not, execute a difference judgment operation; A second judgment module for performing a difference judgment operation: judging whether the difference between the (i + 1)-th loss value and the i-th loss value is less than a preset threshold. If so, update the initial change detection result to the final change detection result of the building; if not, execute the optimization operation to update the i-th loss value and the (i + 1)-th loss value, and return to the difference judgment operation; Wherein, the decoding module includes: Inputting the fine change information into N groups of deconvolution modules to obtain an initial output feature map; Subtracting the initial output feature map from the feature map corresponding to the initial change information to obtain a feature difference map; Converting the feature difference map according to N groups of deconvolution modules to obtain a set of change detection results, and determining the N-th detection result in the set of change detection results as the initial change detection result; The obtaining the i-th loss value by evaluating the loss of the initial change detection result according to the combined loss function includes: Obtaining the result truth value of the training data according to the two-phase images before and after; Resampling the result truth value by using bilinear interpolation to obtain a set of truth value images; Sequentially corresponding the set of truth value images and the set of change detection results respectively, and adding them into the combined loss function to calculate a set of losses; Weighting the set of losses to obtain the i-th loss value.

7. A terminal device, Characterized in that, It includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the remote sensing image building change detection method according to any one of claims 1 to 5.

8. A computer-readable storage medium, Characterized in that, The computer-readable storage medium includes a stored computer program. Wherein, when the computer program runs, it controls the device where the computer-readable storage medium is located to execute the remote sensing image building change detection method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Remote sensing image building change detection method and device, equipment and storage medium

    CN113901877A