Remote sensing image change detection method based on deep convolutional network

By performing linear transformation of remote sensing image features and pyramid feature extraction, combined with the change detection method of deep convolutional network, the problem of insufficient change detection accuracy and accuracy in high-resolution images is solved, and higher detection accuracy and robustness are achieved.

CN116403123BActive Publication Date: 2025-06-06BEIJING DATA INTELLIGENCE INFORMATION TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202310455203.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-25
Publication Date
2025-06-06
Estimated Expiration
2043-04-25

AI Technical Summary

Technical Problem

The existing CNN-based remote sensing image change detection method is easily affected by the "change of unchanged" and "change of unchanged" between the two phases of images in high-resolution images, resulting in insufficient information exchange and insufficient accuracy and accuracy of detection results.

Method used

By performing different linear transformations on features, characterizing the similarity of different focus, and constructing a remote sensing image change detection method based on deep convolutional networks, including pyramid feature extraction, linear mapping and vector internal product calculation to generate differential features with stronger expression capabilities.

Benefits of technology

It improves the accuracy and accuracy of change detection, enhances the robustness of feature extraction, reduces false detection, and improves the reliability of detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116403123B_ABST
    Figure CN116403123B_ABST
Patent Text Reader

Abstract

The present invention provides a remote sensing image change detection method based on a deep convolutional network, which relates to the field of change detection, including: S1 obtaining a pre-phase image and a post-phase image of a target area, and pre-processing and data enhancement of the pre-phase image and the post-phase image; S2 constructing a change detection model and pre-training the change detection model, the change detection model includes a feature extraction network, a change contrast network, and an output network; S3 performing pyramid feature extraction on the pre-phase image and the post-phase image through the feature extraction network to obtain a first feature and a second feature; S4 inputting the first feature and the second feature into the change contrast network, performing linear mapping respectively, obtaining a first feature vector and a second feature vector, calculating the vector inner product of the first feature vector and the second feature vector to obtain a difference feature; S5 inputting the difference feature into the output network, and processing the difference feature to obtain a change detection result. The present invention can improve the accuracy and precision of change detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of change detection, and in particular to a remote sensing image change detection method based on a deep convolutional network. Background Art

[0002] The earth is constantly changing due to natural and human factors. Identifying these changes is important for understanding the environment in which we live. Remote sensing change detection is the process of obtaining dynamic changes of surface objects or phenomena by observing the same area at different times. Since the acquisition of multi-temporal remote sensing images, the use of remote sensing images to identify changes has become an important application in the fields of environmental monitoring, land management, disaster assessment, etc. Remote sensing images have the characteristics of wide coverage, fast revisit cycle and high resolution. With the improvement of image quality and resolution, while providing a rich data source for change detection, it also puts higher requirements on change detection algorithms.

[0003] Change detection research began in the 1970s, with the goal of identifying changes of interest and removing interfering irrelevant changes. In the early stages of research, researchers used simple algebraic methods to identify the probability of change in large areas from low-resolution images. With the improvement of image resolution and the advancement of machine learning technology, transformation methods, classification methods, and object-oriented methods have begun to be applied. These methods have achieved good results on medium and low-resolution images, and can provide relatively reliable change detection results semi-automatically or automatically. With the improvement of image resolution, the ability to express details of images is enhanced. On the one hand, the granularity of identifying objects is enhanced; on the other hand, the phenomenon of different spectra of the same object and the same spectra of different objects is increased, resulting in a large number of false detections. Therefore, in view of the characteristics of high-resolution images, robust feature extraction methods and change detection methods are needed to achieve effective feature extraction, detect effective changes and suppress the interference of invalid changes. The rise of convolutional neural networks (CNNs), especially fully convolutional neural networks, has provided important tools for change detection methods. At present, many methods have introduced convolutional neural networks into remote sensing image analysis and change detection, and have achieved results far exceeding traditional machine learning methods.

[0004] Existing CNN-based change detection methods can be divided into single-stream methods, dual-stream methods, and multi-model integration methods. The single-stream method fuses two-period images or their transformations and directly uses a deep learning model for change detection. This method tends to ignore the spatiotemporal information between the two-period images. The dual-stream method is the most common change detection structure, which is further divided into twin structures, structures based on transfer learning, and post-classification structures. Among them, the method based on the twin structure is the current mainstream method. Since change detection often involves two or more periods, the deep twin network is the most common basic feature extraction module. However, most of the existing methods design change information extraction methods from the feature level, such as subtraction, cascade, attention, and recursive neural networks, and directly use the characteristics of these operations to extract changes. Simple subtraction and cascade operations are used to fully explore the relationship between the two-period images, which makes the algorithm susceptible to the influence of "change seems unchanged" and "unchanged seems changed" between the two-period remote sensing images, and cannot make the information between the two-period images interoperable. The content of the mined remote sensing data is not rich enough, resulting in the change detection results still not reaching the required accuracy and precision. Summary of the invention

[0005] In order to solve the above technical problems, the present invention performs different linear transformations on the features to characterize the similarities with different emphases, thereby obtaining features with stronger expressive power, and achieving higher accuracy and precision when performing change detection.

[0006] In order to achieve the above technical objectives, the present invention provides a remote sensing image change detection method based on a deep convolutional network, comprising:

[0007] S1 obtains the pre-phase image and the post-phase image of the target area, and performs pre-processing and data enhancement on the pre-phase image and the post-phase image;

[0008] S2 builds a change detection model and performs pre-training. The change detection model includes a feature extraction network, a change comparison network, and an output network.

[0009] S3 performs pyramid feature extraction on the front phase image and the back phase image respectively through the feature extraction network to obtain the first feature and the second feature;

[0010] S4 inputs the first feature and the second feature into the change comparison network, performs linear mapping respectively, obtains the first feature vector and the second feature vector, calculates the vector inner product of the first feature vector and the second feature vector to obtain the difference feature;

[0011] S5 inputs the difference features into the output network and processes the difference features to obtain the change detection results.

[0012] In one embodiment of the present invention, step S4 includes:

[0013] S41 linearly maps the first feature and the second feature, transforming them into multiple linear spaces in pairs, to obtain multiple first feature vectors and second feature vectors;

[0014] S42 calculates the vector inner product of the first eigenvector and the second eigenvector in the same linear space to obtain similar features of the linear space;

[0015] S43 fuses similar features of all linear spaces to obtain difference features.

[0016] In an embodiment of the present invention, the plurality of linear spaces are not identical, and the similarity feature of each linear space represents the similarity between the first eigenvector and the second eigenvector in the linear space.

[0017] In one embodiment of the present invention, when the change detection model is pre-trained, the weights are initialized to a random normal distribution with a unit matrix added thereto.

[0018] In one embodiment of the present invention, the pre-training process of the change detection model is as follows:

[0019] Get a dataset with real labels and divide it into training samples and validation samples;

[0020] The AdamW optimizer is used in combination with the cosine annealing algorithm to optimize the learning rate. The initial change detection model is trained based on the training samples to obtain the predicted labels of the training samples. The model weights are adjusted using the validation samples until the model converges.

[0021] During the training process, a loss function is introduced to calculate the gap between the predicted label and the true label.

[0022] In one embodiment of the present invention, the true label is automatically calculated according to a labeling method, and the labeling method includes:

[0023] The dataset consists of multiple image pairs of previous and next phases. The image bands of the image pairs are superimposed, and then the first three principal component images are extracted using the principal component analysis method, and then superpixel segmentation is performed on them to obtain multiple superpixel blocks.

[0024] Extract features from superpixel blocks, including spectral features, texture features, and spatial features;

[0025] Taking superpixel blocks as units, the difference value of each feature is calculated, that is, the spectral feature difference value, texture feature difference value, and spatial feature difference value between the corresponding two superpixel blocks are calculated for each image pair;

[0026] The spectral feature difference value, texture feature difference value, and spatial feature difference value are combined into a vector as the difference vector of the corresponding superpixel block. The difference vector is modulo and normalized to obtain the change value of the superpixel block. The value range is [0,1]. The threshold is set to classify the superpixel block into two categories: changed and unchanged.

[0027] According to the classification results of the superpixel blocks, the classification results of each pixel in the dataset are obtained as the true label of the dataset.

[0028] In one embodiment of the present invention, the loss function includes:

[0029] The cross entropy loss is calculated by predicting the label and the real label, where the cross entropy loss function is used to measure the similarity between the predicted change pixels and the real change pixels;

[0030] The Dice loss is calculated by predicting the label and the true label. The Dice loss function is used to alleviate the problem of imbalance in the number of changed and unchanged samples in the change detection process.

[0031] Add the cross entropy loss function and the Dice loss function to get the total loss function.

[0032] In one embodiment of the present invention, step S3 includes:

[0033] The front phase image and the back phase image are respectively input into the feature extraction network to construct a pyramid, and multi-scale features are extracted according to the scale space of the pyramid, and the first feature and the second feature are obtained after output;

[0034] Among them, the feature extraction network includes a 4*4 convolutional layer, a first residual block, a second residual block, a third residual block, and a fourth residual block. Among them, the first residual block includes 3 residual units, the second residual block includes 3 residual units, the third residual block includes 9 residual units, and the fourth residual block includes 3 residual units. Each residual unit contains a 7*7 depth separation convolution layer, a normalization layer, a 1*1 convolution layer, and an activation function layer. Shortcuts are set between the residual units for connection.

[0035] In one embodiment of the present invention, step S5 includes:

[0036] S51 output network includes 1*1 convolution layer, feature concatenation layer, and MLP layer;

[0037] S52 adjusts the channel through a 1*1 convolution layer, resamples and concatenates the difference features, and obtains the change features;

[0038] S53 performs pixel-level change determination on the change feature through MLP, and outputs the change detection result.

[0039] In one embodiment of the present invention, step S53 includes:

[0040] The value range of the pixel value in the change feature is set to [0,1], and the classification threshold is set. The pixels corresponding to the value range greater than the classification threshold are marked as changed, and the pixels corresponding to the value range less than the classification threshold are marked as unchanged, thereby obtaining the change detection result.

[0041] The beneficial effects of the present invention are:

[0042] (1) Since pyramid features can retain image features at all levels, the accuracy of change detection results can be increased by extracting pyramid features for change detection.

[0043] (2) By linearly mapping the feature vectors to multiple different linear spaces to obtain the vector inner product, different vector inner products are used to represent the similarity between two feature vectors with different focuses. This can obtain more accurate similarity values ​​and improve the accuracy of change detection.

[0044] (3) This application adopts a new labeling method based on superpixel blocks. The difference values ​​of multiple features are calculated based on superpixel blocks, and the data set is labeled according to the classification results of the superpixel blocks. The obtained real labels are highly accurate, and the labor cost is reduced, reducing the consumption of manpower, material resources and other resources.

[0045] (4) When the change detection model in this application starts pre-training, the weights are initialized to a random normal distribution plus a unit matrix. This setting can ensure that the initial weights are reasonable. The corresponding unit matrix initialization weights are to maintain the original spatial calculation inner product, so that the extracted features can have a better clustering effect on different objects. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Various other advantages and benefits will become apparent to those of ordinary skill in the art by reading the detailed description of the preferred embodiments below. The accompanying drawings are only for the purpose of illustrating the preferred embodiments and are not to be considered as limiting the present invention. Moreover, the same reference symbols are used throughout the accompanying drawings to represent the same components. In the accompanying drawings:

[0047] Figure 1 is a flow chart of a method according to an embodiment of the present invention;

[0048] Figure 2 It is a technical processing flow chart of an embodiment of the present invention;

[0049] Figure 3 This is a training result data diagram of an embodiment of the present invention;

[0050] Figure 4is a processing flow chart of a feature extraction network according to an embodiment of the present invention;

[0051] Figure 5 A processing flow chart of a change comparison network according to an embodiment of the present invention;

[0052] Figure 6 The figure is a processing flow chart of the output network according to an embodiment of the present invention. DETAILED DESCRIPTION

[0053] The specific implementation of the present invention is further described in detail below in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present invention, but are not intended to limit the scope of the present invention.

[0054] See also Figure 1 and Figure 2 The present invention provides a remote sensing image change detection method based on a deep convolutional network, comprising:

[0055] S1 obtains a pre-phase image and a post-phase image of a target area, and pre-processes the pre-phase image and the post-phase image.

[0056] Remote sensing imaging is very susceptible to external factors such as changes in sensor attitude, satellite platform movement, earth curvature, terrain undulations, optical system distortion, etc., which may cause the captured remote sensing images to be distorted, offset, squeezed, stretched, and other types of geometric distortion relative to the actual ground position. Before using high-resolution remote sensing images for change detection, the remote sensing images must first be preprocessed as necessary and fully as possible. For different types of remote sensing images, the preprocessing process may vary. The preprocessing process used in this embodiment is as follows:

[0057] 1. Orthorectification: Orthorectification is the process of correcting the spatial and geometric distortion of an image to generate a multi-center projection plane orthorectified image. This implementation adopts a rational polynomial coefficient model and combines it with a digital elevation model to implement the orthorectification process.

[0058] 2. Image registration: Image registration refers to the process of matching and superimposing multi-temporal images acquired at different times, with different imaging devices or under different acquisition conditions into a unified coordinate system. The specific steps are as follows:

[0059] Establish a reference system; use one of the two phase images as the reference image to establish a reference coordinate system; use the other image as the image to be registered to establish a coordinate system for the image to be registered. You can choose any one of the images as the reference image.

[0060] Select tie points; extract feature points using the Forstner corner operator, filter feature points using a fitted global transformation geometry model using a first-order polynomial, and match tie points using a cross-correlation algorithm.

[0061] Establish a transformation model; using the relationship between the connection points in the two images, the parameters of the transformation model used for image registration can be determined.

[0062] Geometric transformation and resampling: Based on the transformation model, the image to be registered is geometrically transformed and resampled to obtain the final registration result. This embodiment uses three 1*1 convolutions for resampling and a polynomial model for geometric transformation.

[0063] 3. Histogram matching: Through histogram matching, the color differences of multi-temporal remote sensing images can be corrected to reduce the impact of color on the accuracy of change detection.

[0064] 4. Contrast-limited adaptive histogram equalization: In order to further enhance the contrast of local images and highlight the changes in local features, this embodiment uses contrast-limited adaptive histogram equalization to enhance local images. After processing, the details of the remote sensing images are clearer, the local features are more obvious, and the tones of multi-temporal remote sensing images are more consistent.

[0065] 5. Median filtering: Median filtering is a type of statistical sorting filter that uses the median of the grayscale levels in a pixel's neighborhood to replace the pixel's value. After this processing, some abrupt details are better handled, and the local features of the object become smoother and more orderly.

[0066] S2 builds a change detection model and performs pre-training. The change detection model includes a feature extraction network, a change comparison network, and an output network.

[0067] Specifically, the pre-training process of the change detection model is as follows:

[0068] (1) Obtain a dataset with true labels and divide the dataset into training samples and validation samples.

[0069] In one embodiment of the present invention, after obtaining the data set, data enhancement may be performed on the data set. The data enhancement strategies include Mixup, Cutmix, RandomErasing, and the like.

[0070] In this embodiment, the true label is automatically calculated according to the labeling method.

[0071] Specifically, the annotation methods include:

[0072] The dataset consists of multiple image pairs of previous and next phases. The image bands of the image pairs are superimposed, and then the first three principal component images are extracted using the principal component analysis method, which are then segmented into superpixels to obtain multiple superpixel blocks.

[0073] Feature extraction is performed on superpixel blocks, including spectral features, texture features, and spatial features.

[0074] Taking superpixel blocks as units, the difference value of each feature is calculated, that is, for each group of image pairs, the spectral feature difference value, texture feature difference value, and spatial feature difference value between the corresponding two superpixel blocks are calculated.

[0075] The spectral feature difference value, texture feature difference value, and spatial feature difference value are combined into a vector as the difference vector of the corresponding superpixel block. The difference vector is modulo and normalized to obtain the change value of the superpixel block with a value range of [0,1]. A threshold is set to classify the superpixel block into two categories: changed and unchanged.

[0076] According to the classification results of the superpixel blocks, the classification results of each pixel in the dataset are obtained as the true label of the dataset.

[0077] This application adopts a new automatic labeling method based on superpixel blocks. The difference values ​​of multiple features are calculated based on superpixel blocks, and the data set is labeled according to the classification results of the superpixel blocks. The obtained real labels have high accuracy, reduce labor costs, and reduce the consumption of manpower, material resources.

[0078] Let’s take a specific example to illustrate the labeling process:

[0079] The dataset contains image pairs of the previous and next phases, and the previous phase image set is denoted as {A i}, and the image set of the later phase is recorded as {B i}, where A i and B i As a set of image pairs.

[0080] A i and B i The bands are superimposed, and the first three principal component images are extracted using the principal component analysis method. These principal component images are then segmented into superpixels. i and B i The segmentation method is the same as that of A, and multiple super pixel blocks are obtained, among which A i The set of superpixel blocks is denoted as {A k}, B i The set of superpixel blocks is denoted as {B k}.

[0081] For k} and {B k} respectively perform feature extraction, the extracted features include spectral feature I, texture feature M, and spatial feature N, and obtain {A k} spectral feature set {A k -I}, texture feature set {Ak -M}, spatial feature set {A k -N}, and {B k} of the spectral feature set {B k -I}, texture feature set {B k -M}, spatial feature set {B k -N}.

[0082] Calculate {A k -I} and {B k -I}, {A k -M} and {B k -M}, {A k -N} and {B k -N}, specifically {A k -I} and {B k -I} as an example:

[0083] {A k -I} refers to the superpixel block set {A k}, where A k Yes A i The kth superpixel block has a spectral feature of A k -I, correspondingly, B k It is B i The kth superpixel block has a spectral feature of B k -I, it should be made clear that A k With B k are two superpixel blocks with the same position in the images of the previous and next phases. When calculating, we can calculate A k -I and B k -I's Euclidean distance to get A k With B k The spectral feature difference values ​​between the super-pixel blocks are obtained based on the Euclidean distance.

[0084] Specifically, {A k -M} and {B k -M}, {A k -N} and {B k -N} are calculated in the same way as above, and the texture feature difference value and spatial feature difference value between the corresponding two super-pixel blocks are obtained.

[0085] For every two corresponding superpixel blocks, such as A k With B k , its spectral feature difference value, texture feature difference value, and spatial feature difference value are combined into a vector to obtain A k With B kThe difference vector is normalized to a value in the range [0,1] after modulo, as A k With B k The change value is used to characterize A k With B k Whether these two superpixel blocks have changed is determined by: k As a reference superpixel block, determine B k Whether it changes, set a threshold, such as 0.5, if the change value is in [0,0.5), determine B k For the unchanged superpixel block, if the change value is in [0.5,1], it is determined that B k is the change superpixel block.

[0086] The label of each pixel is obtained according to the judgment result of the superpixel block, for example, B k is judged to be unchanged, then B k Each pixel in is marked as unchanged, i.e., B k The label of the pixel set is "unchanged", otherwise it is labeled "changed". After labeling all the pixels, the true label of the data set can be obtained.

[0087] (2) The AdamW optimizer is combined with the cosine annealing algorithm to optimize the learning rate. The initial change detection model is trained based on the training samples to obtain the predicted labels of the training samples. The model weights are adjusted using the validation samples until the model converges.

[0088] During the training process, a loss function is introduced to calculate the gap between the predicted label and the true label.

[0089] Specifically, the loss function includes:

[0090] The cross entropy loss is calculated by predicting the label and the real label, where the cross entropy loss function is used to measure the similarity between the predicted change pixels and the real change pixels;

[0091] The Dice loss is calculated by predicting the label and the true label. The Dice loss function is used to alleviate the problem of imbalance in the number of changed and unchanged samples in the change detection process.

[0092] Add the cross entropy loss function and the Dice loss function to get the total loss function.

[0093] Specifically, the AdamW optimizer formula is as follows:

[0094]

[0095] in, and The two hyperparameters are the exponential decay rates of the first-order moment estimate and the second-order moment estimate of the gradient, which are used to distribute weights and the influence of the square of the gradient; represents the gradient of the step at time t; and They are the first-order moment estimation and second-order moment estimation of the gradient respectively; t represents the time step; and Respectively and The paranoid correction results; represents the value of the optimizer output parameter at time t; λ is the weight decay factor; Represents the learning rate; ε represents a very small number to prevent the denominator from being zero.

[0096] The formula for the cosine annealing algorithm is as follows:

[0097]

[0098] in, is the learning rate increased at the current time t, is the minimum learning rate set, is the maximum learning rate set, is the Epoch number of the current training, It is the maximum epoch number that can be set.

[0099] The training parameters are set as: is 0.9, The value of the training epoch is 0.999, the batch size is 4096, the number of epochs during training is 300, the weight decay number is set to 0.05, the learning rate decay method uses the cosine annealing algorithm, and the maximum learning rate is 0.004. During training, warmup can be used to warm up the learning rate to the maximum value, and then decay from the maximum value. The number of epochs for warmup is 20, and the warmup schedule uses linear transformation to adjust the learning rate. The entire training is forward propagated, and then the training parameters are reversely propagated to each node to adjust the weight parameters of each process.

[0100] In this embodiment, when the change detection model is pre-trained, the weights are initialized to a random normal distribution with a unit matrix added. This setting can ensure that the initial weights are reasonable. The corresponding unit matrix initialization weights are to maintain the original spatial calculation inner product, so that the extracted features can have a better clustering effect on different objects.

[0101] During training, cross entropy loss and Dice loss are introduced to evaluate the performance of the model during training. The formula of the cross entropy loss function is:

[0102]

[0103] In the formula, x represents the value after LogSoftMax operation, class represents the real category, and j represents the predicted category.

[0104] The formula for the Dice loss function is as follows:

[0105]

[0106] in, represents the true label of pixel d, represents the predicted label of pixel d, and D is the total number of pixels.

[0107] The total loss function is the sum of cross entropy loss and Dice loss.

[0108] The training results are evaluated based on mF1 to determine the quality of the model training. mF1 refers to the average value of the F1 score. F1 is an indicator of model accuracy. Its maximum value is 1 and its minimum value is 0. The larger the value, the better the model. The formula for F1 is as follows:

[0109]

[0110] Among them, precision refers to the accuracy rate and recall refers to the recall rate.

[0111]

[0112]

[0113] During training, the true labels used are of two categories, one for changes and the other for unchanged. Among them, TP refers to the number of samples whose predicted changed labels are consistent with the true labels of the samples, FP refers to the number of samples whose unchanged true labels are predicted as changed labels, and FN refers to the number of samples whose changed true labels are predicted as unchanged labels.

[0114] In a specific embodiment of the present invention, the feature extraction network includes a 4*4 convolution layer, a first residual block, a second residual block, a third residual block, and a fourth residual block according to the data flow direction, wherein the first residual block includes 3 residual units, the second residual block includes 3 residual units, the third residual block includes 9 residual units, and the fourth residual block includes 3 residual units, each residual unit includes a 7*7 depth separation convolution layer, a normalization layer, a 1*1 convolution layer, and an activation function layer, wherein the normalization layer uses an LN layer, the activation function layer uses a GELU function, and shortcuts are set between residual units for connection; the change contrast network uses a linear mapping layer and a feature fusion layer; the output network includes 3 1*1 convolution layers, a feature splicing layer, and an MLP layer. The training steps of the change detection model are: obtaining a data set, dividing the data set into training samples and verification samples, wherein the number ratio of training samples to verification samples is 7:3, and both training samples and verification samples are remote sensing images of two phases with real labels. The training evaluation method uses mF1. Figure 3 This is the training result of this embodiment.

[0115] It can be seen from the data of the training results that the change detection model of the present invention has good performance and its accuracy is higher than that of the model in the prior art. This shows that the change detection model of the present invention can successfully use multiple linear spaces to linearly map features, and obtain feature similarities that characterize different emphases, so that the features can be expressed more accurately, thereby improving the accuracy and precision of change detection.

[0116] S3 performs pyramid feature extraction on the front phase image and the back phase image through the feature extraction network to obtain the first feature and the second feature.

[0117] See also Figure 4 Specifically, step S3 includes:

[0118] The front phase image and the back phase image are respectively input into the feature extraction network to construct a pyramid, and multi-scale features are extracted according to the scale space of the pyramid, and the first feature and the second feature are obtained after output;

[0119] Among them, the feature extraction network includes a 4*4 convolutional layer, a first residual block, a second residual block, a third residual block, and a fourth residual block. Among them, the first residual block includes 3 residual units, the second residual block includes 3 residual units, the third residual block includes 9 residual units, and the fourth residual block includes 3 residual units. Each residual unit contains a 7*7 depth separation convolution layer, a normalization layer, a 1*1 convolution layer, and an activation function layer. Shortcuts are set between the residual units for connection.

[0120] In the residual unit of this embodiment, the normalization layer adopts the LN layer, the activation function layer adopts the GELU function, there are one 7*7 depth separation convolution layer and two 1*1 convolution layers.

[0121] Optionally, in one embodiment of the present invention, the feature extraction is described by taking the previous phase image as an example:

[0122] The previous phase image is input into the feature extraction network, and the first feature map is extracted according to the 4*4 convolutional layer, which is denoted as X1. Then, four residual blocks are input in sequence to extract four feature maps, which are denoted as X2, X3, X4, and X5. The scale of each feature map is different, and the scale of the feature map increases in sequence according to the data flow.

[0123] A pyramid framework is constructed, where X5 is the highest-level pyramid feature. X5 is upsampled and concatenated with X4 to obtain the second-level pyramid feature. The second-level pyramid feature is also upsampled and concatenated with X3 to obtain the third-level pyramid feature. The third-level pyramid feature is upsampled and concatenated with X2 to obtain the fourth-level pyramid feature. The fourth-level pyramid feature is upsampled and concatenated with X1 to obtain the fifth-level pyramid feature.

[0124] Arrange the highest-level pyramid features, the second-level pyramid features, the third-level pyramid features, the fourth-level pyramid features, and the fifth-level pyramid features in order to form a feature pyramid, namely the first feature.

[0125] The extraction process of the second feature is the same as that of the first feature and will not be repeated here.

[0126] It should be noted that the pyramid features in the present invention include both low-level features that describe detail information and high-level features that describe semantic information. When constructing the feature pyramid, through layer-by-layer upsampling and splicing operations, the accuracy of each feature can be retained to the greatest extent, and the description content of each layer of the pyramid features can be improved, which is beneficial to the subsequent change detection processing.

[0127] S4 inputs the first feature and the second feature into the change comparison network, performs linear mapping respectively, obtains the first feature vector and the second feature vector, and calculates the vector inner product of the first feature vector and the second feature vector to obtain the difference feature.

[0128] See also Figure 5 Specifically, step S4 includes:

[0129] S41 linearly maps the first feature and the second feature, transforming them into multiple linear spaces in pairs, to obtain multiple first feature vectors and second feature vectors;

[0130] S42 calculates the vector inner product of the first eigenvector and the second eigenvector in the same linear space to obtain similar features of the linear space;

[0131] S43 fuses similar features of all linear spaces to obtain difference features.

[0132] In this embodiment, the multiple linear spaces are not identical, and the similarity feature of each linear space represents the similarity between the first eigenvector and the second eigenvector in the linear space.

[0133] In a specific embodiment of the present invention, the implementation process of step S4 is as follows:

[0134] First, the first feature includes the highest level pyramid feature, the second level pyramid feature, the third level pyramid feature, the fourth level pyramid feature, and the fifth level pyramid feature, which are denoted as A1, A2, A3, A4, and A5, where A5 is the feature of the largest scale. The second feature also includes the highest level pyramid feature, the second level pyramid feature, the third level pyramid feature, the fourth level pyramid feature, and the fifth level pyramid feature, which are denoted as B1, B2, B3, B4, and B5, where B5 is the feature of the largest scale.

[0135] In this embodiment, there are three linear spaces, that is, the first feature and the second feature are mapped by linear mapping, and the dimensions of these three linear spaces are different. The first linear space maps the first feature and the second feature to a c-dimensional feature vector, the second linear space maps the first feature and the second feature to a d-dimensional feature vector, and the third linear space maps the first feature and the second feature to a c+d-dimensional feature vector, c<d, c+d is less than the original dimension of the first feature and the second feature.

[0136] In the first linear space, the inner product of the c-dimensional first eigenvector and the c-dimensional second eigenvector is calculated to obtain the first similarity feature C1.

[0137] In the second linear space, the inner product of the d-dimensional first eigenvector and the d-dimensional second eigenvector is calculated to obtain the second similarity feature C2.

[0138] In the third linear space, the inner product of the c+d-dimensional first eigenvector and the c+d-dimensional second eigenvector is calculated to obtain the third similarity feature C3.

[0139] The inner product calculation formula is as follows:

[0140]

[0141]

[0142]

[0143] Among them, i=1,2,3,4,5, j is the dimension of the linear space, E1 and E2 represent the tensor product of A and B, W represents the parameter of the linear mapping, h and w are the height and width of the feature map respectively, and C is the similarity feature.

[0144] The first similar feature C1, the second similar feature C2, and the third similar feature C3 are fused to form a difference feature D.

[0145]

[0146] This application calculates the vector inner product by linearly mapping the feature vector to multiple different linear spaces. The method for calculating the vector inner product is based on the number of channels, the dimension of the linear space, etc., and the tensor product of the pyramid features of the two images in the previous and next phases is first calculated, and then the tensor product is summed in terms of the number of channels to obtain the vector inner product between the two features in the linear space. Using different vector inner products to characterize the similarity between two feature vectors with different focuses can obtain more accurate similarity values ​​to improve the accuracy of change detection.

[0147] S5 inputs the difference features into the output network and processes the difference features to obtain the change detection results.

[0148] See also Figure 6 Specifically, step S5 includes:

[0149] The S51 output network includes a 1*1 convolutional layer, a feature concatenation layer, and an MLP layer.

[0150] S52 After adjusting the channel through the 1*1 convolution layer, the difference features are resampled and feature concatenated to obtain the change features. Specifically, in this embodiment, there are 3 1*1 convolution layers.

[0151] S53 performs pixel-level change determination on the change feature through MLP, and outputs a change detection result. Specifically, in this embodiment, the MLP layer may include three 1*1 convolutional layers.

[0152] Wherein, step S53 comprises:

[0153] The range of pixel values ​​in the change feature is set to [0, 1], a classification threshold is set, pixels corresponding to a range greater than the classification threshold are marked as changed, and pixels corresponding to a range less than the classification threshold are marked as unchanged, to obtain a change detection result. In this embodiment, the classification threshold is 0.5.

[0154] The above implementation modes are only used to illustrate the present invention, but not to limit the present invention. Ordinary technicians in the relevant technical field can make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, all equivalent technical solutions also belong to the scope of the present invention. The patent protection scope of the present invention should be defined by the claims.

Claims

1. A remote sensing image change detection method based on deep convolutional network, It is characterized in that include: S1 obtains the pre-phase image and the post-phase image of the target area, and performs pre-processing and data enhancement on the pre-phase image and the post-phase image; S2 builds a change detection model and performs pre-training. The change detection model includes a feature extraction network, a change comparison network, and an output network; S3 performs pyramid feature extraction on the front phase image and the back phase image respectively through the feature extraction network to obtain the first feature and the second feature; S4 inputs the first feature and the second feature into a change comparison network, performs linear mapping respectively, obtains a first feature vector and a second feature vector, and calculates the vector inner product of the first feature vector and the second feature vector to obtain a difference feature; S5 inputs the difference features into the output network and processes the difference features to obtain the change detection results; Step S4 includes: S41 linearly maps the first feature and the second feature, transforming them into multiple linear spaces in pairs, to obtain multiple first feature vectors and second feature vectors; S42 calculates the vector inner product of the first eigenvector and the second eigenvector in the same linear space to obtain similar features of the linear space; S43 fuses similar features of all linear spaces to obtain difference features; The linear mapping includes: mapping the first feature and the second feature into a c-dimensional feature vector in a first linear space, mapping the first feature and the second feature into a d-dimensional feature vector in a second linear space, and mapping the first feature and the second feature into a c+d-dimensional feature vector in a third linear space, wherein c<d, and c+d is less than the original dimension of the first feature and the second feature; The multiple linear spaces are not identical, and the similarity feature of each linear space represents the similarity between the first eigenvector and the second eigenvector in the linear space.

2. The method according to claim 1, It is characterized in that When pre-training the change detection model, the weights are initialized to a random normal distribution plus an identity matrix.

3. The method according to claim 1, It is characterized in that The pre-training process of the change detection model is as follows: Get a dataset with real labels and divide it into training samples and validation samples; The AdamW optimizer is used in combination with the cosine annealing algorithm to optimize the learning rate. The initial change detection model is trained based on the training samples to obtain the predicted labels of the training samples. The model weights are adjusted using the validation samples until the model converges. During the training process, a loss function is introduced to calculate the gap between the predicted label and the true label.

4. The method according to claim 3, It is characterized in that The true label is automatically calculated based on the annotation method, which includes: The dataset consists of multiple image pairs of previous and next phases. The image bands of the image pairs are superimposed, and then the first three principal component images are extracted using the principal component analysis method, and then superpixel segmentation is performed on them to obtain multiple superpixel blocks. Extract features from superpixel blocks, including spectral features, texture features, and spatial features; Taking superpixel blocks as units, the difference value of each feature is calculated, that is, the spectral feature difference value, texture feature difference value, and spatial feature difference value between the corresponding two superpixel blocks are calculated for each image pair; The spectral feature difference value, texture feature difference value, and spatial feature difference value are combined into a vector as the difference vector of the corresponding superpixel block. The difference vector is modulo and normalized to obtain the change value of the superpixel block. The value range is [0,1]. The threshold is set to classify the superpixel block into two categories: changed and unchanged. According to the classification results of the superpixel blocks, the classification results of each pixel in the dataset are obtained as the true label of the dataset.

5. The method according to claim 3, It is characterized in that The loss functions include: The cross entropy loss is calculated by predicting the label and the real label, where the cross entropy loss function is used to measure the similarity between the predicted change pixels and the real change pixels; The Dice loss is calculated by predicting the label and the true label. The Dice loss function is used to alleviate the problem of imbalance in the number of changed and unchanged samples in the change detection process. Add the cross entropy loss function and the Dice loss function to get the total loss function.

6. The method according to claim 1, It is characterized in that Step S3 includes: The front phase image and the back phase image are respectively input into the feature extraction network to construct a pyramid, and multi-scale features are extracted according to the scale space of the pyramid, and the first feature and the second feature are obtained after output; Among them, the feature extraction network includes a 4*4 convolutional layer, a first residual block, a second residual block, a third residual block, and a fourth residual block. Among them, the first residual block includes 3 residual units, the second residual block includes 3 residual units, the third residual block includes 9 residual units, and the fourth residual block includes 3 residual units. Each residual unit contains a 7*7 depth separation convolution layer, a normalization layer, a 1*1 convolution layer, and an activation function layer. Shortcuts are set between the residual units for connection.

7. The method according to claim 1, It is characterized in that Step S5 includes: The S51 output network includes a 1*1 convolutional layer, a feature concatenation layer, and an MLP layer; S52 adjusts the channel through a 1*1 convolution layer, resamples and concatenates the difference features, and obtains the change features; S53 performs pixel-level change determination on the change feature through MLP, and outputs the change detection result.

8. The method according to claim 7, It is characterized in that Step S53 includes: The value range of the pixel value in the change feature is set to [0,1], and the classification threshold is set. The pixels corresponding to the value range greater than the classification threshold are marked as changed, and the pixels corresponding to the value range less than the classification threshold are marked as unchanged, thereby obtaining the change detection result.

Citation Information

Patent Citations

  • DCNN-based dual-temporal remote sensing image change detection method

    CN112131968A

  • Seismic facies recognition model training method and device and seismic facies prediction method and device

    CN114239655A

  • Aerial photography insulator spontaneous explosion defect detection model, method and equipment and storage medium

    CN114677357A