Building change detection method based on depth separable double-time-phase image converter

By adopting a method based on depth-separable dual-time phase image converter in building change detection, the problem of poor learning and representation ability in the prior art is solved, and more efficient feature extraction and change detection is achieved, reducing the error rate and improving the accuracy and recognition efficiency.

CN119992340APending Publication Date: 2025-05-13CHONGQING INST OF GEOLOGY & MINERAL RESOURCES
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510373989.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing dual-time phase transformers have poor learning and representation capabilities in building change detection, and the feature extraction effect is not obvious, resulting in high detection error rate and low accuracy rate and recognition efficiency.

Method used

The building change detection method based on the depth-separable dual-time phase image converter is adopted, including building a backbone network based on ResNet101 for feature map extraction, using a dual-time phase image converter to process the feature map of the two time phases, and combining the multi-scale jump structure of the residual module and the improved prediction head module, the output of the dual-time phase image converter is processed to generate the final change detection prediction result.

Benefits of technology

It improves the feature extraction effect of building change detection, enhances the network's ability to capture long-distance dependencies, reduces the detection error rate, and improves the accuracy and recognition efficiency of detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992340A_ABST
    Figure CN119992340A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of remote sensing image change detection, in particular to a building change detection method based on a depth separable double-time-phase image converter, and the method comprises the steps: a model construction step: constructing a change detection model for building change detection, and comprises the steps: S1, constructing a backbone network based on ResNet101, and carrying out the feature map extraction; s2, processing the characteristic patterns of the two time phases by adopting a double-time-phase image converter; and S3, in combination with a multi-scale jump structure of the residual module, an improved prediction head module is adopted to process the output of the dual-time-phase image converter, and a final change detection prediction result is generated. According to the scheme, the problems that an existing double-time-phase converter is poor in learning and representation capacity and not obvious in feature extraction effect can be solved, the building change detection error rate is reduced, and the accuracy and recognition efficiency of a detection result are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of remote sensing image change detection, and in particular to a building change detection method based on a depth separable dual-phase image converter. Background Art

[0002] Building change detection in remote sensing images has important application value in urban planning, land management, environmental monitoring, and disaster response. Traditional building change detection methods are mainly based on image processing and machine learning techniques, such as feature extraction, classifiers, and change detection algorithms; these methods often rely on manually designed features and rules, have limitations for complex building change scenarios, and require a lot of manual intervention and adjustment.

[0003] In recent years, the rapid development of deep learning technology has provided new opportunities and challenges for building change detection. Deep learning algorithms can achieve end-to-end building change detection by learning features and patterns from a large amount of labeled data, thereby reducing dependence on artificial feature design and being able to handle more complex and abstract change situations.

[0004] In the deep learning framework, dual-temporal remote sensing images are used to detect and analyze building changes, among which the dual-temporal image transformer is an algorithm for building change detection. However, the feature extraction network of the traditional dual-temporal image transformer is shallow, which limits the learning and representation capabilities of the network. It has certain limitations in processing more complex building change scenes and is not sensitive enough to the complex change patterns in the building change detection task. Summary of the invention

[0005] The present invention aims to provide a building change detection method based on a deeply separable dual-phase image converter, which can solve the problems of poor learning and representation capabilities and unclear feature extraction effects of existing dual-phase converters, so as to reduce the error rate of building change detection and improve the accuracy of detection results and recognition efficiency.

[0006] The present invention provides the following basic scheme: a building change detection method based on a depth-separable dual-phase image converter, comprising the following contents:

[0007] Model construction steps: Build a change detection model for building change detection, including: S1, build a backbone network based on ResNet101 to extract feature maps;

[0008] S2, using a dual-phase image converter to process the feature maps of the two phases;

[0009] S3. Combined with the multi-scale jump structure of the residual module, an improved prediction head module is used to process the output of the dual-phase image transformer to generate the final change detection prediction result.

[0010] Furthermore, it also includes:

[0011] Data preparation steps: construct a building change detection dataset and perform image preprocessing;

[0012] The specific process is as follows:

[0013] We selected the LEVIR-CD dataset, a public dataset for change detection, for experiments;

[0014] The original images in the LEVIR-CD dataset are normalized and standardized, and the result of the normalization is:

[0015]

[0016] Among them, x norm is the normalized image value, min(x) is the minimum value of the image, max(x) is the maximum value of the image, x i is the image value to be processed;

[0017] The result of the normalization process is:

[0018]

[0019] Among them, S is the standardized image value, x i is the image value to be processed, μ is the mean of the image, σ is the standard deviation of the image, and N is the number of image pixels;

[0020] The LEVIR-CD dataset is divided into training set, validation set and test set to construct a building change detection dataset.

[0021] Furthermore, the network architecture of the backbone network includes:

[0022] Convolutional layer, used for initial feature extraction of input images;

[0023] The maximum pooling layer is used to downsample the initial features; the residual module is used to extract features of different scales from the downsampled results, where the residual module adopts a multi-scale jump structure, including multiple stages, each stage includes multiple residual blocks, and extracts features of different scales in turn;

[0024] The global average pooling layer is used to perform an average operation on each feature channel and convert the features of each channel into a single value;

[0025] The fully connected layer is used to map the output of the global average pooling layer to the final feature vector to obtain the feature map of the dual-phase building.

[0026] Further, the S2 includes:

[0027] S201, input feature graphs of two time phases, extract each local area in each feature graph as a discrete semantic symbol Tokens, and obtain the global semantic information set T = {T 1 ,T 2}, capturing semantic information and structural features;

[0028] S202, using an encoder to encode the global semantic information set T;

[0029] The multi-head self-attention module of the encoder includes the calculation of query, key and value, which is calculated as follows:

[0030] Q=T (l-1) W q ;

[0031] K=T (l-1) W k ;

[0032] V=T (l-1) W v ;

[0033] Among them, l is the length of semantic information, W q , W k and W v are the learning parameters of the multi-head self-attention module;

[0034] S203, connecting the encoder outputs of the dual-phase image through a decoder to form a vector H={H1,H2}, reducing the dimension of the connected vector through linear transformation, and repeatedly stacking multiple decoders, where the decoder structure is consistent with the encoder;

[0035] Among them, S202 includes:

[0036] S20201. After the input is linearly transformed, the attention weight is obtained by calculating the correlation between the query and the key. The attention weight is multiplied by the value and weighted summed to obtain the final self-attention output, which is expressed as:

[0037]

[0038] Among them, σ() is the Sigmoid activation function;

[0039] S20202. Calculate the results of multi-head self-attention:

[0040] MultiHead(Q,K,V)=Concat(head1,head2,...,head h )W O ;

[0041] Among them, head i =Att(QW Q i ,KW K i ,VW V i ), W O is the output projection matrix;

[0042] S20203. Perform layer normalization on the output of the multi-head self-attention mechanism:

[0043]

[0044] Among them, μ is the mean of the image, σ is the standard deviation of the image;

[0045] S20204. Perform linear transformation on the input vector and use nonlinear activation function to capture complex features:

[0046] MLP(x) = ReLU(xW1+b1)W2+b2;

[0047] Among them, MLP is a multi-layer perceptron, ReLU is a nonlinear activation function, and W1, b1, W2 and b2 are learning parameters.

[0048] Further, the S3 includes:

[0049] S301, extracting features from the output of the decoder using depthwise separable convolution;

[0050] S302, using the ReLU activation function to process the output of the depthwise separable convolution to obtain feature A after adding nonlinear characteristics:

[0051] A = ReLU(C);

[0052] S303: Using the non-local information residual module, extract features through deep separable convolution to obtain feature maps A at different scales i ;

[0053] S304: Combine the decoder result Z and the features at different scales Perform skip connection fusion to obtain the output value X of the non-local information residual module:

[0054]

[0055] S305: For the output value X, use the non-local weight matrix M ij , calculate the feature map and obtain the convolution feature Y:

[0056]

[0057] Among them, A ij is the non-local weight matrix, X ij is the pixel value of the feature map at pixel i, j;

[0058] S306, convert the convolution feature Y into output dimension Y' through a linear layer:

[0059] Y'=WY+b;

[0060] Among them, W and b are learning parameters;

[0061] S307, using Sigmoid activation function to generate probability prediction of changes

[0062]

[0063] Further, the S301 includes:

[0064] Perform depthwise convolution:

[0065]

[0066] Among them, Y i,j,m is the pixel value of the i,j position of the m-th channel of the output feature map, X i=d-1,j=d'-1,m is the pixel value at the i,j position of the m-th channel of the input feature map, K d,d',m is the parameter of the d,d'th position of the mth convolution kernel;

[0067] The result after depth convolution is convolved point by point using a 1×1 convolution kernel:

[0068]

[0069] Among them, Z i,j,n is the pixel value at the i,j position of the nth channel of the output feature map, Y i,j,m is the pixel value of the i,j position of the mth channel of the output feature map of the depth convolution, L m,n are the parameters of point-wise convolution.

[0070] Furthermore, it also includes:

[0071] Model training steps: Use the building change detection dataset to train and test the change detection model, output a change detection model that meets the requirements, and perform building change detection.

[0072] Furthermore, it also includes: a ground road auxiliary identification step: using a change detection model to identify, perform change detection identification, generate detection and identification results, filter out incomplete detection results, and perform supplementary identification through ground roads.

[0073] Furthermore, the ground road auxiliary identification step includes:

[0074] S401, dividing the remote sensing image to be identified into a plurality of remote sensing image blocks;

[0075] S402, taking the remote sensing image block to be identified as the input of the change detection model, and the change detection model performs identification to obtain the identification result;

[0076] S403, integrating the recognition results of the remote sensing image blocks to form the recognition result of the remote sensing image to be recognized; since the optical satellite remote sensing image is divided into a plurality of remote sensing image blocks to be recognized, the recognition results of the remote sensing image blocks to be recognized are integrated;

[0077] S404, converting the recognition result of the remote sensing image into a vector layer, calculating the coordinates of the center point of each patch surface contour, and obtaining coordinate information;

[0078] S405. Based on the coordinate information, the center point and the road information are superimposed for spatial analysis to determine the spatial position of the image patch and the road. If the positions intersect, the proximity does not exceed a preset distance value, and the image patch is an incomplete image patch, then the image patch is determined to be a remote sensing image of a building; wherein the image patch is a building change contour image patch result identified by a change detection model.

[0079] Furthermore, the method further includes: a shadow auxiliary recognition step:

[0080] According to the time and longitude and latitude of the remote sensing image acquisition, when the brightness is within the preset brightness range, the shadows of the edge buildings in the building complex where the shadows are preset are identified;

[0081] Specifically include:

[0082] According to the acquisition time and longitude and latitude of the remote sensing image, the altitude and azimuth of the sun are calculated to determine the theoretical direction of the shadow;

[0083] According to the spectral characteristics of remote sensing images, segmentation or classification methods are used to extract the shadow area of ​​remote sensing images;

[0084] According to the relationship between the position of the building area and the corresponding shadow area and the azimuth of the sun, the disappeared building detection is performed to obtain the suspected building area;

[0085] For the shadow area, the information of the shadow area on the remote sensing image is enhanced through the transformation formula, and the information of the shadow area is restored;

[0086] Use a smoothing operator to smooth the edges of the shadow area;

[0087] For the same-direction edges of suspected building areas, the existence and size of shadows are identified by comparing with historical shadow-free data.

[0088] Beneficial effects: The change detection model for building change detection constructed in this scheme consists of multiple modules, including: a backbone network, a dual-phase image transformer, and an improved prediction head module;

[0089] First, the deep backbone network ResNet101 is used to extract feature maps, which improves the feature extraction effect of dual-temporal building change detection;

[0090] Then, a dual-phase image transformer is used to process the feature maps of the two phases, and each local area in each image is extracted as a discrete semantic symbol, which increases the efficiency of capturing the semantic information and structural features of the image. Through the encoder-decoder structure, the spatiotemporal information is effectively processed, and the dependency between distant pixels in the image is captured, so as to more accurately identify the changed area.

[0091] Finally, through the improved prediction head module combined with the non-local information residual module, the global correlation modeling of the entire feature map is performed, so that the network can better capture long-distance dependencies and improve the perception of global contextual information; the deep separable convolution is used to generate the original-size segmented image from the output of the dual-phase image converter, which improves the image change detection recognition rate and improves the computational efficiency of the network; the improved prediction head module is combined with the non-local information residual module, and the deep separable convolution is used to generate the original-size segmented image from the output of the dual-phase image converter to obtain the final change detection result. Its advantage is that the global correlation modeling of the entire feature map is performed without considering the spatial position between pixels, so that the network can better capture long-distance dependencies and improve the perception of global contextual information.

[0092] Compared with the traditional dual-phase transformer, the change detection model of this scheme has obvious feature extraction effect. It models the global correlation of the entire feature graph, enables the network to better capture long-distance dependencies, and improves the perception of global contextual information. It effectively solves the problems of poor learning and representation capabilities and unclear feature extraction effects of existing dual-phase transformers, reduces the error rate of building change detection, and improves the accuracy of detection results and recognition efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0093] Figure 1It is a flow chart of an embodiment of a method for detecting building changes based on a depth-separable dual-phase image converter according to the present invention;

[0094] Figure 2 It is a structural schematic diagram of a backbone network in an embodiment of a building change detection method based on a depth-separable dual-phase image converter of the present invention;

[0095] Figure 3 It is a structural schematic diagram of a dual-phase image converter in an embodiment of a building change detection method based on a depth-separable dual-phase image converter of the present invention;

[0096] Figure 4 It is a structural schematic diagram of a prediction head module in an embodiment of a building change detection method based on a depth-separable dual-phase image converter of the present invention. DETAILED DESCRIPTION

[0097] The following is further described in detail through specific implementation methods:

[0098] Embodiment 1

[0099] This embodiment is basically as shown in the attached Figure 1 As shown: A building change detection method based on a deep separable dual-phase image transformer includes the following contents:

[0100] Data preparation steps: construct a building change detection dataset and perform image preprocessing;

[0101] The specific process is as follows:

[0102] A public dataset for change detection is selected for experiments. The public dataset for change detection selected in this embodiment is the LEVIR-CD dataset, where the LEVIR-CD dataset is a LEVIR building change detection dataset. LEVIR-CD contains 637 ultra-high resolution (VHR, 0.5m / pixel) Google Earth image patch pairs with a size of 1024×1024 pixels and a time span of 5 to 14 years. The bit images have significant land use changes, especially the growth of buildings. LEVIR-CD covers various types of buildings, such as villas, high-rise apartments, small garages, and large warehouses.

[0103] The original images in the LEVIR-CD dataset are normalized and standardized, and the result of the normalization is:

[0104]

[0105] Among them, x norm is the normalized image value, min(x) is the minimum value of the image, max(x) is the maximum value of the image, x iis the image value to be processed;

[0106] The result of the normalization process is:

[0107]

[0108] Among them, S is the standardized image value, x i is the image value to be processed, μ is the mean of the image, σ is the standard deviation of the image, and N is the number of image pixels;

[0109] The LEVIR-CD dataset is divided into a training set, a validation set, and a test set to construct a building change detection dataset; in this embodiment, the LEVIR-CD dataset is divided into 7120 training sets, 1024 validation sets, and 2048 test sets to establish a building change detection dataset; the training set is used for training the change detection model constructed subsequently; the validation set is used to evaluate the performance of the model during the training process and help select the optimal hyperparameter combination, such as learning rate and regularization coefficient. Through the evaluation on the validation set, it can be found early whether the model is overfitting. If the performance of the model on the validation set decreases, it means that the model may be overly dependent on the training data. At this time, it is necessary to adjust the hyperparameters or adopt regularization and other means; the validation set can also be used for model selection to determine which model performs best on the validation set, so as to select the optimal model for final evaluation; the test set is used to verify the training results;

[0110] Model building steps: Build a change detection model for building change detection;

[0111] The change detection model includes: a backbone network, a dual-phase image transformer and a prediction head module;

[0112] The specific construction process is as follows:

[0113] S1. Build a backbone network based on ResNet101 (Residual Neural Network 101) to extract feature maps to capture the semantic information of the image;

[0114] The network architecture of the backbone network is shown in the attached figure. Figure 2 As shown, including:

[0115] The convolution layer is used to extract initial features from the input image. In this implementation, a 7x7 convolution layer is used to extract initial features from the dual-phase building image as the input image. The input image is an image preprocessed in the LEVIR-CD dataset, that is, a preprocessed LEVIR-CD dataset image. As a standard remote sensing change detection dataset, LEVIR-CD usually requires preprocessing steps (such as registration, image alignment, normalization, cropping, etc.) to eliminate noise and solve temporal difference problems (such as illumination changes, sensor differences, etc.). The preprocessed image is more in line with the model input requirements (such as fixed size, standardized pixel value range), ensuring that the model can effectively extract features.

[0116] The maximum pooling layer is used to downsample the initial features. In this embodiment, the initial features are downsampled through a 3x3 maximum pooling layer.

[0117] A residual module is used to extract features of different scales from the downsampling results, wherein the residual module adopts a multi-scale jump structure, including multiple stages, each stage including multiple residual blocks, and sequentially extracts features of different scales; in this embodiment, features are extracted through four stages (residual modules), each stage including multiple residual blocks (convolutional layers) for extracting features of different scales;

[0118] The global average pooling layer is used to perform an average operation on each feature channel and convert the features of each channel into a single value;

[0119] The fully connected layer is used to map the output of the global average pooling layer to the final feature vector to obtain the feature map of the dual-phase building. The above features and the initial features exist in the form of feature maps. From the initial convolutional layer to the residual module, the resolution and number of channels of the feature map are constantly changing, but the spatial structure information is always maintained.

[0120] S2, using a dual-phase image converter to process the feature graphs of the two phases, such as Figure 3 As shown;

[0121] The specific process is as follows:

[0122] S201, input feature graphs of two time phases, extract each local area in each feature graph as a discrete semantic symbol Tokens, and obtain the global semantic information set T = {T 1 ,T 2}, capturing semantic information and structural features;

[0123] S202, using an encoder to encode the global semantic information set T;

[0124] The multi-head self-attention module of the encoder includes the calculation of query, key and value, which is calculated as follows:

[0125] Q=T (l-1) W q ;

[0126] K=T (l-1) W k ;

[0127] V=T (l-1) W v ;

[0128] Among them, l is the length of semantic information, W q , W k and W v are the learning parameters of the multi-head self-attention module;

[0129] Specifically, the specific processing process of S202 is as follows:

[0130] S20201. After the input is linearly transformed, the attention weight is obtained by calculating the correlation between the query and the key. The attention weight is multiplied by the value and weighted summed to obtain the final self-attention output, which can be expressed as:

[0131]

[0132] Among them, σ() is the Sigmoid activation function;

[0133] S20202. Calculate the results of multi-head self-attention:

[0134] MultiHead(Q,K,V)=Concat(head1,head2,...,head h )W O ;

[0135] Among them, head i =Att(QW Q i ,KW K i ,VW V i ), W O is the output projection matrix;

[0136] S20203. Perform layer normalization on the output of the multi-head self-attention mechanism:

[0137]

[0138] Where μ is the mean of the image, σ is the standard deviation of the image, and x is the output feature of the multi-head self-attention mechanism (Multi-HeadAttention), that is, the output of S20202;

[0139] S20204. Perform linear transformation on the input vector and use nonlinear activation function to capture complex features:

[0140] MLP(x) = ReLU(xW1+b1)W2+b2;

[0141] Among them, MLP is a multi-layer perceptron, ReLU is a nonlinear activation function, W1, b1, W2 and b2 are learning parameters; MLP(x) is a submodule in the encoder, which is used to further extract high-level semantic features. The role of MLP(x) is to perform nonlinear transformation on the features after layer normalization, enhance the expression ability of the model, and capture more complex feature patterns;

[0142] S203, connecting the encoder outputs of the dual-phase image through a decoder to form a vector H={H1,H2}, reducing the dimension of the connected vector through linear transformation, and repeatedly stacking multiple decoders, where the decoder structure is consistent with the encoder;

[0143] The role of linear transformation is to map the input data to a new feature space by adjusting the weight and bias of the input features, thereby enhancing the expressiveness of the model. The formula of linear transformation is: Y = XW + b, X is the input feature matrix; W is the weight matrix, which is used to linearly map the input features; b is the bias vector, which is used to adjust the baseline value of the output features; Y is the output feature matrix after linear transformation;

[0144] Multiple decoders are stacked repeatedly. The decoder structure is consistent with the encoder. Multiple decoders are connected in sequence, and the output of the previous encoder is used as the input of the next encoder, so that data can be transmitted and processed layer by layer. When stacking decoder layers, residual connections are usually introduced to add the output of the previous encoder to the output of the current encoder to prevent gradient disappearance and accelerate training. Before processing the output data of the encoder, each decoder layer usually performs layer normalization to stabilize the training process and improve model performance.

[0145] Repeatedly stacking multiple decoders, with the decoder structure consistent with the encoder, can enhance the feature extraction capability and improve the model performance. Stacking decoder layers can enhance the model's expressiveness, enabling it to better capture deep patterns in the data, thereby improving the overall performance; it can also achieve multi-level information integration. Each decoder layer will further process the input data. Stacking multiple decoder layers can achieve multi-level information integration, making the final output more accurate. S3. Combined with the multi-scale jump structure of the residual module, the improved prediction head module is used to process the output of the dual-phase image transformer to generate the final change detection prediction result, such as Figure 4 As shown; in this embodiment, the output of the decoder is processed;

[0146] The specific process is as follows:

[0147] S301, using depthwise separable convolution to extract features from the output of the decoder, and obtaining calculated results, including:

[0148] First, perform a depthwise convolution:

[0149]

[0150] Among them, Y i,j,m is the pixel value of the i,j position of the m-th channel of the output feature map, X i=d-1,j=d'-1,m is the pixel value at the i,j position of the m-th channel of the input feature map, K d,d',m is the parameter of the d,d'th position of the mth convolution kernel;

[0151] Then, the result after the depth convolution is convolved point by point using a 1×1 convolution kernel:

[0152]

[0153] Among them, Z i,j,n is the pixel value at the i,j position of the nth channel of the output feature map, Y i,j,m is the pixel value of the i,j position of the mth channel of the output feature map of the depth convolution, L m,n is the parameter of point-by-point convolution;

[0154] S302, using the ReLU activation function to process the output of the depthwise separable convolution to obtain feature A after adding nonlinear characteristics:

[0155] A = ReLU(C);

[0156] S303: Using the non-local information residual module, extract features through deep separable convolution to obtain feature maps A at different scales i ;

[0157] S304: In order to maintain the model performance and further improve the feature extraction capability of the building edge contour, the decoder result Z and the features at different scales are combined. Perform skip connection fusion to obtain the output value X of the non-local information residual module:

[0158]

[0159] In this embodiment, the output A2 in S402 is input into the second depth-wise separable convolution and the output A3 is activated, and the input A1 in S402, the output A2 in S402, and A3 are fused;

[0160] S305: For the output value X, use the non-local weight matrix M ij , calculate the feature map and obtain the convolution feature Y;

[0161] The non-local weight matrix allows the model to model the global correlation of the entire building feature map without considering the spatial position between pixels. The calculation process is as follows:

[0162]

[0163] Among them, A ij is the non-local weight matrix, X ij is the pixel value of the feature map at pixel i, j;

[0164] S306, convert the convolution feature Y into output dimension Y' through a linear layer:

[0165] Y'=WY+b;

[0166] Among them, W and b are learning parameters;

[0167] S307, using Sigmoid activation function to generate probability prediction of changes

[0168]

[0169] Model training steps: Use the building change detection dataset to train and test the change detection model of the S1-S3 combination, output a change detection model that meets the requirements, and perform building change detection;

[0170] The training parameters of the change detection model are set as shown in Table 1:

[0171] Table 1: Model training parameter settings

[0172]

[0173] The change detection model is trained using the training set according to the set parameters;

[0174] Use the validation set to validate the trained change detection model;

[0175] The trained change detection model is tested using the test set, and the corresponding evaluation indicators are obtained, as shown in Table 2;

[0176] Table 2: Model evaluation metrics

[0177]

[0178] The training results of the change detection model are evaluated by F1 score, recall rate, precision, accuracy, Kappa index and IOU evaluation indicators. The experimental results are shown in Table 2.

[0179] If the evaluation index meets the preset evaluation index range, a change detection model that meets the requirements is output to perform building change detection;

[0180] The experimental results show that the invention can realize building change detection using the building change detection dataset LEVIR-CD, which has important economic and application value.

[0181] The change detection model for building change detection constructed in this scheme consists of multiple modules, including: backbone network, dual-phase image transformer and improved prediction head module;

[0182] First, the deep backbone network ResNet101 is used to extract feature maps, which improves the feature extraction effect of dual-temporal building change detection;

[0183] Then, a dual-phase image transformer is used to process the feature maps of the two phases, and each local area in each image is extracted as a discrete semantic symbol, which increases the efficiency of capturing the semantic information and structural features of the image. Through the encoder-decoder structure, the spatiotemporal information is effectively processed, and the dependency between distant pixels in the image is captured, so as to more accurately identify the changed area.

[0184] Finally, through the improved prediction head module combined with the non-local information residual module, the global correlation modeling of the entire feature map is performed, so that the network can better capture long-distance dependencies and improve the perception of global contextual information; the deep separable convolution is used to generate the original-size segmented image from the output of the dual-phase image converter, which improves the image change detection recognition rate and improves the computational efficiency of the network; the improved prediction head module is combined with the non-local information residual module, and the deep separable convolution is used to generate the original-size segmented image from the output of the dual-phase image converter to obtain the final change detection result. Its advantage is that the global correlation modeling of the entire feature map is performed without considering the spatial position between pixels, so that the network can better capture long-distance dependencies and improve the perception of global contextual information.

[0185] Compared with the traditional dual-phase transformer, the change detection model of this scheme has obvious feature extraction effect. It models the global correlation of the entire feature graph, enables the network to better capture long-distance dependencies, and improves the perception of global contextual information. It effectively solves the problems of poor learning and representation capabilities and unclear feature extraction effects of existing dual-phase transformers, reduces the error rate of building change detection, and improves the accuracy of detection results and recognition efficiency.

[0186] Embodiment 2

[0187] This embodiment is basically the same as the above embodiment, except that:

[0188] It also includes: a ground road auxiliary identification step: using a change detection model to identify, perform change detection identification, generate detection and identification results, filter out incomplete detection results, and perform supplementary identification through ground roads; that is, using a change detection model to identify, perform change detection identification, obtain building change contour patch results, filter out incomplete detection results obtained due to cloud and tree canopy occlusion, and perform supplementary identification through ground roads;

[0189] In rural areas, buildings are usually low and have a simple relationship with roads, mainly near the main roads or at the end of the branches of the main roads. This layout feature can effectively assist in identification, that is, by analyzing the spatial relationship between buildings and roads to identify and compensate for obscured buildings. The specific process is as follows:

[0190] S401, dividing the remote sensing image to be identified into a plurality of remote sensing image blocks;

[0191] S402, taking the remote sensing image block to be identified as the input of the change detection model, and the change detection model performs identification to obtain the identification result; wherein the remote sensing image is subjected to the same image preprocessing as the LEVIR-CD dataset to ensure input consistency; remote sensing images have problems such as sensor differences, atmospheric interference, and phase changes, and preprocessing (such as normalization and registration) can reduce the impact of these interferences on the model; if the preprocessing includes cropping the image to a fixed size, the subsequent input must also maintain the same size, otherwise the design of the network's fully connected layer or global pooling layer will cause dimensional mismatch;

[0192] S403, integrating the recognition results of the remote sensing image blocks to form the recognition result of the remote sensing image to be recognized; since the optical satellite remote sensing image is divided into a plurality of remote sensing image blocks to be recognized, the recognition results of the remote sensing image blocks to be recognized are integrated;

[0193] S404, converting the recognition result of the remote sensing image into a vector layer, calculating the coordinates of the center point of each patch surface contour, and obtaining coordinate information;

[0194] S405. Based on the coordinate information, the center point is overlapped with the road information for spatial analysis to determine the spatial position of the image spot and the road. If the positions intersect, the proximity does not exceed the preset distance value, and the image spot is an incomplete image spot, then the image spot is determined to be a remote sensing image of a building; wherein the image spot is a result of a building change contour image spot identified by a change detection model; in this embodiment, the preset distance value is 5 meters;

[0195] Specifically, in rural areas, buildings are usually close to roads. Therefore, if the center point of the patch intersects with the location information of the road and the distance does not exceed a certain range (for example, according to the Highway Safety Protection Regulations, the distance is not less than 5 meters for rural roads), the patch can be considered to be adjacent to the road. For patches identified as incomplete, if their location is adjacent to the road and considering the characteristics of rural buildings (such as small area, irregular shape, and small spacing), the patch can be determined to be a building remote sensing image. This method is particularly suitable for rural areas because the relationship between rural buildings and roads is relatively simple, and buildings are usually not far from the road.

[0196] Embodiment 3

[0197] This embodiment is basically the same as the above embodiment, except that:

[0198] Shadow-assisted recognition steps:

[0199] According to the time and longitude and latitude of the remote sensing image acquisition, when the brightness is within the preset brightness range, the shadows of the edge buildings in the building complex where the shadows are preset are identified; wherein the edge buildings where the shadows are preset are the edge buildings where the shadows should appear theoretically at the corresponding time and longitude and latitude;

[0200] For the same-direction edges of suspected buildings (which may be regular patterns), identify whether there is a shadow and the size of the shadow (there is historical data on whether there is no shadow), and use the shadow to supplement the change detection and identification;

[0201] Specifically, in remote sensing images, there is a close relationship between buildings and shadows; the presence of buildings blocks sunlight, thus forming shadows on the ground, and this relationship can be used to assist in identifying buildings, especially in the presence of regular patterns (which may even lead to misidentification of sports fields, deep pits, etc.);

[0202] The specific identification process is as follows:

[0203] According to the acquisition time and longitude and latitude of the remote sensing image, the altitude and azimuth of the sun are calculated to determine the theoretical direction of the shadow;

[0204] According to the spectral characteristics of remote sensing images, segmentation or classification methods are used to extract the shadow area of ​​remote sensing images. The significant feature of shadows on images is that their pixel grayscale values ​​are relatively low. Grayscale histogram threshold segmentation methods can be used to extract shadows, or color transformation models can be used to highlight the shadow area in the image in an invariant color space.

[0205] According to the relationship between the position of the building area and the corresponding shadow area and the azimuth of the sun, the disappeared building detection is performed to obtain the suspected building area; for example, if an area shows shadow features in the expected shadow direction, then this area may be caused by a building;

[0206] Combined with the shadow probability constraint, the shadows of tall buildings in water bodies can be extracted more accurately. This method can reduce the false detection rate and missed detection rate to less than 6%, and the overall classification accuracy and Kappa coefficient can reach more than 0.9.

[0207] For shadow areas, the information of shadow areas on remote sensing images is enhanced through transformation formulas, and the information of shadow areas is restored to reduce the brightness and chromaticity differences between shadow areas and surrounding areas.

[0208] The smoothing operator is used to smooth the edge of the shadow area. At the edge of the shadow area, due to sunlight refraction, mixed pixels and building umbra, the gray value on the image is larger than that in the center of the shadow area. Therefore, after compensation, the edge will appear brighter. For this reason, the smoothing operator is used to smooth the edge.

[0209] For the same-direction edges of suspected building areas, the existence and size of shadows are identified by comparing with historical shadow-free data. For example, compared with historical data, two buildings A and B are adjacent, and in fact the shadow of A is cast on B. If there is no shadow data of A in history, then the suspected building area is A, and it can be identified that it has no shadow.

[0210] This approach can help distinguish between shadow changes due to terrain changes and shadows due to the presence of buildings.

[0211] The above is only an embodiment of the present invention. The common sense such as the known specific structure and characteristics in the scheme is not described in detail here. The ordinary technicians in the relevant field know all the common technical knowledge in the technical field of the invention before the application date or priority date, can know all the existing technologies in the field, and have the ability to apply the conventional experimental means before that date. The ordinary technicians in the relevant field can improve and implement this scheme in combination with their own abilities under the enlightenment given by this application. Some typical known structures or known methods should not become obstacles for ordinary technicians in the relevant field to implement this application. It should be pointed out that for those skilled in the art, without departing from the structure of the present invention, several deformations and improvements can be made, which should also be regarded as the protection scope of the present invention, which will not affect the effect of the implementation of the present invention and the practicality of the patent. The protection scope required by this application shall be based on the content of its claims, and the specific implementation methods and other records in the specification can be used to interpret the content of the claims.

Claims

1. A building change detection method based on a deep separable dual-phase image transformer, characterized in that: It includes the following: Model construction steps: Build a change detection model for building change detection, including: S1, build a backbone network based on ResNet101 to extract feature maps; S2, using a dual-phase image converter to process the feature maps of the two phases; S3. Combined with the multi-scale jump structure of the residual module, an improved prediction head module is used to process the output of the dual-phase image transformer to generate the final change detection prediction result.

2. The building change detection method based on the depth-separable dual-phase image converter according to claim 1 is characterized in that: Also includes: Data preparation steps: Construct a building change detection dataset and perform image preprocessing, including: Select a public dataset for change detection for experiments; The original images in the public dataset for change detection are normalized and standardized, and the results of the normalization are: Among them, x norm is the normalized image value, min(x) is the minimum value of the image, max(x) is the maximum value of the image, x i is the image value to be processed; The result of the normalization process is: Among them, S is the standardized image value, x i is the image value to be processed, μ is the mean of the image, σ is the standard deviation of the image, and N is the number of image pixels; The public change detection dataset is divided into training set, validation set and test set to construct a building change detection dataset.

3. The building change detection method based on a depth-separable dual-phase image converter according to claim 1, characterized in that: The network architecture of the backbone network includes: Convolutional layer, used for initial feature extraction of input image; the image after image preprocessing; The maximum pooling layer is used to downsample the initial features; the residual module is used to extract features of different scales from the downsampled results, where the residual module adopts a multi-scale jump structure, including multiple stages, each stage includes multiple residual blocks, and extracts features of different scales in turn; The global average pooling layer is used to perform an average operation on each feature channel and convert the features of each channel into a single value; The fully connected layer is used to map the output of the global average pooling layer to the final feature vector to obtain the feature map of the dual-phase building.

4. The building change detection method based on a depth-separable dual-phase image converter according to claim 1, characterized in that: The S2 comprises: S201, input feature graphs of two time phases, extract each local area in each feature graph as a discrete semantic symbol Tokens, and obtain the global semantic information set T = {T 1 ,T 2 }, capturing semantic information and structural features; S202, using an encoder to encode the global semantic information set T; The multi-head self-attention module of the encoder includes the calculation of query, key and value, which is calculated as follows: Q=T (l-1) W q ; K=T (l-1) W k ; V=T (l-1) W v ; Among them, l is the length of semantic information, W q , W k and W v are the learning parameters of the multi-head self-attention module; S203, connecting the encoder outputs of the dual-phase image through a decoder to form a vector H={H1,H2}, reducing the dimension of the connected vector through linear transformation, and repeatedly stacking multiple decoders, where the decoder structure is consistent with the encoder; Among them, S202 includes: S20201. After the input is linearly transformed, the attention weight is obtained by calculating the correlation between the query and the key. The attention weight is multiplied by the value and weighted summed to obtain the final self-attention output, which is expressed as: Among them, σ() is the Sigmoid activation function; S20202. Calculate the results of multi-head self-attention: MultiHead(Q,K,V)=Concat(head1,head2,...,head h )W O ; Among them, head i =Att(QW Q i ,KW K i ,VW V i ), W O is the output projection matrix; S20203. Perform layer normalization on the output of the multi-head self-attention mechanism: Among them, μ is the mean of the image, σ is the standard deviation of the image; S20204. Perform linear transformation on the input vector and use nonlinear activation function to capture complex features: MLP(x) = ReLU(xW1+b1)W2+b2; Among them, MLP is a multi-layer perceptron, ReLU is a nonlinear activation function, and W1, b1, W2 and b2 are learning parameters.

5. The building change detection method based on the depth-separable dual-phase image converter according to claim 4 is characterized in that: The S3 includes: S301, extracting features from the output of the decoder using depthwise separable convolution; S302, using the ReLU activation function to process the output of the depthwise separable convolution to obtain feature A after adding nonlinear characteristics: A = ReLU(C); S303: Using the non-local information residual module, extract features through deep separable convolution to obtain feature maps A at different scales i ; S304: Combine the decoder result Z and the features at different scales Perform skip connection fusion to obtain the output value X of the non-local information residual module: S305: For the output value X, use the non-local weight matrix M ij , calculate the feature map and obtain the convolution feature Y: Among them, A ij is the non-local weight matrix, X ij is the pixel value of the feature map at pixel i, j; S306, convert the convolution feature Y into output dimension Y' through a linear layer: Y'=WY+b; Among them, W and b are learning parameters; S307, using Sigmoid activation function to generate probability prediction of changes 6. The building change detection method based on a depth-separable dual-phase image converter according to claim 5, characterized in that: The S301 includes: Perform depthwise convolution: Among them, Y i,j,m is the pixel value at the i,j position of the m-th channel of the output feature map, X i=d-1,j=d'-1,m is the pixel value at the i,j position of the m-th channel of the input feature map, K d,d',m is the parameter of the d,d'th position of the mth convolution kernel; The result after depth convolution is convolved point by point using a 1×1 convolution kernel: Among them, Z i,j,n is the pixel value at the i,j position of the nth channel of the output feature map, Y i,j,m is the pixel value of the i,j position of the mth channel of the output feature map of the depth convolution, L m,n are the parameters of point-wise convolution.

7. The building change detection method based on a depth-separable bi-temporal image converter according to claim 1, characterized in that: Also includes: Model training steps: Use the building change detection dataset to train and test the change detection model, output a change detection model that meets the requirements, and perform building change detection.

8. The building change detection method based on a depth-separable dual-phase image converter according to claim 1, characterized in that: Also includes: The steps of auxiliary recognition of ground roads are as follows: using change detection model recognition, performing change detection recognition, generating detection recognition results, filtering out incomplete detection results, and performing supplementary recognition through ground roads.

9. The building change detection method based on the depth-separable bi-temporal image converter according to claim 8 is characterized in that: The ground road auxiliary identification step includes: S401, dividing the remote sensing image to be identified into a plurality of remote sensing image blocks; S402, taking the remote sensing image block to be identified as the input of the change detection model, and the change detection model performs identification to obtain the identification result; S403, integrating the recognition results of the remote sensing image blocks to form the recognition result of the remote sensing image to be recognized; since the optical satellite remote sensing image is divided into a plurality of remote sensing image blocks to be recognized, the recognition results of the remote sensing image blocks to be recognized are integrated; S404, converting the recognition result of the remote sensing image into a vector layer, calculating the coordinates of the center point of each patch surface contour, and obtaining coordinate information; S405. Based on the coordinate information, the center point and the road information are superimposed for spatial analysis to determine the spatial position of the image patch and the road. If the positions intersect, the proximity does not exceed a preset distance value, and the image patch is an incomplete image patch, then the image patch is determined to be a remote sensing image of a building; wherein the image patch is a building change contour image patch result identified by a change detection model.

10. The building change detection method based on a depth-separable bi-temporal image converter according to claim 1, characterized in that: Also includes: Shadow auxiliary identification steps: According to the time and longitude and latitude of the remote sensing image acquisition, when the brightness is within the preset brightness range, the shadows of the edge buildings in the building complex where the shadows are preset are identified, including: According to the acquisition time and longitude and latitude of the remote sensing image, the altitude and azimuth of the sun are calculated to determine the theoretical direction of the shadow; According to the spectral characteristics of remote sensing images, segmentation or classification methods are used to extract the shadow area of ​​remote sensing images; According to the relationship between the position of the building area and the corresponding shadow area and the azimuth of the sun, the disappeared building detection is performed to obtain the suspected building area; For the shadow area, the information of the shadow area on the remote sensing image is enhanced through the transformation formula, and the information of the shadow area is restored; Use a smoothing operator to smooth the edges of the shadow area; For the same-direction edges of suspected building areas, the existence and size of shadows are identified by comparing with historical shadow-free data.

Citation Information

Cited By

  • Dual-time-phase remote sensing image change detection method, device, equipment and medium

    CN121033670A

  • Remote sensing image change prediction method and device and computer equipment

    CN121437428A