Building damage result detection method and device based on time sequence knowledge guidance
By acquiring medium- and high-resolution images to generate multi-temporal image sequences and extracting temporal feature vectors, the problem of insufficient high-resolution image samples is solved, enabling efficient and accurate prediction of building damage detection.
Patent Information
- Application Number
- CN202510920584.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-04
- Publication Date
- 2025-11-28
AI Technical Summary
In building damage detection, high-resolution satellite images are difficult to obtain and the number of samples is small, resulting in poor damage classification performance of deep learning models and making it difficult to achieve accurate building damage detection and assessment.
By acquiring multiple medium-resolution and high-resolution images of a building, downsampling is performed to generate medium-resolution images. Multi-temporal image sequences are generated based on the image acquisition time, temporal feature vectors are extracted, and a trained deep learning model is used to predict building damage.
This improved the richness of sample information, ensured the accuracy of the model's output predictions, and enabled efficient and accurate predictions in building damage detection.
Smart Images

Figure CN121033652A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present specification relate to the technical field of image processing, and particularly relate to a building damage result detection method based on time sequence knowledge guidance. BACKGROUND
[0002] When buildings in some areas are damaged due to a specific event, it is usually desirable to detect and evaluate the damage of buildings in the area. At present, satellite technology is mature, and city damage identification through satellite images has become a feasible technical means. However, for building damage detection, high-resolution satellite images are required, and such high-precision satellite images are difficult to obtain in the civil field, and the number of samples is very small. Therefore, the deep learning model trained by relying on high-resolution satellite images has poor performance in war damage classification.
[0003] It can be seen that, in the case of building damage detection which requires high resolution of satellite images but has a small number of image samples, how to train a deep learning model with ideal performance to accurately detect and evaluate the damage of city buildings in the event of a specific event is a technical problem that needs to be solved. SUMMARY
[0004] Therefore, the embodiments of the present specification provide a building damage result detection method based on time sequence knowledge guidance. One or more embodiments of the present specification also relate to a building damage result detection device based on time sequence knowledge guidance, a computing device, a computer-readable storage medium, and a computer program to solve the technical defects in the prior art.
[0005] According to a first aspect of the embodiments of the present specification, a building damage result detection method based on time sequence knowledge guidance is provided, comprising:
[0006] Obtaining a plurality of first medium-resolution images and a plurality of high-resolution images of a building, and performing down-sampling processing on the plurality of high-resolution images to generate corresponding second medium-resolution images;
[0007] Based on the image acquisition time of each image, processing the plurality of first medium-resolution images and the second medium-resolution images to generate a corresponding multi-temporal image sequence of the building;
[0008] By extracting features from the images in the multi-temporal image sequence, a corresponding time sequence feature vector is obtained;
[0009] Inputting the time sequence feature vector into a trained deep learning model for processing to obtain a prediction result of whether the building is damaged.
[0010] Optionally, the processing of the plurality of first medium-resolution images and the second medium-resolution image based on the image acquisition time of each image generates a multi-temporal image sequence corresponding to the building, including:
[0011] Stacking the plurality of first medium-resolution images and the second medium-resolution image based on the image acquisition time of each image, and dividing the stacking result into a plurality of grids that do not overlap with each other;
[0012] Generating a dual-temporal image corresponding to the building based on the first medium-resolution image and / or the second medium-resolution image contained in each grid, and generating a multi-temporal image corresponding to the building based on the dual-temporal image.
[0013] Optionally, the processing of the plurality of first medium-resolution images and the second medium-resolution image based on the image acquisition time of each image generates a multi-temporal image sequence corresponding to the building, including:
[0014] According to the sorting result of n images in a target grid, extracting the 1st target image and the i th target image in the target grid, and stacking the 1st target image and the i th target image according to the image acquisition time to generate a dual-temporal image corresponding to the building, wherein the target grid is each of the plurality of grids, n is a positive integer greater than or equal to 2, and i is a positive integer greater than 1 and less than or equal to n;
[0015] Stacking each dual-temporal image according to the image acquisition time corresponding to the i th target image contained in each dual-temporal image to generate a multi-temporal image sequence corresponding to the building.
[0016] Optionally, the feature extraction of the images in the multi-temporal image sequence to obtain a corresponding time sequence feature vector includes:
[0017] Determining a first dual-temporal image containing a first medium-resolution image in the multi-temporal image sequence;
[0018] Processing the first medium-resolution image in the first dual-temporal image according to the band value of the second medium-resolution image and through an upsampling block;
[0019] Feature extraction of the second medium-resolution image and the processed first medium-resolution image to obtain a corresponding time sequence feature vector.
[0020] Optionally, the feature extraction of the images in the multi-temporal image sequence to obtain a corresponding time sequence feature vector includes:
[0021] Determine the second biphase image in the multiphase image sequence, which is generated by stacking two second medium-resolution images;
[0022] The second dual-temporal image is convolved, and the convolution output is standardized using a batch normalization algorithm and a linear rectification unit.
[0023] By extracting features from the standardized image, the corresponding temporal feature vector is obtained.
[0024] Optionally, the step of extracting features from the standardized image to obtain the corresponding temporal feature vector includes:
[0025] Initial image features are obtained by extracting features from the standardized image using residual blocks.
[0026] The initial image features are processed by the self-attention mechanism of the first deep learning network, the processing result is normalized, and the normalization result is standardized to form a target feature vector, wherein the target feature vector is composed of 36 basic units of 128 dimensions.
[0027] The dimension of each basic unit is compressed by the fully connected layer of the first deep learning network, the compressed basic units are concatenated, and the concatenation result is determined as the corresponding temporal feature vector.
[0028] Optionally, the method for detecting building damage results based on time-series knowledge further includes:
[0029] Acquire multiple historical images of the building, wherein the multiple historical images include high-resolution images and medium-resolution images;
[0030] The high-resolution image is downsampled to generate a corresponding medium-resolution image, and a historical dual-temporal image of the building is constructed based on each medium-resolution image;
[0031] Damage result labels corresponding to some historical dual-temporal images are determined, and the historical dual-temporal images are stacked based on the image acquisition time of each medium-resolution image to generate a historical multi-temporal image sequence corresponding to the building.
[0032] The second deep learning model is trained based on the historical multi-temporal image sequence and the damage result label to obtain the trained deep learning model.
[0033] According to a second aspect of the embodiments of this specification, a building damage result detection device based on time-series knowledge is provided, comprising:
[0034] The acquisition module is configured to acquire multiple first medium-resolution images and multiple high-resolution images of the building, and to perform downsampling processing on the multiple high-resolution images to generate corresponding second medium-resolution images;
[0035] The processing module is configured to process the plurality of first medium-resolution images and second medium-resolution images based on the image acquisition time of each image to generate a multi-temporal image sequence corresponding to the building;
[0036] The extraction module is configured to obtain the corresponding temporal feature vector by performing feature extraction on the images in the multi-temporal image sequence;
[0037] The prediction module is configured to input the time-series feature vector into the trained deep learning model for processing to obtain a prediction result of whether the building is damaged.
[0038] According to a third aspect of the embodiments of this specification, a computing device is provided, comprising:
[0039] Memory and processor;
[0040] The memory is used to store computer-executable instructions, and the processor is used to execute any one of the steps of the time-series knowledge-guided building damage result detection method described in the computer-executable instructions.
[0041] According to a fourth aspect of the embodiments of this specification, a computer-readable storage medium is provided that stores computer-executable instructions, which, when executed by a processor, implement the steps of any one of the time-series knowledge-guided building damage result detection methods.
[0042] According to a fifth aspect of the embodiments of this specification, a computer program is provided, wherein when the computer program is executed in a computer, it causes the computer to perform the steps of the above-described method for detecting building damage results based on time-series knowledge.
[0043] The method for detecting building damage based on temporal knowledge provided in this specification involves acquiring multiple first medium-resolution images and multiple high-resolution images of a building, and downsampling the high-resolution images to generate corresponding second medium-resolution images. Based on the image acquisition time of each image, the multiple first and second medium-resolution images are processed to generate a multi-temporal image sequence corresponding to the building. Feature extraction is performed on the images in the multi-temporal image sequence to obtain corresponding temporal feature vectors. These temporal feature vectors are then input into a trained deep learning model for processing to obtain a prediction result indicating whether the building is damaged. Generating a multi-temporal image sequence improves the richness of information contained in the samples, thereby ensuring the accuracy of the prediction results output by the trained deep learning model when using the temporal feature vectors generated from the multi-temporal image sequence to predict building damage. Attached Figure Description
[0044] Figure 1 This is a flowchart of a method for detecting building damage results based on time-series knowledge, provided in one embodiment of this specification.
[0045] Figure 2 This is a schematic diagram of a building damage detection device based on time-series knowledge, provided in one embodiment of this specification.
[0046] Figure 3 This is a structural block diagram of a computing device provided in one embodiment of this specification. Detailed Implementation
[0047] Many specific details are set forth in the following description to provide a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.
[0048] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this specification. The singular forms “a,” “described,” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.
[0049] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."
[0050] This specification provides a method for detecting building damage results based on time-series knowledge. This specification also relates to a device for detecting building damage results based on time-series knowledge, a computing device, a computer-readable storage medium, and a computer program, which will be described in detail in the following embodiments.
[0051] Figure 1 A flowchart is shown of a method for detecting building damage results based on time-series knowledge according to an embodiment of this specification, which specifically includes the following steps.
[0052] Step 102: Acquire multiple first medium-resolution images and multiple high-resolution images of the building, and perform downsampling processing on the multiple high-resolution images to generate corresponding second medium-resolution images.
[0053] The images in the embodiments of this specification may be satellite images.
[0054] Specifically, a medium-resolution image of the building can be obtained, or if the obtained image contains a high-resolution image, the high-resolution image can be downsampled to a medium-resolution image.
[0055] The purpose of downsampling high-resolution images to medium-resolution images is as follows: Currently, most war damage detection methods rely on high-resolution imagery, but the acquisition of high-resolution imagery is extremely limited, especially during wartime when it is not publicly available. Therefore, using medium-resolution images, which have more accessibility, as the basis for war damage detection and assessment has become a current research hotspot. However, the accuracy of medium-resolution images is lower than that of high-resolution images, so relying on medium-resolution images for war damage detection and assessment is not easy. This specification addresses this problem by generating dual-temporal images (i.e., dual-temporal image patch samples) from medium-resolution images, and generating multi-temporal image sequences (i.e., multi-temporal patch sample sequences) from dual-temporal image patch samples.
[0056] In addition, the resolution of the high-resolution images in the embodiments of this specification can be: 0.5m resolution of the Earth Imaging Satellite (Worldview satellite); the resolution of the medium-resolution images can be 10m resolution of the Worldview satellite after downsampling or 10m resolution of the Sentinel-2 satellite.
[0057] The first medium-resolution image has a resolution of 10m from the Sentinel-2 satellite, and the second medium-resolution image has a resolution of 10m from the Worldview satellite after downsampling from the 0.5m resolution of the Worldview satellite.
[0058] Step 104: Based on the image acquisition time of each image, process the multiple first medium-resolution images and the second medium-resolution images to generate a multi-temporal image sequence corresponding to the building.
[0059] In one optional implementation, the step of processing the plurality of first medium-resolution images and second medium-resolution images based on the image acquisition time of each image to generate a multi-temporal image sequence corresponding to the building includes:
[0060] Based on the image acquisition time of each image, the multiple first medium-resolution images and the second medium-resolution images are stacked, and the stacking result is divided into multiple non-overlapping grids;
[0061] Based on the first medium-resolution image and / or the second medium-resolution image contained in each grid, a dual-temporal image corresponding to the building is generated, and a multi-temporal image corresponding to the building is generated based on the dual-temporal image.
[0062] Furthermore, the step of generating a dual-temporal image corresponding to the building based on a first medium-resolution image and / or a second medium-resolution image contained in each grid, and generating a multi-temporal image corresponding to the building based on the dual-temporal image, includes:
[0063] According to the sorting result of n images in the target grid, the first target image and the i-th target image in the target grid are extracted, and the first target image and the i-th target image are stacked according to the image acquisition time to generate a dual-temporal image corresponding to the building. Here, the target grid is each of the multiple grids, n is a positive integer greater than or equal to 2, and i is a positive integer greater than 1 and less than or equal to n.
[0064] Based on the image acquisition time corresponding to the i-th target image contained in each dual-temporal image, the dual-temporal images are stacked to generate a multi-temporal image sequence corresponding to the building.
[0065] Specifically, after acquiring high-resolution and medium-resolution images of buildings, and reducing the resolution of the high-resolution images to medium resolution through downsampling, the acquired medium-resolution images and the medium-resolution images generated by downsampling can be stacked sequentially according to the image acquisition time of each image, and the stacking result can be divided into several non-overlapping grids.
[0066] Furthermore, when each grid contains n images, images of each building before and after the conflict can be extracted from each grid in a "1 to n-1" manner. "1" represents a pre-war image before the conflict, and "n-1" represents n-1 post-war images after the conflict. Based on the image acquisition time corresponding to the pre-war image and the n-1 post-war images, the pre-war image and the n-1 post-war images are stacked sequentially according to the image acquisition time to form a dual-temporal image patch sample associated with the building, so as to reflect the damage changes of buildings in the area contained in the same grid in the time dimension.
[0067] With each grid containing 10 images, the first image in each grid can be identified as a pre-war image before the conflict, and the second to tenth images can be identified as nine post-war images after the conflict. Then, the first image is stacked with the second to tenth images in order of image acquisition time to form nine dual-temporal image patch samples associated with the building.
[0068] In practical applications, there are many ways to divide the grid. For example, the extract_patches_2d function in the sklearn.feature_extraction.image library can be used to divide a stacked image sequence into several non-overlapping grids. At the same time, this function can be used to perform batch image operations on the grids.
[0069] After generating the dual-temporal image patch samples corresponding to the building, the dual-temporal image patch samples can be stacked according to the image acquisition time corresponding to the post-war images contained in each dual-temporal image patch sample to generate a multi-temporal patch sample sequence corresponding to the building.
[0070] Step 106: Extract features from the images in the multi-temporal image sequence to obtain the corresponding temporal feature vector.
[0071] In one optional implementation, the step of extracting features from the images in the multi-temporal image sequence to obtain the corresponding temporal feature vector includes:
[0072] The first biphase image in the multiphase image sequence is identified as containing a first medium-resolution image;
[0073] Based on the band values of the second medium-resolution image, the first medium-resolution image in the first dual-temporal image is processed through an upsampling block;
[0074] By extracting features from the second medium-resolution image and the processed first medium-resolution image, the corresponding temporal feature vectors are obtained.
[0075] Specifically, since a multi-temporal image sequence is composed of multiple stacked bi-temporal images, and each bi-temporal image is composed of two stacked medium-resolution images, the bands required for different medium-resolution images may be different. For example, the first medium-resolution image may require both 10m and 20m bands, while the second medium-resolution image may only contain the 10m band. In this case, it is necessary to process the first medium-resolution image (multi-scale feature fusion) to convert the 20m band in the first medium-resolution image into the 10m band, so that the bands of the processed first medium-resolution image are consistent.
[0076] Based on this, the embodiments of this specification can first determine a first dual-temporal image containing a first medium-resolution image in a multi-temporal image sequence, and then process the first medium-resolution image in the first dual-temporal image using a multi-scale feature fusion method.
[0077] In practical applications, the accuracy of the first medium-resolution image can be 10m resolution from the Sentinel-2 satellite. There are 6 20m bands and 4 10m bands in the Sentinel-2 satellite image.
[0078] Specifically, the multi-scale feature fusion method employs an upsampling deconvolution strategy. The upscaling block consists of a 2×2 deconvolution layer and a 1×1 convolutional layer. The deconvolution layer enlarges the 3×3 pixel 20m band image to 6×6 pixels, and the convolutional layer then processes the deconvolution result, suppressing redundant features or noise information introduced by the upsampling feature map enlargement. This strategy only changes the spatial resolution matching feature map; the number of feature channels in the image remains constant at 12 to maintain resolution consistency.
[0079] This specification describes an embodiment that fuses 10m and 20m band images using a multi-scale feature fusion method. Specifically, it utilizes an upscaling block to maintain the same size between the low-band and high-band images, then concatenates the images to form a fused image. Features are then extracted from the fused image to obtain a corresponding temporal feature vector, which can be used as input to the subsequent first deep learning model. Upsampling operations are used to upscale the 20m band to the 10m band to form the fused image. This process, known as multi-scale feature fusion, aims to better capture features from each band, thereby enhancing subsequent processing.
[0080] In another optional implementation, the step of extracting features from the images in the multi-temporal image sequence to obtain the corresponding temporal feature vector includes:
[0081] Determine the second biphase image in the multiphase image sequence, which is generated by stacking two second medium-resolution images;
[0082] The second dual-temporal image is convolved, and the convolution output is standardized using a batch normalization algorithm and a linear rectification unit.
[0083] By extracting features from the standardized image, the corresponding temporal feature vector is obtained.
[0084] Furthermore, the step of extracting features from the standardized image to obtain the corresponding temporal feature vector includes:
[0085] Initial image features are obtained by extracting features from the standardized image using residual blocks.
[0086] The initial image features are processed by the self-attention mechanism of the first deep learning network, the processing result is normalized, and the normalization result is standardized to form a target feature vector, wherein the target feature vector is composed of 36 basic units of 128 dimensions.
[0087] The dimension of each basic unit is compressed by the fully connected layer of the first deep learning network, the compressed basic units are concatenated, and the concatenation result is determined as the corresponding temporal feature vector.
[0088] Specifically, as mentioned earlier, since the multi-temporal image sequence is composed of multiple stacked dual-temporal images, and each dual-temporal image is composed of two stacked medium-resolution images, when both medium-resolution images are of the same second medium resolution, pixel embedding can be performed on these two medium-resolution images. Specifically, convolution operation is performed on the two medium-resolution images, and the convolution output is standardized using batch normalization algorithm and linear rectifier unit. Deeper features are further extracted through two residual networks, and average pooling is performed in conjunction with the temporal attention module.
[0089] Specifically, a convolution operation is performed on two medium-resolution images. This involves using convolution kernels as weight matrices to perform local dot product operations on the feature maps (receptive fields) of the two medium-resolution images, accumulating the results, and then inputting them into their respective output channels. The initial convolution embedding of each pixel converts the number of channels from 20 to 64, and the initial convolution uses 64 1×1 convolution kernels.
[0090] After the convolution operation is completed, normalization can be used to calculate the mean and variance of the convolution feature map. Then, the normalization function is substituted to make the mean of the sample feature values "0" and the variance "1", thus forming a standardized output.
[0091] Specifically, after the convolution operation, to prevent the newly obtained feature data from falling into the saturation region when applied to the ReLU activation function, which would result in a small gradient during network training, this embodiment uses Batch Normalization (BN) to ensure that the feature map after convolution has a mean of 0 and a variance of 1. Then, the pixel values of all points in each channel are counted, and the mean and variance are calculated. The normalized pixel value is obtained by subtracting the mean from the pixel value and dividing by the variance. This normalized pixel value is then applied to the ReLU activation function, retaining positive pixel values and assigning 0 to negative pixel values. This ensures the sparsity and positive distribution of the activation output, giving the convolution output non-linear characteristics.
[0092] Furthermore, the method for extracting deep features from the normalized processing result using two residual networks is as follows: Each residual block in the residual network contains two convolutional layers. These two convolutional layers are used to extract features from the normalized image to obtain the initial image features. In the embodiments of this specification, the residual network solves the gradient vanishing and exploding problems during model training through skip connections.
[0093] After obtaining the initial image features, these features are input into the first deep learning network (Transformer network) to process the initial image features using a pixel-based Transformer network, generating the target feature vector, i.e., Tokens T. 2 A sequence in the form of ×36, Tokens T 2 The expression for the ×36 form sequence is: T B×36×128 Where Token is the basic unit, B represents the batch size, that is, multiple target feature vectors will be input into the trained deep learning model in batches of B, 36 is the length of each feature map after flattening, that is, each feature map contains 36 basic units after flattening, and 128 is the dimension of each Token.
[0094] In the embodiments of this specification, the Transformer network processes the initial image features as follows: Location information is added to the time series corresponding to the initial image features; a self-attention mechanism is used to process the initial image features, forming a global time dependency and focusing on key change information; residual connections and normalization are performed to connect the attention output and the original input, stabilizing the training process; a multilayer perceptron is used to enhance nonlinear representation capabilities and fuse temporal context information; finally, Tokens T is formed through feature normalization. 1×36 format sequence.
[0095] The process of image processing of initial image features using a self-attention mechanism is as follows: spatial information is preserved by adding positional encoding to each basic unit; the initial image features are processed using a self-attention mechanism, and Q, K, and V are calculated using a learnable weight matrix and attention weights are calculated to reflect the semantic association between each basic unit; multi-head attention is used to capture feature patterns of different subspaces in parallel, generate images, and stitch them together for projection.
[0096] In the embodiments described in this specification, when processing the initial image features using a self-attention mechanism, the image is segmented into multiple subspaces, and then each subspace is processed. Additionally, the feature pattern can be a feature value. By comparing the changes in feature values between subspaces of two images, or the magnitude of changes between different subspaces, this indicator is used to highlight the change information of that subspace. Then, based on the change information of each subspace, a large image is stitched and projected to form a larger image, which reflects all the changes before and after the war.
[0097] Furthermore, after the self-attention mechanism, Layer Normalization is used to normalize the data in the sequence dimension. The method is as follows: normalize the feature vector in each input sequence, adjust the mean of each feature vector to 0, and adjust the variance to 1.
[0098] For Tokens T 1 After normalizing the ×36 sequence, the normalization result is standardized to form Tokens T. 2 The sequence is in the form of ×36. Feature normalization (FN) is used here, which is a special case of BN with scaling and translation parameters of 0 and 1. After normalization through self-attention mechanism, the subsequent fully connected layer will compress the semantic tag length of the basic unit from 128 dimensions to 4 dimensions, but the total number of tokens (semantic tags) remains unchanged at 36.
[0099] Finally, each semantic token, compressed to 4 dimensions, is concatenated into a semantic vector of length 144, thereby combining the features of each sample in the multi-temporal patch sample sequence into a temporal feature vector.
[0100] Step 108: Input the time-series feature vector into the trained deep learning model for processing to obtain a prediction result of whether the building is damaged.
[0101] Specifically, after obtaining the temporal feature vector, it can be input into the trained deep learning model for processing to obtain a prediction result of whether the building is damaged. By outputting the classification result of the target domain, the model can accurately identify the damage state of a small number of target domains.
[0102] In one alternative implementation, the trained deep learning model can be obtained by training in the following manner:
[0103] Acquire multiple historical images of the building, wherein the multiple historical images include high-resolution images and medium-resolution images;
[0104] The high-resolution image is downsampled to generate a corresponding medium-resolution image, and a historical dual-temporal image of the building is constructed based on each medium-resolution image;
[0105] Damage result labels corresponding to some historical dual-temporal images are determined, and the historical dual-temporal images are stacked based on the image acquisition time of each medium-resolution image to generate a historical multi-temporal image sequence corresponding to the building.
[0106] The second deep learning model is trained based on the historical multi-temporal image sequence and the damage result label to obtain the trained deep learning model.
[0107] In the embodiments described in this specification, a semi-supervised domain adaptation strategy is used for model training, that is, the model is trained in the source domain using a small number of labeled samples from the source domain and a large number of unlabeled samples from the target domain. Alternatively, a small number of labeled samples from the target domain can be used together with labeled samples from the source domain to participate in the model training.
[0108] For a small number of labeled samples from the source domain, they can be obtained in the following way:
[0109] First, multiple historical images of the building are acquired, including high-resolution and medium-resolution images. The high-resolution images are then downsampled to medium-resolution images. The acquired medium-resolution images and the downsampled medium-resolution images are then stacked sequentially according to their acquisition time, and the stacked results are divided into several non-overlapping grids. Furthermore, if each grid contains n historical images, historical images of each building before and after the conflict can be extracted from each grid in a "1-to-n-1" manner. "1" represents a pre-conflict historical image, and "n-1" represents n-1 post-conflict historical images. Based on the acquisition time of the pre-conflict and post-conflict images, the pre-conflict images are stacked sequentially with the n-1 post-conflict images according to their acquisition time, forming a historical bi-temporal image associated with the building. Finally, historical bi-temporal images showing building damage are marked as positive samples (damaged), and those not marked as negative samples (undamaged). After labeling each historical dual-temporal image generated from the same grid, the historical dual-temporal images can be stacked according to the image acquisition time corresponding to the historical post-war images contained in each historical dual-temporal image to generate a historical multi-temporal image sequence corresponding to the building.
[0110] For a large number of unlabeled samples from the target domain, the same method can be used. The difference is that the historical pre-war images in the target domain are stacked sequentially with n-1 historical post-war images according to the image acquisition time to form a historical bi-temporal image associated with the building. There is no need to label them. Instead, the historical bi-temporal images are stacked according to the image acquisition time of the historical post-war images contained in each historical bi-temporal image to generate a historical multi-temporal image sequence corresponding to the building.
[0111] After generating the historical multi-temporal image sequences corresponding to the source and target domains, feature extraction can be performed on the images in the multi-temporal image sequences to obtain the corresponding temporal feature vectors. These temporal feature vectors can then be input into the second deep learning model to be trained. The specific implementation method for feature extraction from the images in the multi-temporal image sequences to obtain the corresponding temporal feature vectors can be found in step 106, and will not be repeated here.
[0112] The second deep learning model to be trained includes a TCD (Trajectory Consistency Distillation) module. The TCD module is a temporal convolutional decoder containing two one-dimensional convolutional layers (kernel size of 3 and padding of 1). In addition, the TCD module processes the output of the one-dimensional convolutional layers through a linear rectified function (ReLU).
[0113] The aforementioned steps extract temporal feature vectors (including a 144-dimensional semantic vector) from historical multi-temporal image sequences. These temporal feature vectors are then input into the second deep learning model to be trained, where the TCD module performs temporal convolution to fuse local temporal context information. The resulting feature vector (after temporal convolution) is the input to the last layer of the TCD module. This feature vector consists of a 144-dimensional semantic vector plus temporal local context information. The input feature vector to the last layer of the TCD module is then divided into patch features f. Specifically, the feature vector can be segmented into independent time-step features based on the temporal dimension, with each time step corresponding to one f.
[0114] After generating patch features f, the MMD (domain adaptation) loss can be calculated based on the source domain features and the target domain features to align the patch features f, thereby reducing the distribution differences between different domains and making the features of the source domain and the target domain more similar.
[0115] In the embodiments of this specification, the method for calculating the domain adaptation loss between the source domain and the target domain is as follows: A Gaussian kernel matrix is calculated between the samples in the source domain and the target domain using a Gaussian kernel function. Specifically, the samples in the source domain and the target domain are concatenated into a merge matrix, and the square of the Euclidean distance between every two samples in the merge matrix is calculated. A kernel matrix is calculated independently for each bandwidth. The kernel matrix is then divided into four types according to the number of samples in the source and target domains (source / source, target / target, source / target, target / source). The kernel matrices are normalized to obtain their respective similarities, and the sum of these similarities is the loss.
[0116] The method for aligning patch features f is as follows: based on the feature distribution differences obtained from the domain loss results, a sample prototype labeled with sample features is constructed, and the compactness of similar features in spherical space is constrained to enhance the feature alignment effect. Feature distribution alignment is achieved by optimizing parameter settings through backpropagation gradient descent.
[0117] After aligning the patch features f, the TCD module can be used to generate the prediction result CS′ as an intermediate feature, which can then be used together with the ground truth label to participate in subsequent loss comparison and optimization.
[0118] The TCD module generates the prediction result CS′ based on the patch features f as follows: the temporal feature vector is convolved by the TCD module to extract patch features, then the MMD loss of the patch features in the source domain and the target domain is calculated to achieve feature alignment and reduce distribution differences, then the distance between the expected features of the two distributions in the feature space is calculated to evaluate their similarity, and then the corresponding prediction result is output based on the patch features through a fully connected layer.
[0119] In this embodiment, after dividing the input feature vector of the last layer of the TCD module into patch features f, the local receptive field of the TCD is used to simulate and analyze the temporal trend of building damage status. Based on the analysis results, a gradual feature of local temporal context is added to the feature at each time step. The temporal feature vector is processed by the TCD module to capture temporal changes and generate a classification probability vector. Then, the output features of the TCD module are compressed into a single predicted classification result by a multilayer perceptron. The above operations encode each historical dual-temporal image in the historical multi-temporal image sequence into a semantic vector of size 144. The semantic vectors are ordered sequentially in the time dimension. The TCD module compresses the semantic vectors into classification probability vectors. The classification probability vectors participate in the subsequent loss optimization process, thereby improving the model's classification performance.
[0120] Furthermore, after generating the prediction result CS′, a global time constraint is applied to the prediction result CS′ to calculate the loss of the prediction result. The loss calculation method is implemented through the following formulas (1) to (5):
[0121] TotalLoss=CELoss(C′,l)+TTV(CS′) Formula (1)
[0122] In formula (1), TotalLoss represents the final loss calculated on the prediction result CS′.
[0123] CELoss(C′, l) represents the cross-entropy loss, used to calculate the mapping error loss between the damage prediction label C′ and the ground truth value l for each image.
[0124] CELoss(C′, l) can be calculated using the following formula:
[0125]
[0126] Among them, l m It is the m-th component of the ground truth label; C′ m is the probability value of the m-th class predicted by the model; TTV(CS′) is the TTV regularization term, used to measure the deviation between the predicted result CS′ and its damage mode; the relationship between C′ and CS′ is: CS′ is a set of m C′.
[0127] TTV(CS′) is calculated using the following formula (3):
[0128]
[0129] In formula (3), B is the weighting coefficient; B is the batch size. It is the sum of the absolute values of the first-order differences of the predicted sequence CS′, representing the number of transitions between the two states in the predicted sequence (the building damage state is irreversible in the time dimension, that is, the number of transitions between the building damage state does not exceed 1).
[0130] Calculated using the following formula (4):
[0131]
[0132] In formula (4), j represents the j-th time. and These represent the predicted values of a single sample in the sequence at times j+1 and j, respectively; for Length; g p represents the grid; p represents the grid number.
[0133] The piecewise constant indicates that regularization applies a global constraint only to samples with significant time variations.
[0134] It is calculated using the following formula (5):
[0135]
[0136] In formula (5), The predicted value represents the initial state of the sample.
[0137] In this embodiment, the labeled sample features are constructed according to region (source domain or target domain) and state (damaged or undamaged). The respective sample feature prototypes are calculated and normalized to unit vectors, so that all sample feature prototypes fall on a unit hypersphere. The distance between each feature prototype and other feature prototypes is calculated based on the distance between vectors on the sphere to measure feature similarity. Based on whether they come from images of the same category, the model can map patches from two domains with the same category to the feature space and extract discriminative features, thus achieving supervised contrastive learning.
[0138] Furthermore, to verify the accuracy of the model's detection results, specific event data from relevant news and investigative reports can be used to verify the accuracy of the predictions regarding building damage. A two-way fixed effects regression model can be used to analyze the correlation between the damage predictions and the actual bombing events, thus assessing the reliability of the predictions. The method for analyzing damage predictions using a two-way fixed effects regression model is briefly described below:
[0139] A dataset of damaged buildings at multiple time points and locations was constructed. The model comprehensively considers regional and temporal fixed effects, eliminates some influencing features, uses damage prediction as the independent variable and actual bombing data as the dependent variable, establishes a regression equation, performs significance tests on the regression coefficients, and evaluates the model's explanatory power using model fit and error term analysis.
[0140] Furthermore, for interpreting Grad-CAM heatmaps, Gradient Weighted Class Activation Mapping (Grad-CAM) can be used to generate heatmaps. Grad-CAM calculates the importance of each feature map to the predicted class label by predicting the output label through backpropagation and computing gradients. The importance of these channels is then summed with their weights to generate a heatmap, where color represents the contribution of each region to the classification result. This heatmap visually demonstrates the model's decision-making process, helping to understand which image regions contribute most to the prediction of building damage and highlighting the extent of damaged areas.
[0141] Regarding the construction of urban destruction maps in urban destruction impact assessment, after verifying the spatial distribution and scale of severely damaged areas in the city, an urban destruction map is constructed based on the prediction results, and it is spatially overlaid with population and critical infrastructure (such as hospitals, schools, etc.) from international data sources (worldPop, OpenStreetMap, GlobalMLBuildingFootprints).
[0142] The quantitative assessment of indicators in urban damage impact assessment involves evaluating the impact of specific events on the city by statistically analyzing indicators such as the number of damaged buildings, population, school-age population, and population seeking medical assistance in the damaged area.
[0143] This specification's embodiments address the technical problem of poor performance in deep learning models trained for building damage detection due to insufficient numbers of high-resolution battle-damaged sample images by generating historical multi-temporal image sequences. Specifically, generating historical multi-temporal image sequences enhances the richness of information contained in the samples, thereby improving the accuracy of the model's output. Furthermore, using the input features of the last layer of the TCD module as the classification basis enables the deep learning model to learn more detailed classification features, significantly improving the model's classification performance.
[0144] The method for detecting building damage based on temporal knowledge provided in this specification involves acquiring multiple first medium-resolution images and multiple high-resolution images of a building, and downsampling the high-resolution images to generate corresponding second medium-resolution images. Based on the image acquisition time of each image, the multiple first and second medium-resolution images are processed to generate a multi-temporal image sequence corresponding to the building. Feature extraction is performed on the images in the multi-temporal image sequence to obtain corresponding temporal feature vectors. These temporal feature vectors are then input into a trained deep learning model for processing to obtain a prediction result indicating whether the building is damaged. Generating a multi-temporal image sequence improves the richness of information contained in the samples, thereby ensuring the accuracy of the prediction results output by the trained deep learning model when using the temporal feature vectors generated from the multi-temporal image sequence to predict building damage.
[0145] Corresponding to the above method embodiments, this specification also provides an embodiment of a building damage result detection device guided by time-series knowledge. Figure 2 This specification illustrates a schematic diagram of a building damage detection device based on time-series knowledge, according to one embodiment of this specification. Figure 2 The device includes:
[0146] The acquisition module 202 is configured to acquire multiple first medium-resolution images and multiple high-resolution images of the building, and to perform downsampling processing on the multiple high-resolution images to generate corresponding second medium-resolution images;
[0147] Processing module 204 is configured to process the plurality of first medium-resolution images and second medium-resolution images based on the image acquisition time of each image to generate a multi-temporal image sequence corresponding to the building;
[0148] Extraction module 206 is configured to obtain corresponding temporal feature vectors by performing feature extraction on images in the multi-temporal image sequence;
[0149] The prediction module 208 is configured to input the time-series feature vector into the trained deep learning model for processing to obtain a prediction result of whether the building is damaged.
[0150] Optionally, the processing module 204 is further configured to:
[0151] Based on the image acquisition time of each image, the multiple first medium-resolution images and the second medium-resolution images are stacked, and the stacking result is divided into multiple non-overlapping grids;
[0152] Based on the first medium-resolution image and / or the second medium-resolution image contained in each grid, a dual-temporal image corresponding to the building is generated, and a multi-temporal image corresponding to the building is generated based on the dual-temporal image.
[0153] Optionally, the processing module 204 is further configured to:
[0154] According to the sorting result of n images in the target grid, the first target image and the i-th target image in the target grid are extracted, and the first target image and the i-th target image are stacked according to the image acquisition time to generate a dual-temporal image corresponding to the building. Here, the target grid is each of the multiple grids, n is a positive integer greater than or equal to 2, and i is a positive integer greater than 1 and less than or equal to n.
[0155] Based on the image acquisition time corresponding to the i-th target image contained in each dual-temporal image, the dual-temporal images are stacked to generate a multi-temporal image sequence corresponding to the building.
[0156] Optionally, the extraction module 206 is further configured to:
[0157] The first biphase image in the multiphase image sequence is identified as containing a first medium-resolution image;
[0158] Based on the band values of the second medium-resolution image, the first medium-resolution image in the first dual-temporal image is processed through an upsampling block;
[0159] By extracting features from the second medium-resolution image and the processed first medium-resolution image, the corresponding temporal feature vectors are obtained.
[0160] Optionally, the extraction module 206 is further configured to:
[0161] Determine the second biphase image in the multiphase image sequence, which is generated by stacking two second medium-resolution images;
[0162] The second dual-temporal image is convolved, and the convolution output is standardized using a batch normalization algorithm and a linear rectification unit.
[0163] By extracting features from the standardized image, the corresponding temporal feature vector is obtained.
[0164] Optionally, the extraction module 206 is further configured to:
[0165] Initial image features are obtained by extracting features from the standardized image using residual blocks.
[0166] The initial image features are processed by the self-attention mechanism of the first deep learning network, the processing result is normalized, and the normalization result is standardized to form a target feature vector, wherein the target feature vector is composed of 36 basic units of 128 dimensions.
[0167] The dimension of each basic unit is compressed by the fully connected layer of the first deep learning network, the compressed basic units are concatenated, and the concatenation result is determined as the corresponding temporal feature vector.
[0168] Optionally, the time-series knowledge-guided building damage detection device further includes: a training module configured to:
[0169] Acquire multiple historical images of the building, wherein the multiple historical images include high-resolution images and medium-resolution images;
[0170] The high-resolution image is downsampled to generate a corresponding medium-resolution image, and a historical dual-temporal image of the building is constructed based on each medium-resolution image;
[0171] Damage result labels corresponding to some historical dual-temporal images are determined, and the historical dual-temporal images are stacked based on the image acquisition time of each medium-resolution image to generate a historical multi-temporal image sequence corresponding to the building.
[0172] The second deep learning model is trained based on the historical multi-temporal image sequence and the damage result label to obtain the trained deep learning model.
[0173] The above is a schematic scheme of a building damage result detection device based on time-series knowledge in this embodiment. It should be noted that the technical solution of this building damage result detection device based on time-series knowledge and the technical solution of the building damage result detection method based on time-series knowledge described above belong to the same concept. For details not described in detail in the technical solution of the building damage result detection device based on time-series knowledge, please refer to the description of the technical solution of the building damage result detection method based on time-series knowledge described above.
[0174] Figure 3 A structural block diagram of a computing device 300 according to one embodiment of this specification is shown. The components of the computing device 300 include, but are not limited to, a memory 310 and a processor 320. The processor 320 is connected to the memory 310 via a bus 330, and a database 350 is used to store data.
[0175] The computing device 300 also includes an access device 340, which enables the computing device 300 to communicate via one or more networks 360. Examples of these networks include a Public Switched Telephone Network (PSTN), a Local Area Network (LAN), a Wide Area Network (WAN), a Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 340 may include one or more of any type of wired or wireless network interface (e.g., a Network Interface Card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) interface, a Wi-MAX interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, and so on.
[0176] In one embodiment of this specification, the aforementioned components of the computing device 300 and Figure 3 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 3 The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.
[0177] The computing device 300 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or PCs. The computing device 300 can also be a mobile or stationary server.
[0178] The processor 320 is used to execute the following computer-executable instructions, which, when executed by the processor, implement the steps of the above-mentioned time-series knowledge-guided building damage result detection method.
[0179] The above is a schematic diagram of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the above-described method for detecting building damage results based on time-series knowledge belong to the same concept. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solution of the above-described method for detecting building damage results based on time-series knowledge.
[0180] An embodiment of this specification also provides a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement the steps of the above-described method for detecting building damage results based on timing knowledge.
[0181] The above is an illustrative scheme of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium belongs to the same concept as the technical solution of the building damage result detection method based on time-series knowledge described above. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solution of the building damage result detection method based on time-series knowledge described above.
[0182] An embodiment of this specification also provides a computer program, wherein when the computer program is executed in a computer, it causes the computer to perform the steps of the above-described method for detecting building damage results based on time-series knowledge.
[0183] The above is an illustrative scheme of a computer program according to this embodiment. It should be noted that the technical solution of this computer program belongs to the same concept as the technical solution of the above-described method for detecting building damage results based on time-series knowledge. For details not described in detail in the technical solution of the computer program, please refer to the description of the technical solution of the above-described method for detecting building damage results based on time-series knowledge.
[0184] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0185] The computer instructions include computer program code, which may be in the form of source code, object code, executable file, or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately added to or subtracted according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.
[0186] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments in this specification are not limited to the described order of actions, because according to the embodiments in this specification, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments in this specification.
[0187] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0188] The preferred embodiments disclosed above are merely illustrative of this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the embodiments described herein. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the embodiments, thereby enabling those skilled in the art to better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.
Claims
1. A method for detecting building damage results based on time-series knowledge, comprising: Acquire multiple first medium-resolution images and multiple high-resolution images of the building, and perform downsampling processing on the multiple high-resolution images to generate corresponding second medium-resolution images; Based on the image acquisition time of each image, the multiple first medium-resolution images and the second medium-resolution images are processed to generate a multi-temporal image sequence corresponding to the building; By extracting features from the images in the multi-temporal image sequence, the corresponding temporal feature vectors are obtained; The time-series feature vector is input into the trained deep learning model for processing to obtain a prediction result of whether the building is damaged.
2. The method for detecting building damage results based on temporal knowledge according to claim 1, wherein processing the plurality of first medium-resolution images and the second medium-resolution images based on the image acquisition time of each image to generate a multi-temporal image sequence corresponding to the building includes: Based on the image acquisition time of each image, the multiple first medium-resolution images and the second medium-resolution images are stacked, and the stacking result is divided into multiple non-overlapping grids; Based on the first medium-resolution image and / or the second medium-resolution image contained in each grid, a dual-temporal image corresponding to the building is generated, and a multi-temporal image corresponding to the building is generated based on the dual-temporal image.
3. The method for detecting building damage results based on temporal knowledge as described in claim 2, wherein generating a dual-temporal image corresponding to the building based on a first medium-resolution image and / or a second medium-resolution image contained in each grid, and generating a multi-temporal image corresponding to the building based on the dual-temporal image, comprises: According to the sorting result of n images in the target grid, the first target image and the i-th target image in the target grid are extracted, and the first target image and the i-th target image are stacked according to the image acquisition time to generate a dual-temporal image corresponding to the building. Here, the target grid is each of the multiple grids, n is a positive integer greater than or equal to 2, and i is a positive integer greater than 1 and less than or equal to n. Based on the image acquisition time corresponding to the i-th target image contained in each dual-temporal image, the dual-temporal images are stacked to generate a multi-temporal image sequence corresponding to the building.
4. The method for detecting building damage results based on temporal knowledge according to claim 1, wherein the step of extracting features from the images in the multi-temporal image sequence to obtain the corresponding temporal feature vector includes: The first biphase image in the multiphase image sequence is identified as containing a first medium-resolution image; Based on the band values of the second medium-resolution image, the first medium-resolution image in the first dual-temporal image is processed through an upsampling block; By extracting features from the second medium-resolution image and the processed first medium-resolution image, the corresponding temporal feature vectors are obtained.
5. The method for detecting building damage results based on temporal knowledge according to claim 1, wherein the step of extracting features from the images in the multi-temporal image sequence to obtain the corresponding temporal feature vector includes: Determine the second biphase image in the multiphase image sequence, which is generated by stacking two second medium-resolution images; The second dual-temporal image is convolved, and the convolution output is standardized using a batch normalization algorithm and a linear rectification unit. By extracting features from the standardized image, the corresponding temporal feature vector is obtained.
6. The method for detecting building damage results based on temporal knowledge according to claim 5, wherein the step of extracting features from the standardized image to obtain the corresponding temporal feature vector includes: Initial image features are obtained by extracting features from the standardized image using residual blocks. The initial image features are processed by the self-attention mechanism of the first deep learning network, the processing result is normalized, and the normalization result is standardized to form a target feature vector, wherein the target feature vector is composed of 36 basic units of 128 dimensions. The dimension of each basic unit is compressed by the fully connected layer of the first deep learning network, the compressed basic units are concatenated, and the concatenation result is determined as the corresponding temporal feature vector.
7. The method for detecting building damage results based on time-series knowledge according to claim 1 further includes: Acquire multiple historical images of the building, wherein the multiple historical images include high-resolution images and medium-resolution images; The high-resolution image is downsampled to generate a corresponding medium-resolution image, and a historical dual-temporal image of the building is constructed based on each medium-resolution image; Damage result labels corresponding to some historical dual-temporal images are determined, and the historical dual-temporal images are stacked based on the image acquisition time of each medium-resolution image to generate a historical multi-temporal image sequence corresponding to the building. The second deep learning model is trained based on the historical multi-temporal image sequence and the damage result label to obtain the trained deep learning model.
8. A building damage detection device guided by time-series knowledge, comprising: The acquisition module is configured to acquire multiple first medium-resolution images and multiple high-resolution images of the building, and to perform downsampling processing on the multiple high-resolution images to generate corresponding second medium-resolution images; The processing module is configured to process the plurality of first medium-resolution images and second medium-resolution images based on the image acquisition time of each image to generate a multi-temporal image sequence corresponding to the building; The extraction module is configured to obtain the corresponding temporal feature vector by performing feature extraction on the images in the multi-temporal image sequence; The prediction module is configured to input the time-series feature vector into the trained deep learning model for processing to obtain a prediction result of whether the building is damaged.
9. A computing device, comprising: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, they implement the steps of the building damage result detection method based on time-series knowledge as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement the steps of the time-series knowledge-guided building damage result detection method according to any one of claims 1 to 7.
Citation Information
Patent Citations
High-resolution optical remote sensing image building change detection method
CN115471467A
Method and program for time series image analysis
KR101628723B1