A method, system, electronic device and medium for change detection of remote sensing images
Through the twin nearest neighbor sliding window Transformer network model, the remote sensing image features are extracted and differential features are fused by using the encoder and decoder structure, which solves the problem of inefficiency in remote sensing image change detection and achieves efficient and accurate change detection.
Patent Information
- Application Number
- CN202310258534.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-14
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2043-03-14
AI Technical Summary
The prior art is difficult to efficiently and accurately identify real change information in remote sensing image change detection, and it is inefficient when facing a large amount of data and calculation amount.
The twin nearest neighbor sliding window Transformer network model is used to extract feature maps of different resolutions through the encoder and decoder structure, combine the self-attention mechanism and differential feature fusion, and use the loss value to update the model parameters to improve detection accuracy.
The efficiency and accuracy of remote sensing image change detection are improved, the calculation amount and training cost are reduced, and the practicality and detection speed of the model are enhanced.
Smart Images

Figure CN116343034B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image detection, and particularly to a method, a system, an electronic device and a medium for change detection of remote sensing images. Background Art
[0002] Change detection of remote sensing images is an important method for analyzing surface changes. This method mainly obtains change information of the surface area by analyzing remote sensing images of the same surface area at different times. When analyzing the change information of remote sensing images, due to complex scene conditions and different imaging conditions within the surface area, the same surface area may exhibit different spectral characteristics at different times, resulting in some irrelevant changes in the remote sensing images. For example, image changes caused by seasonal changes, building shadows, atmospheric changes, lighting condition changes or other irrelevant change conditions. Such irrelevant changes in the images affect the true change detection of remote sensing images. Therefore, current change detection of remote sensing images focuses on identifying the true change information of remote sensing images to obtain accurate surface change information.
[0003] However, currently, when performing the change detection task of remote sensing images, it is difficult for the detection personnel to efficiently and accurately identify the true change information of remote sensing images in the face of a large amount of data volume and calculation volume. Summary of the Invention
[0004] In view of the above-mentioned disadvantages of the prior art, the purpose of the present invention is to provide a method, a system, an electronic device and a medium for change detection of remote sensing images, and the present invention can improve the efficiency and accuracy of change detection of remote sensing images.
[0005] To achieve the above object and other related objects, the present invention provides a method for change detection of remote sensing images, including:
[0006] Obtain a preset training image set and a plurality of change label maps, wherein the training image set includes a plurality of pre-temporal remote sensing images and the corresponding post-temporal remote sensing images of each of the pre-temporal remote sensing images, and each change label map is used to indicate the image change data between each pre-temporal remote sensing image and the corresponding post-temporal remote sensing image;
[0007] Perform image preprocessing on all the image data in the training image set to generate an input image set;
[0008] Establish a change detection network model, wherein the change detection network model includes an encoder structure and a decoder structure;
[0009] Input a certain pre-temporal remote sensing image and the corresponding post-temporal remote sensing image in the input image set into the encoder structure to output a plurality of pre-temporal feature maps and corresponding post-temporal feature maps with different resolutions;
[0010] Based on the decoder structure, perform fusion processing on the difference features between each pair of the previous-temporal feature maps and the corresponding subsequent-temporal feature maps to generate a change prediction image;
[0011] Based on the loss value between the change prediction image and the corresponding change label map, update the parameters of the change detection network model to establish a trained change detection network model;
[0012] Input a preset previous-temporal test image and the corresponding subsequent-temporal test image into the trained change detection network model to output a target change map.
[0013] In an embodiment of the present invention, the step of performing image preprocessing on all image data in the training image set to generate an input image set includes:
[0014] Perform cropping processing on all image data in the training image set and the corresponding change label maps;
[0015] Perform preprocessing operations on all the cropped image data to generate an input image set, where the preprocessing operations include grayscale processing, geometric transformation processing, and image enhancement processing.
[0016] In an embodiment of the present invention, the step of inputting a certain previous-temporal remote sensing image and the corresponding subsequent-temporal remote sensing image in the input image set into the encoder structure to output multiple previous-temporal feature maps and corresponding subsequent-temporal feature maps with different resolutions includes:
[0017] Based on the encoder structure, perform block segmentation processing on a certain previous-temporal remote sensing image and the corresponding subsequent-temporal remote sensing image in the input image set to generate a previous-temporal segmented image and a subsequent-temporal segmented image;
[0018] Perform linear mapping on the previous-temporal segmented image and the subsequent-temporal segmented image to adjust the image dimensions of the previous-temporal segmented image and the subsequent-temporal segmented image;
[0019] At each encoding stage of the encoder structure, perform downsampling operations on the previous-temporal segmented image and the subsequent-temporal segmented image after dimension adjustment;
[0020] Perform window segmentation operations on the downsampled previous-temporal segmented image and the subsequent-temporal segmented image;
[0021] Perform self-attention mechanism operations within the window on the previous-temporal segmented image and the subsequent-temporal segmented image after window segmentation;
[0022] Perform a shifted window attention operation on the pre-temporal chunk image and the post-temporal chunk image after the self-attention mechanism operation to generate each pair of pre-temporal feature maps and corresponding post-temporal feature maps with different resolutions corresponding to each encoding stage.
[0023] In an embodiment of the present invention, the step of performing a self-attention mechanism operation within the window on the pre-temporal chunk image and the post-temporal chunk image after window segmentation includes:
[0024] Perform sampling processing on the pre-temporal chunk image and the post-temporal chunk image after window segmentation to generate sample vectors;
[0025] Based on the sample vectors, perform a linear mapping on a preset initial neighbor affinity matrix to generate a target neighbor affinity matrix, where the initial neighbor affinity matrix is expressed as φ q and φ k represent linear mapping, z represents the input vector, and z ∈ R N×d , q ∈ R N×d , k ∈ R N×d , N represents the number of input images, d represents the dimension of the vector, K represents the inner product function, and the target neighbor affinity matrix is expressed as q l ∈ R l×d , k l ∈ R l×d .
[0026] In an embodiment of the present invention, the step of performing a fusion process on the difference features between each pre-temporal feature map and the corresponding post-temporal feature map based on the decoder structure to generate a change prediction image includes:
[0027] Based on the difference module of the decoder structure, perform a difference feature extraction operation on each pair of pre-temporal feature maps and corresponding post-temporal feature maps with different resolutions to obtain multiple difference feature maps with different resolutions;
[0028] Perform a channel number conversion process on the multiple difference feature maps to unify the channel numbers of the multiple difference feature maps;
[0029] Perform a fusion process on the multiple difference feature maps with unified channel numbers to generate a fusion feature map;
[0030] Perform a two-dimensional transposed convolution operation on the fusion feature map to generate the upsampled fusion feature map;
[0031] Based on the multi-layer perceptron layer, process the upsampled fusion feature map to generate a change prediction image.
[0032] In an embodiment of the present invention, the step of updating the parameters of the change detection network model based on the loss value between the change prediction image and the corresponding change label map to establish a trained change detection network model includes:
[0033] Obtaining the label values of all pixel points of the change prediction image and the corresponding change label map;
[0034] Based on a preset loss function, calculating the loss value between the change prediction image and the corresponding change label map, where the loss function is expressed as N represents the number of all pixel points in the change prediction image, y i represents the label value of the i-th pixel point in the change label map, and p i represents the probability that the i-th pixel point in the change prediction image is predicted as a positive class;
[0035] Updating the parameters of the change detection network model based on the loss value.
[0036] In an embodiment of the present invention, after the step of updating the parameters of the change detection network model based on the loss value between the change prediction image and the corresponding change label map to establish a trained change detection network model, the following is further included:
[0037] Inputting a preset test image set into the change detection network model to output a test change map, where the test image set includes a pre-temporal test image and a post-temporal test image;
[0038] Comparing the test change map with the test label map corresponding to the test image set to obtain the intersection over union (IoU) of the change detection network model, where the intersection over union is expressed as IoU = (area i ∩area j ) / (area i ∪area j ), area i represents the area of the true change region in the test label map, area j represents the area of the predicted change region in the test change map, and IoU represents the intersection over union of the change regions of the test label map and the test change map;
[0039] Performing performance analysis on the change detection network model based on the intersection over union.
[0040] The present invention also provides a change detection system for remote sensing images, including:
[0041] A data acquisition module, configured to obtain a preset training image set and a plurality of change label maps, wherein the training image set includes a plurality of pre-temporal remote sensing images and the corresponding post-temporal remote sensing images for each of the pre-temporal remote sensing images, and each change label map is used to indicate the image change data between each pre-temporal remote sensing image and the corresponding post-temporal remote sensing image;
[0042] A data processing module, configured to perform image preprocessing on all image data in the training image set to generate an input image set;
[0043] A model establishment module, configured to establish a change detection network model, wherein the change detection network model includes an encoder structure and a decoder structure;
[0044] An encoding structure module, configured to input a certain pre-temporal remote sensing image and the corresponding post-temporal remote sensing image in the input image set into the encoder structure to output a plurality of pre-temporal feature maps with different resolutions and the corresponding post-temporal feature maps;
[0045] A decoding structure module, configured to perform fusion processing on the difference features between each pair of the pre-temporal feature maps and the corresponding post-temporal feature maps based on the decoder structure to generate a change prediction image;
[0046] A model training module, configured to update the parameters of the change detection network model based on the loss value between the change prediction image and the corresponding change label map to establish a trained change detection network model;
[0047] A data detection module, configured to input a preset pre-temporal image to be measured and the corresponding post-temporal image to be measured into the trained change detection network model to output a target change map.
[0048] The present invention further provides an electronic device, which includes:
[0049] One or more processors;
[0050] A storage device, configured to store one or more programs, and when the one or more programs are executed by the one or more processors, the electronic device implements the above-mentioned method for detecting changes in remote sensing images.
[0051] The present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor of a computer, the computer implements the above-mentioned method for detecting changes in remote sensing images.
[0052] As described above, the present invention provides a method, a system, an electronic device and a medium for detecting changes in remote sensing images, and the present invention can improve the efficiency and accuracy of detecting changes in remote sensing images. Brief Description of the Drawings
[0053] To describe the technical solutions of the embodiments of the present invention more clearly, the following briefly introduces the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0054] Figure 1 is a flowchart of the change detection method for the remotely sensed images shown in this application;
[0055] Figure 2 is a schematic diagram of the overall structure of the change detection network model of this application;
[0056] Figure 3 is the window segmentation strategy before window self-attention of this application;
[0057] Figure 4 is the window division strategy of the shifted window self-attention operation of this application;
[0058] Figure 5 are two consecutive twin nearest neighbor sliding window Transformer modules of this application;
[0059] Figure 6 is of this application Figure 1 The flowchart of step S20 in;
[0060] Figure 7 is of this application Figure 1 The flowchart of step S40 in;
[0061] Figure 8 is of this application Figure 7 The flowchart of step S45 in;
[0062] Figure 9 is of this application Figure 1 The flowchart of step S50 in;
[0063] Figure 10 is of this application Figure 1 The flowchart of step S60 in;
[0064] Figure 11 is a block diagram of the change detection system for the remotely sensed images shown in this application;
[0065] Figure 12 Shows a schematic diagram of the structure of a computer system of an electronic device suitable for implementing the embodiments of this application. Detailed Description of the Embodiments
[0066] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0067] Please refer to Figure 1-12 . It should be noted that the illustrations provided in this embodiment only schematically illustrate the basic concept of the present invention. Therefore, only the components related to the present invention are shown in the illustrations, rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and ratio of each component in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.
[0068] As Figure 1 shown, the embodiment of the present invention provides a method for change detection of remote sensing images, which can be applied to the actual detection of remote sensing images. This method can be to prepare in advance multiple pre-temporal remote sensing images representing the images before change, and multiple post-temporal remote sensing images representing the images after change, as a training image set. And, a change label map for identifying the changes between each pair of pre-temporal remote sensing images and post-temporal remote sensing images can be prepared in advance. Further, this method can establish a change detection network model for detection. This change detection network model can be a Transformer network model based on a siamese network, a nearest neighbor affinity matrix, and a sliding window method. This network model can be further expressed as a siamese nearest neighbor sliding window Transformer network model. After inputting the images in the training image set into the network model, the encoder part of the change detection network model can extract pre-temporal feature maps and corresponding post-temporal feature maps with different resolutions at each encoding stage. And the decoder structure of the change detection network model can extract the difference features between the pre-temporal feature maps and post-temporal feature maps at each stage. By fusing the difference features at each stage, a change prediction image can be output. Finally, the loss value between the change prediction image and the change label map can be calculated to update the parameters of the change detection network model. Eventually, a trained change detection network model can be obtained. The trained change detection network model can quickly and accurately perform change detection on remote sensing images.
[0069] Please refer to Figure 1 shown, the method for change detection of remote sensing images provided by the present invention may include the following steps:
[0070] Step S10: Obtain a preset training image set and multiple change label maps. Among them, the training image set includes multiple pre-temporal remote sensing images and the corresponding post-temporal remote sensing images for each of the pre-temporal remote sensing images. Each of the change label maps is used to indicate the image change data between each pre-temporal remote sensing image and the corresponding post-temporal remote sensing image.
[0071] Step S20: Perform image preprocessing on all the image data in the training image set to generate an input image set.
[0072] Step S30: Establish a change detection network model, where the change detection network model includes an encoder structure and a decoder structure.
[0073] Step S40: Input a certain pre-temporal remote sensing image and the corresponding post-temporal remote sensing image in the input image set into the encoder structure to output multiple pre-temporal feature maps and corresponding post-temporal feature maps with different resolutions.
[0074] Step S50: Based on the decoder structure, perform fusion processing on the difference features between each pair of the pre-temporal feature maps and the corresponding post-temporal feature maps to generate a change prediction image.
[0075] Step S60: Based on the loss value between the change prediction image and the corresponding change label map, update the parameters of the change detection network model to establish a trained change detection network model.
[0076] Step S70: Input a preset pre-temporal image to be measured and the corresponding post-temporal image to be measured into the trained change detection network model to output a target change map.
[0077] In an embodiment of the present invention, when performing step S10, that is, obtaining a preset training image set and multiple change label maps. Specifically, the training image set may include multiple pre-temporal remote sensing images and the corresponding post-temporal remote sensing images for each of the pre-temporal remote sensing images. Each pre-temporal remote sensing image may represent a remote sensing image before the change of the surface area, and each post-temporal remote sensing image may represent a remote sensing image after the change of the surface area. Each of the change label maps can be used to indicate the image change data between each pre-temporal remote sensing image and the corresponding post-temporal remote sensing image. For example, the training image set may contain M pre-temporal remote sensing images X = {X1, X2,..., X m ,…,X M}, and M corresponding post-temporal remote sensing images Y = {Y1, Y2,..., Y m ,…,Y M}. At the same time, M change label maps Z = {Z1, Z2,..., Zm ,…,Z M}. The change label map can be a black-and-white binary map. Specifically, the black part in the change label map can represent the background part, and the white part can represent the changed part between the pre-temporal remote sensing image and the corresponding post-temporal remote sensing image. Further, a verification image set can be preset to verify the test situation of the change detection network model during model training and adjust the hyperparameters of the change detection network model based on the test situation. Also, a test image set can be preset to test the performance of the change detection network model trained by the training image set.
[0078] Please refer to Figure 6 As shown, in an embodiment of the present invention, when performing step S20, that is, performing image preprocessing on all image data in the training image set to generate an input image set. Specifically, step S20 may include the following steps:
[0079] Step S21: Crop all image data in the training image set and the corresponding change label map;
[0080] Step S22: Perform preprocessing operations on all the cropped image data to generate an input image set, where the preprocessing operations include grayscale processing, geometric transformation processing, and image enhancement processing.
[0081] In an embodiment of the present invention, when performing step S21, that is, cropping all image data in the training image set and the corresponding change label map. Specifically, since the current mainstream public data sets include data sets of sizes 1024×1024, 512×512, 256×256, or other sizes. To meet the use of the vast majority of data sets, the input picture size used by the change detection network model can be set to 256×256. Further, the input picture size can also be set to other sizes, which are not limited herein. It should be noted that for image data with too large a size, corresponding image cropping operations need to be performed to ensure that the image data can be normally input into the network model.
[0082] In one embodiment of the present invention, when performing step S22, all the image data after cropping processing is preprocessed to generate an input image set. Specifically, the preprocessing operation may include grayscale processing, geometric transformation processing, image enhancement processing or other image preprocessing operations. Grayscale processing may refer to a method of changing the grayscale value of each pixel in the input image point by point according to a preset target condition and a certain transformation relationship. Through grayscale processing, the image quality can be improved and the display effect of the image can be made clearer. Geometric transformation processing may refer to mapping the coordinates in the original image to the new coordinate positions in the new image, which does not change the pixel values of the original image but only changes the geometric positions where the pixels are located, and is used to correct the random errors generated during image acquisition. Through image enhancement processing, the image features can be enhanced.
[0083] In one embodiment of the present invention, when performing step S30, a change detection network model is established. The change detection network model can be represented as a Siamese Nearest Neighbor Sliding Window Transformer model. Specifically, the change detection network model may include an encoder structure 10 and a decoder structure 11. Each stage of the encoder structure 10 can output a pre-temporal feature map of the pre-temporal remote sensing image and a post-temporal feature map of the corresponding post-temporal remote sensing image; the decoder structure 11 can be used to fuse the difference features between the pre-temporal feature map and the corresponding post-temporal feature map of each stage to generate a change prediction image. Further, when training this change detection network model, an Adaptive Moment Estimation AdamW optimizer can be used to update the parameters during model training. The initial learning rate of this optimizer can be set to 6×10 -5 Or other learning rate values, which are not limited here. Further, the weight decay of the optimizer can be set to 0.01 or other values, which are not limited here. At the same time, the default settings can be adopted for random horizontal flipping, random rescaling within the ratio range [0.5, 2.0], and random photometric distortion. The change detection network model can adopt a random depth with a ratio of 0.2.
[0084] In one embodiment of the present invention, when performing step S40, a certain pre-temporal remote sensing image 21 in the input image set and the corresponding post-temporal remote sensing image 22 are input into the encoder structure 10 to output multiple pre-temporal feature maps and corresponding post-temporal feature maps with different resolutions. Specifically, this encoder structure 10 can be used to encode the pre-temporal remote sensing image 21 and the corresponding post-temporal remote sensing image 22, and obtain pre-temporal feature maps with different sizes at multiple resolutions corresponding to the pre-temporal remote sensing image 21, and post-temporal feature maps with different sizes at multiple resolutions corresponding to the post-temporal remote sensing image 22. In such feature maps, there are not only rough features with high resolution but also fine-grained features with low resolution. It should be noted that the initial size of the input image can be H×W×3, and the image size output by the encoder of this change detection network model at each stage can be calculated based on the formula for calculation. Specifically, the value of i can be i∈{1,2,3,4}, and C i+1 >C i . Four feature maps can be obtained during the encoding process. The difference features of each pair of pre-temporal feature maps and post-temporal feature maps at each stage can be obtained through the difference module. Finally, the difference features at each stage are processed by a multi-layer perceptron, feature fusion, etc. in the decoder structure 11 to output a change prediction map.
[0085] Please refer to Figure 7 As shown, in one embodiment of the present invention, when performing step S40, a certain pre-temporal remote sensing image in the input image set and the corresponding post-temporal remote sensing image are input into the encoder structure 10 to output multiple pre-temporal feature maps and corresponding post-temporal feature maps with different resolutions. Specifically, step S40 may include the following steps:
[0086] Step S41: Based on the encoder structure, perform block segmentation processing on a certain pre-temporal remote sensing image and the corresponding post-temporal remote sensing image in the input image set to generate a pre-temporal segmented image and a post-temporal segmented image;
[0087] Step S42: Perform linear mapping on the pre-temporal segmented image and the post-temporal segmented image to adjust the image dimensions of the pre-temporal segmented image and the post-temporal segmented image;
[0088] Step S43: At each encoding stage of the encoder structure, perform downsampling operations on the pre-temporal segmented image and the post-temporal segmented image after dimension adjustment;
[0089] Step S44: Perform window segmentation operations on the pre-temporal segmented image and the post-temporal segmented image after downsampling;
[0090] Step S45: Perform self-attention mechanism operations within the window on the pre-temporal segmented images and the post-temporal segmented images after window segmentation;
[0091] Step S46: Perform shifted window attention operations on the pre-temporal segmented images and the post-temporal segmented images after the self-attention mechanism operations to generate each pair of pre-temporal feature maps and corresponding post-temporal feature maps with different resolutions corresponding to each encoding stage.
[0092] In an embodiment of the present invention, when executing step S41, that is, based on the encoder structure 10, perform block segmentation processing on a certain pre-temporal remote sensing image 21 and the corresponding post-temporal remote sensing image 22 in the input image set to generate pre-temporal segmented images and post-temporal segmented images. Specifically, after the preprocessing operation on the training image set, the image can be transformed into a size of H×W×3. First, a block segmentation operation can be performed on a certain pre-temporal remote sensing image 21 and the corresponding post-temporal remote sensing image 22 after preprocessing. Specifically, the block segmentation operation can be to divide the input picture with a size of H×W×3 into small blocks of (4, 4), and the size of the picture after block segmentation can be It should be noted that in order to achieve non-overlapping convolution according to the size of a block to divide the image into multiple non-overlapping blocks, it is necessary to ensure that the size of the convolution kernel is equal to the size of the block, and the stride of the convolution is equal to the size of the block.
[0093] In an embodiment of the present invention, when executing step S42, that is, perform linear mapping on the pre-temporal segmented images and the post-temporal segmented images to adjust the image dimensions of the pre-temporal segmented images and the post-temporal segmented images. Specifically, through linear mapping, the pre-temporal segmented images and the post-temporal segmented images can be mapped to 96 dimensions, so that the image becomes Dimensions. During this process, the size of the input image can be (256, 256, 3), the size of the block can be (4, 4). After the image is further divided into blocks, it can be obtained as (64, 64, 48), and then through linear mapping, it can be obtained as (64, 64, 96). At this time, the flattening operation provided by the tensor can be performed on the image data to flatten the tensor into (4096, 96). The tensor at this time is extremely similar to the sentence in natural language processing. Among them, 96 can be equivalent to the dimension of a word, and 4096 can be equivalent to the number of words. Through the above operations, the image of (256, 256, 3) can be converted into a tensor of (4096, 96) that can be further input into the encoder. Since the Transformer network model is generally used in natural language processing, its processing of text input is to convert each word in the sentence into a token, that is, a sentence consists of multiple tokens. And the above operations can convert the two-dimensional image into a one-dimensional vector, so that it can be further input into the encoder of the Transformer network model.
[0094] In an embodiment of the present invention, when performing step S43, that is, in each encoding stage of the encoder structure 10, a downsampling operation is performed on the pre-temporal block image and the post-temporal block image after dimension adjustment. Specifically, after processing the image into blocks one by one, the next step is to perform a block fusion operation on the image. This operation is used to downsample the image by a factor of two at the end of each stage. While reducing the image resolution, it can also adjust the number of channels, forming a hierarchical operation design, and at the same time saving the amount of computation. Specifically, when performing the downsampling operation by a factor of two, in each row and column direction of the image, elements can be selected at an interval of 1 unit, and then spliced together as an entire tensor, and then unfolded. At this time, since the number of rows and columns of the image is reduced by a factor of two each, the channel dimension of the image can become 4 times the original. Further, the channel dimension is adjusted to 2 times the original through a fully connected layer. In this way, the downsampling operation on the original image is achieved.
[0095] Specifically, for example, in the above downsampling process, for a tensor of size 4*4*1, if elements are selected at an interval of one unit in both the row and column directions, 4 tensors of 2*2*1 can be obtained. Then these four tensors are spliced into a tensor (2*2*4) according to the channel dimension, and then the channel dimension is converted to 2 through a fully connected layer. Finally, the size of the obtained tensor is (2*2*2). Therefore, the transformation from size (H, W, C) to size (H / 2, W / 2, 2*C) is achieved.
[0096] In one embodiment of the present invention, when performing step S44, a window segmentation operation is performed on the downsampled pre-temporal block image and the post-temporal block image. Specifically, after performing staged downsampling, window segmentation and window restoration operations can be performed to implement window self-attention operations. Specifically, before performing the window self-attention operation, a window segmentation operation 30 can be performed, as Figure 3 shown, to divide multiple blocks into one window. After performing the self-attention operation within the window, a window restoration operation is performed to restore the picture to its normal block form. It should be further explained that the tensor size of a normal picture is (B, H, W, C), where B is the batch size, that is, the number of pictures obtained at one time during training. Dividing the window can mean dividing the original picture into multiple windows of window_size, which can be reflected in the tensor size by dividing the tensor of (B, H, W, C) into (num_windows*B, window_size, window_size, C). Restoring the window can mean performing the inverse operation of dividing the window, that is, restoring multiple windows to an entire picture, which can be reflected in the tensor size by dividing the tensor of (num_windows*B, window_size, window_size, C) into (B, H, W, C).
[0097] It should be further explained that for traditional Transformer network models, for example, the attention mechanism in the ViT network model performs global self-attention on the entire original picture, and this attention operation requires a large amount of computation. The change detection network model proposed by the present invention divides the image into small windows one by one before performing the attention mechanism. Each time the calculation only needs to perform the attention mechanism within the window, reducing the amount of calculation, improving the training and detection speed of the model, reducing the training cost of the model, and enabling it to be better applied to actual detection. It is worth noting that since the change detection network model of the present invention only performs self-attention in small windows each time, the receptive field of feature extraction will become smaller. Therefore, it is necessary to expand the receptive field of feature extraction in the subsequent process.
[0098] Please refer to Figure 8 shown. In one embodiment of the present invention, when performing step S45, a self-attention mechanism operation within the window is performed on the pre-temporal block image and the post-temporal block image after window segmentation. Specifically, step S45 may include the following steps:
[0099] Step S451: Perform sampling processing on the pre-temporal block image and the post-temporal block image after window segmentation to generate sample vectors.
[0100] Step S452: Based on the sample vectors, perform a linear mapping on the preset initial neighbor affinity matrix to generate a target neighbor affinity matrix.
[0101] In an embodiment of the present invention, when performing step S451 and step S452, specifically, after performing the window segmentation operation on the image, the self-attention mechanism operation within the window can be continued. The self-attention mechanism used in the present invention is different from the traditional attention mechanism. This self-attention mechanism operation does not need to calculate the similarity between high-dimensional vectors, but obtains the target neighbor affinity matrix through a more simple and efficient method. The main method of this attention mechanism is to map the high-dimensional representation vector z of the image to a low-dimensional coding space.
[0102] Regarding the self-attention mechanism of the present invention, it should be noted first that the initial neighbor affinity matrix can be expressed as where φ q (·) and φ k (·) can represent two linear mappings, through which the input z ∈ R N×d can be mapped into q, k ∈ R N×d , N can represent the number of input images, and d can represent the dimension of the vector. K(·, ·) can represent a typical inner product function. For the method of transforming the initial neighbor affinity matrix into the target neighbor affinity matrix, specifically, after vectorizing the image, the input of the attention mechanism can be expressed as z ∈ R N×d . Then randomly sample l samples from the input, for example, z l ∈ R l×d . Then use three different linear mappings to map the input z into q, k ∈ R N×d , and at the same time use the same three mappings to map z l into q l and k l matrices, where q l , k l ∈ R l×d . Then use q l and k l matrices to map the original q and k to the l-dim space, that is where can represent the similarity between the characterization vector l ∈ {1,..., N} and the boundary motivator j ∈ {1,..., l}. After the above derivation, the initial neighbor affinity matrix can be transformed into the target neighbor affinity matrix. The target neighbor affinity matrix can be expressed as Based on the above process, the algorithm complexity of obtaining the affinity matrix is significantly reduced from O(N 2 d) to O(N 2(l), and during the testing process, the value of d is much smaller than the value of l, which greatly reduces the complexity and computational amount of the algorithm, improves the training and detection speed of the model, reduces the cost of model training, and improves the practicability of the model.
[0103] In an embodiment of the present invention, when performing step S46, that is, performing a shifted window attention operation on the pre-temporal block image and the post-temporal block image after the self-attention mechanism operation to generate each pair of pre-temporal feature maps and corresponding post-temporal feature maps with different resolutions corresponding to each encoding stage. As Figure 5 shown, they are two consecutive twin nearest neighbor sliding window Transformer modules. Specifically, after performing the self-attention operation within each window, the computational amount of the model is greatly reduced. However, since there is no interaction between the information of each window and its surrounding windows, the receptive field is greatly reduced. Therefore, it is necessary to expand the receptive field of feature extraction. Specifically, expanding the receptive field of feature extraction can be as Figure 4 shown to perform the shifted window attention operation 31. This operation can first transform the originally divided window, move it half the window size towards the upper left corner of the image, and fill the extra part in the upper left corner to the lower right corner, and then perform self-attention according to the current window. At this time, the content included in the window is to combine the information of the previous window and its adjacent window, so that the current window and the surrounding windows can perform information interaction, thereby expanding the receptive field of feature extraction. This method reduces the computational amount by performing self-attention within the window, and at the same time, the shifted window attention operation can obtain a larger receptive field and improve the detection accuracy.
[0104] In an embodiment of the present invention, when performing step S50, that is, based on the decoder structure 11, fusing the difference features between each pre-temporal feature map and the corresponding post-temporal feature map to generate a change prediction image 23. Specifically, the decoder structure 11 can predict the change map by aggregating the difference features of each pair of pre-temporal feature maps and post-temporal feature maps in each stage. The decoder part extracts the difference features between the pre-temporal feature map and the post-temporal feature map in each stage. The difference feature maps after the difference extraction in the first three stages of the decoder are linearly interpolated and upsampled by a factor of two, and the upsampled difference feature maps are fused with the difference feature maps of the next stage, so as to fuse the difference features of high resolution and low resolution and obtain finer-grained features of the image. The decoder structure 11 can include a multi-layer perceptron and upsampling in the first part, splicing and fusion in the second part, and splicing, classification, and peer-to-peer nearest neighbor normalization exponential function in the third part.
[0105] Please refer to Figure 9As shown, in one embodiment of the present invention, when performing step S50, that is, based on the decoder structure 11, the difference features between each of the pre-temporal feature maps and the corresponding post-temporal feature maps are fused to generate a change prediction image 23. Specifically, step S50 may include the following steps:
[0106] Step S51: Based on the difference module of the decoder structure, perform a difference feature extraction operation on each pair of pre-temporal feature maps and corresponding post-temporal feature maps with different resolutions to obtain a plurality of difference feature maps with different resolutions;
[0107] Step S52: Perform a channel number conversion process on the plurality of difference feature maps to unify the channel numbers of the plurality of difference feature maps;
[0108] Step S53: Perform a fusion process on the plurality of difference feature maps with unified channel numbers to generate a fusion feature map;
[0109] Step S54: Perform a two-dimensional transposed convolution operation on the fusion feature map to generate the upsampled fusion feature map;
[0110] Step S55: Based on the multi-layer perceptron layer, process the upsampled fusion feature map to generate a change prediction image.
[0111] In one embodiment of the present invention, when performing step S51, that is, based on the difference module 12 of the decoder structure 11, a difference feature extraction operation is performed on each pair of pre-temporal feature maps and corresponding post-temporal feature maps with different resolutions to obtain a plurality of difference feature maps with different resolutions. Specifically, first, in the four stages of the encoder part, the difference features of the pre-temporal feature map and the post-temporal feature map of each stage can be extracted through the difference module 12 to obtain difference features with different resolutions and different sizes. This difference feature extraction process can be based on a siamese network structure. The pre-change image passes through one network, and the post-change image passes through another network, and then the difference features of each stage are extracted through the difference module 12.
[0112] It should be noted that the difference module 12 may include two-dimensional convolution (Conv2D), rectified linear unit (ReLU), and batch normalization (BN). Specifically, the difference module 12 can be expressed as where, and can represent the i-th layer of the pre-temporal feature map and the corresponding post-temporal feature map, and cat can represent tensor concatenation. This difference module 12 does not simply calculate and the absolute difference value, but learns the optimal distance metric for each scale during training to obtain better change detection effects.
[0113] In one embodiment of the present invention, when step S52 is executed, that is, channel number conversion processing is performed on the multiple difference feature maps to unify the channel numbers of the multiple difference feature maps. Specifically, after difference extraction is performed on the pre-temporal feature maps and the corresponding post-temporal feature maps at each stage, multiple difference feature maps with different resolutions can be obtained. Then, the channel numbers of all the difference feature maps can be converted into a unified channel number word embedding dimension through a linear layer. For example, the word embedding dimension can be set to 256, which is the input and output size of the image. Finally, each dimension is upsampled to a size of H / 4×W / 4. The specific process can be expressed as, where C ebd can represent the embedding dimension, that is, the above-mentioned word embedding dimension.
[0114] In one embodiment of the present invention, when step S53 is executed, that is, fusion processing is performed on the multiple difference feature maps with unified channel numbers to generate a fused feature map. Specifically, after the channel number unification operation, tensors with the same channel number corresponding to four different scales can be obtained. Then, through concatenation, a tensor of four times C ebd can be obtained. Since this tensor is obtained by fusing four difference maps with different scales, it fuses the rough features of high resolution and the fine-grained features of low resolution. Finally, through a multi-layer perceptron layer, the tensor of four times C ebd can be converted into a tensor of one time C ebd that is, the 256 size of the final output image. The specific process can be expressed as,
[0115] In one embodiment of the present invention, when steps S54 and S55 are executed, specifically, in the final upsampling process, using a two-dimensional transposed convolution with S = 4 and K = 3, the fused feature map can be upsampled to a size of H×W. Finally, the upsampled fused feature map is processed through a multi-layer perceptron layer, and a change mask image with a resolution of H×W×n cls can be predicted, that is, the change prediction image 23. Where n cls can be 2, which represents the number of categories in the image. In the change detection process of the present invention, n cls = 2 can represent the two categories of change and no change. The specific process can be expressed as, where ConvTranspose2D represents the transposed convolution. It should be noted that by using the simple multi-layer perceptron layer in the above decoder part to replace the original convolutional network, change detection can be quickly completed, thereby reducing the complexity of the model and improving the detection efficiency of the model.
[0116] For the equivalent neighbor normalized exponential function (Softmax) in the decoder structure 11, it should be noted that the original normalized exponential function (Softmax) aggregates all samples, but the significant presence of irrelevant samples will have a negative impact on the final calculation. In addition, in addition to the negative impact on the final output representation, the computational complexity of the representation aggregation is O(N 2 d) Since the input size N is large, the computational burden is also large. Therefore, the peer-neighbor normalized exponential function Softmax (RNS) can be used to force the sparsification of the few relevant attention weights with the peer-neighbor mask. For example, if two images are neighbors of each other in the feature space, they are likely to be related. To this end, the neighbor affinity matrix Compute a top-k neighbor mask M k , by focusing on the first k affinity values of each row, and setting the first k largest attention weights s of each row in the A matrix to 1, and the rest to 0. The formula can be expressed as, The neighbor mask M can be calculated by this formula. ij =M k °M kT For each element M ij , if i and j are each other's first k neighbors, the value will be set to 1, otherwise it is 0. By adding this mask M to the conventional normalized exponential function (Softmax), sparse attention that only occurs in the neighbors is achieved to increase attention to more relevant images. The calculation formula of the peer neighbor normalized exponential function Softmax (RNS) can be expressed as Since most of the attention values are set to zero, the aggregation in the calculation formula is more concentrated and more robust. Since there is no need to add the representation of zero weights, the time complexity of feature aggregation is reduced from O(N 2 d) is reduced to O(Nkd).
[0117] In one embodiment of the present invention, for step S60, based on the loss value between the change prediction image 23 and the corresponding change label map, the change detection network model is parameter updated to establish a trained change detection network model. Specifically, after the previous tense image to be tested and the corresponding post-tense image to be tested are input into the change detection network model, loss data can be obtained. The loss data can be used to update the parameters of the change detection network model through the gradient back propagation algorithm to complete the training.
[0118] See also Figure 10As shown, in an embodiment of the present invention, when step S60 is executed, that is, based on the loss value between the change prediction image and the corresponding change label map, the parameters of the change detection network model are updated to establish a trained change detection network model. Specifically, step S60 may include the following steps:
[0119] Step S61, obtain the label values of all pixel points of the change prediction image and the corresponding change label map;
[0120] Step S62, calculate the loss value between the change prediction image and the corresponding change label map based on a preset loss function;
[0121] Step S63, update the parameters of the change detection network model based on the loss value.
[0122] In an embodiment of the present invention, when steps S61 and S62 are executed, specifically, the loss function can be expressed as where N can represent the number of all pixel points in the change prediction image, y i can represent the label value of the i-th pixel point in the change label map, and p i can represent the probability that the i-th pixel point in the change prediction image is predicted as the positive class. Through the above loss function, the loss value between the change prediction image and the corresponding change label map can be calculated.
[0123] In an embodiment of the present invention, when step S63 is executed, that is, based on the loss value, the parameters of the change detection network model are updated. Specifically, the gradient backpropagation algorithm can be used to update the parameters of the change detection network model with the loss value to complete the training.
[0124] In an embodiment of the present invention, after step S60, that is, after the step of updating the parameters of the change detection network model based on the loss value between the change prediction image and the corresponding change label map to establish a trained change detection network model, the following steps may further be included:
[0125] Step S64, input a preset test image set into the change detection network model to output a test change map, where the test image set includes a pre-temporal test image and a post-temporal test image;
[0126] Step S65, compare the test change map with the test label map corresponding to the test image set to obtain the intersection over union (IoU) of the change detection network model, where the intersection over union is expressed as IoU = (area i ∩area j ) / (areai ∪ area j ), area i represents the area of the true change region in the test label map, area j represents the area of the predicted change region in the test change map, and IoU is the intersection over union of the change regions of the test label map and the test change map;
[0127] Step S66, perform performance analysis on the change detection network model based on the intersection over union.
[0128] In an embodiment of the present invention, when performing step S64, that is, inputting a preset test image set into the change detection network model to output a test change map, where the test image set includes a pre-temporal test image and a post-temporal test image. Specifically, a preset test image set can be used for testing. For example, Q can be used to test the pre-temporal test image and the corresponding post-temporal test image. Q pre-temporal test images can be represented as Q post-temporal test images can be represented as where, can represent the q-th pre-temporal test image, can represent the q-th post-temporal test image.
[0129] In an embodiment of the present invention, when performing step S65 and step S66, specifically, the test image can be input into the trained change detection network model. During testing, the image propagates forward in the network model. The network model extracts features and differences of the image according to the parameters obtained from the previous training, and then fuses the difference feature maps of multiple stages to obtain a test change map. Finally, the test change map predicted by the model is compared with the corresponding test label map, and the quality of the model is judged according to the comparison result.
[0130] Specifically, the above comparison process can analyze the accuracy of the change detection network model based on the intersection over union. The calculation process of the intersection over union can be expressed as IoU = (area i ∩ area j ) / (area i ∪ area k ). Where, area i can represent the area of the true change region in the test label map, area j can represent the area of the predicted change region in the test change map, and IoU can represent the intersection over union of the change regions of the test label map and the test change map.
[0131] In one embodiment of the present invention, for step S70, the preset pre-temporal image to be measured and the corresponding post-temporal image to be measured are input into the trained change detection network model to output a target change map. Specifically, after the change detection network model finishes detection, the preset pre-temporal image to be measured and the corresponding post-temporal image to be measured can be input into the trained change detection network model. During detection, the image undergoes forward propagation in the network model. The change detection network model extracts features and differences of the picture according to the parameters obtained from previous training, and then fuses the difference feature maps of multiple stages to obtain the final target change map. Thus, fast and accurate change detection can be achieved.
[0132] It can be seen that in the above solution, the Transformer network model is applied to the change detection task, and the full-image self-attention is changed to window-based self-attention, greatly reducing the computational amount of the model, thereby improving the training and detection speed of the model and making it better applicable to actual use. The network model of the present invention does not require a large amount of data for training, reducing the training cost. And the present invention uses a sliding window, performs self-attention within the window, and at the same time transforms and moves the window, and then calculates self-attention. This sliding window method enables good interaction between the current window and its surrounding windows, thereby obtaining a larger receptive field during feature extraction, greatly improving the detection accuracy while reducing the computational amount. At the same time, the attention mechanism used in the present invention obtains a neighbor affinity matrix in a more efficient manner, maps the high-dimensional representation vector z to a low-dimensional coding space, reduces the complexity of attention calculation, and improves the running efficiency of the model. Further, in the model architecture of the present invention, through a hierarchical structure of multiple stages, feature maps of multiple different sizes are obtained, realizing the flexibility of modeling. In the decoder part of the network model of the present invention, a multi-layer perceptron is used to replace the convolutional network, greatly reducing the model complexity while improving the model efficiency.
[0133] Please refer to Figure 11 As shown, the present invention also provides a change detection system for remote sensing images, which corresponds one-to-one with the change detection method in the above embodiment. The change detection system may include a data acquisition module 101, a data processing module 102, a model establishment module 103, an encoding structure module 104, a decoding structure module 105, a model training module 106, and a data detection module 107.
[0134] In one embodiment of the present invention, the data acquisition module 101 can be used to obtain a preset training image set and a plurality of change label maps, wherein the training image set includes a plurality of pre-temporal remote sensing images and the corresponding post-temporal remote sensing images of each of the pre-temporal remote sensing images, and each of the change label maps is used to indicate the image change data between each of the pre-temporal remote sensing images and the corresponding post-temporal remote sensing images;
[0135] In one embodiment of the present invention, the data processing module 102 can be used to perform image preprocessing on all the image data in the training image set to generate an input image set. Specifically, the data processing module 20 can specifically be used to perform cropping processing on all the image data in the training image set and the corresponding change label maps; perform preprocessing operations on all the cropped image data to generate an input image set, wherein the preprocessing operations include grayscale processing, geometric transformation processing, and image enhancement processing.
[0136] In one embodiment of the present invention, the model establishment module 103 can be used to establish a change detection network model, wherein the change detection network model includes an encoder structure and a decoder structure.
[0137] In one embodiment of the present invention, the encoding structure module 104 can be used to input a certain pre-temporal remote sensing image in the input image set and the corresponding post-temporal remote sensing image into the encoder structure to output a plurality of pre-temporal feature maps and corresponding post-temporal feature maps with different resolutions. Specifically, the encoding structure module can specifically be used to perform block segmentation processing on a certain pre-temporal remote sensing image in the input image set and the corresponding post-temporal remote sensing image based on the encoder structure to generate a pre-temporal segmented image and a post-temporal segmented image; perform a linear mapping on the pre-temporal segmented image and the post-temporal segmented image to adjust the image dimensions of the pre-temporal segmented image and the post-temporal segmented image; perform a downsampling operation on the dimension-adjusted pre-temporal segmented image and the post-temporal segmented image at each encoding stage of the encoder structure; perform a window segmentation operation on the downsampled pre-temporal segmented image and the post-temporal segmented image; perform a self-attention mechanism operation within the window on the window-segmented pre-temporal segmented image and the post-temporal segmented image; perform a shifted window attention operation on the pre-temporal segmented image and the post-temporal segmented image after the self-attention mechanism operation to generate each pair of pre-temporal feature maps and corresponding post-temporal feature maps with different resolutions corresponding to each encoding stage.
[0138] In one embodiment of the invention, the encoding structure module 104 may also be specifically configured to perform sampling processing on the pre-temporal segmented images and the post-temporal segmented images after window segmentation to generate sample vectors; and perform a linear mapping on a preset initial neighbor affinity matrix based on the sample vectors to generate a target neighbor affinity matrix, where the initial neighbor affinity matrix is expressed as φ q and φ k represents a linear mapping, z represents an input vector, and z ∈ R N×d ,q ∈ R N×d ,k ∈ R N×d ,N represents the number of input images, d represents the dimension of the vector, K represents an inner product function, and the target neighbor affinity matrix is expressed as q l ∈ R l×d ,k l ∈ R l×d 。
[0139] In one embodiment of the present invention, the decoding structure module 105 may be configured to perform a fusion process on the difference features between each pre-temporal feature map and the corresponding post-temporal feature map based on the decoder structure to generate a change prediction image; specifically, the decoding structure module 50 may be specifically configured to perform a difference feature extraction operation on each pair of pre-temporal feature maps and the corresponding post-temporal feature maps with different resolutions based on the difference module of the decoder structure to obtain a plurality of difference feature maps with different resolutions; perform a channel number conversion process on the plurality of difference feature maps to unify the channel numbers of the plurality of difference feature maps; perform a fusion process on the plurality of difference feature maps with unified channel numbers to generate a fusion feature map; perform a two-dimensional transposed convolution operation on the fusion feature map to generate the upsampled fusion feature map; and perform a process on the upsampled fusion feature map based on a multi-layer perceptron layer to generate a change prediction image.
[0140] In one embodiment of the present invention, the model training module 106 may be configured to update the parameters of the change detection network model based on the loss value between the change prediction image and the corresponding change label map to establish a trained change detection network model; specifically, the model training module 60 may be specifically configured to obtain the label values of all pixel points of the change prediction image and the corresponding change label map; calculate the loss value between the change prediction image and the corresponding change label map based on a preset loss function, where the loss function is expressed as N represents the number of all pixel points in the change prediction image, y i represents the label value of the i-th pixel point in the change label map, p iRepresents the probability that the i-th pixel in the predicted change image is predicted as the positive class; based on the loss value, update the parameters of the change detection network model.
[0141] In an embodiment of the present invention, the model training module 106 may further be specifically configured to input a preset test image set into the change detection network model to output a test change map, where the test image set includes a pre-temporal test image and a post-temporal test image; compare the test change map with the test label map corresponding to the test image set to obtain the intersection over union (IoU) of the change detection network model, where the IoU data is expressed as IoU = (area i ∩area j ) / (area i ∪area j ), area i represents the area of the true change region in the test label map, area j represents the area of the predicted change region in the test change map, and IoU represents the intersection over union of the change regions of the test label map and the test change map; perform performance analysis on the change detection network model based on the IoU.
[0142] In an embodiment of the present invention, the data detection module 107 may be configured to input a preset pre-temporal image to be measured and the corresponding post-temporal image to be measured into the trained change detection network model to output a target change map.
[0143] It should be noted that the remote sensing image change detection system provided in the above embodiments belongs to the same concept as the remote sensing image change detection method provided in the above embodiments. The specific ways in which each module and unit perform operations have been described in detail in the method embodiments and will not be repeated here. In practical applications, the remote sensing image change detection system provided in the above embodiments can allocate the above functions to different functional modules as needed, that is, divide the internal structure of the device into different functional modules to complete all or part of the functions described above, and this is not limited here either.
[0144] An embodiment of the present application further provides an electronic device, including: one or more processors; a storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, the electronic device implements the remote sensing image change detection method provided in each of the above embodiments.
[0145] Figure 12 Shows a schematic structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application. It should be noted that Figure 12The computer system of the electronic device shown is only an example and should not impose any limitations on the functions and scope of use of the embodiments of the present application.
[0146] As Figure 12 shown, the computer system includes a central processing unit (CPU) 1201, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1202 or the program loaded from the storage section 1208 into the random access memory (RAM) 1203, such as executing the method described in the above embodiments. In the RAM 1203, various programs and data required for system operation are also stored. The CPU 1201, ROM 1202, and RAM 1203 are connected to each other via a bus 1204. An input / output (I / O) interface 1205 is also connected to the bus 1204.
[0147] The following components are connected to the I / O interface 1205: an input section 1206 including a keyboard, a mouse, etc.; an output section 1207 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 1208 including a hard disk, etc.; and a communication section 1209 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 1209 performs communication processing via a network such as the Internet. A drive 1210 is also connected to the I / O interface 1205 as needed. A removable medium 1211, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1210 as needed so that a computer program read from it can be installed into the storage section 1208 as needed.
[0148] Specifically, according to the embodiments of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments of the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication section 1209, and / or installed from the removable medium 1211. When the computer program is executed by the central processing unit (CPU) 1201, various functions defined in the system of the present application are executed.
[0149] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. Among them, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0150] Another aspect of the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor of a computer, the computer is caused to execute the change detection method of the remote sensing image as described above. The computer-readable storage medium may be included in the electronic device described in the above embodiments, or may exist separately and not be assembled into the electronic device.
[0151] The above embodiments merely exemplarily illustrate the principles and effects of the present invention, rather than limiting the present invention. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or changes completed by those with ordinary knowledge in the technical field without departing from the spirit and technical idea disclosed by the present invention should still be covered by the claims of the present invention.
Claims
1. A method for change detection of remote sensing images, characterized in that, Including: Obtain a preset training image set and multiple change label maps, where the training image set includes multiple pre-temporal remote sensing images and the corresponding post-temporal remote sensing images for each of the pre-temporal remote sensing images, and each change label map is used to indicate the image change data between each pre-temporal remote sensing image and the corresponding post-temporal remote sensing image; Perform image preprocessing on all the image data in the training image set to generate an input image set; Establish a change detection network model, where the change detection network model includes an encoder structure and a decoder structure; Based on the encoder structure, perform block segmentation on a pre-temporal remote sensing image and the corresponding post-temporal remote sensing image in the input image set to generate a pre-temporal segmented image and a post-temporal segmented image; Perform a linear mapping on the pre-temporal segmented image and the post-temporal segmented image to adjust the image dimensions of the pre-temporal segmented image and the post-temporal segmented image; At each encoding stage of the encoder structure, perform a downsampling operation on the pre-temporal segmented image and the post-temporal segmented image after dimension adjustment; Perform a window segmentation operation on the pre-temporal segmented image and the post-temporal segmented image after downsampling; Perform a self-attention mechanism operation within the window on the pre-temporal segmented image and the post-temporal segmented image after window segmentation; Perform a shifted window attention operation on the pre-temporal segmented image and the post-temporal segmented image after the self-attention mechanism operation to generate each pair of pre-temporal feature maps and corresponding post-temporal feature maps with different resolutions at each encoding stage; Based on the difference module of the decoder structure, perform a difference feature extraction operation on each pair of pre-temporal feature maps and corresponding post-temporal feature maps with different resolutions to obtain multiple difference feature maps with different resolutions; Perform a channel number conversion process on the multiple difference feature maps to unify the channel numbers of the multiple difference feature maps; Perform a fusion process on the multiple difference feature maps with unified channel numbers to generate a fusion feature map; Perform a two-dimensional transposed convolution operation on the fusion feature map to generate an upsampled fusion feature map; Based on a multi-layer perceptron layer, process the upsampled fusion feature map to generate a change prediction image; Based on the loss value between the change prediction image and the corresponding change label map, update the parameters of the change detection network model to establish a trained change detection network model; Input a preset pre-temporal image to be measured and the corresponding post-temporal image to be measured into the trained change detection network model to output a target change map.
2. The change detection method for remote sensing images according to claim 1, wherein The step of performing image preprocessing on all the image data in the training image set to generate an input image set includes: Perform a cropping process on all the image data in the training image set and the corresponding change label maps; Perform a preprocessing operation on all the image data after the cropping process to generate an input image set, where the preprocessing operation includes grayscale processing, geometric transformation processing, and image enhancement processing.
3. The change detection method for remote sensing images according to claim 1, wherein, The steps of performing self-attention mechanism operations within the window on the pre-temporal segmented images and the post-temporal segmented images after window segmentation include: Performing sampling processing on the pre-temporal segmented images and the post-temporal segmented images after window segmentation to generate sample vectors; Based on the sample vectors, perform a linear mapping on a preset initial neighbor affinity matrix to generate a target neighbor affinity matrix, where the initial neighbor affinity matrix is expressed as , and represent linear mapping, represents the input vector, and , , , N represents the number of input images, d represents the dimension of the vector, K represents the inner product function, and the target neighbor affinity matrix is expressed as q l ∈R l×d , .
4. The change detection method for remote sensing images according to claim 1, characterized in that The steps of updating the parameters of the change detection network model based on the loss value between the change prediction image and the corresponding change label map to establish a trained change detection network model include: Obtaining the label values of all pixel points of the change prediction image and the corresponding change label map; Calculate the loss value between the change prediction image and the corresponding change label map based on a preset loss function, where the loss function is expressed as , represents the number of all pixel points in the change prediction image, represents the label value of the i-th pixel point in the change label map, represents the probability that the i-th pixel point in the change prediction image is predicted as the positive class; Updating the parameters of the change detection network model based on the loss value.
5. The change detection method for remote sensing images according to claim 1, characterized in that, After the steps of updating the parameters of the change detection network model based on the loss value between the change prediction image and the corresponding change label map to establish a trained change detection network model, it further includes: Inputting a preset test image set into the change detection network model to output a test change map, where the test image set includes pre-temporal test images and post-temporal test images; Compare the test change map with the test label map corresponding to the test image set to obtain the intersection over union (IoU) of the change detection network model, where the intersection over union (IoU) is expressed as , represents the area of the true change region in the test label map, represents the area of the predicted change region in the test change map, and IoU represents the intersection over union of the change regions between the test label map and the test change map; Performing performance analysis on the change detection network model based on the area intersection over union.
6. A change detection system for remote sensing images, characterized in that, Including: A data acquisition module for obtaining a preset training image set and multiple change label maps, where the training image set includes multiple pre-temporal remote sensing images and the post-temporal remote sensing images corresponding to each pre-temporal remote sensing image, and each change label map is used to indicate the image change data between each pre-temporal remote sensing image and the corresponding post-temporal remote sensing image; A data processing module for performing image preprocessing on all image data in the training image set to generate an input image set; A model establishment module for establishing a change detection network model, where the change detection network model includes an encoder structure and a decoder structure; An encoding structure module for, based on the encoder structure, performing block segmentation processing on a certain pre-temporal remote sensing image and the corresponding post-temporal remote sensing image in the input image set to generate pre-temporal segmented images and post-temporal segmented images; performing linear mapping on the pre-temporal segmented images and the post-temporal segmented images to adjust the image dimensions of the pre-temporal segmented images and the post-temporal segmented images; performing downsampling operations on the pre-temporal segmented images and the post-temporal segmented images with adjusted dimensions at each encoding stage of the encoder structure; performing window segmentation operations on the downsampled pre-temporal segmented images and the post-temporal segmented images; performing self-attention mechanism operations within the window on the pre-temporal segmented images and the post-temporal segmented images after window segmentation; performing shifted window attention operations on the pre-temporal segmented images and the post-temporal segmented images after self-attention mechanism operations to generate each pair of pre-temporal feature maps and corresponding post-temporal feature maps with different resolutions corresponding to each encoding stage; A decoding structure module is used to perform differential feature extraction operations on each pair of pre-temporal feature maps and corresponding post-temporal feature maps with different resolutions based on the differential module of the decoder structure to obtain multiple differential feature maps with different resolutions; perform channel number conversion processing on the multiple differential feature maps to unify the channel numbers of the multiple differential feature maps; perform fusion processing on the multiple differential feature maps with unified channel numbers to generate a fused feature map; perform a two-dimensional transposed convolution operation on the fused feature map to generate the upsampled fused feature map; and process the upsampled fused feature map based on a multi-layer perceptron layer to generate a change prediction image. A model training module is used to update the parameters of the change detection network model based on the loss value between the change prediction image and the corresponding change label map to establish a trained change detection network model. A data detection module is used to input a preset pre-temporal image to be measured and the corresponding post-temporal image to be measured into the trained change detection network model to output a target change map.
7. An electronic device, characterized in that, The electronic device includes: One or more processors; A storage device for storing one or more programs, which when executed by the one or more processors, cause the electronic device to implement the remote sensing image change detection method according to any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that, A computer program is stored thereon, which when executed by a processor of a computer, causes the computer to execute the remote sensing image change detection method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Novel ultra-high-definition remote sensing image change detection method based on AFFPN
CN112818818A
Multi-classification change detection method and system for multi-scale fusion dual-temporal remote sensing image
CN115035334A