A dual-temporal remote sensing image change detection method, model construction method and device

By using the depth residual network of the extrusion-excitation module and the pyramid attention module in the two-time phase remote sensing image change detection, combined with the composite loss function and the Adam optimizer, the problem of poor detection of targets for different sizes in complex backgrounds is solved, the recognition ability of changing areas and invariant areas is improved, and the impact of positive and negative samples is reduced.

CN114494870BActive Publication Date: 2025-05-30SHANDONG UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210073167.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-21
Publication Date
2025-05-30
Estimated Expiration
2042-01-21

AI Technical Summary

Technical Problem

In the dual-time phase remote sensing image change detection, the detection effect of different size changes targets under complex backgrounds is poor, the recognition ability of the changing areas and the unchanging areas is insufficient, and the training effect is reduced due to the imbalance between positive and negative samples of the training data set.

Method used

A depth residual network model with an extrusion-excitation module is used to construct a dual-time image remote sensing image feature extractor, and a pyramid attention module is constructed to calculate the correlation between feature image pixel pairs, which is weighted to calculate the feature map. The model is trained using the composite loss function and the Adam optimizer to improve the model's ability to identify areas with varying sizes and invariant areas.

Benefits of technology

The model's detection ability of different sizes, especially those with smaller sizes, is enhanced, and the ability to identify changing areas and unchanged areas is reduced, the impact of positive and negative samples imbalance on training effects is reduced, and the model's prediction effect in complex backgrounds is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114494870B_ABST
    Figure CN114494870B_ABST
Patent Text Reader

Abstract

The present invention discloses a dual-temporal remote sensing image change detection method, a model construction method and a device, belonging to the technical field of remote sensing image processing. A deep residual network model with a squeeze-and-excitation module is used to construct a dual-temporal remote sensing image feature extractor. The feature extractor integrates the rich semantic information of high-level features and the rich detailed information of low-level features. By introducing the squeeze-and-excitation module to weight the information of each channel, the model pays more attention to important features, improving the effect of feature extraction. The pyramid attention module has made improvements in the way of cutting images for detecting change targets of different sizes. The pyramid attention module cuts the feature map into several groups of feature submaps with overlapping edge pixels of different sizes, and uses the common attention algorithm to process the feature submaps, improving the detection ability of the model for different-size change regions, especially for small-size change regions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of remote sensing image processing, and particularly to a method for detecting changes in dual-temporal remote sensing images, a method for constructing a model, and a device. Background Art

[0002] In recent years, due to the ability to conveniently and accurately obtain surface information around the world, remote sensing images have been widely used. As an important branch in the field of remote sensing image research, change detection has been studied for decades. Currently, change detection tasks generally include the following aspects: land use change detection, forest and vegetation change detection, urban expansion change detection, disaster monitoring and assessment such as earthquakes and forest fires, etc. Change detection tasks usually involve vast surface areas. Manual execution is time-consuming and laborious. The introduction of change detection algorithms based on deep learning has greatly saved the human, material, and financial resources of researchers.

[0003] The existing change detection methods for dual-temporal remote sensing images mainly include the following several types:

[0004] I. Traditional methods: First, preprocess two dual-temporal remote sensing images, then obtain the difference map of the two images through methods such as change vector analysis and wavelet fusion, and finally process the difference map to generate a binary change detection image. Traditional methods rely to a large extent on the experience judgment of engineers and long-term debugging of algorithms. The cost of implementing traditional methods is relatively high.

[0005] II. Pixel-based change detection methods: Extract the change information by processing the spectral information of pixel pairs at the same positions of two dual-temporal remote sensing images. This method has the retention of detailed information that other methods do not have, and is simple and easy to implement, and is widely used. However, this method has poor robustness to some interference factors, such as changes in illumination angle and intensity, registration errors, etc.; and does not fully explore the spatial position information of each pixel point and its neighboring pixel points.

[0006] III. Object-based change detection algorithms: First, segment the remote sensing image, and then perform change detection on the objects generated by the segmentation. This method makes up for the defects of pixel-based change detection methods to a certain extent. Image segmentation pays more attention to the spatial information between pixel points, and the robustness of this method to noise is enhanced. However, due to the need to improve the integrity and accuracy of existing segmentation technologies, this method also needs to be further improved.

[0007] The above-mentioned several types of methods have poor detection effects on change targets of different sizes in complex backgrounds; the recognition ability for change regions and unchanged regions in remote sensing images needs to be improved; and the problem of reduced training effects caused by the imbalance between positive and negative samples in the training dataset cannot be effectively solved. Summary of the Invention

[0008] The present invention provides a method for detecting changes in dual-temporal remote sensing images, a method for constructing a model, and a device, aiming to solve the problems in the prior art that the method for detecting changes in dual-temporal remote sensing images has poor detection effects on changing targets of different sizes in complex backgrounds, poor recognition capabilities for changing areas and unchanged areas, and reduced training effects caused by unbalanced positive and negative samples in the training dataset.

[0009] The specific technical solutions provided by the present invention are as follows:

[0010] On the one hand, a method for constructing a dual-temporal remote sensing image change detection model provided by the present invention includes:

[0011] Construct a feature extractor for dual-temporal remote sensing images using a deep residual network model with a squeeze-and-excitation module added;

[0012] Construct a pyramid attention module for calculating the correlation between pixels of dual-temporal feature images and weighted calculation of the feature map with the correlation as the weight, wherein a co-attention algorithm is nested inside the pyramid attention module;

[0013] Use a preset remote sensing image dataset, perform image offset, brightness, and contrast transformations on the preset remote sensing image dataset, and then train a dual-temporal remote sensing image change detection model composed of the feature extractor for dual-temporal remote sensing images and the pyramid attention module using a composite loss function. The dual-temporal remote sensing image change detection model is used to obtain a binary change detection image based on dual-temporal remote sensing images.

[0014] Optionally, constructing the feature extractor for dual-temporal remote sensing images is specifically:

[0015] Retain the first convolutional layer to the fifth convolutional layer of the deep residual network model and delete the subsequent global pooling layer, fully connected layer, and activation function layer as the basic network of the feature extractor for dual-temporal remote sensing images;

[0016] Add a 1x1 convolutional layer after the second convolutional layer of the deep residual network model, add a 1x1 convolutional layer and an upsampling layer after the third convolutional layer and the fourth convolutional layer respectively, and add a global average pooling layer, a 1x1 convolutional layer, an upsampling layer, and a concatenation layer after the output of the fifth convolutional layer;

[0017] Add a squeeze-and-excitation module after the concatenation layer to form the feature extractor for dual-temporal remote sensing images as a whole, wherein the squeeze-and-excitation module is used for weighted learning of each channel of the feature map.

[0018] Optionally, the specific process of constructing the pyramid attention module is:

[0019] Construct a feature sub - map by cutting the feature map output by the feature extractor of the dual - temporal remote - sensing image based on the adjacent sub - map edge pixel overlap mechanism;

[0020] Adopt the co - attention algorithm for prediction for each pair of feature sub - maps to obtain an attention sub - map;

[0021] Construct an image stitching module for stitching the attention sub - maps. Among them, the size of the stitched attention feature map is equal to the size of the image before cutting, and the stitching result of the overlapping pixel region is equal to the result of weighted summation of the prediction results of the corresponding pixel regions of the two adjacent attention sub - maps where it is located with equal weights;

[0022] Construct a cascaded convolution module composed of 1 cascaded layer and 1 1x1 convolution layer. The cascaded convolution module is used to fuse the stitched attention feature map, and the generated attention feature map after fusion is equal to the size of the feature map before inputting into the pyramid attention module.

[0023] Optionally, adopting the co - attention algorithm for prediction for each pair of feature sub - maps to obtain an attention sub - map, specifically:

[0024] Multiply the two sub - maps included in a pair of dual - temporal sub - maps after transforming them to a preset size respectively to obtain a correlation feature map. Among them, each element of the correlation feature map represents the correlation degree between two pixel points in the two sub - maps;

[0025] Use the correlation feature map as weights to multiply with the two sub - maps respectively to obtain two weighted feature maps, and fuse the weighted feature maps with the two sub - maps to obtain an attention sub - map.

[0026] Optionally, during the process of training the dual - temporal remote - sensing image change detection model using the composite loss function, use the Adam optimizer to optimize the model parameters.

[0027] Optionally, the composite loss function is:

[0028]

[0029] where h and w respectively represent the height and width of the attention feature map; y ij represents the value of the label image, n u and n c respectively represent the total number of unchanged pixel point pairs and the total number of changed pixel point pairs in the attention feature map. Posdist and Negdist are used to expand the distance of the changed pixel point pairs with too small distance and shrink the distance of the unchanged pixel point pairs with too large distance respectively; Posdiff and Negdiff are used to make the model pay more attention to the learning of difficult - to - distinguish samples.

[0030] Optionally, the total number of unchanged pixel pairs and the total number of changing pixel pairs in the attention feature map are calculated using the following formulas respectively:

[0031]

[0032]

[0033] where h and w represent the height and width of the attention feature map respectively; y ij represents the value of the label image, and n u , n c represent the total number of unchanged pixel pairs and the total number of changing pixel pairs in the attention feature map respectively.

[0034] Optionally, Posdist ij = max{m - D ij , 0}

[0035] Negdist ij = max{D ij - τ, 0}

[0036]

[0037]

[0038] where D ij represents the value at the position of the i-th row and j-th column in the distance map D, α and β are weighting coefficients, τ and m are thresholds, α > β > 1, and m > τ.

[0039] On the other hand, the present invention also provides a device for constructing a dual-temporal remote sensing image change detection model, including:

[0040] A feature extractor construction module configured to construct a dual-temporal remote sensing image feature extractor using a deep residual network model with a squeeze-and-excitation module added;

[0041] A pyramid attention construction module configured to construct a pyramid attention module for calculating the correlation between pixels of the dual-temporal feature image and weighting the feature map with the correlation as the weight, wherein a co-attention algorithm is nested inside the pyramid attention module;

[0042] The model training module is configured to adopt a preset remote sensing image dataset, perform image offset, brightness, and contrast transformation on the preset remote sensing image dataset, and then use a composite loss function to train a dual-temporal remote sensing image change detection model composed of the dual-temporal remote sensing image feature extractor and the pyramid attention module. The dual-temporal remote sensing image change detection model is used to obtain a binary change detection image according to the dual-temporal remote sensing image.

[0043] In another aspect, the present invention also provides a method for detecting changes in dual-temporal remote sensing images. The dual-temporal remote sensing image change detection model constructed by the above method is used to detect changes in dual-temporal remote sensing images, and a binary change detection image is output.

[0044] The beneficial effects of the present invention are as follows:

[0045] The present invention provides a method, a model construction method, and a device for detecting changes in dual-temporal remote sensing images. A dual-temporal remote sensing image feature extractor is constructed using a deep residual network model with a squeeze-and-excitation module. The feature extractor integrates the rich semantic information of high-level features and the rich detailed information of low-level features. By introducing the squeeze-and-excitation module to weight the information of each channel, the model pays more attention to important features, improving the effect of feature extraction. The pyramid attention module improves the way of cutting images for detecting change targets of different sizes. The pyramid attention module cuts the feature map into several groups of feature submaps with overlapping edge pixels of different sizes, and uses a common attention algorithm to process the feature submaps, improving the model's detection ability for different-sized change regions, especially small-sized change regions. The common attention algorithm calculates the correlation between features in the dual-temporal feature map and weights the feature map with the correlation as the weight, improving the model's recognition ability for change regions and unchanged regions. In the model training stage, a composite loss function and an Adam optimizer are used to optimize the model. The composite loss function reduces the impact of positive and negative sample imbalance on the training effect in a weighted manner, increasing the learning intensity of the model for difficult-to-separate samples, thereby improving the prediction effect of the model in complex backgrounds. The Adam optimizer can ensure the stable change of model parameters within a certain range during the training process. Description of the Drawings

[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0047] Figure 1Flow chart of a method for constructing a dual - temporal remote sensing image change detection model according to an embodiment of the present invention;

[0048] Figure 2 Structural schematic diagram of a dual - temporal remote sensing image change detection model according to an embodiment of the present invention;

[0049] Figure 3 Structural schematic diagram of a feature extractor according to an embodiment of the present invention;

[0050] Figure 4 Schematic diagram of image cutting using an adjacent sub - map edge pixel overlapping mechanism according to an embodiment of the present invention;

[0051] Figure 5 Structural schematic diagram of a pyramid attention module according to an embodiment of the present invention;

[0052] Figure 6 Structural schematic diagram of a co - attention module according to an embodiment of the present invention;

[0053] Figure 7 Schematic diagram of the detection result of a method for detecting changes in a dual - temporal remote sensing image according to an embodiment of the present invention. Detailed implementation manners

[0054] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0055] The terms "including" and "having" in the specification and claims of the present invention, and any variations thereof, are intended to cover non - exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0056] Next, a method, a model construction method and a device for detecting changes in a dual - temporal remote sensing image according to an embodiment of the present invention will be described in detail with reference to Figures 1 to 7 shown in, a method for constructing a dual - temporal remote sensing image change detection model provided by an embodiment of the present invention includes:

[0057] Referring to Figure 1 、 Figure 2 、 Figure 3 and Figure 4 shown in, a method for constructing a dual - temporal remote sensing image change detection model provided by an embodiment of the present invention includes:

[0058] Step 100: Construct a dual-temporal remote sensing image feature extractor using a deep residual network model incorporated with a squeeze-and-excitation module.

[0059] The deep residual network model selected in the embodiment of the present invention is the ResNet-18 residual network model, that is, the feature extractor is implemented based on the ResNet-18 residual network model. The ResNet-18 model structure is as follows: The first convolutional layer conv1 includes a 7×7 convolutional layer and a 3×3 global max pooling layer. After conv1, there are four convolutional layers, namely conv2_x, conv3_x, conv4_x, and conv5_x, with the BasicBlock residual network as the main body, followed by a global pooling layer and a fully connected layer. The feature extractor constructed in the embodiment of the present invention deletes, modifies, and adds some sub-structures of ResNet-18, so that the improved structure can be successfully applied to the dual-temporal remote sensing image feature extraction task. The extracted feature maps have both rich semantic information and detailed information, and a channel weighting mechanism is used to increase the learning intensity of the model for more important features.

[0060] Specifically, step 100 includes the following execution process:

[0061] (1) Retain the first convolutional layer to the fifth convolutional layer of the deep residual network model, and delete the subsequent global pooling layer, fully connected layer, and activation function layer as the basic network of the dual-temporal remote sensing image feature extractor.

[0062] Among them, a 5-layer ResNet-18 residual network model is passed in as the basic network of the feature extractor model. Only retain the five convolutional layers of conv1, conv2_x, conv3_x, conv4_x, and conv5_x of ResNet-18, and delete the subsequent global pooling layer, fully connected layer, and softmax activation function.

[0063] (2) Add a 1x1 convolutional layer after the second convolutional layer of the deep residual network model, add a 1x1 convolutional layer and an upsampling layer after the third convolutional layer and the fourth convolutional layer respectively, and add a global average pooling layer, a 1x1 convolutional layer, an upsampling layer, and a concatenation layer after the output of the fifth convolutional layer.

[0064] After the conv2_x convolutional layer of the ResNet-18 residual network model, a 1×1 convolutional layer is added. After the conv3_x and conv4_x convolutional layers, a 1×1 convolutional layer and an upsampling layer are added respectively. After the output of conv5_x, a global average pooling layer, a 1×1 convolutional layer, and an upsampling layer are added to process the output feature maps of the four modules. A concatenation layer is added to perform data fusion on the four feature maps in the channel dimension.

[0065] (3) After the concatenation layer, a squeeze-and-excitation module is added to form a dual-temporal remote sensing image feature extractor. The squeeze-and-excitation module is used to perform weighted learning on each channel of the feature map.

[0066] Reference Figure 1 、 Figure 3 and Figure 4 As shown in, assume that two dual-temporal remote sensing images are input into the dual-temporal remote sensing image change detection model of the embodiment of the present invention. The feature extractor will process the two images separately, and the respective processing processes are as follows: The image first passes through the conv1 convolutional layer, and then passes through the conv2_x, conv3_x, conv4_x, and conv5_x convolutional layers in sequence. The feature extractor performs global average pooling, convolution, and upsampling on the four feature maps with different feature levels generated by the four convolutional layers. The feature maps are fused in the concatenation layer and then processed by the squeeze-and-excitation module (SE module layer), so that the model performs weighted learning on the features according to the importance of different features. The feature map generated after processing is the output of the feature extractor. Furthermore, two feature maps are obtained after the two dual-temporal remote sensing images are processed by the feature extractor.

[0067] The conv2_x, conv3_x, conv4_x, and conv5_x convolutional layers with different depths of the feature extractor will learn different levels of features of the remote sensing image. The low-level features contain richer detail information and higher resolution, but have less semantic information and more noise. The high-level features have rich semantic features, but have poor perception ability for details. The feature extractor needs to fuse different levels of features to give play to the advantages of different levels of features. The global average pooling layer is applied after the conv5_x module, which can simplify the parameters, reduce the calculation amount, and also achieve the purpose of dimensionality reduction. The SE module layer weights the information of each channel, increases the learning intensity of the model for more important features, and improves the effect of feature extraction. It is an effective semantic information supplement for the model.

[0068] Step 200: Construct a pyramid attention module for calculating the correlation between pixels of the dual-temporal feature image and weighted calculation of the feature map with the correlation as the weight.

[0069] By constructing a pyramid attention module, the image is cut into regions of different sizes for prediction. There is some pixel overlap between adjacent sub-images during cutting. A co-attention module is nested in the pyramid attention algorithm: the co-attention algorithm calculates the correlation between pixels of the dual-temporal feature image and weights the feature map with the correlation as the weight.

[0070] If two images contain the same target, then this target will exhibit relatively similar features in the two images. Therefore, the dual-temporal remote sensing image change detection model of the embodiments of the present invention can increase the discrimination ability for changed and unchanged targets by mining the correlation between the features of the dual-temporal images, and achieve this by designing a co-attention module. In addition, the pixel sizes of each changed target are different, and the prediction effect of the change detection algorithm on small targets is often not satisfactory. Existing methods generally solve this problem by cutting the image at different scale levels, but image cutting will inevitably cause some changed targets to be cut into two parts located on two adjacent sub-images respectively, thus affecting the prediction effect. The image cutting method of the dual-temporal remote sensing image change detection model of the embodiments of the present invention can not only ensure the locality of the sub-images, but also ensure that most of the changed targets can be included in one sub-image with a suitable size, that is, to a large extent, it guarantees the integrity of the changed targets in the sub-images.

[0071] Specifically, the specific process of constructing the pyramid attention module includes the following four steps:

[0072] Step 1: Construct a feature sub-image by cutting the feature map output by the feature extractor of the dual-temporal remote sensing image based on the adjacent sub-image edge pixel overlap mechanism.

[0073] By constructing an image cutting module, the feature map obtained by the feature extractor can be evenly cut into s×s sub-images (s∈{1, 2, 4, 8}); among them, the image cutting adopts the adjacent sub-image edge pixel overlap mechanism: for two adjacent sub-images at the edge, the pixel range extending 5 to 10 pixel points from the adjacent edge into the two images respectively is shared by the two sub-images; after the two dual-temporal images are cut respectively, the cut sub-images are recombined into several groups of dual-temporal feature sub-images;

[0074] Suppose a feature map with a size of 256×256 is input into the image cutting module. After cutting, four groups of feature sub-images will be generated. The first group of feature sub-images contains one sub-image with a size of 256×256, which is the same as the feature map before cutting.

[0075] The second group of feature sub - graphs contains four sub - graphs. The cutting module first cuts the original feature graph into 2×2 sub - graphs, and the size of each sub - graph is 128×128. However, considering that the image cutting line may pass through some change targets with relatively small sizes, in order to ensure the integrity of the change targets in the sub - graphs to a greater extent and to satisfy that there are several overlapping pixels at the edges of adjacent sub - graphs, the right edge and the bottom edge of the upper - left - corner sub - graph after cutting are respectively expanded by ten pixel points to the right and downwards. Similarly, the two edges adjacent to other sub - graphs of the upper - right, lower - left, and lower - right sub - graphs are also expanded by ten pixel points. Finally, the size of the obtained feature sub - graphs is 138×138, as Figure 5 shown.

[0076] The third group of feature sub - graphs contains 16 sub - graphs, and the size of each sub - graph is 74×74. In order to make the sizes of all sub - graphs in this group of images the same, if the upper and lower edges or the left and right edges of a sub - graph are adjacent to other images at the same time, then this sub - graph should be expanded by five pixel points in each of these two directions. If only one of the upper and lower edges or the left and right edges is adjacent to other sub - graphs, then this sub - graph is expanded by ten pixel points in this direction.

[0077] Similarly, the fourth group of feature sub - graphs contains 64 sub - graphs, and the size of each sub - graph is 42×42. The image cutting module will cut the two double - temporal feature graphs respectively to obtain two large groups of sub - graphs, and each large group of sub - graphs contains four groups of sub - graphs with different sizes as described above. For the sake of easy expression, two sub - graphs with the same size and position in the two large groups of sub - graphs are called a pair of double - temporal feature sub - graphs.

[0078] Step 2: Use the co - attention algorithm for each pair of feature sub - graphs to obtain the attention sub - graph.

[0079] Specifically, the two sub - graphs included in a pair of double - temporal sub - graphs are respectively transformed to a preset size and then multiplied to obtain a correlation feature graph. Among them, each element of the correlation feature graph represents the correlation degree between two pixel points in the two sub - graphs; the correlation feature graph is used as a weight to multiply with the two sub - graphs respectively to obtain two weighted feature graphs, and the weighted feature graphs are fused with the two sub - graphs to obtain the attention sub - graph.

[0080] For each pair of double - temporal feature sub - graphs, use the co - attention module for prediction to obtain the attention sub - graph; among them, the two sub - graphs included in a pair of double - temporal feature sub - graphs are respectively transformed to a preset size and then multiplied to obtain a correlation feature graph; each element of the correlation feature graph represents the correlation degree between two pixel points in the two sub - graphs;

[0081] Suppose the sizes of both pairs of feature sub - graphs are \(h\times w\). The size of one of the feature sub - graphs is transformed into \(N\times1\), and the other is transformed into \(1\times N\) (\(N = h\times w\)). After matrix multiplication, the size of the correlation feature map is \(N\times N\), where each element corresponds to the similarity of two pixel points from the two sub - graphs respectively. Multiply the correlation feature map as weights with the two sub - graphs respectively to obtain two weighted feature maps, and fuse the weighted feature maps with the two sub - graphs to obtain the attention sub - graph.

[0082] The role of the co - attention module is to obtain the correlation between two features in two multi - temporal remote sensing images. Since the features corresponding to the invariant targets in the multi - temporal images are highly similar, weighting the feature maps according to the correlation between features will help improve the model's ability to distinguish between the changed areas and the invariant areas.

[0083] Step 3: Construct an image stitching module for stitching the attention sub - graphs. Among them, the size of the stitched attention feature map is equal to the size of the image before cutting. The stitching result of the overlapping pixel area is equal to the result of weighting the prediction results of the corresponding pixel areas of the two adjacent attention sub - graphs where it is located with equal weights.

[0084] Reference Figure 5 and Figure 6 As shown in

[0085] and

[0086] re - stitch the attention sub - graphs by constructing an image stitching module. The size of the stitched attention feature map is equal to the size of the image before cutting. For the overlapping pixel area, its stitching result is equal to the result of weighting the prediction results of the corresponding pixel areas of the two adjacent attention sub - graphs where it is located with equal weights. For example, for a \(2\times2\) sub - graph, the pixel points in rows 1 - 118 and columns 118 - 138 of the original image are respectively in the upper - left sub - graph and the upper - right sub - graph. Then in this area: stitching result=(prediction result of the corresponding area of the upper - left sub - graph+prediction result of the corresponding area of the upper - right sub - graph) / 2. And for the pixel points in rows 118 - 138 and columns 118 - 138, they are included in four sub - graphs respectively. Then the stitching result of the pixel points in this area is equal to the average value of the prediction results of the corresponding areas of the four sub - graphs.

[0087] Reference Figure 6 and Figure 7As shown, the concatenated attention feature map is fused by constructing a cascaded convolutional module, which consists of a concatenation layer and a 1×1 convolutional layer; the size of the attention feature map generated after fusion is equal to that of the feature map before inputting into the pyramid module.

[0088] Reference Figure 1 、 Figure 5 and Figure 6 As shown, the pyramid attention module helps the dual-temporal remote sensing image change detection model of the embodiment of the present invention to identify small-sized change regions, the overlapping pixel mechanism increases the integrity and accuracy of the identification of change targets at the edge positions of sub-images, and the co-attention module helps to improve the discrimination ability of the dual-temporal remote sensing image change detection model for change regions and unchanged regions.

[0089] Step 300: Adopt a preset remote sensing image dataset, perform image offset, brightness, and contrast transformations on the preset remote sensing image dataset, and then use a composite loss function to train the dual-temporal remote sensing image change detection model composed of the dual-temporal remote sensing image feature extractor and the pyramid attention module.

[0090] The dual-temporal remote sensing image change detection model provided by the embodiment of the present invention is used to obtain a binary change detection image based on dual-temporal remote sensing images. By performing data augmentation on the training dataset, the robustness of the model to interference factors such as illumination intensity differences, illumination angle differences, and registration differences is improved; by designing a composite loss function, the impact of positive and negative sample imbalance on the training effect is reduced, and the learning intensity of the model for difficult-to-separate samples is increased, so that the dual-temporal remote sensing image change detection model can achieve better results in more complex change detection tasks; the Adam optimizer can ensure that the parameters of the dual-temporal remote sensing image change detection model change within a certain range in each iteration process, improving the stability of the dual-temporal remote sensing image change detection model during the training process.

[0091] Specifically, based on the publicly available remote sensing dataset, preprocess the dual-temporal remote sensing images and label images of the dataset to construct a training set and a test set for the dual-temporal remote sensing image change detection model of the embodiment of the present invention. Among them, the publicly available remote sensing dataset includes the SZTAKI dataset and the LEVIR-CD dataset.

[0092] Among them, the steps for preprocessing the dual-temporal remote sensing images and label images in the dataset include but are not limited to: performing random image offsets of no more than 5 pixel points and random image rotations with an angle of no more than 15° on the dual-temporal remote sensing images and label images, and making random changes of no more than 10% to the brightness and contrast of the dual-temporal remote sensing images. The data augmentation method involved in the preprocessing is beneficial to improving the robustness of the model to noise: image offset helps to increase the robustness of the model to the registration error of dual-temporal images; the changes in brightness and contrast help to increase the adaptability of the model to changes in light intensity and light angle.

[0093] Use the composite loss function as the loss function of the dual-temporal remote sensing image change detection model in the embodiments of the present invention, and use the Adam optimizer to optimize the model parameters to obtain the trained dual-temporal remote sensing image change detection model. Among them, according to the formula Calculate the Euclidean distance between pixel pairs at the same position in the two attention feature maps to obtain the distance map D, where D ij represents the value at the position of the i-th row and j-th column in the distance map D, and y ij represent the values in the two label images. The distance map D can measure the similarity between features, and the regions where pixel pairs with larger distances are located are more likely to be the changed regions.

[0094] Calculate the composite loss function L according to the following formula:

[0095]

[0096] Among them, h and w respectively represent the height and width of the attention feature map; y ij represents the value of the label image, and n u and n c respectively represent the total number of unchanged pixel pairs and the total number of changed pixel pairs in the attention feature map. Posdist and Negdist respectively enlarge the distance of the changed pixel pairs with too small distances and shrink the distance of the unchanged pixel pairs with too large distances; Posdiff and Negdiff are used to make the model pay more attention to the learning of difficult-to-separate samples.

[0097] y ij represents the value of the label image. The value of the pixel corresponding to the changed region is 1, and the value of the pixel corresponding to the unchanged region is 0. Through different values of y ij , the loss function is divided into two parts, respectively corresponding to the set of changed pixel points with a label of 1 and the set of unchanged pixel points with a label of 0. n u and n crespectively represent the total number of invariant pixel pairs and the total number of changing pixel pairs in the attention feature map, and are calculated according to the following formula:

[0098]

[0099]

[0100] where n u and n c respectively represent the total number of invariant pixel pairs and the total number of changing pixel pairs in the attention feature map, y ij represents the value of the binary label map, and h and w respectively represent the height and width of the attention feature map.

[0101] Related research shows that the training effect of convolutional neural networks is sensitive to the balance of training sample categories. More balanced samples help improve the prediction effect of the model, while unbalanced samples will cause the model's prediction to be biased towards a certain category. To overcome the problem of a large gap in the number of positive and negative samples in the training set that may occur in practical applications, the composite loss function calculates the total number n u and n c of the two types of samples, and uses them to weight the losses of the two types of samples respectively. The weight of the positive samples is The weight of the negative samples is

[0102] Posdist and Negdist perform the following tasks: expand the distance of changing pixel pairs with too small a distance and shrink the distance of invariant pixel pairs with too large a distance. Posdist and Negdist are calculated according to the following formula:

[0103] Posdist ij = max{m - D ij , 0}

[0104] Negdist ij = max{D ij - τ, 0}

[0105] where D ij represents the value at the position of the i-th row and j-th column in the distance map D, α and β are weighting coefficients, T and m are thresholds, α > β > 1, and m > τ.

[0106] Posdist and Negdist implement a mechanism similar to Contrastive loss: for positive samples with a predicted distance less than the threshold m, this loss function expands their distances; for negative samples with a predicted distance greater than the threshold τ, their distances are shrunk; for a large portion of samples with good prediction results, their losses are ignored, reducing the computational amount. This mechanism helps to expand the discrimination boundary between classes, making the model's prediction for adversarial samples more accurate.

[0107] Posdiff and Negdiff accomplish the following tasks: for each pair of pixel points, according to their classes and the corresponding distances in the distance map D, their difficulty levels of discrimination are divided into three categories, and different weights are given to each category for learning: more difficult samples are given greater weights, ordinary samples are next, and the losses of easy samples are ignored. They are calculated according to the following formula:

[0108]

[0109]

[0110] where D ij represents the value at the position of the i-th row and j-th column in the distance map D, α and β are weighting coefficients, τ and m are thresholds, α > β > 1, and m > τ.

[0111] Posdiff and Negdiff can make the dual-temporal remote sensing image change detection model of the embodiments of the present invention pay more attention to the learning of difficult samples. The positive samples and negative samples are respectively divided by two thresholds m and τ, and the division criterion is the difficulty level of the samples, and the difficulty level is reflected by the corresponding distance of the pixel points in the distance map D. Negative samples with too large predicted distances and positive samples with too small predicted distances are more difficult to distinguish. For difficult samples, larger weights are given to them in the loss function, for samples with moderate difficulty, their weights are relatively small, and for a large number of easy samples, their losses are ignored. The Adam optimizer can dynamically adjust the learning rate of each parameter using the first-order moment estimate and second-order moment estimate of the gradient. The learning rate can achieve a stable change within a clear range after each iteration.

[0112] The present invention provides a method for constructing a dual-temporal remote sensing image change detection model. A deep residual network model with a squeeze-and-excitation module is used to construct a feature extractor for dual-temporal remote sensing images. The feature extractor integrates the rich semantic information of high-level features and the rich detailed information of low-level features. By introducing the squeeze-and-excitation module to weight the information of each channel, the model pays more attention to important features, improving the effect of feature extraction. The pyramid attention module improves the way of cutting images for detecting change targets of different sizes. The pyramid attention module cuts the feature map into several groups of feature submaps with overlapping edge pixels of different sizes, and processes the feature submaps using the common attention algorithm, improving the model's detection ability for different-size change regions, especially small-size change regions. The common attention algorithm calculates the correlation between features in the dual-temporal feature maps and weights the feature maps with the correlation as the weight, improving the model's recognition ability for change regions and unchanged regions. In the model training stage, a composite loss function and an Adam optimizer are used to optimize the model. The composite loss function reduces the impact of positive and negative sample imbalance on the training effect in a weighted manner, increasing the learning intensity of the model for difficult-to-separate samples, thereby improving the prediction effect of the model in complex backgrounds. The Adam optimizer can ensure the stable change of model parameters within a certain range during the training process.

[0113] Based on the same inventive concept, an embodiment of the present invention further provides a device for constructing a dual-temporal remote sensing image change detection model for executing the above method for constructing a dual-temporal remote sensing image change detection model. The device for constructing a dual-temporal remote sensing image change detection model includes:

[0114] A feature extractor construction module configured to construct a feature extractor for dual-temporal remote sensing images using a deep residual network model with a squeeze-and-excitation module;

[0115] A pyramid attention construction module configured to construct a pyramid attention module for calculating the correlation between pixels of dual-temporal feature images and weighting the feature map with the correlation as the weight, wherein the common attention algorithm is nested inside the pyramid attention module;

[0116] A model training module configured to use a preset remote sensing image data set, perform image offset, brightness, and contrast transformations on the preset remote sensing image data set, and then train a dual-temporal remote sensing image change detection model composed of the feature extractor for dual-temporal remote sensing images and the pyramid attention module using the composite loss function. The dual-temporal remote sensing image change detection model is used to obtain a binary change detection image based on dual-temporal remote sensing images.

[0117] Based on the same inventive concept, an embodiment of the present invention further provides a method for dual-temporal remote sensing image change detection. The method uses the dual-temporal remote sensing image change detection model constructed by the above method to perform change detection on dual-temporal remote sensing images and outputs a binary change detection image.

[0118] Specifically, the method for dual-temporal remote sensing image change detection inputs the dual-temporal remote sensing images into the trained temporal remote sensing image change detection model using the above method. Through a threshold segmentation mechanism, the final binary change detection image is obtained. By calculating the Euclidean distance between pixel pairs at the same position in the two attention feature maps, a distance map D is obtained; traverse each pixel point D in the distance map D ij , a fixed threshold θ is selected. If D ij > θ, then this pixel point is predicted to have changed, otherwise it is predicted to have not changed. Refer to Figure 7 As described above, a binary change detection image is generated. The corresponding value of the pixel point that has changed in the binary change detection image is 1, and the corresponding value of the pixel point that has not changed is 0.

[0119] Obviously, those skilled in the art can make various changes and modifications to the embodiments of the present invention without departing from the spirit and scope of the embodiments of the present invention. Thus, if these modifications and variations of the embodiments of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention also intends to include these changes and modifications.

Claims

1. A method for constructing a dual - temporal remote - sensing image change detection model, characterized in that, the method for constructing a dual - temporal remote - sensing image change detection model includes: using a deep residual network model with a squeeze - and - excitation module to construct a dual - temporal remote - sensing image feature extractor; constructing a pyramid attention module for calculating the correlation between pixels of dual - temporal feature images and weighted - calculating the feature map with the correlation as the weight, wherein a co - attention algorithm is nested inside the pyramid attention module; using a preset remote - sensing image dataset, performing image offset, brightness, and contrast transformation on the preset remote - sensing image dataset, and then training a dual - temporal remote - sensing image change detection model composed of the dual - temporal remote - sensing image feature extractor and the pyramid attention module using a composite loss function, and the dual - temporal remote - sensing image change detection model is used to obtain a binary change detection image according to the dual - temporal remote - sensing images; Specifically, constructing the dual - temporal remote - sensing image feature extractor is as follows: Retaining the first convolutional layer to the fifth convolutional layer of the deep residual network model and deleting the subsequent global pooling layer, fully - connected layer, and activation function layer as the basic network of the dual - temporal remote - sensing image feature extractor; Adding a 1x1 convolutional layer after the second convolutional layer of the deep residual network model, adding a 1x1 convolutional layer and an up - sampling layer after the third convolutional layer and the fourth convolutional layer respectively, and adding a global average pooling layer, a 1x1 convolutional layer, an up - sampling layer, and a concatenation layer after the output of the fifth convolutional layer; Adding a squeeze - and - excitation module after the concatenation layer to form the dual - temporal remote - sensing image feature extractor as a whole, wherein the squeeze - and - excitation module is used for weighted learning of each channel of the feature map; The specific process of constructing the pyramid attention module is as follows: Constructing a feature sub - map by cutting the feature map output by the dual - temporal remote - sensing image feature extractor based on an adjacent sub - map edge pixel overlap mechanism; Using the co - attention algorithm to perform prediction on each pair of feature sub - maps to obtain an attention sub - map; Constructing an image stitching module for stitching the attention sub - maps, wherein the size of the stitched attention feature map is equal to the size of the image before cutting, and the stitching result of the overlapping pixel region is equal to the result of weighted averaging of the prediction results of the corresponding pixel regions of the two adjacent attention sub - maps where it is located with equal weights; Constructing a cascaded convolutional module composed of 1 cascading layer and 1 1x1 convolutional layer, and the cascaded convolutional module is used for fusing the stitched attention feature map, and the generated attention feature map after fusion has the same size as the feature map before inputting into the pyramid attention module.

2. The method for constructing a dual - temporal remote - sensing image change detection model according to claim 1, characterized in that, using the co - attention algorithm to perform prediction on each pair of feature sub - maps to obtain an attention sub - map, specifically: Multiplying two sub - maps included in a pair of dual - temporal sub - maps after transforming them to a preset size respectively to obtain a correlation feature map, where each element of the correlation feature map represents the correlation degree between two pixel points in the two sub - maps; Multiply the relevant feature maps with the two sub - maps as weights respectively to obtain two weighted feature maps, and fuse the weighted feature maps with the two sub - maps to obtain an attention sub - map.

3. The method for constructing a dual - temporal remote sensing image change detection model according to claim 1, wherein, during the process of training the dual - temporal remote sensing image change detection model using the composite loss function, the Adam optimizer is used to optimize the model parameters.

4. The method for constructing a dual - temporal remote sensing image change detection model according to claim 1, wherein, the composite loss function is: where h and w respectively represent the height and width of the attention feature map; represents the value of the label image, , respectively represent the total number of unchanged pixel pairs and the total number of changed pixel pairs in the attention feature map. Posdist and Negdist respectively enlarge the distance of changed pixel pairs with too small a distance and shrink the distance of unchanged pixel pairs with too large a distance; Posdiff and Negdiff are used to make the model pay more attention to the learning of difficult-to-separate samples.

5. The method for constructing a dual - temporal remote sensing image change detection model according to claim 4, wherein, the total number of unchanged pixel - point pairs and the total number of changed pixel - point pairs in the attention feature map are calculated using the following formulas respectively: where h and w respectively represent the height and width of the attention feature map; represents the value of the label image, , respectively represent the total number of unchanged pixel pairs and the total number of changed pixel pairs in the attention feature map.

6. The method for constructing a dual - temporal remote sensing image change detection model according to claim 5, wherein, Among them, represents the value at the position of the i-th row and j-th column in the distance graph D, α and β are weighting coefficients, τ and m are thresholds, .

7. A device for constructing a dual - temporal remote sensing image change detection model, wherein, the device for constructing a dual - temporal remote sensing image change detection model includes: A feature extractor construction module, configured to construct a dual - temporal remote sensing image feature extractor using a deep residual network model with a squeeze - and - excitation module added; A pyramid attention construction module, configured to construct a pyramid attention module for calculating the correlation between pixels of dual - temporal feature images and weighting the feature maps with the correlation as the weight, wherein the co - attention algorithm is nested inside the pyramid attention module; A model training module, configured to use a preset remote sensing image data set, perform image offset, brightness, and contrast transformations on the preset remote sensing image data set, and then use the composite loss function to train the dual - temporal remote sensing image change detection model composed of the dual - temporal remote sensing image feature extractor and the pyramid attention module. The dual - temporal remote sensing image change detection model is used to obtain a binary change detection image according to the dual - temporal remote sensing images; wherein, the specific process of constructing the dual - temporal remote sensing image feature extractor is: Retain the first convolutional layer to the fifth convolutional layer of the deep residual network model and delete the subsequent global pooling layer, fully - connected layer, and activation function layer as the basic network of the dual - temporal remote sensing image feature extractor; Add a 1x1 convolutional layer after the second convolutional layer of the deep residual network model, add a 1x1 convolutional layer and an up - sampling layer after the third convolutional layer and the fourth convolutional layer respectively, and add a global average pooling layer, a 1x1 convolutional layer, an up - sampling layer, and a concatenation layer after the output of the fifth convolutional layer; Add a squeeze - and - excitation module after the concatenation layer to form the dual - temporal remote sensing image feature extractor as a whole, wherein the squeeze - and - excitation module is used for weighted learning of each channel of the feature map; The specific process of constructing the pyramid attention module is: Construct a feature sub - map by cutting the feature map output by the dual - temporal remote sensing image feature extractor based on the adjacent sub - map edge pixel overlap mechanism; Use the co - attention algorithm to perform prediction for each pair of feature sub - maps to obtain an attention sub - map; Construct an image stitching module for stitching the attention subgraphs, wherein the size of the stitched attention feature map is equal to the size of the image before cutting, and the stitching result of the overlapping pixel region is equal to the result of weighted prediction results of the corresponding pixel regions of the two adjacent attention subgraphs where it is located with equal weights; Construct a cascaded convolution module composed of 1 cascaded layer and 1 1x1 convolution layer, and the cascaded convolution module is used to fuse the stitched attention feature map, and the generated attention feature map after fusion is equal to the size of the feature map before inputting into the pyramid attention module.

8. A method for change detection of dual-temporal remote sensing images Characterized in that The method for change detection of dual-temporal remote sensing images uses the dual-temporal remote sensing image change detection model constructed by any one of claims 1 to 6 to perform change detection on dual-temporal remote sensing images and outputs a binary change detection image.

Citation Information

Patent Citations

  • Target detection improved algorithm based on feature pyramid network and attention mechanism

    CN111914917A

  • Double-time-phase high-resolution remote sensing image change detection algorithm

    CN112577473A