Lightweight Optical Remote Sensing Image Change Detection Method Based on Efficient Attention

By introducing FOCUS module, depth residual block, efficient attention module and multi-scale feature fusion module into the change detection network, a lightweight change detection network is designed, which solves the problems of large parameters and slow inference speed of existing deep learning change detection methods, and achieves more efficient change detection performance.

CN115713529BActive Publication Date: 2025-06-24HENAN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211524552.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-30
Publication Date
2025-06-24
Estimated Expiration
2042-11-30

AI Technical Summary

Technical Problem

When the existing deep learning change detection methods based on CNN and Transformer are implemented, the number of parameters is large and the inference speed is slow, making it difficult to apply to large-scale remote sensing image processing, industrial fields or applications that require real-time performance.

Method used

A lightweight optical remote sensing image change detection method based on efficient attention is proposed. Through the FOCUS module, depth residual block, efficient attention module and multi-scale feature fusion module, an end-to-end lightweight change detection network (LCDNet) is designed to achieve fewer parameters and faster inference speed.

Benefits of technology

It has achieved better change detection performance, with the parameter volume of only 0.88MB and the inference speed of 4.75ms, which can effectively solve the application difficulties of high-performance change detection algorithms in industrial fields or applications requiring real-time performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115713529B_ABST
    Figure CN115713529B_ABST
Patent Text Reader

Abstract

The present invention discloses a lightweight optical remote sensing image change detection method based on efficient attention, comprising the following steps: First, preprocess the optical remote sensing image to obtain a corresponding change label map; then cut it to obtain training samples; then concatenate the two-temporal images and pass them through the FOCUS module, the deep residual block, and the lightweight attention module to obtain refined feature maps of different scales. Subsequently, use the multi-scale feature fusion module to aggregate the obtained multiple feature maps to generate a change map; after the training is completed, save all the parameter information of the model; finally, input the preprocessed sample to be measured into the change detection model, and calculate and output the detection result map. The solution of the present invention uses the FOCUS downsampling layer, the deep residual convolution block that can expand the receptive field, the efficient attention mechanism, and the multi-scale feature fusion module to extract the change region with fewer parameters and computational complexity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of optical remote sensing image change detection, and particularly to a lightweight optical remote sensing image change detection method based on efficient attention. Background Art

[0002] Remote sensing image change detection is a technology that uses two or more remote sensing images acquired at different times in the same area for comparative analysis to obtain change information. Different fields have different definitions of change, such as agricultural surveys, forest monitoring, urban expansion, and disaster assessment. In recent years, with the rapid development of satellite remote sensing technology and computer vision technology, remote sensing image change detection has become an active research topic.

[0003] With the rapid development of computer technology and the continuous increase of high-resolution optical remote sensing image datasets, many deep learning-based change detection methods have been proposed by domestic and foreign scholars. Deep learning-based change detection methods have non-linear features and excellent feature extraction capabilities, predicting pixel classification maps and highly semantic abstract spatial contexts from raw images, blurring the boundaries between traditional pixel-based and object-based methods, and enabling better understanding of complex scenes. Deep learning-based change detection methods do not require image preprocessing, which not only reduces manual intervention but also avoids errors caused by preprocessing. Therefore, the use of deep learning-based remote sensing image change detection methods in solving remote sensing image change detection problems has increased exponentially. Currently, the mainstream deep learning-based change detection methods can be roughly divided into two categories: change detection methods based on Convolutional Neural Network (CNN) and change detection methods based on Transformer. CNN is widely used in deep learning due to its powerful feature learning ability. In recent years, many CNN-based change detection methods [1]-[7] have been proposed. The self-attention mechanism [8] has been widely applied in the field of natural language processing to find the correlation between different parts of the input. Vision Transformer [9] and Swin Transformer

[10] introduce the self-attention mechanism into the field of computer vision and improve it. Networks based on the self-attention mechanism and Transformer model the global distance of the input feature map through non-local self-attention. Based on this, many scholars have introduced it into the change detection field to obtain better change detection performance [11-13].

[0004] However, there are still some problems with deep learning change detection methods based on CNN and Transformer. Simple models [1-2,7] have small parameter counts and fast inference speeds, but their change detection performance is low and cannot meet the requirements for accurately identifying change regions. Complex models [3-6,11-13] use more modules, larger structures, and more complex training processes, resulting in a significant improvement in change detection ability. However, they also lead to high model parameter counts and slow inference speeds, limiting their application to large-scale remote sensing image processing, industrial fields, or applications requiring real-time performance. Summary of the Invention

[0005] The object of the present invention is to provide a lightweight optical remote sensing image change detection method based on efficient attention, which can achieve better change detection performance with a parameter count of 0.88 MB and an inference speed of 4.75 ms.

[0006] The technical solution adopted by the present invention is as follows:

[0007] A. Orthorectify, image register, image stretch, and image numerical normalization preprocessing are sequentially performed on the dual-temporal optical remote sensing images to obtain remote sensing images with consistent data distributions.

[0008] B. The updated parts in the remote sensing images are marked for the preprocessed dual-temporal optical remote sensing images obtained in step A to obtain corresponding change label maps.

[0009] C. The label maps obtained in step B and the preprocessed dual-temporal optical remote sensing images obtained in step A are cut to the same size to obtain training samples.

[0010] D. The dual-temporal remote sensing images in the training samples are concatenated.

[0011] E. The concatenated image pairs are downsampled through the FOCUS module, and the downsampled feature maps are input into the Depthwise Residual Block (DRB) for encoding to extract feature maps related to the change regions.

[0012] F. The feature maps obtained in step E are input into the Efficient Attention Module (EAM) to refine the feature maps.

[0013] G. The features of different scales obtained in step F are input into the Multiscale Feature Fusion Module (MFFM) to obtain the final feature map X.

[0014] H. Input the final fused feature X into the prediction head composed of 1×1 convolutions ( Figure 2 (c)) to obtain the predicted change map of the dual-temporal image of the training sample;

[0015] I. Combine binary cross-entropy loss and Dice loss to form a hybrid loss function to calculate the loss between the predicted change map of the dual-temporal image of the training sample obtained in step H and the corresponding label map;

[0016] J. After the training is completed, save both the weight parameters and hyperparameters of the trained change detection model;

[0017] K. After the pre- and post-temporal remote sensing images to be detected are successively preprocessed by orthorectification, image registration, image stretching, and image numerical normalization, then cut them with the same size to obtain the samples to be detected;

[0018] L. Input the samples to be detected into the change detection model obtained in step J, and calculate and output the predicted change map of the samples to be detected.

[0019] The present invention takes the change detection of optical remote sensing images as the application background. Aiming at the problem that the existing change detection methods are difficult to balance the change detection performance and the number of model parameters, a lightweight optical remote sensing image change detection method based on efficient attention (Lightweight Change Detection Network, LCDNet) is proposed. This detection method can achieve better change detection performance and faster inference speed with fewer parameters. Specifically, aiming at the above problems, the present invention designs an end-to-end lightweight change detection method. The present invention designs from four aspects: downsampling layer, convolution method, attention mechanism, and feature fusion to meet the requirements of fewer parameters and higher change detection performance. The present invention sets the downsampling layer at the beginning rather than the end of each encoding layer to reduce the number of model parameters. Using the downsampling layer at the beginning of the encoding layer will cause the loss of some feature information. Therefore, the present invention introduces the FOCUS module widely used in the field of object detection to solve this problem. The FOCUS module can ensure that the information is not lost and achieve the feature Figure 2Downsampling by a factor of two. In the present invention, depthwise (DW) convolutions with large convolutional kernels are used in the network to not only expand the receptive field but also significantly compress the parameters and computational complexity. To enable the network to pay more attention to the changing regions and improve the change detection performance of the network, the present invention designs an efficient attention module (EAM). The EAM sums the channel dimension weights obtained by fast one-dimensional convolutions and the spatial dimension weights obtained by single-layer two-dimensional convolutions and redistributes the weights to retain the correlation between the channel and spatial features. The present invention also designs a multi-scale feature fusion module (MFFM) that only uses effective feature streams during feature fusion. The MFFM can effectively fuse multi-scale features with a simple structure and a low number of parameters. Compared with traditional algorithms, the solution of the present invention can achieve better change detection performance with fewer parameters and a faster inference speed, and can effectively solve the problem that high-performance change detection algorithms are difficult to apply to industrial fields or applications requiring real-time performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0021] Figure 1 It is a schematic flowchart of the present invention.

[0022] Figure 2 It is a structural diagram of LCDNet of the present invention.

[0023] Figure 3 It is a diagram of the efficient attention module of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0024] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, not all of them. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention belong to the scope of protection of the present invention.

[0025] As Figure 1 shown, the present invention includes the following steps:

[0026] A. Orthorectify, image register, image stretch, and numerically normalize the dual-temporal optical remote sensing images in sequence to obtain remote sensing images with consistent data distributions;

[0027] B. Label the updated parts in the pre - processed dual - temporal optical remote - sensing image obtained in step A (mainly including vegetation changes, newly built urban buildings, suburban expansion, foundations before construction, and road expansion, etc.) to obtain the corresponding change label map;

[0028] C. Cut the change label map obtained in step B and the pre - processed dual - temporal optical remote - sensing image obtained in step A with the same size to obtain training samples;

[0029] D. Since the change detection task can be regarded as segmenting the changed areas in the dual - temporal images, the remote - sensing images before and after the change in the training samples can be concatenated as a whole;

[0030] E. Encode the concatenated remote - sensing images obtained in step D. The input feature map first undergoes down - sampling operations to reduce the number of model parameters. In the first two encoder layers, the FOCUS module is used to achieve 2 - fold down - sampling while ensuring that the information of the feature map is not lost. In the last two decoder layers, MaxPool2d with a convolution kernel of 3 and a stride of 2 is used for 2 - fold down - sampling. Since the feature maps of the shallow network contain more detailed texture information, the FOCUS module is only used in the first two encoding layers in the present invention. The disadvantage of the standard CNN is that the use of fixed small convolution kernels in the network leads to a limited receptive field. To overcome this problem, recent work focuses on using larger convolution kernels to expand the receptive field. Therefore, in the present invention, deep convolution layers with convolution kernels of 3×3 and 5×5 are stacked, which not only expands the receptive field but also significantly compresses the parameters and computational amount. To prevent the phenomenon of network degradation as the number of network layers increases, a 1×1 convolution is also performed on the down - sampled feature map to form residuals; Figure 2

[0031] F. Input the feature map F obtained by the encoder in step E into the designed efficient attention module (EAM) to further refine the feature map extracted in step E, thereby improving the change detection performance of the network. The efficient attention module is as shown in Figure 3 . Use the average pooling layer for the feature map F to generate aggregated vectors of size C×1×1 and 1×H×W (C is the number of channels, H and W are the height and width of the feature map). Apply a one - dimensional convolution with a convolution kernel of 3 to the aggregated vector of C×1×1 to obtain an attention map in the channel dimension. Apply a two - dimensional convolution with a convolution kernel of 7 to the aggregated vector of 1×H×W to obtain an attention map in the spatial dimension. Expand the above two attention feature maps to C×H×W, add them, and re - distribute the weights to obtain an attention map M(F) that retains the correlation between the channel and spatial features. Multiply the input feature map element - by - element with the attention map M(F) to obtain the refined feature map F':

[0032] ​M(F) = σ(C1D3(AvgPool(F)) + C2D7(AvgPool(F)))

[0033]

[0034] where F represents the feature map obtained after passing through the deep residual block, AvgPool(·) represents the average pooling operation, C1D3(·) represents the one-dimensional convolution with a kernel size of 3, C2D7(·) represents the two-dimensional convolution with a kernel size of 7, and σ represents the sigmoid function. represents element-wise multiplication, and F' represents the refined feature map after weighting;

[0035] G. In order to effectively fuse the feature maps of different scales extracted in step F, the present invention proposes a multi-scale feature fusion module (MFFM), as shown in Figure 2 (b). The acquisition of the feature map is mathematically described as follows:

[0036] X1 = C(F4)

[0037] X2 = C(C(X1, F3))

[0038] X3 = C(C(X1, X2, F2))

[0039] X4 = C(X2, X3, F1)

[0040] X = X1 + X2 + X3 + X4

[0041] where the function C(·) represents the convolution operation using the convolution block ( Figure 2 (d)) composed of 1×1 convolution and 3×3 DW convolution. X1, X2, X3, and X4 respectively represent the feature maps obtained after passing through four layers of deep residual blocks, and F1, F2, F3, and F4 represent the refined feature maps obtained after the feature maps X1, X2, X3, and X4 pass through the efficient attention module. The final fused feature X is obtained by upsampling and adding the feature maps X1, X2, X3, and X4. Since only the effective feature stream and the designed convolution block are used for feature fusion, the proposed MFFM can effectively fuse multi-scale features with a simple structure and a low number of parameters;

[0042] H. Input the final fused feature X into the prediction head ( Figure 2 (c)) composed of 1×1 convolution to obtain the predicted change map of the dual-temporal image of the training sample;

[0043] I. Combine the cross-entropy loss commonly used in the binary classification task with the loss that can alleviate the background imbalance in the samples bce+L dice Calculate the loss between the predicted change map of the dual-temporal image of the training samples obtained in calculation step H and the corresponding label map, where y i,j represents the probability that the pixel point (i, j) in the corresponding label map is a changed pixel, represents the probability that the pixel point (i, j) in the predicted change map is a changed pixel, and n and m respectively represent the width and height at the pixel level of the image;

[0044] J. After the training is completed, save both the weight parameters and hyperparameter information of the trained change detection model;

[0045] K. After the pre- and post-temporal remote sensing images to be detected are successively subjected to orthorectification, image registration, image stretching, and image numerical normalization preprocessing, cut them with the same size to obtain the samples to be detected;

[0046] L. Input the sample to be detected into the change detection model saved in step J, and calculate and output the predicted change map of the sample to be detected.

[0047] In the present invention, in order to solve the problem that existing high-performance change detection methods have a large number of parameters, slow inference speed, and are difficult to be deployed in industrial fields or applications requiring real-time performance, a FOCUS module used at the beginning of the encoder, a depth residual block that can expand the receptive field used in the encoder, a high-efficiency attention module for refining the features extracted by the encoder, and a multi-scale fusion module that can make full use of feature information at different scales are utilized. The present invention is solved by using a FOCUS downsampling layer, a depth residual convolution block, an efficient attention mechanism, and a multi-scale feature fusion module. Among them, the FOCUS downsampling layer is set at the beginning of the decoding layer to reduce the number of parameters of the model on the premise of ensuring that information is not lost. The depth residual convolution block expands the receptive field and greatly compresses the parameters and computational amount through DW convolutions with convolution kernels of 3×3 and 5×5. The efficient attention mechanism can refine the feature map with a slight increase in the number of parameters and improve the performance of network change detection. The multi-scale feature fusion module only uses effective feature streams and can effectively fuse multi-scale features with a relatively low number of parameters.

[0048] The present invention conducts experiments on a change detection (CDD) dataset [1] containing various change types. In order to verify the effectiveness of the proposed LCDNet, the following thirteen advanced remote sensing image change detection methods are selected for comparison with the method of the present invention, and a brief introduction is made to them.

[0049] FC-EF (Fully Convolutional-Early Fusion)[1] is proposed based on the U-Net architecture, where the bi-temporal images are concatenated into multi-band images for input, and skip connections are used to gradually transfer multi-scale features from the encoder to the decoder to recover spatial information. FC-Siam-conc (Fully Convolutional-Siamese-Concatenation)[1], as a variant of the FC-EF model, uses a Siamese encoder to extract the features of bi-temporal images, and then concatenates the features of the same level from the encoder to the decoder. Different from FC-Siam-conc, FC-Siam-diff (Fully Convolutional-Siamese-Difference)[1] has another type of skip connection in the FC-EF model, which transfers the absolute difference between bi-temporal features. CDNet[2] is used for the research of street scene change detection. It consists of a contraction block and an expansion block, and obtains a change map through a softmax layer. DDCNN (Difference-enhancement Dense-attention Convolutional Neural Network)[3] simplifies UNet++. When fusing features, it combines the dense attention method, using high-level features to guide the selection of low-level features to retain the texture and detail information of the change area. DSIFN (Deeply Supervised Image Fusion Network)[4] uses channel attention and spatial attention to cross-utilize the feature maps obtained by the VGG16 pre-trained model multiple times at multiple scales for effective fusion to obtain a change map more accurately. SNUNet-CD (Siamese NestedUNet-Change Detection)[5] combines the Siamese network with the UNet++ network and uses the Ensemble Channel Attention Module (ECAM) to fuse the feature maps obtained from the backbone network at multiple semantic levels, thus suppressing localization errors and semantic voids. RDP-Net (Region Detail Preserving Network)[6] is a network based on ConvMixer, which can obtain good change detection performance with only a small number of parameters. In this method, a training method of learning detail information from easy to difficult and an edge loss that focuses on the network boundary details are proposed. LSNet (Lightweight Siamese network)[7] replaces the standard convolution with depthwise separable dilated convolution.LSNet_denseFPN uses the denseFPN (dense Feature Pyramid Network) proposed by SNUNet-CD to fuse multi-scale features. LSNet_diffFPN proposes diffFPN (difference Feature Pyramid Network) based on denseFPN. LSNet_diffFPN eliminates redundant dense connections and only retains effective feature streams during siamese feature fusion to compress parameters and computational complexity. STANet (Spatial–Temporal Attention neural Network)

[11] inputs the global features extracted by the ResNet18 network into the self-attention mechanism module and captures long-term spatio-temporal correlations to learn better representations. DASNet (DualAttentive fully convolutional Siamese Networks)

[12] applies the attention mechanism to the siamese network. BIT (Bitemporal Image Transformer)

[13] expresses the bitemporal image as several semantics and uses the transformer encoder to model the context in the spatio-temporal based on the compact semantics. The semantics are fed back to the pixel space for refining the original features through the transformer decoder.

[0050] Table Ⅰ shows the comparative experiments conducted on the CDD dataset. Precision (P), Recall (R), F1 Score (F1), and Intersection over Union (IoU) are used to quantitatively evaluate the performance of the methods involved. The number of parameters (Params), floating-point operations per second (FLOPs), and inference speed (Inference time) are used to measure the computational complexity and efficiency of the methods involved. The calculation of the precision, recall, F1 score, and intersection over union metrics is as follows:

[0051]

[0052]

[0053]

[0054]

[0055] Among them, true positive (TP) represents the number of unchanged pixels correctly detected, false positive (FP) represents the number of unchanged pixels not predicted, and false negative (FN) represents the number of changed pixels not predicted. Precision represents the probability that all detected pixels have changed. Recall represents the probability that all changed pixels are correctly detected. F1 is the harmonic mean of precision and recall, which can balance conflicts by considering both precision and recall simultaneously. IoU is the overlapping area between the predicted changed pixels and the changed pixels divided by the union area between them.

[0056] Table Ⅰ Comparative experiments conducted on the CDD dataset

[0057]

[0058]

[0059] It can be seen from the data in the above table that compared with other existing remote sensing image change detection methods, the proposed scheme of the present invention improves 0.56% in F1 and 1.01% in IoU on the CDD dataset. The model proposed by the present invention's scheme achieves state-of-the-art change detection performance with only 0.88MB of parameter quantity, 2.20GB of FLOPs, and an inference speed of 4.75ms. The present invention's scheme achieves the best performance on the CDD dataset and can identify the changed area with a lower parameter quantity and a faster speed.

[0060] To solve the problems of large computational complexity and slow inference speed existing in the prior art, the present invention constructs an end-to-end network architecture called the lightweight change detection network (LCDNet). LCDNet greatly compresses the parameters and computational complexity through the FOCUS downsampling module, depth residual module, and multi-scale feature fusion module. To realize the ability of the network to refine features and focus on the changed area, the present invention constructs an efficient attention module based on channel attention and spatial attention. To realize the effective fusion of feature maps at different scales, the present invention constructs a simple and effective multi-scale feature fusion module.

[0061] The proposed LCDNet of the present invention enables the model to achieve a faster inference speed on the premise of ensuring good change detection performance through the FOCUS module, depth residual block, and lightweight attention. The proposed multi-scale feature fusion module can make full use of the information of feature maps at each scale and only use effective feature streams when performing feature fusion, thereby more efficiently and quickly realizing the extraction of the changed area.

[0062] The references in the invention are as follows:

[0063] [1] Daudt R C, Le Saux B, Boulch A. Fully convolutional siamese networks for change detection[C] / / 2018 25th IEEE International Conference on Image Processing(ICIP), 2018: 4063 - 4067.

[0064] [2] Alcantarilla P F, Stent S, Ros G, et al. Street - view change detection with deconvolutional networks[J]. Autonomous Robots, 2018, 42(7): 1301 - 1322.

[0065] [3] Peng X, Zhong R, Li Z, et al. Optical remote sensing image change detection based on attention mechanism and image difference[J]. IEEE Transactions on Geoscience and Remote Sensing, 2020, 59(9): 7296 - 7307.

[0066] [4] Zhang C, Yue P, Tapete D, et al. A deeply supervised image fusion network for change detection in high resolution bi - temporal remote sensing images[J]. ISPRS Journal of Photogrammetry and Remote Sensing, 2020, 166: 183 - 200.

[0067] [5] Fang S, Li K, Shao J, et al. SNUNet - CD: A densely connected Siamese network for change detection of VHR images[J]. IEEE Geoscience and Remote Sensing Letters, 2021, 19: 1 - 5.

[0068] [6] Chen H, Pu F, Yang R, et al. RDP-Net: Region detail preserving network for change detection[J]. arXiv 2022, arXiv:2202.09745.

[0069] [7] Liu B, Chen H, Wang Z. LSNet: Extremely Light-weight siamese network for change detection in remote aensing image[J]. arXiv 2022, arXiv:2201.09156.

[0070] [8] Vaswani A, Shazeer N, Parmar N, et al. Attention is all you need[J]. Advances in neural information processing systems, 2017, 30.

[0071] [9] Dosovitskiy A, Beyer L, Kolesnikov A, et al. An image is worth 16x16 words: Transformers for image recognition at scale[J]. arXiv 2020, arXiv:2010.11929.

[0072]

[10] Liu Z, Lin Y, Cao Y, et al. Swin transformer: Hierarchical vision transformer using shifted windows[C] / / Proceedings of the IEEE / CVF International Conference on Computer Vision. 2021:10012-10022.

[0073]

[11] Chen H,Shi Z.A spatial-temporal attention-based method and a newdataset for remote sensing image change detection[J].Remote Sensing,2020,12(10):1662.

[0074]

[12] Chen J,Yuan Z,Peng J,et al.DASNet:Dual attentive fullyconvolutional siamese networks for change detection in high-resolutionsatellite images[J].IEEE Journal of Selected Topics in Applied EarthObservations and Remote Sensing,2020,14:1194-1206.

[0075]

[13] Chen H,Qi Z,Shi Z.Remote sensing image change detection withtransformers[J].IEEE Transactions on Geoscience and Remote Sensing,2021,60:1-14.

[0076]

[14] N.Bourdis,D.Marraud,and H.Sahbi,“Constrained optical flow foraerial image change detection,”in Proc.IEEE Int.Geosci.Remote Sens.Symp.,2011,4176–4179.

[0077] In the description of the present invention, it should be noted that for orientation terms, such as the terms "center", "horizontal", "vertical", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", etc., the indicated orientation and positional relationships are based on the orientation or positional relationships shown in the drawings. This is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and should not be construed as limiting the specific protection scope of the present invention.

[0078] It should be noted that the terms "first", "second", etc. in the description and claims of this application are used to distinguish similar objects and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances for the embodiments of the present application described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0079] Note that the above is only the preferred embodiment of the present invention and the application of technical principles. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and various obvious changes, re-adjustments, and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in more detail through the above embodiments, the present invention is not limited to the specific embodiments described herein. Without departing from the concept of the present invention, more other effective embodiments can also be included, and the scope of the present invention is determined by the scope of the appended claims.

Claims

1. A lightweight optical remote sensing image change detection method based on efficient attention, characterized in that: It includes the following steps: A. Orthorectify, image register, image stretch, and perform image numerical normalization preprocessing on the dual-temporal optical remote sensing images in sequence, so as to obtain dual-temporal optical remote sensing images with consistent data distribution; B. Label the updated parts in the remotely sensed images of the dual-temporal optical remote sensing images obtained in step A to obtain the corresponding change label map; C. Cut the change label map obtained in step B and the preprocessed dual-temporal optical remote sensing images obtained in step A with the same size to obtain training samples; D. Concatenate the dual-temporal remotely sensed images in the training samples obtained in step C; E. Downsample the concatenated image pairs through the FOCUS module, and then input the downsampled feature maps into a depth residual block composed of DW convolutions with a convolution kernel size of 3×3 and 5×5 for encoding, so as to extract feature maps related to the changed areas; F. Input the feature maps obtained in step E into an efficient attention module to refine the feature maps; G. Input the features at different levels obtained in step F into a multi-scale feature fusion module to obtain the final feature map X; H. Input the final fused feature X obtained in step G into a prediction head composed of 1×1 convolutions to obtain the predicted change map of the dual-temporal images of the training samples; I. Combine the binary cross-entropy loss and the Dice loss to form a hybrid loss function to calculate the loss between the predicted change map of the dual-temporal images of the training samples obtained in step H and the corresponding label map; J. After the training is completed, save the weight parameters and hyperparameter information of the trained change detection model; K. Perform orthorectification, image registration, image stretching, and image numerical normalization preprocessing on the pre- and post-temporal remotely sensed images to be detected in sequence, and then cut them with the same size to obtain samples to be detected; L. Input the samples to be detected into the change detection model obtained in step J, and calculate and output the predicted change map of the samples to be detected.

2. The lightweight optical remote sensing image change detection method based on efficient attention according to claim 1, wherein: The specific steps of step F are as follows: First, use the average pooling layer on the feature map F to generate aggregation vectors of size C×1×1 and 1×H×W, where C is the number of channels, and H and W are the height and width of the feature map; then apply one-dimensional convolution to the aggregation vector of C×1×1 to obtain the attention map in the channel dimension, and apply two-dimensional convolution to the aggregation vector of 1×H×W to obtain the attention map in the spatial dimension; expand the attention maps in the channel dimension and the spatial dimension to C×H×W, add them up and reassign weights to obtain the final attention map M(F); finally, multiply M(F) element-wise with the input feature map: M(F) = σ(C1D3(AvgPool(F)) + C2D7(AvgPool(F))), where F represents the input feature map, AvgPool(·) represents the average pooling operation, C1D3(·) represents one-dimensional convolution with a convolution kernel size of 3, C2D7(·) represents two-dimensional convolution with a convolution kernel size of 7, σ represents the sigmoid function, represents element-wise multiplication, and F' represents the refined feature map after weighting.

3. The lightweight optical remote sensing image change detection method based on efficient attention according to claim 1, characterized in that: The final fused feature X described in step G is obtained by upsampling and element-wise adding the four-scale feature maps X1, X2, X3, and X4 with different network depths. Among them, the four-scale feature maps are: X1 = C(F4), X2 = C(C(X1, F3)), X3 = C(C(X1, X2, F2)), X4 = C(X2, X3, F1), and the final fused feature X: X = X1 + X2 + X3 + X2. Here, the function C(·) represents the convolution operation using a convolution block composed of 1×1 convolution and 3×3 DW convolution. X1, X2, X3, and X4 respectively represent the feature maps obtained through four layers of depth residual blocks, and F1, F2, F3, and F4 represent the refined feature maps obtained after the feature maps X1, X2, X3, and X4 pass through the efficient attention module.

4. The lightweight optical remote sensing image change detection method based on efficient attention according to claim 1, wherein: The hybrid loss function described in Step I is \(L = L bce +L dice , where the binary cross-entropy loss Dice loss y i,j represents the probability that the pixel (i, j) in the corresponding label map is a changing pixel, represents the probability that the pixel (i, j) in the predicted change map is a changing pixel, and n and m respectively represent the width and height at the pixel level of the image.

Citation Information

Patent Citations

  • Remote sensing image change detection method based on twinborn multi-scale difference feature fusion

    CN113420662A

  • Remote sensing image typical surface feature extraction method based on multi-task attention mechanism

    CN114694031A