Lightweight remote sensing change detection method and system based on dual-temporal remote sensing images

By using lightweight feature extraction backbone RSShuffle, LightSEPP and Transformer encoder-decoder technologies in remote sensing change detection, the existing methods are solved by the problem of limited application and insufficient detection accuracy on devices with limited computing capabilities, and efficient and accurate remote sensing change detection is achieved.

CN118334532BActive Publication Date: 2025-05-13CHANGSHA UNIVERSITY OF SCIENCE AND TECHNOLOGY

Patent Information

Application Number
CN202410501815.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-25
Publication Date
2025-05-13
Estimated Expiration
2044-04-25

AI Technical Summary

Technical Problem

The existing remote sensing change detection methods are limited in applications on devices with limited computing capabilities, and the detection accuracy needs to be improved.

Method used

A lightweight remote sensing change detection method based on dual-time image remote sensing images is designed, using technology such as feature extraction backbone RSShuffle, lightweight space exchange pyramid pooling LightSEPP and Transformer encoder-decoder to extract shallow and deep features, and generate change maps through prediction of the head.

Benefits of technology

This method realizes efficient remote sensing change detection on devices with limited computing capabilities, reducing the amount of parameters and calculation complexity, and improving the detection accuracy. It is suitable for terminal equipment such as drones and satellites.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118334532B_ABST
    Figure CN118334532B_ABST
Patent Text Reader

Abstract

The present invention discloses a lightweight remote sensing change detection method based on dual-time remote sensing images, including: building a remote sensing change detection network, inputting a data set with dual-time remote sensing graphics to train the remote sensing change detection network, inputting dual-time remote sensing graphics, and outputting a predicted change map of the predicted change area through the trained remote sensing change detection network, wherein the remote sensing change detection network includes a feature extraction backbone RSShuffle, LightSEPP, a Transformer encoder-decoder and a prediction head. The present invention solves the problem that the current RSCD method is limited in practical application on devices with limited computing power due to too many parameters and high complexity, and the detection accuracy needs to be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of remote sensing change detection, and in particular to a lightweight remote sensing change detection method and system based on dual-temporal remote sensing images. Background Art

[0002] Remote sensing change detection (RSCD) refers to the use of satellites, aircraft and other remote sensing platforms to obtain two or more remote sensing images, and analyze the changes in surface features by comparing these images. RSCD is widely used in urban planning, damage assessment, resource management and other fields. By analyzing the changes, it can provide information about land use, vegetation cover, urban expansion, natural disaster impacts, etc., to support decision-making and measures. Therefore, RSCD has attracted more and more attention and research from researchers around the world. Early RSCD relied on hand-crafted features to obtain change results. However, due to the limited representation ability of hand-crafted features and the time-consuming nature of the process, this method is usually only used in simple scenes. Traditional RSCD mainly uses the following six methods: layer arithmetic, post-classification change, direct classification, transformation, change vector analysis (CVA) and hybrid change detection. Among them, layer arithmetic obtains the result of change by directly comparing the features of remote sensing images numerically. This method is simple to implement, but it is difficult to identify multiple changes. Post-classification changes do not require prior radiometric calibration to generate labeled maps to obtain the change process, but any errors in the input map will be directly converted into change maps. Direct classification classifies multi-temporal remote sensing data through a classification stage and directly generates labeled change maps, but due to the complexity of image time series construction, the training data set of this method is difficult to construct. The transformation method finds change features by emphasizing the differences between remote sensing images, but this method is challenging when locating multiple changes. CVA gives the size and position of the change by calculating the difference vector between units, but this method is prone to produce the same change direction and magnitude in different change themes. Hybrid change detection finds change features through two steps: change location and change identification. Due to the limitations of sensor technology, the resolution of early remote sensing images is usually low, and traditional RSCD can meet the needs. However, with the advancement of sensor technology, high-resolution remote sensing images have become easier to obtain. For high-resolution remote sensing images, the texture features of the ground are richer. However, the traditional RSCD method obtains change features through manual feature selection, which is limited by manually set parameters and the understanding of data features, and often performs poorly in high-resolution remote sensing images.

[0003] In recent years, with the continuous development of deep learning technology, many researchers have applied deep learning technology to RSCD in order to obtain better performance on various public datasets. They choose to add complex feature enhancement modules (self-attention mechanism, pyramid pooling, etc.) to the network. Thanks to the powerful feature extraction capabilities of convolutional neural networks (CNN) and Transformer, better results have been achieved than traditional methods. Through deep learning technology, richer and more abstract feature representations can be learned, thereby improving the performance of change detection. There are two main methods for RSCD based on deep learning: two-stage solution and single-stage solution. The two-stage solution first extracts features from the two input images, and then determines the changed area by comparing the feature maps of two time points based on the feature extraction. This method clearly defines the feature extraction stage, which can better capture the relevant information of the image, but requires additional steps to extract features, which increases the computational complexity. The single-stage solution combines feature extraction and change detection into one stage. This method directly integrates the bi-temporal feature information to generate change results. Since one feature extraction stage is reduced, this method usually requires more training data. Therefore, although the RSCD method based on deep learning has improved the effect of the model to a certain extent, the designed network is usually complex in structure and has many parameters, which will greatly increase the complexity and number of parameters of the model, thus hindering the deployment of the change detection model on terminal devices such as drones and satellites; and the training and inference process requires a lot of computing resources, which is very time-consuming; making it necessary to develop a lightweight RSCD model. In addition, most RSCD networks only focus on the features of each temporal state in the dual temporal state, while ignoring the interaction between the two temporal features, which makes the detection accuracy need to be improved. Therefore, in view of the limited computing power of terminal devices such as drones and satellites, it is necessary to design a lightweight detection algorithm with low computing cost, few model parameters and high detection accuracy. Summary of the invention

[0004] 1. Technical issues to be resolved

[0005] Based on the above problems, the present invention provides a lightweight detection method for remote sensing change detection, which solves the problem that the current RSCD method is limited in practical application on devices with limited computing power due to too many parameters and high complexity, and the detection accuracy needs to be improved.

[0006] (II) Technical solution

[0007] Based on the above technical problems, the present invention provides a lightweight remote sensing change detection method for dual-temporal remote sensing images, comprising:

[0008] S1. Build a remote sensing change detection network;

[0009] S2, inputting a data set with dual-time remote sensing graphics to train the remote sensing change detection network;

[0010] S3, inputting a dual-time remote sensing image, and outputting a predicted change map of the predicted change area through the trained remote sensing change detection network; the remote sensing change detection network includes:

[0011] S11, extracting shallow bi-temporal features through the feature extraction backbone RSShuffle;

[0012] S12, enhance the feature extraction capability through LightSEPP and exchange the bi-temporal features;

[0013] S13, extract deep information through Transformer encoder-decoder and decode it into change area mask;

[0014] S14. Generate a predicted change graph by predicting the head.

[0015] Furthermore, the feature extraction backbone RSShuffle includes: subjecting the input feature map to a convolution with a convolution kernel of 3×3, three downsamplings, a convolution with a convolution kernel of 1×1, an attention mechanism ShuffleAttention, two upsamplings, and connecting the output of the second downsampling with the output residual of the first upsampling after passing through the RSConv module, and connecting the output of the first downsampling with the output residual of the second upsampling after passing through the RSConv module, and then outputting, and the multiples of upsampling and downsampling are both 2.

[0016] Furthermore, the RSConv includes two branches, the first branch performs channel-by-channel convolution and point-by-point convolution in sequence, the second branch performs standard convolution with a convolution kernel of 1×1, and finally the two branches are connected in a residual manner.

[0017] Furthermore, the LightSEPP includes: inputting feature maps of two temporal states; first performing a channel separation operation on the input feature maps, splitting a feature x on the channel dimension into four features x1, x2, x3, and x4; performing maximum pooling on the four features, with convolution kernels of 3×3, 5×5, 7×7, and 9×9 respectively, and splicing the obtained features together; convolving the spliced ​​features with a convolution kernel of 1×1, and then adding the feature values ​​of feature x after convolution with a convolution kernel of 3×3; finally, performing channel mixing on the feature maps after the feature maps of the two temporal states are added.

[0018] Furthermore, the channel mixing includes: alternately dividing the feature maps of the two temporal states into a first half and a second half respectively in the channel dimension; alternately synthesizing the feature map of the first half of the first temporal state with the feature map of the second half of the second temporal state in the channel dimension to form a mixed feature map of the first temporal state and then outputting it; alternately synthesizing the feature map of the second half of the first temporal state with the feature map of the first half of the second temporal state in the channel dimension to form a mixed feature map of the second temporal state and then outputting it.

[0019] Furthermore, a Token generator is provided before the Transformer encoder-decoder, which is used to generate Token sets X1 and X2 of size L×C from the input bi-temporal feature maps F1 and F2 of size H×W×C, where H is the feature map height, W is the width, C is the number of channels, and L is the size of Tokens.

[0020] Furthermore, the Transformer encoder-decoder adopts BIT.

[0021] Furthermore, the prediction header includes: calculating a change probability map P:

[0022] P=Softmax(g(X))=Softmax(g(|X1-XD)

[0023] Among them, given two decoded feature maps X1 and X2, X is the absolute value of the subtraction of the two temporal feature maps X1 and X2, and g is the change classifier, which consists of two 3×3×2 convolutional layers.

[0024] The present invention also discloses a lightweight remote sensing change detection system based on dual-time remote sensing images, comprising:

[0025] at least one processor; and at least one memory in communication with the processor, wherein:

[0026] The memory stores program instructions that can be executed by the processor, and the processor can execute the method by calling the program instructions.

[0027] The present invention also discloses a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer instructions, wherein the computer instructions enable the computer to execute the method.

[0028] (III) Beneficial effects

[0029] The above technical solution of the present invention has the following advantages:

[0030] (1) The present invention adopts a lightweight feature extraction backbone RSShuffle to extract shallow bi-temporal features, abandons complex structures, adds a lightweight attention mechanism ShuffleAttention to obtain more feature information, and adopts a lightweight convolution RSConv to improve the model feature extraction capability, so that a small number of parameters can be used to fully learn the bi-temporal remote sensing image features, thereby being suitable for remote sensing change detection on devices with limited computing power;

[0031] (2) The present invention uses lightweight spatial exchange pyramid pooling LightSEPP to achieve channel separation, splicing and feature addition of feature maps of each temporal state, realize the fusion of local and global features, reduce the number of parameters, and effectively integrate the dual-temporal features through channel mixing of dual-temporal features and deep learning of dual-temporal scale, space and channel features, thereby further applying it to remote sensing change detection on devices with limited computing power and improving detection accuracy;

[0032] (3) The present invention uses a token generator to generate high-level semantic tokens from the features extracted by the convolutional neural network backbone architecture, and uses a transformer for encoding and decoding to achieve global feature learning. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] The features and advantages of the present invention will be more clearly understood by referring to the accompanying drawings, which are schematic and should not be construed as limiting the present invention in any way. In the accompanying drawings:

[0034] Figure 1 This is an overall structural diagram of the remote sensing change detection network LightRSCDNet according to an embodiment of the present invention;

[0035] Figure 2 A network structure comparison diagram of RSShuffle, a feature extraction backbone of an embodiment of the present invention, and shuffleNetv2, a background technology;

[0036] Figure 3 A network structure diagram of RSConv in RSShuffle according to an embodiment of the present invention;

[0037] Figure 4 This is a network structure diagram of LightSEPP according to an embodiment of the present invention;

[0038] Figure 5 A schematic diagram of channel mixing in LightSEPP according to an embodiment of the present invention;

[0039] Figure 6 This is a schematic diagram of the location of a Token generator in an embodiment of the present invention;

[0040] Figure 7A network structure diagram of a Transformer encoder-decoder according to an embodiment of the present invention;

[0041] Figure 8 This is a visual comparison diagram of the effects of the LightRSCDNet network of an embodiment of the present invention and the background technology network. DETAILED DESCRIPTION

[0042] The specific implementation of the present invention is further described in detail below in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present invention, but are not intended to limit the scope of the present invention.

[0043] In recent years, more and more lightweight networks have been developed and widely used in tasks such as target detection, segmentation, and classification. These networks have excellent performance, and many RSCD researchers have extended them in the field of RSCD. Liu et al. designed a lightweight change detection network ELW_CDNet based on shufflenetv2 and separable self-attention. Li et al. designed a lightweight change detection network A2Net based on MobileNetv2 and neighbor aggregation module. Many researchers have also designed lightweight feature extraction modules and integrated them into existing encoders and decoders to achieve lightweight while improving model performance. Although the existing lightweight backbone network already has good performance, combined with the characteristics of RSCD, there is still a lot of room for improvement.

[0044] Embodiments of the present invention include:

[0045] S1. Build a remote sensing change detection network;

[0046] S2, inputting a data set with dual-time remote sensing graphics to train the remote sensing change detection network;

[0047] S3, inputting a dual-temporal remote sensing image, and outputting a predicted change map of the predicted change area through the trained remote sensing change detection network.

[0048] Among them, in step S1, the remote sensing change detection network is a lightweight network named LightRSCDNet. We designed a lightweight feature extraction backbone RSShuffle, which abandons the complex structure and uses a small number of parameters to fully learn the dual-temporal remote sensing image features. We also designed a lightweight spatial exchange pyramid pooling (LightSEPP) to effectively fuse dual-temporal features. Through multi-scale pooling and dual-temporal feature exchange operations, we deeply learned the scale, space and channel characteristics of dual-temporal images. The specific structure is as follows:

[0049] LightRSCDNet is a standard twin network structure. Its network structure is as follows Figure 1 As shown. LightRSCDNet consists of three parts: feature extraction, Transformer encoder-decoder, and prediction head. Feature extraction is used to extract shallow information, including a shared backbone network RSShuffle and a lightweight SPP network LightSEPP. RSShuffle is improved from shuffleNetv2 and is used to extract shallow dual-temporal features; LightSEPP is used to enhance feature extraction capabilities and exchange dual-temporal features so that each encoder contains dual-temporal features. The Transformer encoder-decoder is used to extract deep information and decode it into a change region mask, where the Token generator extracts the semantic tags required by the Transformer encoder from the shallow feature map information. The prediction head generates a predicted change map through the high-level semantic features of the Transformer decoder. Overall, LightRSCDNet integrates modules such as feature extraction, Transformer encoding and decoding, and change map prediction through a twin network structure to achieve a multi-level, global understanding and expression of temporal features.

[0050] S11, extract shallow bi-temporal features through the lightweight feature extraction backbone RSShuffle;

[0051] RSShuffle is improved from shuffleNetv2. ShuffleNetv2 is a lightweight network structure with excellent performance. RSShuffle network is obtained by improving ShuffleNetv2. Figure 2 This is a comparison chart of RSShuffle and ShuffleNetv2. Figure 2 Figure (a) is RSShuffle. Figure 2Figure (b) is ShuffleNetv2. ShuffleNetv2 has 7 stages, the first stage downsamples by 4, the second to fourth stages downsample by 2, the fifth stage is a 1×1 convolution, the sixth stage is a 7×7 global pooling, and the seventh stage is the FC layer for classification. In RSShuffle, the maximum pooling in the first stage and the global pooling in the sixth stage are canceled to reduce the loss of spatial details. In order to obtain more feature information, we add a lightweight attention mechanism ShuffleAttention after the fifth stage. For details of ShuffleAttention, see Q.-L.Zhang and Y.-B.Yang, "Sa-net: Shuffle attention for deep convolutional neural networks," in ICASSP 2021-2021IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2021: IEEE, pp. 2235-2239. ShuffleAttention efficiently combines the two attention mechanisms of spatial attention mechanism and channel attention mechanism under low complexity. We also constructed a lightweight convolution RSConv and added it to RSShuffle through residual connection to improve the model's feature extraction capability. Finally, through two upsamplings, i.e. bilinear interpolation, we obtained an output feature map with a downsampling factor of 4, avoiding a significant reduction in spatial information. RSShuffle has fewer parameters and stronger feature extraction capabilities than shuffleNet2.

[0052] Therefore, the network structure of RSShuffle is as follows Figure 2 As shown in Figure (a), the input feature map is sequentially subjected to a convolution with a convolution kernel of 3×3, three downsamplings, a convolution with a convolution kernel of 1×1, an attention mechanism ShuffleAttention, and two upsamplings. The output of the second downsampling is connected to the output residual of the first upsampling after passing through the RSConv module, and the output of the first downsampling is connected to the output residual of the second upsampling after passing through the RSConv module. The numbers of upsampling 6 and upsampling 7 in the figure indicate the number of stages, and the multiples of upsampling and downsampling are both 2.

[0053] The network structure of RSConv is as follows Figure 3 Input a F1×F2×C in feature map, and generate a F1×F2×C outThe feature map of its standard convolution parameters and calculation amount are P Conv and C Conv , where D k1 ×D k2 ×C out is the convolution and size.

[0054] P Conv =D k1 ·D k2 ·C out ·C in

[0055] C Conv =F1·F2·D k1 ·D k2 ·C out ·C in

[0056] RSConv consists of two branches. The first branch first performs channel-by-channel convolution and point-by-point convolution, and the second branch performs 1×1 standard convolution. Finally, the two branches are connected through the residual method. The number of parameters and the amount of calculation of RSConv are P RSConv and C RSConv .

[0057] P RSConv =D k1 ·D k2 ·C in +2C out ·C in

[0058] C RSConv =F1·F2·D k1 ·D k2 ·C in +2F1·F2·C out ·C in

[0059] According to the above formula, using RSConv instead of 3×3 ordinary convolution in the second stage of RSShuffle reduces the number of parameters by 4-5 times, the amount of calculation is also reduced by 4-5 times, and the accuracy remains basically unchanged.

[0060]

[0061]

[0062] S12, enhance the feature extraction capability through LightSEPP and exchange the bi-temporal features;

[0063] The main function of LightSEPP is to achieve the interaction of channel features while realizing the fusion of local features and global features at the feature map level. The network structure of LightSEPP is as follows Figure 4 As shown in the figure, the input is a feature map of two temporal states. In order to reduce the amount of calculation, a channel separation operation is first performed on the input feature map, and a feature x in the channel dimension is split into four features x1, x2, x3, and x4; then the four features are subjected to maximum pooling, and the convolution kernels are 3×3, 5×5, 7×7, and 9×9 respectively, and the obtained features are spliced ​​together; the spliced ​​features are convolved with a convolution kernel of 1×1, and then added to the feature value of feature x after convolution with a convolution kernel of 3×3; finally, the feature maps of the two temporal states are added and the feature maps are subjected to a channel mixing.

[0064] The structure of channel mixing is as follows Figure 5 As shown in the figure, half of the bi-temporal features are exchanged in channel mixing. This module alternately divides the feature maps of the two temporal states into the first half and the second half in the channel dimension; the feature map of the first half of the first temporal state and the feature map of the second half of the second temporal state are alternately synthesized into a mixed feature map of the first temporal state in the channel dimension and then output; the feature map of the second half of the first temporal state and the feature map of the first half of the second temporal state are alternately synthesized into a mixed feature map of the second temporal state in the channel dimension and then output. Half of the input bi-temporal features are alternately exchanged in the channel dimension to interact the channel information.

[0065] In general, LightSEPP obtains multiple branches by means of channel separation. The branches are pooled with different convolution kernels and connected in parallel with the original features, thereby reducing the number of parameters. Finally, the spatial features of the two temporal states are extracted in the shallow layer and semi-exchanged to fuse the dual temporal features. Through operations such as channel separation, pooling, convolution, feature addition, and channel interaction, the fusion of local and global features is achieved, and the interaction between channel features is promoted. This helps to improve the model's perception and understanding of temporal features.

[0066] S13, extract deep information through Transformer encoder-decoder and decode it into change area mask;

[0067] Before the Transformer encoder-decoder, the shallow feature information obtained from feature extraction is mapped to the Token set required by the Transformer through the Token generator, such as Figure 6As shown. Let F1 and F2 be the input bi-temporal feature maps, with a size of H×W×C, where H is the feature map height, W is the width, and C is the number of channels. The token sets X1 and X2 are obtained through the Token generator, with a size of L×C, where L is the size of tokens. Subsequently, the obtained Token sets are input into the Transformer encoder and decoder.

[0068] Transformer encoder and decoder structure is as follows Figure 7 As shown. The Transformer encoder and decoder here use BIT, see H.Chen, Z.Qi, and Z.Shi, "Remote sensing image change detection with transformers," IEEE Transactions on Geoscience and RemoteSensing, vol.60, pp.1-14, 2021. The self-attention mechanism is mainly used. In the encoder work, it first receives the input sequence and maps it into a vector representation of a series of Tokens. This mapping process uses a multi-head self-attention mechanism, which allows the model to focus on different parts of other Tokens in the input sequence when processing each Token, which helps to capture global information and establish contextual relevance. At the same time, in feature extraction and abstraction, multiple attention heads work in parallel, and each head learns a different focus direction. This enables the model to better understand the correlation between different positions in the input sequence, thereby extracting more expressive features.

[0069] The Transformer decoder also generates the target sequence and receives the output of the encoder. Through the attention mechanism, it associates the generated sequence with the output of the encoder to better generate the next token. At the same time, the decoder also combines the output of the encoder to integrate the timing information. This helps the model understand the timing relationship of the input sequence, especially when dealing with sequence generation tasks.

[0070] The Transformer encoder and decoder use the self-attention mechanism to capture the global relationship between the input sequence and the generated sequence, making the model perform well in sequence processing tasks. Compared with the traditional recurrent neural network (RNN), this structure effectively avoids the long dependency problem and improves the expressiveness and adaptability of the model.

[0071] S14. Generate a predicted change graph by predicting the head.

[0072] The function of the prediction head is to generate a change map. Given two decoded feature maps X1 and X2, the prediction head obtains the change probability map P through the following formula. Among them, g is a change classifier, which consists of two 3×3×2 convolutional layers. In particular, X is the absolute value of the subtraction of the two temporal feature maps X1 and X2. The calculation of the change probability map P is implemented by the Softmax function. Specifically, its formula is as follows:

[0073] P=Softmax(g(X))=Softmax(g(|X1-X2|)).

[0074] This embodiment selects change detection datasets to verify the performance of LightRSCDNet. These datasets include LEVIR building Change Detection dataset (LEVIR-CD), Wuhan University BuildingCD (WHU-CD), Deeply supervised image fusion network CD (DSIFN-CD), etc. Due to the differences in the sizes of these datasets, we uniformly crop them to 256×256. For each dataset, we randomly divide the samples into three parts: training set, validation set, and test set. Among them, the training set is used for model fitting, the validation set is used to adjust the model during training, and the test set is used to evaluate the generalization performance of the model. In LEVIR-CD, the number of training sets, validation sets, and test sets is 7,120, 1,024, and 2,048, respectively. In order to optimize the training process, we selected stochastic gradient descent (SGD) as the training optimizer, where the momentum is set to 0.9 and the weight decay is set to 0.0005. The model was trained for 400 iterations on the LEVIR-CD dataset, with an initial learning rate set to 0.01 and a batch size of 8.

[0075] This example uses five accuracy indicators, namely precision, recall, intersection-over-union, overall accuracy, and F1 score, and three complexity indicators, namely the number of parameters, number of floating-point operations (FLOPs), and model size, to comprehensively evaluate the performance of the detection algorithm. The proposed LightRSCDNet network is compared with nine state-of-the-art RSCD methods, including three fully convolutional methods, FC-EF, FC-Siam-Di, and FC-Siam-Conc, and six transformer-based methods, namely DTCDSCN, STANet, IFNet, SNUNet, BIT, and VcT. Taking the WHU-CD dataset as an example, as shown in Table 1, for the accuracy index, the values ​​of LightRSCDNet in the five indicators of precision (P), recall (R), F1 score, intersection-over-union ratio, and overall accuracy are 91.60%, 89.16%, 90.36%, 82.42%, and 99.03%, respectively. Compared with the recent VcT algorithm, F1 score, intersection-over-union ratio, and overall accuracy are all advantageous. Compared with the BIT algorithm, our algorithm is superior to BIT in most indicators. Specifically, our algorithm is 2.36%, 1.05%, 1.74%, and 0.11% higher than the BIT algorithm in precision, F1 score, intersection-over-union ratio, and overall accuracy, respectively. In terms of complexity indicators, our number of parameters and floating-point operations are both optimal, which are 0.53M and 2.51G, respectively. In summary, compared with networks such as DTCDSCN, STANet, and SNUNet, our extracted network LightRSCDNet also has obvious advantages and superiority.

[0076] Table 1 Quantitative comparison with other SOTA models on the LEVIR-CD dataset

[0077]

[0078]

[0079] In summary, the above results fully verify that our algorithm LightRSCDNet has excellent change detection capabilities, and thanks to the low number of parameters and floating-point operations of the algorithm, it is possible to deploy it on satellites, aircraft and other equipment. The proposed LightRSCDNet network inference results are visualized and compared with the advanced network visualization differences, such as Figure 8 As shown, for better visualization, we use white, black, red, and green to represent true positives, true negatives, false positives, and false negatives, respectively.

[0080] Finally, it should be noted that the above method can be converted into software program instructions, which can be implemented by using a system including a processor and a memory, or by computer instructions stored in a non-transitory computer-readable storage medium. The above integrated unit implemented in the form of a software functional unit can be stored in a computer-readable storage medium. The above software functional unit is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to perform some steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, ROM), random access memory (Random Access Memory, RAM), disk or optical disk and other media that can store program codes.

[0081] In summary, the above-mentioned lightweight remote sensing change detection method and system based on dual-temporal remote sensing images have the following beneficial effects:

[0082] (1) The present invention adopts a lightweight feature extraction backbone RSShuffle to extract shallow bi-temporal features, abandons complex structures, adds a lightweight attention mechanism ShuffleAttention to obtain more feature information, and adopts a lightweight convolution RSConv to improve the model feature extraction capability, so that a small number of parameters can be used to fully learn the bi-temporal remote sensing image features, thereby being suitable for remote sensing change detection on devices with limited computing power;

[0083] (2) The present invention uses lightweight spatial exchange pyramid pooling LightSEPP to achieve channel separation, splicing and feature addition of feature maps of each temporal state, realize the fusion of local and global features, reduce the number of parameters, and effectively integrate the dual-temporal features through channel mixing of dual-temporal features and deep learning of dual-temporal scale, space and channel features, thereby further applying it to remote sensing change detection on devices with limited computing power and improving detection accuracy;

[0084] (3) The present invention uses a token generator to generate high-level semantic tokens from the features extracted by the convolutional neural network backbone architecture, and uses a transformer for encoding and decoding to achieve global feature learning.

[0085] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the embodiments of the present invention are described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations shall fall within the scope defined by the appended claims.

Claims

1. A lightweight remote sensing change detection method based on dual-temporal remote sensing images, characterized in that: include: S1. Build a remote sensing change detection network; S2, inputting a data set with dual-time remote sensing graphics to train the remote sensing change detection network; S3, inputting a dual-temporal remote sensing image, and outputting a predicted change map of a predicted change area through the trained remote sensing change detection network; The remote sensing change detection network includes: S11. Extract shallow bi-temporal features through the feature extraction backbone RSShuffle; the feature extraction backbone RSShuffle includes: subjecting the input feature map to a convolution with a convolution kernel of 3×3, three downsamplings, a convolution with a convolution kernel of 1×1, an attention mechanism ShuffleAttention, two upsamplings, and connecting the output of the second downsampling to the output residual of the first upsampling after passing through the RSConv module, and connecting the output of the first downsampling to the output residual of the second upsampling after passing through the RSConv module, and then outputting, and the multiples of upsampling and downsampling are both 2; the RSConv includes two branches, the first branch performs channel-by-channel convolution and point-by-point convolution in sequence, and the second branch performs a standard convolution with a convolution kernel of 1×1, and finally the two branches are connected in a residual manner; S12. Enhance the feature extraction capability through LightSEPP and exchange the dual-temporal features; the LightSEPP includes: inputting a feature map of two temporal states; first performing a channel separation operation on the input feature map, splitting a feature x in the channel dimension into four features x1, x2, x3, and x4; performing maximum pooling on the four features, with convolution kernels of 3×3, 5×5, 7×7, and 9×9, respectively, and splicing the obtained features together; after the spliced ​​features are convolved with a convolution kernel of 1×1, they are added to the feature values ​​of the feature x after convolution with a convolution kernel of 3×3; finally, the feature maps of the two temporal states are added and the feature maps are channel mixed; S13, extract deep information through Transformer encoder-decoder and decode it into change area mask; S14. Generate a predicted change graph by predicting the head.

2. The lightweight remote sensing change detection method based on dual-temporal remote sensing images according to claim 1 is characterized in that: The channel mixing includes: alternately dividing the feature maps of the two temporal states into a first half and a second half in the channel dimension respectively; alternately synthesizing the feature map of the first half of the first temporal state with the feature map of the second half of the second temporal state in the channel dimension into a mixed feature map of the first temporal state and then outputting it; alternately synthesizing the feature map of the second half of the first temporal state with the feature map of the first half of the second temporal state in the channel dimension and then outputting it.

3. The lightweight remote sensing change detection method based on dual-temporal remote sensing images according to claim 1 is characterized in that: Before the Transformer encoder-decoder, a Token generator is also provided to convert the input size into Bi-temporal feature map of , The generated size is Token Set and , is the feature map height, is the width, is the number of channels, The size of Tokens.

4. The lightweight remote sensing change detection method based on dual-temporal remote sensing images according to claim 1, characterized in that: The Transformer encoder-decoder adopts BIT.

5. The lightweight remote sensing change detection method based on dual-temporal remote sensing images according to claim 1, characterized in that: The prediction header includes: calculating the change probability map P: Among them, given two decoded feature maps , , Two temporal feature maps , The absolute value of the subtraction, It is a change classifier, consisting of two 3×3×2 convolutional layers.

6. A lightweight remote sensing change detection system based on dual-temporal remote sensing images, characterized in that: include: at least one processor; and at least one memory in communication with the processor, wherein: The memory stores program instructions executable by the processor, and the processor can execute the method according to any one of claims 1 to 5 by calling the program instructions.

7. A non-transitory computer-readable storage medium, characterized in that: The non-transitory computer-readable storage medium stores computer instructions, which cause the computer to execute the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Remote sensing image change detection method and device, computer equipment and storage medium

    CN114022788A

  • A lightweight neural network-based remote sensing change detection method with multi-feature aggregation

    CN114937204A

Cited By

  • Remote sensing image change detection method and system and electronic equipment

    CN121121491A