Remote sensing image change detection method and system based on multi-scale fusion lightweight network
Through multi-scale fusion lightweight network, the problems of insufficient feature expression capabilities and high computational complexity in remote sensing image change detection are solved, efficient multi-scale feature extraction and information fusion are achieved, and the accuracy and efficiency of change detection are improved, and suitable for disaster analysis and emergency response.
Patent Information
- Application Number
- CN202510575500.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2045-05-06
AI Technical Summary
The existing remote sensing image change detection methods lack feature expression capabilities in lightweight models, lack of multi-scale context fusion strategies, complex design of hybrid architecture and high training cost, making it difficult to effectively extract multi-scale features and realize information fusion while reducing the computational complexity.
Multi-scale fusion lightweight networks are adopted, including symmetrical encoding-decoding structural networks, lightweight global dynamic feature extraction modules, lightweight spatio-temporal local feature extraction modules and adaptive feature fusion modules. Through technical means such as large-core convolution decomposition, grouping hole convolution and parameterless attention mechanisms, channel attention mechanisms, etc., multi-directional capture of features, cross-scale semantic association, background interference suppression and multi-level information fusion are achieved.
While reducing the computational complexity, multi-scale features are effectively extracted and information fusion is realized, improving the characteristic discrimination and accuracy of change detection, and is suitable for tasks such as disaster analysis and emergency response.
Smart Images

Figure CN120472318A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer data processing, and in particular to a remote sensing image change detection method and system based on a multi-scale fusion lightweight network. Background Art
[0002] In recent years, the rapid development of remote sensing technology has greatly enhanced our ability to monitor land surface changes, particularly in areas such as urban planning, environmental monitoring, and disaster assessment. Change detection techniques based on multi-temporal remote sensing imagery effectively reveal changing characteristics of land use, vegetation cover, and other areas by comparing surface information over time. With advances in high-resolution remote sensing data and artificial intelligence, change detection technology is developing towards higher precision and intelligence, providing crucial data support.
[0003] Traditional change detection methods include those based on pixels, features, and objects. However, with the continuous development of deep learning technology, deep neural networks have demonstrated significant advantages in feature extraction, driving improvements in change detection performance. However, deep learning-based models still face several challenges: lightweight models sacrifice feature expressiveness for efficiency, multi-scale context fusion methods suffer from insufficient fusion strategies, and hybrid architectures and novel operators are complex to design and expensive to train. Therefore, effectively extracting multi-scale features and achieving information fusion while reducing computational complexity remains a key challenge in change detection. Summary of the Invention
[0004] One of the purposes of the present invention is to provide a remote sensing image change detection method and system based on a multi-scale fusion lightweight network to solve the problems in the background technology.
[0005] An embodiment of the present invention provides a remote sensing image change detection method based on a multi-scale fusion lightweight network, comprising:
[0006] Change detection in dual-temporal remote sensing images using a multi-scale fusion lightweight network.
[0007] Optional, multi-scale fusion lightweight network includes:
[0008] A symmetrical encoder-decoder structure network is used to extract branch features, encode and decode dual-temporal remote sensing images and generate change maps;
[0009] A lightweight global dynamic feature extraction module is used to perform the first processing of the change map based on the large kernel convolution decomposition strategy and the global enhanced multi-scale attention mechanism;
[0010] A lightweight spatiotemporal local feature extraction module for performing secondary processing on the change map based on a grouped dilated convolution strategy and a parameter-free spectral-spatial joint attention mechanism;
[0011] The adaptive feature fusion module is used to process the change graphs after the first processing and the second processing based on the dual pooling strategy and the channel attention mechanism to obtain the change detection results.
[0012] Optionally, when performing branch feature extraction in a symmetric encoder-decoder structure network, two 3×3 convolutions are used to extract shallow feature information.
[0013] During encoding, the symmetric branches use a cross combination of five main modules and four maximum pooling operations to perform downsampling operations to obtain encoding features of different scales, and calculate the absolute value difference between the upper and lower feature maps to obtain the change characteristics of the two branches in the encoding stage;
[0014] During decoding, the main module is used to refine features and calculate the absolute value difference between the upper and lower feature maps. The upsampling operation is used to restore the scale of the features. The high-level difference features are upsampled to the resolution of the previous layer and spliced with the difference features of the corresponding layer in the encoding stage along the channel dimension.
[0015] After four rounds of feature extraction, encoding, and decoding operations, a 1×1 convolution is used to operate on the change features of the last layer to obtain a change map.
[0016] Optionally, the lightweight global dynamic feature extraction module uses a large-core convolution decomposition strategy to process the feature map of the input change map. After channel division, it is divided into four groups of heterogeneous branches according to the proportional coefficient, respectively extracting local fine-grained context features, long-distance dependent horizontal features, vertical global correlation features, and original feature information.
[0017] When processing based on the global enhanced multi-scale attention mechanism, the four groups of branch features are initially fused through splicing. The left branch generates a semantic vector through global context compression, and the right branch captures context information through multi-scale convolution. Finally, the two are added together to obtain the final output change map after the first processing.
[0018] Optionally, the lightweight spatiotemporal local feature extraction module processes the feature map of the input change map based on the grouped dilated convolution strategy. The feature map is grouped by channel, and a dilated convolution with a different dilation rate is applied to each group to capture local detail information at multiple scales.
[0019] When processing based on the parameter-free spectral-spatial joint attention mechanism, the feature map is used as the key and value, and the channel mean is used as the query to calculate the generated attention; the weight is generated by Softmax and multiplied with the feature map to obtain the change map after the second processing.
[0020] Optionally, when the adaptive feature fusion module is processed based on the dual pooling strategy, channel stitching is performed on the input change maps after the first processing and the second processing, global and local features are obtained respectively through global average pooling and global maximum pooling, and feature stitching is performed;
[0021] When processing based on the channel attention mechanism, the spliced features are flattened and dimensionally adjusted, cross-channel interactive learning is performed through one-dimensional convolution, weights are generated through Softmax, and feature calibration is performed using broadcast multiplication to obtain change detection results.
[0022] Optionally, the loss function of the multi-scale fusion lightweight network adopts a binary cross entropy loss function.
[0023] Optionally, the remote sensing image change detection method based on a multi-scale fusion lightweight network also includes:
[0024] Visually output the change detection results; assist users in achieving their change detection goals based on the visually output change detection results.
[0025] Optionally, steps to assist the user include:
[0026] A second spatiotemporal graph for decision support based on a first spatiotemporal graph formed by the user's interaction with the change detection results of the visual output before the most recent trigger moment;
[0027] Based on the second spatiotemporal graph, the user is assisted to continue interacting with the change detection results outputted visually.
[0028] An embodiment of the present invention provides a remote sensing image change detection system based on a multi-scale fusion lightweight network, comprising:
[0029] The remote sensing image change detection module is used to perform change detection on dual-temporal remote sensing images using a multi-scale fusion lightweight network.
[0030] The present invention has achieved the following beneficial effects:
[0031] This paper uses a lightweight global dynamic feature extraction module combined with large kernel convolution decomposition and a multi-scale attention mechanism to capture multi-directional features and establish cross-scale semantic associations. The spatiotemporal local feature extraction module uses grouped dilated convolution and a parameter-free attention mechanism to suppress background interference and optimize local detail representation. The adaptive feature fusion module uses efficient channel attention to dynamically calibrate global and local features, achieving multi-level information adaptive fusion and improving feature discriminability. This effectively extracts multi-scale features and achieves information fusion while reducing computational complexity.
[0032] Other features and advantages of the present invention will be described in the following description, and in part will become apparent from the description, or will be understood by practicing the present invention. The purpose and other advantages of the present invention can be realized and obtained by the structures particularly pointed out in the written description and the accompanying drawings.
[0033] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:
[0035] Figure 1 This is the overall network architecture diagram of the present invention;
[0036] Figure 2 is a pair of remote sensing images at different times on the LEVIR-CD dataset;
[0037] Figure 3 This is the structural diagram of the lightweight global dynamic feature extraction module;
[0038] Figure 4 This is the structural diagram of the lightweight spatiotemporal local feature extraction module;
[0039] Figure 5 This is the structure diagram of the adaptive feature fusion module;
[0040] Figure 6 Schematic diagram for visual comparison of change maps obtained with other methods on the LEVIR-CD dataset;
[0041] Figure 7 Schematic diagram for visual comparison of change maps obtained with other methods on the CDD dataset;
[0042] Figure 8 Schematic diagram for visual comparison of different ablation settings in two datasets;
[0043] Figure 9 Schematic diagram of the visual comparison of different ablations of the LGDM module on two datasets.
[0044] Figure 10 Schematic diagram of the visual comparison of different ablations of the LSTLM module on two datasets. DETAILED DESCRIPTION
[0045] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.
[0046] Example 1:
[0047] The embodiment of the present invention provides a remote sensing image change detection method based on a multi-scale fusion lightweight network. Figure 1 As shown, including:
[0048] Use multi-scale fusion lightweight network to perform change detection on dual-temporal remote sensing images;
[0049] Among them, the multi-scale fusion lightweight network includes:
[0050] A symmetrical encoder-decoder structure network is used to extract branch features, encode and decode dual-temporal remote sensing images and generate change maps;
[0051] A lightweight global dynamic feature extraction module is used to perform the first processing of the change map based on the large kernel convolution decomposition strategy and the global enhanced multi-scale attention mechanism;
[0052] A lightweight spatiotemporal local feature extraction module for performing secondary processing on the change map based on a grouped dilated convolution strategy and a parameter-free spectral-spatial joint attention mechanism;
[0053] The adaptive feature fusion module is used to process the change graphs after the first processing and the second processing based on the dual pooling strategy and the channel attention mechanism to obtain the change detection results.
[0054] This paper uses a lightweight global dynamic feature extraction module combined with large kernel convolution decomposition and a multi-scale attention mechanism to capture multi-directional features and establish cross-scale semantic associations. The spatiotemporal local feature extraction module uses grouped dilated convolution and a parameter-free attention mechanism to suppress background interference and optimize local detail representation. The adaptive feature fusion module uses efficient channel attention to dynamically calibrate global and local features, achieving multi-level information adaptive fusion and improving feature discriminability. This effectively extracts multi-scale features and achieves information fusion while reducing computational complexity.
[0055] The present invention also has broad application prospects and is helpful for tasks such as disaster analysis and emergency response.
[0056] Example 2:
[0057] In this embodiment, when the symmetric encoding-decoding structure network performs branch feature extraction, two 3×3 convolutions are used to extract shallow feature information;
[0058] During encoding, the symmetric branches use a cross combination of five main modules and four maximum pooling operations to perform downsampling operations to obtain encoding features of different scales, and calculate the absolute value difference between the upper and lower feature maps to obtain the change characteristics of the two branches in the encoding stage;
[0059] During decoding, the main module is used to refine features and calculate the absolute value difference between the upper and lower feature maps. The upsampling operation is used to restore the scale of the features. The high-level difference features are upsampled to the resolution of the previous layer and spliced with the difference features of the corresponding layer in the encoding stage along the channel dimension.
[0060] After four rounds of feature extraction, encoding, and decoding operations, a 1×1 convolution is used to operate on the change features of the last layer to obtain a change map.
[0061] The symmetric encoder-decoder structure network is the first part of the model: for the input dual-phase images X1, X2 (such as Figure 2 As shown), two 3×3 convolutions are used to extract shallow feature information as input features in the encoding stage. In the encoding stage, the symmetric branches use a cross combination of 5 main modules and four maximum pooling to downsample the input shallow feature information to obtain encoding features of different scales, and calculate the absolute value difference between the upper and lower feature maps of the upper and lower branches to obtain the change features of the dual branches in the encoding stage. In the decoding stage, the main module is first used to refine the features to obtain refined features, and the absolute value difference between the upper and lower feature maps is calculated. Then, the upsampling operation is used to restore the scale of the final output features of the decoding stage. In this process, in order to suppress redundant noise and strengthen the areas of significant changes, the high-level difference features (the high-level difference features are derived from the absolute value differences calculated after each layer of the decoding stage processes the input feature map) are upsampled to the resolution of the previous layer, and then spliced with the difference features of the corresponding layer in the encoding stage along the channel dimension. After four times of the above upsampling and splicing operations, 1×1 convolution is used to operate on the change features of the last layer to obtain the final change map. The symmetric encoding-decoding structure network corresponds to Figure 1 (a) MSFLNet part in.
[0062] Example 3:
[0063] In this embodiment, the lightweight global dynamic feature extraction module processes the feature map of the input change map based on the large kernel convolution decomposition strategy. After the channel division operation, it is divided into four groups of heterogeneous branches according to the proportional coefficient, respectively extracting local fine-grained context features, long-distance dependent horizontal features, vertical global correlation features, and original feature information;
[0064] When processing based on the global enhanced multi-scale attention mechanism, the four groups of branch features are initially fused through splicing. The left branch generates a semantic vector through global context compression, and the right branch captures context information through multi-scale convolution. Finally, the two are added together to obtain the final output change map after the first processing.
[0065] like Figure 3As shown in the figure, the lightweight global dynamic feature extraction module (LGDM) is the second part of the model. First, it adopts a large kernel convolution decomposition strategy in the horizontal and vertical directions, significantly reducing computational complexity while effectively improving the ability to capture multi-directional features. Second, this module integrates a global enhanced multi-scale attention mechanism. Through adaptive global context compression and expanded convolution branches, it achieves cross-scale dynamic feature fusion, thereby accurately modeling the semantic associations of objects in any orientation.
[0066] Specifically, this module adopts the collaborative design of decomposed large kernel convolution and global enhanced multi-scale attention mechanism, and realizes the effective fusion of global semantic information through feature channel decoupling and dynamic weight allocation. Figure 1 In (b), the MSFLModule performs a 1×1 convolution on the input features to reduce the dimension of the features f', and then there are two branches, one named S1 and the other named S2, both of which have input f'. Input feature map R is a set of real numbers, C, H, and W represent the number of channels, height, and length of the feature respectively. First, the feature is decoupled by channel division operation, and the ratio coefficient α = 1 / 4 is used to divide it into four groups of heterogeneous branches I, II, III, and IV. The features of the four groups of branches are named S I 、S II 、S III 、S IV Among them, branch I uses k×k square depth-wise separable convolution to capture local fine-grained context features and obtains the output f I =DW k×k (S I ), branch II passes through 1×k h Convolution models long-distance dependencies in the horizontal direction to obtain features k h Represents a whole, 1×k h Represents the shape of the convolution kernel, which is used to extract horizontal features. Branch III uses k v ×1 convolution extracts global correlation features in the vertical direction k v Represents a whole, k v ×1 represents the shape of the convolution kernel, which is used to extract vertical features. Branch IV is used to preserve the original feature information and only perform identity mapping. IV Then, the four groups of branch features are preliminarily fused:
[0067] f new1 =[f I ,f II ,f III ,f IV ]
[0068] Among them, [,] represents the concatenation operation. Through the design of the above heterogeneous convolution kernel, LGDM not only makes the features complementary in space, but also greatly reduces the computational complexity. In order to further strengthen the global semantic association, the fusion feature f new1 For further processing, we design a dual-branch Global Enhanced Multi-scale Attention (GEMA) mechanism for collaborative optimization. In this mechanism, there are two branches, both of which are for f new1 The two branches are independent of each other and the input is f new1 , one on the left and one on the right. First, in the left branch, the fused features are globally compressed, and a compact semantic vector G1 is generated through adaptive average pooling (GAP) and 1×1 convolution:
[0069] G1=Conv 1×1 (GAP(f new1 ))
[0070] Among them, G1 encodes the prior information of the global context, Conv 1×1 Represents a 1×1 convolution. In the operation of the right branch, the input feature map f new1 The channel dimension is divided into g subgroups, and the number of channels in each subgroup is C / g. The variable g represents the total number of groups formed after the feature is divided. After that, a convolution with a dilation rate of 3×3 is used to expand the receptive field to capture multi-scale context information, resulting in feature G2, which can be represented by the following process:
[0071]
[0072] Dconv stands for dilated convolution, d stands for dilation rate, and 3×3 stands for dilation with a 3×3 convolution kernel. Subsequently, G2 is transformed into (Depend on Convert to Finally, the global context features G1 and G2' are added together to obtain the final output S1'.
[0073] Example 4:
[0074] In this embodiment, the lightweight spatiotemporal local feature extraction module processes the feature map of the input change map based on the grouped dilated convolution strategy, groups the feature map by channel, and applies dilated convolution with different dilation rates to each group to capture multi-scale local detail information;
[0075] When processing based on the parameter-free spectral-spatial joint attention mechanism, the feature map is used as the key and value, and the channel mean is used as the query to calculate the generated attention; the weight is generated by Softmax and multiplied with the feature map to obtain the change map after the second processing.
[0076] like Figure 4 As shown in the figure, the Lightweight Spatiotemporal Local Feature Extraction Module (LSTLM) is the third component of the model. It uses grouped dilated convolution and feature masking to effectively suppress background interference and enhance the representation of local details. This module innovatively introduces a parameter-free attention mechanism, modeling the joint spectral-spatial relationship through an energy function. It dynamically enhances the response of abnormal regions in multiple bands, achieving an optimal balance between spectral sensitivity and spatial detail without adding additional parameters.
[0077] Specifically, this module adopts the grouped hole convolution strategy, first inputting the feature map After grouping them by channel, each group independently applies a dilated convolution with different dilation rates (e.g., d = {1, 3, 6}) to capture multi-scale local detail information in a hierarchical manner. The above process can be expressed as:
[0078]
[0079] in, Represents the value of the output feature map at channel c, spatial position (h, w), Indicates the value of the input feature map at channel c, spatial position (h+id-1,w+jd-1) is a cyclic multi-scale convolution kernel, R represents a real number set, C represents the number of channels, k×k represents the size of the convolution kernel, i and j represent the spatial position of the convolution kernel, h and w represent the spatial position of the input feature map respectively. In addition, LSTLM also proposes a parameter-free spectral-space joint attention, which combines the feature map f new2 As the key (K) and value (V), and f new2 The channel dimension mean As a query (Q), calculate the generated Attention:
[0080]
[0081] Among them, d K Represents the dimension of the feature. Finally, the attention weight is generated by the Softmax function (Softmax is a function that converts the original score into a probability distribution, thereby indicating the relative importance of different input positions), and then multiplied by S2 to obtain S2' to achieve dynamic enhancement, focusing on amplifying the response of abnormal areas. The process can be expressed as:
[0082] S2′=S2⊙Softmax(Attention)
[0083] Among them, ⊙ is the dot product operation.
[0084] Example 5:
[0085] In this embodiment, when the adaptive feature fusion module is processed based on the dual pooling strategy, channel stitching is performed on the input change maps after the first processing and the second processing, and global and local features are obtained respectively through global average pooling and global maximum pooling, and feature stitching is performed;
[0086] When processing based on the channel attention mechanism, the spliced features are flattened and dimensionally adjusted, cross-channel interactive learning is performed through one-dimensional convolution, weights are generated through Softmax, and feature calibration is performed using broadcast multiplication to obtain change detection results.
[0087] like Figure 5 As shown in Figure 2, the Adaptive Feature Fusion Module (AFFM) is the fourth component of the model. This module leverages an efficient channel attention mechanism to dynamically calibrate the contributions of global and local features, enabling adaptive weighted fusion of multi-level information. This design effectively avoids feature conflicts and significantly improves feature discrimination.
[0088] Specifically, first, the module performs channel concatenation on the input features S1′ and S2′ to obtain S′. Then, dual-path global pooling is performed, that is, global average pooling (GAP) is used to capture the overall distribution characteristics of the features to obtain the feature S gap = GAP(S′); use global maximum pooling (GMP) to focus on significant local features and obtain feature S gmp =GMP(S′). Then, T can be obtained by concatenating the two vectors along the channel dimension:
[0089] T=[S gap ,S gmp ]
[0090] Then, AFFM flattens T and changes its dimension (through the above flattening operation, that is, adjusting the order of data), and obtains T′∈R 2×C Then, we use one-dimensional convolution to perform cross-channel interactive learning, and use the Softmax function to constrain the weights to the [0,1] interval to obtain M:
[0091] M = Softmax(Conv1d(T′)), where Conv1d represents one-dimensional convolution;
[0092] Finally, feature calibration is achieved through broadcast multiplication:
[0093] f c =S′⊙M ↑
[0094] Among them, M ↑ ∈R C×H×Wrepresents the expansion of M to the spatial dimension. This mechanism enhances the robustness of feature representation through a dual-pooling complementary strategy, while its lightweight design avoids the introduction of excessive parameters.
[0095] Example 6:
[0096] In this embodiment, the loss function of the multi-scale fusion lightweight network adopts a binary cross entropy loss function.
[0097] In the change graph generation stage, the binary cross-entropy loss function is used to optimize the network, which can be expressed as:
[0098]
[0099] Among them, N represents the total number of training samples (CDD and LEVIR datasets used for training), Y i represents the true value of the i-th sample, l CD is the binary cross entropy loss function, MSFLNet(X1,X2) represents the attached Figure 1 (a) In the MSFLNet part, X1 and X2 represent two data input to the network, namely images at different phases.
[0100] Example 7:
[0101] In this embodiment, the remote sensing image change detection method based on a multi-scale fusion lightweight network further includes:
[0102] Visually output change detection results; assist users in achieving their change detection goals based on the visually output change detection results;
[0103] The steps of assisting the user include:
[0104] A second spatiotemporal graph for decision support based on a first spatiotemporal graph formed by the user's interaction with the change detection results of the visual output before the most recent trigger moment;
[0105] Based on the second spatiotemporal graph, the user is assisted to continue interacting with the change detection results outputted visually.
[0106] When visually outputting the change detection results, the change detection results are output based on the preset visual output template.
[0107] When determining the trigger moment, obtain the first behavior of the user viewing the change detection result of the visual output at the new moment; if the previous trigger moment is not empty, use the association between the second behavior of the user viewing the change detection result of the visual output at the previous trigger moment and the first behavior, and the first behavior as the trigger factor; otherwise, only use the first behavior as the trigger factor; if the trigger factor matches the preset trigger indication, use the new moment as the trigger moment;
[0108] When deciding on the second space-time graph, call the preset initial space-time graph, map the space-time content from the first space-time graph to the initial space-time graph based on the space-time content mapping table, and use the initial space-time graph after the space-time content mapping is completed as the second space-time graph.
[0109] After change detection is complete, users may wish to achieve their change detection objectives based on the change detection results. For example, they may assess the impact of a disaster in a region based on change detection results from remote sensing images before and after the disaster. In this case, the change detection results are visualized using pre-set visualization output templates that define visualization methods (such as comparison charts and image parameter comparison tables). The system also assists users in viewing the visualized change detection results, helping them achieve their change detection objectives more quickly and improving their work efficiency.
[0110] Specifically, during assistance, a trigger moment is introduced. The user's interaction with the visually output change detection results before this moment serves as the basis for decision-making on assistance. Based on the first spatiotemporal graph (which records the time and viewing behavior of the user viewing different contents in the visually output change detection results before the trigger moment), the second spatiotemporal graph is determined. The second spatiotemporal graph can be used as a reference for the system to assist the user in continuing to interact with the visually output change detection results, helping them achieve their change detection goals as quickly as possible.
[0111] When determining the trigger moment, first determine the trigger factor and pre-set trigger indications that match different trigger factors. For example, if a certain association relationship is the same type of behavior and the first behavior is to start decision analysis, it means that the user has started decision analysis before and is currently starting other decision analysis, and assistance is needed, so a trigger indication is set for it; for example, if the first behavior is the end of decision analysis, the system can assist it in reviewing the decision and optimizing the decision analysis results, and a trigger indication is also set.
[0112] The spatiotemporal content mapping table contains spatiotemporal content and its mapping positions that assist users with different contents. For example, if a decision-making process behavior occurs at a certain spatiotemporal location in the first spatiotemporal diagram, then the corresponding spatiotemporal content is the decision-making knowledge corresponding to the decision-making process behavior (such as decision-making guidance content, etc.), and its mapping position is the spatial position corresponding to the spatiotemporal location on the initial spatiotemporal diagram, and its time dimension parameter is any future moment. When assisting users based on the second spatiotemporal diagram, when they review the spatial position at any future moment, the corresponding decision-making knowledge is output for their reference to optimize the decision-making results. This greatly improves the assistance effect and efficiency.
[0113] Example 8:
[0114] An embodiment of the present invention provides a remote sensing image change detection system based on a multi-scale fusion lightweight network, comprising:
[0115] The remote sensing image change detection module is used to perform change detection on dual-temporal remote sensing images using a multi-scale fusion lightweight network.
[0116] Example 9:
[0117] To effectively and systematically evaluate the proposed method, experimental comparisons and analyses were conducted on two public datasets: LEVIR-CD and CDD. The LEVIR-CD dataset consists of 637 pairs of ultra-high-resolution remote sensing images, each with a size of 1024×1024 pixels and a resolution of 0.5 meters. To maximize GPU memory utilization and avoid overfitting, the images were cropped to form 10,192 image patches of 256×256 pixels. Ultimately, the dataset was divided into three parts: 7,120, 1,024, and 2,048 image patches for training, validation, and testing, respectively. The CDD dataset contains 11 pairs of multispectral images acquired from Google Earth, documenting seasonal changes in the same area. The resolution ranges from 0.03 meters to 1 meter. The CDD dataset contains 16,000 images of 256×256 pixels, of which 10,000, 3,000, and 3,000 image patches are used for training, validation, and testing, respectively.
[0118] Regarding implementation details, MSFLNet uses the PyTorch framework and is trained on an NVIDIA GeForce RTX 4090 GPU. Using the Adam optimizer, we set the learning rate to 0.0001, the weight decay parameter to 0.0005, the number of epochs to 200, and the batch size to 16. We evaluated the performance of MSFLNet using precision, recall, F1 score, and the number of model parameters. Results for all compared methods are directly cited from their original papers.
[0119] To evaluate the performance of the method, precision, recall, and F1 score are used to quantitatively evaluate the method. Table 1 shows the performance comparison of different algorithms on the LEVIR-CD dataset.
[0120] Table 1: Performance comparison with state-of-the-art methods on the LEVIR-CD dataset
[0121]
[0122]
[0123] As can be seen from Table 1, the performance of the present invention is superior to other advanced methods. The reasons are as follows:
[0124] (1) Early fully convolutional Siamese networks achieved relatively basic recognition tasks through the interaction of shallow features, but their coarse feature modeling limited the accuracy improvement. With the introduction of the attention mechanism, semantic consistency was significantly enhanced by modeling spatiotemporal dependencies. However, the number of parameters of these methods increased significantly, limiting their efficiency in practical deployment.
[0125] (2) The lightweight model attempts to strike a balance between accuracy and computational cost, but in complex scenarios, it still faces the problem of insufficient multi-scale feature fusion.
[0126] (3) The hybrid structure-based methods achieve the complementarity of global and local features by coordinating convolutional neural networks and Transformers, but the significant increase in the complexity of these models places higher demands on computing resources. Unlike the above methods, our proposed MSFLNet shows significant advantages in terms of accuracy and parameter balance. The results show that MSFLNet achieves an F1 score of 91.39% with only 2.78M parameters. Compared with HANet of the same parameter scale, its performance is improved by 1.11%, and its parameter efficiency is 1.09 times that of the latter. It is worth noting that compared with the MambaBCD series of methods with more than 20M parameters, MSFLNet achieves higher accuracy while maintaining computational efficiency. At the same time, MSFLNet achieves a high balance between recall (91.75%) and precision (91.04%), which shows that it can effectively reduce pseudo-changes and missed detections.
[0127] like Figure 6 The visual comparison results shown also show that, first, MSFLNet uses the LGDM module to establish contextual associations between multi-level features, effectively avoiding false positives. Second, MSFLNet uses the LSTLM module to suppress background interference and enhance the effective representation of local details, thereby reducing false negatives. Finally, the introduction of AFFM allows MSFLNet to adaptively fuse multi-level information, improving overall discriminability. Therefore, compared to other methods, MSFLNet's transformation results are more accurate.
[0128] Table 2 gives the performance comparison of different algorithms on the CDD dataset.
[0129] Table 2: Performance comparison with state-of-the-art methods on the CDD dataset
[0130]
[0131]
[0132] From the experimental results in Table 2 and Figure 7 The visual comparison results clearly demonstrate the superior performance of our proposed method. Specifically, MSFLNet is able to extract more detailed change information from bi-temporal images, resulting in a better F1 score than other methods. As analyzed above, this performance improvement is attributed to the Local Global Difference Map (LGDM), Long Short-Term Temporal Learning Module (LSTLM), and Adaptive Feature Fusion Module (AFFM) introduced in our method.
[0133] In order to verify the effectiveness of each module in the model proposed in this invention, we conducted several sets of ablation experiments to evaluate the effectiveness of the components of our invention (MSFLNet). First, we removed the lightweight global dynamic feature extraction module (LGDM) alone, which limits the capture of global context information. Then, we removed the lightweight spatiotemporal local feature extraction module (LSTLM) module from the network, which limits the capture of local and spatial detail information. Finally, we also deleted the adaptive feature fusion module (AFFM) module and replaced it with a simple splicing operation, which ignores the effective role of feature adaptive fusion. By implementing the above settings, we can obtain three different network structures: MSFLNet without LGDM, MSFLNet without LSTLM, and MSFLNet without AFFM. In order to promote fair performance comparison, all networks are trained under the same parameter settings. The results of the ablation experiments on the two datasets are shown in Table 3, and the comparison results are shown in Table 3. Figure 8 The data in the table and the visual comparison results show the importance of LGDM, LSTLM and AFFM.
[0134] Table 3: Ablation comparison of different settings on two datasets
[0135]
[0136] In addition, we also tested the rationality of each component in the LGDM module. First, we removed the decomposed large kernel convolution from the LGDM module and replaced it with ordinary convolution only, which means that we no longer enhance multi-directional feature capture by decomposing large kernel convolution, which makes it difficult for the model to capture multi-directional feature information. Second, we removed the GEMA module from the LGDM module, which means that we will not perform cross-scale dynamic feature fusion, nor can we introduce adaptive global context branches to capture long-distance dependencies, which makes the model unable to better model global semantic information. We use Network1 to indicate the use of decomposed large kernel convolution, Network2 to indicate the use of GEMA, and Network3 to indicate the use of both. The results of the ablation experiments on the two datasets are shown in Table 4, and the comparison results are shown in Table 4. Figure 9 As shown in the figure, white represents true positives, black represents true negatives, red represents false positives, and blue represents false negatives. The data in the table and the visual comparison results show that decomposition of large kernel convolution and GEMA are indispensable for achieving fine change recognition at the pixel level.
[0137] Table 4: Comparison of ablation for the LGDM module
[0138]
[0139] Finally, we evaluate the effectiveness of cross-circular dilated convolution in the LSTLM module. We use Network1 to represent the use of ordinary convolution in the LSTLM module, and Network2 to represent the use of cross-circular dilated convolution. From Table 5 and Figure 10 The advantages of cross-circular dilated convolution can be clearly seen from the experimental results and visualization comparison. Figure 10 In the figure, white represents true positives, black represents true negatives, red represents false positives, and blue represents false negatives. That is, on both datasets, using a traditional CNN instead of cross-circular dilated convolutions reduced the F1 score by 1.57% and 2.29%, respectively. This demonstrates that using cross-circular dilated convolutions in the LSTLM module can obtain rich information from different receptive fields and effectively capture local features, which is particularly important for extracting subtle variations.
[0140] Table 5: Ablation comparison for LSTLM module
[0141]
[0142] In light of this, we propose a Multi-Scale Fusion Lightweight Network (MSFLNet) based on rich information extraction, aiming to address the limitations of existing change detection methods in global and local feature extraction, multi-scale information fusion, and computational complexity. First, to enhance the ability to capture multi-directional features and reduce computational complexity, we propose a lightweight global dynamic feature extraction module (LGDM). This module achieves cross-scale dynamic feature fusion through large kernel convolution decomposition and a globally enhanced multi-scale attention mechanism (GEMA), effectively modeling the semantic associations of objects in any orientation. Second, to suppress background interference and optimize spectral-spatial detail information, we design a lightweight spatiotemporal local feature extraction module (LSTLM). This module combines grouped atrous convolution with a parameter-free attention mechanism to significantly enhance the representation of local details. Finally, we propose an adaptive feature fusion module (AFFM). This module utilizes an efficient channel-wise attention mechanism to dynamically calibrate the contributions of global and local features, achieving adaptive weighted fusion of multi-level information, avoiding feature conflicts and improving model performance. Experiments on the CDD and LEVIR-CD datasets verify the effectiveness of MSFLNet in change detection tasks.
[0143] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.
Claims
1. A remote sensing image change detection method based on a multi-scale fusion lightweight network, characterized in that: include: Change detection in dual-temporal remote sensing images using a multi-scale fusion lightweight network.
2. The remote sensing image change detection method based on a multi-scale fusion lightweight network according to claim 1, characterized in that: The multi-scale fusion lightweight network includes: A symmetrical encoder-decoder structure network is used to extract branch features, encode and decode dual-temporal remote sensing images and generate change maps; A lightweight global dynamic feature extraction module is used to perform the first processing of the change map based on the large kernel convolution decomposition strategy and the global enhanced multi-scale attention mechanism; A lightweight spatiotemporal local feature extraction module for performing secondary processing on the change map based on a grouped dilated convolution strategy and a parameter-free spectral-spatial joint attention mechanism; The adaptive feature fusion module is used to process the change graphs after the first processing and the second processing based on the dual pooling strategy and the channel attention mechanism to obtain the change detection results.
3. The remote sensing image change detection method based on a multi-scale fusion lightweight network according to claim 2, characterized in that: When the symmetric encoder-decoder structure network performs branch feature extraction, two 3×3 convolutions are used to extract shallow feature information; During encoding, the symmetric branches use a cross combination of five main modules and four maximum pooling operations to perform downsampling operations to obtain encoding features of different scales, and calculate the absolute value difference between the upper and lower feature maps to obtain the change characteristics of the two branches in the encoding stage; During decoding, the main module is used to refine features and calculate the absolute value difference between the upper and lower feature maps. The upsampling operation is used to restore the scale of the features. The high-level difference features are upsampled to the resolution of the previous layer and spliced with the difference features of the corresponding layer in the encoding stage along the channel dimension. After four rounds of feature extraction, encoding, and decoding operations, a 1×1 convolution is used to operate on the change features of the last layer to obtain a change map.
4. The remote sensing image change detection method based on a multi-scale fusion lightweight network according to claim 2, characterized in that: The lightweight global dynamic feature extraction module uses a large-core convolution decomposition strategy to process the feature map of the input change map. After the channel division operation, it is divided into four groups of heterogeneous branches according to the proportional coefficient. These branches extract local fine-grained contextual features, long-distance dependent horizontal features, vertical global correlation features, and original feature information respectively. When processing based on the global enhanced multi-scale attention mechanism, the four groups of branch features are initially fused through splicing. The left branch generates a semantic vector through global context compression, and the right branch captures context information through multi-scale convolution. Finally, the two are added together to obtain the final output change map after the first processing.
5. The remote sensing image change detection method based on a multi-scale fusion lightweight network according to claim 2, characterized in that: The lightweight spatiotemporal local feature extraction module is based on the grouped dilated convolution strategy. The feature map of the input change map is grouped by channel, and dilated convolution with different dilation rates is applied to each group to capture multi-scale local detail information. When processing based on the parameter-free spectral-spatial joint attention mechanism, the feature map is used as the key and value, and the channel mean is used as the query to calculate the generated attention; the weight is generated by Softmax and multiplied with the feature map to obtain the change map after the second processing.
6. The remote sensing image change detection method based on a multi-scale fusion lightweight network according to claim 2, characterized in that: When the adaptive feature fusion module is based on the dual pooling strategy, it performs channel splicing on the input change maps after the first and second processing, obtains global and local features respectively through global average pooling and global maximum pooling, and performs feature splicing; When processing based on the channel attention mechanism, the spliced features are flattened and dimensionally adjusted, cross-channel interactive learning is performed through one-dimensional convolution, weights are generated through Softmax, and feature calibration is performed using broadcast multiplication to obtain change detection results.
7. The remote sensing image change detection method based on a multi-scale fusion lightweight network according to claim 1, characterized in that: The loss function of the multi-scale fusion lightweight network adopts the binary cross entropy loss function.
8. The remote sensing image change detection method based on a multi-scale fusion lightweight network according to claim 1, characterized in that: Also includes: Visual output of change detection results; Assist users to achieve their change detection goals based on the change detection results of visual output.
9. The remote sensing image change detection method based on a multi-scale fusion lightweight network according to claim 8, characterized in that: Steps to assist the user include: A second spatiotemporal graph for decision support based on a first spatiotemporal graph formed by the user's interaction with the change detection results of the visual output before the most recent trigger moment; Based on the second spatiotemporal graph, the user is assisted to continue interacting with the change detection results outputted visually.
10. A remote sensing image change detection system based on a multi-scale fusion lightweight network, characterized in that: include: The remote sensing image change detection module is used to perform change detection on dual-temporal remote sensing images using a multi-scale fusion lightweight network.
Citation Information
Patent Citations
Remote sensing image change detection method based on spatial-spectral feature fusion network
CN114359723A
Dual-branch multi-scale dynamic local convolution attention method based on remote sensing change detection
CN118736416A
Remote sensing data change detection system
CN119478691A
Lightweight remote sensing image change detection method based on binary neural network and large kernel stripe convolution
CN119672295A
Cited By
Transform-based lightweight remote sensing geographic image change detection method and system
CN121789066A