A method and system for detecting changes in remote sensing images based on multi-scale fusion lightweight networks
By using a multi-scale fusion lightweight network, the problems of insufficient feature representation ability and high computational complexity in remote sensing image change detection are solved, achieving efficient multi-scale feature extraction and information fusion, and improving the accuracy and efficiency of change detection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2026-03-10
AI Technical Summary
Existing remote sensing image change detection methods suffer from insufficient feature representation capabilities in lightweight models, inadequate multi-scale context fusion strategies, complex hybrid architecture designs, and high training costs, making it difficult to effectively extract multi-scale features and achieve information fusion while reducing computational complexity.
A multi-scale fusion lightweight network is adopted, including a symmetric encoder-decoder structure network, a lightweight global dynamic feature extraction module, a lightweight spatiotemporal local feature extraction module, and an adaptive feature fusion module. Through techniques such as large kernel convolution decomposition, grouped dilated convolution and parameterless attention mechanism, and channel attention mechanism, multi-directional feature capture, cross-scale semantic association and information fusion are achieved.
While reducing computational complexity, it effectively extracts multi-scale features and achieves information fusion, improving the accuracy and efficiency of change detection. It can better capture multi-directional features and suppress background interference, optimize local detail representation, and achieve multi-level adaptive information fusion.
Smart Images

Figure CN120472318B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer data processing technology, and in particular to a method and system for detecting changes in remote sensing images based on a multi-scale fusion lightweight network. Background Technology
[0002] In recent years, the rapid development of remote sensing technology has greatly enhanced the ability to monitor land surface changes, particularly in fields such as urban planning, environmental monitoring, and disaster assessment. Change detection technology based on multi-temporal remote sensing imagery effectively reveals changes in land use, vegetation cover, and other characteristics by comparing land surface information from different periods. With advancements in high-resolution remote sensing data and artificial intelligence, change detection technology is developing towards higher precision and intelligence, providing crucial data support.
[0003] Traditional change detection methods include pixel-based, feature-based, and object-based approaches. However, with the continuous development of deep learning technology, deep neural networks have shown significant advantages in feature extraction, driving improvements in the performance of change detection tasks. Nevertheless, deep learning-based models still face several challenges: lightweight models sacrifice feature representation capabilities to improve efficiency; multi-scale context fusion methods suffer from insufficient fusion strategies; and hybrid architectures and novel operators are complex to design and have high training costs. Therefore, effectively extracting multi-scale features and achieving information fusion while reducing computational complexity remains a key challenge in change detection. Summary of the Invention
[0004] One of the objectives of this invention is to provide a method and system for detecting changes in remote sensing images based on a multi-scale fusion lightweight network, in order to solve the problems in the background art.
[0005] This invention provides a remote sensing image change detection method based on a multi-scale fusion lightweight network, comprising:
[0006] Change detection is performed on dual-temporal remote sensing images using a multi-scale fusion lightweight network.
[0007] Optional, multi-scale fusion lightweight networks include:
[0008] A symmetric encoder-decoder network is used to extract branch features, encode, and decode dual-temporal remote sensing images to generate change maps.
[0009] A lightweight global dynamic feature extraction module is used to perform the first processing on the change map based on a large kernel convolution decomposition strategy and a global enhanced multi-scale attention mechanism.
[0010] A lightweight spatiotemporal local feature extraction module is used to perform secondary processing on the change map based on a grouped dilated convolution strategy and a parameterless spectral-space joint attention mechanism.
[0011] The adaptive feature fusion module is used to process the change maps after the first and second processing based on the dual pooling strategy and channel attention mechanism to obtain change detection results.
[0012] Optionally, when performing branch feature extraction in a symmetric encoder-decoder network, two 3×3 convolutions are used to extract shallow feature information respectively;
[0013] During encoding, the symmetric branch uses a cross-combination of 5 main modules and 4 max pooling to perform downsampling operations, obtaining encoded features at different scales, and calculates the absolute difference between the upper and lower feature maps to obtain the change features of the two branches in the encoding stage.
[0014] During decoding, the main module is used to refine the features and calculate the absolute difference between the upper and lower feature maps. The upsampling operation is used to restore the scale of the features, upsampling the high-level difference features to the resolution of the previous level, and concatenating them with the difference features of the corresponding level in the encoding stage along the channel dimension.
[0015] After four feature extraction, encoding, and decoding operations, a 1×1 convolution is used to operate on the changing features of the last layer to obtain the change map.
[0016] Optionally, when the lightweight global dynamic feature extraction module processes the input feature map based on the large kernel convolution decomposition strategy, it divides the feature map of the input change map into four heterogeneous branches according to the scaling factor after channel partitioning operation, and extracts local fine-grained context features, long-distance dependent horizontal features, vertical global correlation features and original feature information respectively.
[0017] When processing based on the global enhanced multi-scale attention mechanism, the four branches of features are initially fused by concatenation. The left branch generates a semantic vector through global context compression, and the right branch captures contextual information through multi-scale convolution. Finally, the two are added together to obtain the final output, the transformation map after the first processing.
[0018] Optionally, when the lightweight spatiotemporal local feature extraction module processes the input feature map based on the grouped dilated convolution strategy, it groups the feature map of the change map by channel, and applies dilated convolution with different dilation rates to each group to capture multi-scale local detail information.
[0019] When processing based on the parameterless spectral-space joint attention mechanism, the feature map is used as the key and value, and the channel mean is used as the query to calculate and generate attention; weights are generated through Softmax and multiplied with the feature map to obtain the change map after the second processing.
[0020] Optionally, when the adaptive feature fusion module processes the input after the first and second processing of the change map, it performs channel stitching, obtains global and local features through global average pooling and global max pooling respectively, and then performs feature stitching.
[0021] When processing based on the channel attention mechanism, the spliced features are flattened and their dimensions are adjusted. Cross-channel interactive learning is performed through one-dimensional convolution, weights are generated through Softmax, and feature calibration is performed using broadcast multiplication to obtain change detection results.
[0022] Optionally, the loss function for the multi-scale fusion lightweight network can be the binary cross-entropy loss function.
[0023] Optionally, remote sensing image change detection methods based on multi-scale fusion lightweight networks also include:
[0024] Visualize the change detection results; assist users in achieving their change detection objectives based on the visualized change detection results.
[0025] Optional steps to assist the user include:
[0026] A first spatiotemporal graph is formed based on the interaction between the user and the change detection results of the visual output before the most recent trigger time; a second spatiotemporal graph is used for decision support.
[0027] Based on the second spatiotemporal graph, it assists users in continuing to interact with the change detection results of the visual output.
[0028] This invention provides a remote sensing image change detection system based on a multi-scale fusion lightweight network, comprising:
[0029] The remote sensing image change detection module is used to detect changes in dual-temporal remote sensing images using a multi-scale fusion lightweight network.
[0030] The present invention has achieved the following beneficial effects:
[0031] This invention employs a lightweight global dynamic feature extraction module combined with large kernel convolution decomposition and a multi-scale attention mechanism to capture multi-directional features and establish cross-scale semantic associations. A spatiotemporal local feature extraction module uses grouped dilated convolution and a parameter-free attention mechanism to suppress background interference and optimize local detail representation. An adaptive feature fusion module dynamically calibrates global and local features through efficient channel attention, achieving multi-level information adaptive fusion and improving feature discriminative power. This effectively extracts multi-scale features and achieves information fusion while reducing computational complexity.
[0032] Other features and advantages of the invention will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in the written description and the accompanying drawings.
[0033] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0034] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:
[0035] Figure 1 This is a diagram of the overall network architecture of the present invention;
[0036] Figure 2 Given a pair of remote sensing images from different times on the LEVIR-CD dataset;
[0037] Figure 3 Structure diagram of the lightweight global dynamic feature extraction module;
[0038] Figure 4 This is a structural diagram of the lightweight spatiotemporal local feature extraction module;
[0039] Figure 5 Here is a structural diagram of the adaptive feature fusion module;
[0040] Figure 6 A visual comparison of change maps obtained with other methods on the LEVIR-CD dataset;
[0041] Figure 7 A visual comparison diagram of change maps obtained with other methods on the CDD dataset;
[0042] Figure 8 A visual comparison diagram of different ablation settings across two datasets;
[0043] Figure 9 This is a visual comparison diagram of different ablation methods of the LGDM module on two datasets.
[0044] Figure 10 This is a visual comparison diagram of different ablation methods of the LSTLM module on two datasets. Detailed Implementation
[0045] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.
[0046] Example 1:
[0047] This invention provides a method for detecting changes in remote sensing images based on a multi-scale fusion lightweight network, such as... Figure 1 As shown, it includes:
[0048] A lightweight multi-scale fusion network is used to detect changes in dual-temporal remote sensing images.
[0049] The multi-scale fusion lightweight network includes:
[0050] A symmetric encoder-decoder network is used to extract branch features, encode, and decode dual-temporal remote sensing images to generate change maps.
[0051] A lightweight global dynamic feature extraction module is used to perform the first processing on the change map based on a large kernel convolution decomposition strategy and a global enhanced multi-scale attention mechanism.
[0052] A lightweight spatiotemporal local feature extraction module is used to perform secondary processing on the change map based on a grouped dilated convolution strategy and a parameterless spectral-space joint attention mechanism.
[0053] The adaptive feature fusion module is used to process the change maps after the first and second processing based on the dual pooling strategy and channel attention mechanism to obtain change detection results.
[0054] This invention employs a lightweight global dynamic feature extraction module combined with large kernel convolution decomposition and a multi-scale attention mechanism to capture multi-directional features and establish cross-scale semantic associations. A spatiotemporal local feature extraction module uses grouped dilated convolution and a parameter-free attention mechanism to suppress background interference and optimize local detail representation. An adaptive feature fusion module dynamically calibrates global and local features through efficient channel attention, achieving multi-level information adaptive fusion and improving feature discriminative power. This effectively extracts multi-scale features and achieves information fusion while reducing computational complexity.
[0055] This invention also has broad application prospects and can be helpful for tasks such as disaster analysis and emergency response.
[0056] Example 2:
[0057] In this embodiment, when the symmetric encoder-decoder network extracts branch features, it uses two 3×3 convolutions to extract shallow feature information.
[0058] During encoding, the symmetric branch uses a cross-combination of 5 main modules and 4 max pooling to perform downsampling operations, obtaining encoded features at different scales, and calculates the absolute difference between the upper and lower feature maps to obtain the change features of the two branches in the encoding stage.
[0059] During decoding, the main module is used to refine the features and calculate the absolute difference between the upper and lower feature maps. The upsampling operation is used to restore the scale of the features, upsampling the high-level difference features to the resolution of the previous level, and concatenating them with the difference features of the corresponding level in the encoding stage along the channel dimension.
[0060] After four feature extraction, encoding, and decoding operations, a 1×1 convolution is used to operate on the changing features of the last layer to obtain the change map.
[0061] The symmetric encoder-decoder network is the first part of the model: for the input biphase images X1, X2 (e.g. Figure 2 As shown, two 3×3 convolutions are used to extract shallow feature information, which serves as the input features for the encoding stage. In the encoding stage, the symmetric branch uses a cross-combination of five main modules and four max-pooling operations to downsample the input shallow feature information, obtaining encoded features at different scales. The absolute difference between the upper and lower feature maps of the upper and lower branches is calculated to obtain the variation features of the two branches in the encoding stage. In the decoding stage, the main module is first used to refine the features to obtain refined features, and the absolute difference between the upper and lower feature maps is calculated. Then, upsampling is used to restore the scale of the final output features of the decoding stage. During this process, to suppress redundant noise and enhance significantly changing regions, the high-level difference features (derived from the absolute difference calculated after each layer of the decoding stage processes the input feature map) are upsampled to the resolution of the previous layer, and then concatenated with the difference features of the corresponding layer in the encoding stage along the channel dimension. After four such upsampling and concatenation operations, a 1×1 convolution is used to operate on the variation features of the last layer to obtain the final variation map. The symmetric encoder-decoder network corresponds to... Figure 1 (a) MSFLNet part.
[0062] Example 3:
[0063] In this embodiment, when the lightweight global dynamic feature extraction module processes the feature map of the input change map based on the large kernel convolution decomposition strategy, it divides the feature map into four heterogeneous branches according to the channel partitioning operation, and extracts local fine-grained context features, long-distance dependent horizontal features, vertical global correlation features and original feature information respectively.
[0064] When processing based on the global enhanced multi-scale attention mechanism, the four branches of features are initially fused by concatenation. The left branch generates a semantic vector through global context compression, and the right branch captures contextual information through multi-scale convolution. Finally, the two are added together to obtain the final output, the transformation map after the first processing.
[0065] like Figure 3As shown, the Lightweight Global Dynamic Feature Extraction (LGDM) module is the second part of the model. First, it employs a large kernel convolution decomposition strategy in both the horizontal and vertical directions, which significantly reduces computational complexity while effectively improving the ability to capture features from multiple directions. Second, this module integrates a global enhanced multi-scale attention mechanism, achieving cross-scale dynamic feature fusion through adaptive global context compression and expanded convolution branches, thereby accurately modeling the semantic associations of targets in any direction.
[0066] Specifically, this module adopts a collaborative design of decomposing large kernel convolution and a global enhanced multi-scale attention mechanism, and achieves effective fusion of global semantic information through feature channel decoupling and dynamic weight allocation. Figure 1 In section (b), the MSFLModule performs 1×1 convolution on the input features to reduce the dimensionality of the resulting feature f'. Then, it branches into two branches, one named S1 and the other S2, both with f' as input. (Input feature map) R is the set of real numbers, and C, H, and W represent the number of channels, height, and length of the feature, respectively. First, feature decoupling is performed through channel partitioning, with a scaling factor α = 1 / 4, resulting in four heterogeneous branches: I, II, III, and IV. The features of these four branches are named S. I S II S III S IV Branch I employs a k×k square depthwise separable convolution to capture local fine-grained contextual features, yielding the output f. I =DW k×k (S I Branch II then passes through 1×k h Convolution models long-range dependencies in the horizontal direction to obtain features. k h Represents a whole, 1×k h This represents the shape of the convolution kernel, used to extract horizontal features. Branch III uses k... v ×1 convolution extracts global association features in the vertical direction. k v To represent a whole, k v ×1 represents the shape of the convolution kernel, used to extract features in the vertical direction. Branch IV, on the other hand, performs only identity mapping to preserve information from the original features, using feature f. IV This indicates that the four sets of branch features are then initially fused.
[0067] f new1 =[f I ,f II ,f III ,f IV ]
[0068] Here, [,] represents the concatenation operation. Through the design of the heterogeneous convolution kernels described above, LGDM not only enables spatial complementarity of features but also significantly reduces computational complexity. To further enhance global semantic association, the fused feature f... new1 Further processing involves designing a dual-branch Global Augmented Multi-Scale Attention (GEMA) mechanism for collaborative optimization. This mechanism has two branches, both of which optimize f. new1 The two branches are independent of each other, and the input is f. new1 One on the left, one on the right. First, in the left branch, global context compression is performed on the fused features, generating a compact semantic vector G1 through adaptive average pooling (GAP) and 1×1 convolution:
[0069] G1 = Conv 1×1 (GAP(f new1 ))
[0070] Among them, G1 encodes prior information about the global context, Conv 1×1 This represents a 1×1 convolution. In the right branch operation, the input feature map f is... new1 The channel dimension is divided into g subgroups, and the number of channels in each subgroup is C / g. The variable g represents the total number of groups formed after the features are divided. Then, a convolution with a dilation rate d of 3×3 is used to expand the receptive field to capture multi-scale contextual information, resulting in feature G2, which can be represented by the following process:
[0071]
[0072] Dconv represents dilated convolution, d represents the dilation rate, and 3×3 indicates dilation using a 3×3 convolution kernel. Subsequently, G2 is transformed into... (Depend on Transform into (Transformation is performed by merging dimensions). Finally, the global context features G1 and G2' are added together to obtain the final output S1'.
[0073] Example 4:
[0074] In this embodiment, when the lightweight spatiotemporal local feature extraction module processes the input feature map based on the grouped dilated convolution strategy, it groups the feature map of the change map by channel, and applies dilated convolution with different dilation rates to each group to capture multi-scale local detail information.
[0075] When processing based on the parameterless spectral-space joint attention mechanism, the feature map is used as the key and value, and the channel mean is used as the query to calculate and generate attention; weights are generated through Softmax and multiplied with the feature map to obtain the change map after the second processing.
[0076] like Figure 4 As shown, the Lightweight Spatiotemporal Local Feature Extraction Module (LSTLM) is the third part of the model. During local feature extraction, it employs grouped dilated convolution and feature masking mechanisms to effectively suppress background interference and enhance local detail representation. This module innovatively introduces a parameter-free attention mechanism, modeling the spectral-spatial joint relationship through an energy function to dynamically enhance the response of anomalous regions in multiple bands, achieving an optimized balance between spectral sensitivity and spatial detail without adding additional parameters.
[0077] Specifically, this module employs a grouped dilated convolution strategy, first processing the input feature map... After grouping them by channel, each group is independently subjected to dilated convolutions with different dilation rates (e.g., d = {1, 3, 6}) to capture multi-scale local detail information in a hierarchical manner. The above process can be represented as:
[0078]
[0079] in, This represents the numerical value of the output feature map at spatial location (h, w) in channel c. This represents the numerical value of the input feature map at spatial location (h+id-1, w+jd-1) in channel c. It is a recurrent multi-scale convolutional kernel, where R represents the set of real numbers, C represents the number of channels, k×k represents the kernel size, i and j represent the spatial positions of the kernel, and h and w represent the spatial positions of the input feature map, respectively. Furthermore, LSTLM proposes a parameterless spectral-space joint attention mechanism, which integrates the feature map f... new2 As key (K) and value (V), simultaneously f new2 channel dimension mean As a query (Q), the attention mechanism is calculated and generated:
[0080]
[0081] Where, d K The dimension representing the feature is used. Finally, attention weights are generated using the Softmax function (Softmax is a function that transforms the original scores into a probability distribution, thus representing the relative importance of different input positions), and then multiplied by S2 to obtain S2', achieving dynamic enhancement and focusing on amplifying the response in anomalous regions. This process can be described as:
[0082] S2′=S2⊙Softmax(Attention)
[0083] Here, ⊙ represents the dot product operation.
[0084] Example 5:
[0085] In this embodiment, when the adaptive feature fusion module processes the input after the first and second processing transformation maps using the dual pooling strategy, it performs channel stitching, obtains global and local features through global average pooling and global max pooling respectively, and then performs feature stitching.
[0086] When processing based on the channel attention mechanism, the spliced features are flattened and their dimensions are adjusted. Cross-channel interactive learning is performed through one-dimensional convolution, weights are generated through Softmax, and feature calibration is performed using broadcast multiplication to obtain change detection results.
[0087] like Figure 5 As shown, the Adaptive Feature Fusion (AFFM) module is the fourth part of the model. This module utilizes an efficient channel attention mechanism to dynamically calibrate the contributions of global and local features, achieving adaptive weighted fusion of multi-level information. This design effectively avoids feature conflicts and significantly improves the discriminative ability of features.
[0088] Specifically, firstly, the module concatenates the input features S1′ and S2′ to obtain S′. Then, it performs dual-path global pooling, that is, it uses global average pooling (GAP) to capture the overall distribution characteristics of the features, to obtain feature S. gap =GAP(S′); Global max pooling (GMP) is used to focus on salient local features to obtain feature S. gmp =GMP(S′). Then, by concatenating the two vectors along the channel dimension, we can obtain T:
[0089] T = [S gap ,S gmp ]
[0090] Subsequently, AFFM flattens T and changes its dimensions (by adjusting the data order through the flattening operation) to obtain T′∈R. 2×C Then, cross-channel interactive learning is performed through one-dimensional convolution, and the weights are constrained to the [0,1] interval using the Softmax function to obtain M:
[0091] M = Softmax(Conv1d(T′)), where Conv1d represents one-dimensional convolution;
[0092] Finally, feature calibration is achieved through broadcast multiplication:
[0093] f c =S′⊙M ↑
[0094] Among them, M ↑ ∈R C×H×WThis indicates that M is extended to the spatial dimension. This mechanism enhances the robustness of feature representation through a dual-pooling complementary strategy, while the lightweight design avoids introducing too many parameters.
[0095] Example 6:
[0096] In this embodiment, the loss function of the multi-scale fusion lightweight network adopts the binary cross-entropy loss function.
[0097] During the transformation graph generation stage, the network is optimized using the binary cross-entropy loss function, which can be expressed as:
[0098]
[0099] Where N represents the total number of training samples (CDD and LEVIR datasets used for training), and Y... i Let l represent the true value of the i-th sample. CD Let MSFLNet(X1,X2) be the binary cross-entropy loss function, representing the additional cross-entropy loss function. Figure 1 In the (a) MSFLNet part, X1 and X2 represent two data inputs to the network, namely images at different times.
[0100] Example 7:
[0101] In this embodiment, the remote sensing image change detection method based on a multi-scale fusion lightweight network further includes:
[0102] Visualize the change detection results; assist users in achieving their change detection objectives based on the visualized change detection results.
[0103] The steps to assist users include:
[0104] A first spatiotemporal graph is formed based on the interaction between the user and the change detection results of the visual output before the most recent trigger time; a second spatiotemporal graph is used for decision support.
[0105] Based on the second spatiotemporal graph, it assists users in continuing to interact with the change detection results of the visual output.
[0106] When visualizing the change detection results, the change detection results are output based on a preset visualization output template.
[0107] When determining the trigger time, obtain the first behavior of the change detection result of the user viewing the visualization output at the new time; if the previous trigger time is not empty, use the relationship between the second behavior and the first behavior of the change detection result of the user viewing the visualization output at the previous trigger time, and the first behavior as the trigger factor; otherwise, only the first behavior is used as the trigger factor; if the trigger factor matches a preset trigger indication, use the new time as the trigger time.
[0108] When deciding on the second spatiotemporal map, the preset initial spatiotemporal map is invoked. Based on the spatiotemporal content mapping table, spatiotemporal content is mapped from the first spatiotemporal map to the initial spatiotemporal map. The initial spatiotemporal map after the spatiotemporal content mapping is completed is used as the second spatiotemporal map.
[0109] After change detection is completed, users may want to use the results to achieve their desired outcome, such as assessing the impact of a disaster on a region based on changes in remote sensing images before and after the disaster. In this case, the change detection results are visualized using a pre-defined visualization template with predefined visualization methods (such as comparison charts or image parameter comparison tables). The system also provides assistance when users view the visualized change detection results, helping them achieve their change detection goals more quickly and improving their work efficiency.
[0110] Specifically, during assistance, a trigger point is introduced. The interactions between the user and the change detection results in the visual output prior to this trigger point serve as the basis for decision-making regarding assistance. A second spatiotemporal graph is then determined based on a first spatiotemporal graph (which records the time and viewing behavior of the user when viewing different content in the change detection results in the visual output before the trigger point). The second spatiotemporal graph can be used by the system to guide how to assist the user in continuing to interact with the change detection results in the visual output, helping them achieve their change detection objectives as quickly as possible.
[0111] When determining the trigger time, first determine the trigger factor and pre-set trigger indicators that match different trigger factors. For example, if a certain relationship is the same type of behavior and the first behavior is to start decision analysis, it means that the user started decision analysis before and is now starting other decision analysis, so assistance is needed and a trigger indicator is set for them. Another example is that if the first behavior is to end decision analysis, the system can assist them in reviewing the decision and optimizing the decision analysis results, and a trigger indicator is also set.
[0112] The spatiotemporal content mapping table contains different types of spatiotemporal content that provide assistance to users, along with their corresponding mapping locations. For example, if a decision-making process occurs at a certain spatiotemporal location in the first spatiotemporal map, the corresponding spatiotemporal content is the decision knowledge (such as decision guidance content) associated with that decision-making process. Its mapping location is the spatial position corresponding to that location on the initial spatiotemporal map, and its time dimension parameter is any future moment. Therefore, when assisting users based on the second spatiotemporal map, when the user revisits that spatial position at any future moment, the corresponding decision knowledge is output for their reference to optimize the decision-making result. This significantly improves the effectiveness and efficiency of the assistance.
[0113] Example 8:
[0114] This invention provides a remote sensing image change detection system based on a multi-scale fusion lightweight network, comprising:
[0115] The remote sensing image change detection module is used to detect changes in dual-temporal remote sensing images using a multi-scale fusion lightweight network.
[0116] Example 9:
[0117] To effectively and systematically evaluate the proposed method, experiments were conducted and compared on two public datasets, LEVIR-CD and CDD. The LEVIR-CD dataset consists of 637 pairs of ultra-high-resolution remote sensing images, each with a size of 1024×1024 pixels and a resolution of 0.5 meters. To maximize GPU memory utilization and avoid overfitting, the images were cropped to form 10192 image patches of 256×256 pixels each. Finally, the dataset was divided into three parts: 7120, 1024, and 2048 image patches for training, validation, and testing, respectively. The CDD dataset contains 11 pairs of multispectral images obtained from Google Earth, recording seasonal variations in the same region, with resolutions ranging from 0.03 meters to 1 meter. The CDD dataset contains 16,000 images of 256x256 pixels each, with 10,000, 3,000, and 3,000 image patches used for training, validation, and testing, respectively.
[0118] Regarding implementation details, MSFLNet uses the PyTorch framework and is trained on an NVIDIA GeForce RTX 4090 graphics processing unit. In training with the Adam optimizer, we set the learning rate to 0.0001, the weight decay parameter to 0.0005, the epoch to 200, and the batch size to 16. We evaluate the performance of MSFLNet using precision, recall, F1 score, and the number of model parameters. Results for all comparison methods are directly cited from their original papers.
[0119] To evaluate the performance of the method, precision, recall, and F1 score were used for quantitative evaluation. Table 1 shows the performance comparison of different algorithms on the LEVIR-CD dataset.
[0120] Table 1: Performance comparison with state-of-the-art methods on the LEVIR-CD dataset
[0121]
[0122]
[0123] As can be seen from Table 1, the performance of this invention is superior to other advanced methods. The reasons are as follows:
[0124] (1) Early fully convolutional Siamese networks achieved basic recognition tasks through the interaction of shallow features, but their coarse feature modeling limited accuracy improvement. With the introduction of attention mechanisms, semantic consistency was significantly enhanced by modeling spatiotemporal dependencies. However, the number of parameters in these methods increased dramatically, limiting their efficiency in practical deployments.
[0125] (2) Lightweight models attempt to balance accuracy and computational cost, but in complex scenarios, they still face the problem of insufficient multi-scale feature fusion.
[0126] (3) Hybrid-structure-based methods achieve complementarity of global and local features through collaborative convolutional neural networks and Transformers, but the significant increase in model complexity places higher demands on computational resources. Unlike the methods mentioned above, our proposed MSFLNet demonstrates a significant advantage in terms of accuracy and parameter balance. Results show that MSFLNet achieves an F1 score of 91.39% with only 2.78M parameters, a 1.11% performance improvement compared to HANet with the same parameter size, and its parameter efficiency is 1.09 times that of the latter. Notably, compared to the MambaBCD series of methods with over 20M parameters, MSFLNet achieves higher accuracy while maintaining computational efficiency. Furthermore, MSFLNet achieves a high balance between recall (91.75%) and precision (91.04%), indicating its ability to effectively reduce spurious changes and missed detections.
[0127] like Figure 6 The visualization comparison results also show that, firstly, MSFLNet utilizes the LGDM module to establish contextual associations for multi-level features, effectively avoiding false positives; secondly, MSFLNet suppresses background interference and enhances the effective representation of local details through the LSTLM module, thereby reducing false negatives; finally, due to the introduction of AFFM, MSFLNet can achieve adaptive fusion of multi-level information, improving overall discriminability. Therefore, compared with other methods, MSFLNet's results are more accurate.
[0128] Table 2 shows a performance comparison of different algorithms on the CDD dataset.
[0129] Table 2: Performance comparison with state-of-the-art methods on the CDD dataset
[0130]
[0131]
[0132] From the experimental results in Table 2 and Figure 7 The visualization comparison results clearly reveal the superior performance of the proposed method. Specifically, MSFLNet is able to extract more detailed variation information from bi-temporal images, thus achieving a better F1 score than other methods. As analyzed above, this performance improvement is attributed to the Local-Global Difference Mapping (LGDM), Long Short-Term Learning Module (LSTLM), and Adaptive Feature Fusion Module (AFFM) introduced in our method.
[0133] To verify the effectiveness of each module in the proposed model, we conducted several ablation experiments to evaluate the effectiveness of the components of our invention (MSFLNet). First, we removed the Lightweight Global Dynamic Feature Extraction (LGDM) module, which limits the capture of global contextual information. Then, we removed the Lightweight Spatiotemporal Local Feature Extraction (LSTLM) module from the network, which limits the capture of local and spatial detail information. Finally, we removed the Adaptive Feature Fusion (AFFM) module and replaced it with a simple concatenation operation, which ignores the effective role of adaptive feature fusion. By implementing the above settings, we can obtain three different network structures: MSFLNet without LGDM, MSFLNet without LSTLM, and MSFLNet without AFFM. To promote fair performance comparison, all networks were trained under the same parameter settings. The ablation experiment results on the two datasets are shown in Table 3, and the comparison results are as follows: Figure 8 As shown in the table, both the data and the visualization comparison results demonstrate the importance of LGDM, LSTLM, and AFFM.
[0134] Table 3: Ablation comparisons under different settings on the two datasets
[0135]
[0136] Furthermore, we tested the rationality of each component in the LGDM module. First, we removed the decomposed large kernel convolution from the LGDM module, replacing it with ordinary convolution. This means we no longer enhance multi-directional feature capture by decomposing the large kernel convolution, making it difficult for the model to capture multi-directional feature information. Second, we removed the GEMA module from the LGDM module. This means we do not perform cross-scale dynamic feature fusion, nor can we introduce adaptive global context branches to capture long-distance dependencies, resulting in the model's inability to better model global semantic information. We use Network1 to represent the use of decomposed large kernel convolution, Network2 to represent the use of GEMA, and Network3 to represent the use of both. The ablation experiment results on the two datasets are shown in Table 4, and the comparison results are as follows: Figure 9 As shown, white represents true positives, black represents true negatives, red represents false positives, and blue represents false negatives. Both the data in the table and the visualization comparison results demonstrate that decomposed large kernel convolution and GEMA are indispensable for achieving pixel-level fine-grained change recognition tasks.
[0137] Table 4: Ablation Comparison for LGDM Modules
[0138]
[0139] Finally, we evaluated the effectiveness of cross-cyclic dilated convolution in the LSTLM module. We use Network1 to represent the use of regular convolution in the LSTLM module and Network2 to represent the use of cross-cyclic dilated convolution. (See Table 5 and...) Figure 10 The advantages of cross-cyclic dilated convolution are clearly evident in the experimental results and visualization comparisons. Figure 10 In the diagram, white represents true positives, black represents true negatives, red represents false positives, and blue represents false negatives. That is, on both datasets, using a traditional CNN instead of cross-recurrent dilated convolutions reduced the F1 score by 1.57% and 2.29%, respectively. This indicates that using cross-recurrent dilated convolutions in the LSTLM module can obtain rich information from different receptive fields, effectively capturing local features, which is particularly important for extracting subtle variations.
[0140] Table 5: Ablation Comparison for LSTLM Modules
[0141]
[0142] In view of this, this invention proposes a lightweight multi-scale fusion network (MSFLNet) based on rich information extraction, aiming to address the limitations of existing change detection methods in terms of global and local feature extraction, multi-scale information fusion, and computational complexity. First, to enhance the capture capability of multi-directional features and reduce computational complexity, we propose a lightweight global dynamic feature extraction module (LGDM). This module achieves cross-scale dynamic feature fusion through large kernel convolution decomposition and a global enhanced multi-scale attention mechanism (GEMA), effectively modeling the semantic association of targets in any orientation. Second, to suppress background interference and optimize spectral-spatial detail information, we design a lightweight spatiotemporal local feature extraction module (LSTLM). This module combines grouped dilated convolution with a parameter-free attention mechanism, significantly enhancing the representation capability of local details. Finally, we propose an adaptive feature fusion module (AFFM), which utilizes an efficient channel attention mechanism to dynamically calibrate the contribution of global and local features, achieving adaptive weighted fusion of multi-level information, avoiding feature conflicts and improving model performance. Through experiments on the CDD and LEVIR-CD datasets, we verify the effectiveness of MSFLNet in change detection tasks.
[0143] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.
Claims
1. A remote sensing image change detection method based on a multi-scale fusion lightweight network, characterized in that, The application relates to a multi-scale fusion lightweight network for change detection of double-time-phase remote sensing images. The multi-scale fusion lightweight network comprises: a symmetric encoding-decoding structure network for branch feature extraction, encoding and decoding of the double-time-phase remote sensing images to generate a change map; a lightweight global dynamic feature extraction module for first processing of the change map based on a large-kernel convolution decomposition strategy and a global enhanced multi-scale attention mechanism; a lightweight spatio-temporal local feature extraction module for second processing of the change map based on a grouping hollow convolution strategy and a parameter-free spectral space joint attention mechanism; an adaptive feature fusion module for processing the change map after the first processing and the second processing based on a double-pooling strategy and a channel attention mechanism to obtain a change detection result; when the lightweight global dynamic feature extraction module processes the feature map of the change map based on the large-kernel convolution decomposition strategy, the feature map is divided into four groups of heterogeneous branches according to proportional coefficients through a channel division operation, and local fine-grained context features, long-distance dependent horizontal direction features, vertical direction global correlation features and original feature information are extracted respectively; when the global enhanced multi-scale attention mechanism is used for processing, the four groups of branch features are preliminarily fused through splicing, a semantic vector is generated through global context compression for the left branch, context information is captured through multi-scale convolution for the right branch, and finally the two are added to obtain the change map after the first processing; when the lightweight spatio-temporal local feature extraction module processes the feature map of the change map based on the grouping hollow convolution strategy, the feature map is grouped according to channels, and different dilated convolutions with different dilated rates are applied to each group to capture multi-scale local detail information; when the parameter-free spectral space joint attention mechanism is used for processing, the feature map is taken as a key and a value, a channel mean value is taken as a query, and attention is calculated and generated; weights are generated through Softmax, and are multiplied with the feature map to obtain the change map after the second processing; when the adaptive feature fusion module processes the change map after the first processing and the second processing based on the double-pooling strategy, channel splicing is performed on the input change map, global and local features are obtained through global average pooling and global maximum pooling respectively, and feature splicing is performed; when the channel attention mechanism is used for processing, the spliced features are flattened and dimension-adjusted, cross-channel interaction learning is performed through one-dimensional convolution, weights are generated through Softmax, feature calibration is performed through broadcast multiplication, and the change detection result is obtained. when the symmetric encoding-decoding structure network extracts branch features, two 3*3 convolutions are used to extract shallow feature information respectively; 2. The remote sensing image change detection method based on the multi-scale fusion lightweight network of claim 1, wherein, when encoding, the symmetric branches are subjected to down-sampling operation through five times of main modules and four times of maximum pooling cross combination, different scale encoding features are obtained, and absolute value differences of upper and lower feature maps are calculated to obtain change features of the double branches in the encoding stage; when decoding, the main module is used for refining features, the absolute value differences of the upper and lower feature maps are calculated, the features are restored in scale through up-sampling operation, the high-level difference features are up-sampled to the resolution of the previous level, and the difference features of the corresponding level in the encoding stage are spliced along the channel dimension. After four feature extraction, encoding, decoding operations, the last layer of the change feature is operated using a 1*1 convolution to obtain the change map.
3. The remote sensing image change detection method based on the multi-scale fusion lightweight network of claim 1, wherein, The loss function of the multi-scale fusion lightweight network adopts a binary cross-entropy loss function. 4.The remote sensing image change detection method based on multi-scale fusion lightweight network according to claim 1, wherein, Also includes: Visualize the change detection result; Assist the user to achieve his change detection purpose based on the visualized change detection result.
5. The remote sensing image change detection method based on the multi-scale fusion lightweight network of claim 4, wherein, The steps of assisting the user include: Based on the first spatio-temporal graph formed by the user's interaction with the change detection result of the visualized output before the last triggering time, the second spatio-temporal graph for decision assistance is decided; Based on the second spatio-temporal graph, the user continues to interact with the change detection result of the visualized output.
6. A remote sensing image change detection system based on a multi-scale fusion lightweight network, characterized in that, It includes: A remote sensing image change detection module for detecting changes in double-time-phase remote sensing images using a multi-scale fusion lightweight network; The multi-scale fusion lightweight network includes: A symmetric encoding-decoding structure network for branch feature extraction, encoding and decoding of double-time-phase remote sensing images to generate a change map; A lightweight global dynamic feature extraction module for first processing of the change map based on a large kernel convolution decomposition strategy and a global enhanced multi-scale attention mechanism; A lightweight spatio-temporal local feature extraction module for second processing of the change map based on a grouping hollow convolution strategy and a parameter-free spectral space joint attention mechanism; An adaptive feature fusion module for processing the change map after the first processing and the second processing based on a double-pooling strategy and a channel attention mechanism to obtain a change detection result; When the lightweight global dynamic feature extraction module processes the feature map of the input change map, it is divided into four groups of heterogeneous branches by a channel division operation according to the proportion coefficient, and local fine-grained context features, long-distance dependent horizontal direction features, vertical direction global correlation features and original feature information are extracted respectively; When processing based on the global enhanced multi-scale attention mechanism, the four groups of branch features are preliminarily fused by concatenation, the left branch generates a semantic vector by global context compression, the right branch captures context information by multi-scale convolution, and finally the two are added to obtain the change map after the first processing. When the lightweight spatio-temporal local feature extraction module processes the feature map of the input change map, it is divided into groups by channel, and different dilated convolutions with different dilated rates are applied to each group to capture multi-scale local detail information. When processing based on the parameter-free spectral space joint attention mechanism, the feature map is used as the key and value, and the channel mean is used as the query to calculate the attention; the weight is generated by Softmax, and multiplied with the feature map to obtain the change map after the second processing. When the adaptive feature fusion module processes based on the double-pooling strategy, the first-processed and second-processed change maps are concatenated by channel, and global and local features are obtained by global average pooling and global maximum pooling respectively, and the features are concatenated; When processing based on the channel attention mechanism, the concatenated features are flattened and dimension adjusted, cross-channel interaction learning is performed by one-dimensional convolution, the weight is generated by Softmax, and feature calibration is performed using broadcast multiplication to obtain the change detection result.
Citation Information
Patent Citations
Remote sensing image change detection method based on spatial-spectral feature fusion network
CN114359723A
Dual-branch multi-scale dynamic local convolution attention method based on remote sensing change detection
CN118736416A