Lightweight feature smooth group attention network for remote sensing image change detection

By proposing a lightweight feature smooth group attention network (LAGANet) in remote sensing image change detection, the self-attention and dual-feature fusion module of the group convolution channel are used to solve the problem of low global-local feature extraction efficiency in the existing technology, and more efficient and accurate detection of change areas is achieved.

CN120088616AActive Publication Date: 2025-06-03ANHUI UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510155940.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-12
Publication Date
2025-06-03
Estimated Expiration
2045-02-12

AI Technical Summary

Technical Problem

The prior art is difficult to effectively extract global-local features in remote sensing image change detection, resulting in low detection efficiency of changing areas.

Method used

A lightweight feature smooth group attention network (LAGANet) is proposed, which captures global-local features and reduces attention calculation redundancy by introducing CNN's group convolutional channel self-attention (CGSA) module and dual-feature fusion module.

Benefits of technology

The accuracy and efficiency of remote sensing image change detection is improved, the generalization and calculation efficiency of the model are enhanced, and the characteristics of the changing area can be extracted more accurately.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088616A_ABST
    Figure CN120088616A_ABST
Patent Text Reader

Abstract

The invention discloses a lightweight feature smoothing group attention network for remote sensing image change detection, and relates to the technical field of remote sensing information extraction, first, group convolution channel self-attention is designed, and input features are divided into small windows along channel dimensions; secondly, a double-feature fusion module is designed to filter and reconstruct multi-scale features, so that the difference between convolution kernels is smoothed, and the generalization ability of the model is enhanced; according to the method, the problems of interaction capability limitation and calculation cost of the CNN and Transform in global-local feature extraction are solved, a remote sensing target detection technology with practical value is obtained in application, and application and development of target ground object monitoring in remote sensing images are expected to be practically promoted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of remote sensing information extraction, and particularly to a lightweight feature smoothing group attention network for remote sensing image change detection. Background Art

[0002] Change detection (CD), as an important technical means in earth observation, has been widely applied in various fields such as post-disaster assessment, urban management, and environmental monitoring. It can accurately identify changes in land cover and environmental conditions by analyzing remote sensing images from different periods. With the rapid development of optical sensors, automated CD technology has attracted increasing attention. This method can save a considerable amount of time and cost. Binary change detection (BCD) is a common and important basic method in CD, which can be divided into between-class CD and within-class CD. The former is mainly used to detect changes between all land cover elements, while the latter is specifically used to detect changes related to specific tasks, such as constructing CD.

[0003] In recent years, CD methods for very high resolution (VHR) remote sensing images based on deep learning have been widely applied. The emergence of convolutional neural network (CNN) has stimulated the rapid development of CD technology, leading to the proposal of a series of classic models. Due to the characteristics of weight sharing and local receptive fields of CNN, it has shown excellent capabilities in obtaining rich and abstract local features. Relying on deeper network architectures and larger convolutional kernels, the model can often obtain a wider receptive field. However, this also leads to redundancy in the number of parameters, making model training more challenging. In addition, the increase in the receptive field obtained by this method is limited. Since the advent of Transformer, its ability to establish connections between global features through self-attention modules has been widely applied in the fields of natural language processing and image processing. Some studies have integrated Vision Transformer (ViT) into CD, providing new possibilities for feature extraction and solving the problem that CNN cannot establish global feature correlations. However, due to the quadratic growth of the computational complexity of the self-attention mechanism with the increase in image size, the requirements for hardware in model training are higher.

[0004] How to effectively extract the global-local features of remote sensing image change detection tasks and improve the detection efficiency of change regions is an urgent problem to be solved. To address these challenges, the present invention proposes a lightweight feature smoothing group attention network (LAGANet). In this model, the group convolutional channel self-attention (CGSA) of CNN is introduced to capture global-local features and reduce the redundancy of attention calculation. At the same time, the forward feed network (FFN) of CNN is used to process local information and non-linear transformation. In addition, a dual feature fusion module (DFFM) is designed to filter and reconstruct multi-scale features. This strategy requires few learnable parameters and can effectively smooth the differences between convolutional kernels, thereby enhancing the generalization of the model. For this reason, we propose a lightweight feature smoothing group attention network for remote sensing image change detection. Summary of the Invention

[0005] The object of the present invention is to solve the problems in the background technology, improve the accuracy of remote sensing image target detection, and achieve high-precision remote sensing target detection. Under the guidance of this idea, a lightweight feature smoothing group attention network for remote sensing image change detection is designed. The lightweight feature smoothing group attention network for remote sensing image change detection is characterized by including the following steps:

[0006] Step 1: Collect remote sensing image datasets;

[0007] Step 2: Model construction;

[0008] Step 3: Model implementation and training;

[0009] Step 4: Model accuracy evaluation.

[0010] Preferably, in the first step, the CDD, SECOND, LEVIR-CD, LEVIR-CD+ and WHU-CD datasets are downloaded from the Internet respectively, and the image size is uniformly cropped to 256×256 pixels. The images are enhanced by random flipping, rotation, and scaling operations, and all patches are randomly divided into a training set, a validation set, and a test set according to a 7:1:2 division ratio.

[0011] Preferably, the LAGANet in the second step consists of a dual encoder and a dual decoder. The encoding part consists of a CGSA module and a downsampling module. Depthwise separable convolutions are used instead of standard convolution blocks. The downsampling module consists of multiple groups of depthwise separable convolutional layers, batch normalization layers, and ReLU activation function layers, which are used to extract features of different scales while minimizing computational complexity and information loss;

[0012] In the decoding stage, the decoding part is formed by integrating two sets of depthwise separable convolutional layers, batch normalization layers, and ReLU activation function layers. In the feature fusion stage of the decoder, features from two time periods are first fused, and then, through a filtering and reconstruction process, the differences between convolutional kernels are smoothed and merged with the features of the previous stage.

[0013] Preferably, the model in step two includes a CGSA module, a dual feature fusion module, and a loss function module.

[0014] Preferably, in step two, the multi-head attention mechanism is applied to couple lightweight, inter-channel features. The CGSA module evenly divides the input features into four attention heads in the channel dimension to ensure that each attention head can learn different features. The specific formula is as follows:

[0015] Q, K, V = Conv Q,K,V (w(x i ))

[0016]

[0017]

[0018] In the formula, x i represents the full input features of the i-th block, and w(x i ) represents window segmentation processing; among them, Q, K, and V are obtained by 3 1×1 convolutions; T represents the transpose operation, and d is for dimensionality reduction; is the channel, that is, MHSA. The input window features are divided into n groups according to the number of fronts, and then, self-attention calculations are performed on these n channel features respectively. Finally, the forward propagation τ f perceives and analyzes the local features of the cumulative attention feature map.

[0019] Preferably, in step two, is the feature of the previous time stage, is the feature of the later time stage. The specific formula is as follows:

[0020] P 1 = |F 1 - F 2 |

[0021] P 2 = Conv(SE(Cat(F 1 , F 2 )))

[0022] FF = Cat(σ(P 2 ) × P 1 ) + P 2

[0023] In the formula, P 1 represents the feature enhancement branch, and P 2 represents the feature retention branch. Cat represents the feature concatenation of F 1 and F 2 along the channel dimension. The SE module is used to strengthen the channel features of the input feature map, and 1×1 convolution is used to align the channels. To enhance the interaction between P 1 and P 2 a sigmoid activation function is applied to normalize P 2 . Then, the normalized P 2 is multiplied by the enhanced P 1 features. Finally, the features of the two paths are fused;

[0024]

[0025] In the formula, in the group normalization (GN) operation, the dynamically updated parameter γ is used to record the weights of each feature channel. By comparing the preset empirical coefficient τ, different channels are filtered out. Then, these channels are multiplied by the original input feature F i and the weighted attention feature to obtain the final high-weight and low-weight features. Finally, inter-kernel smoothing is performed through channel concatenation.

[0026] Preferably, in the second step, LossNet is designed to provide comprehensive supervision across different feature levels. The specific formula is as follows:

[0027]

[0028] In the formula, represents the binary cross-entropy loss, which is a commonly used loss function in the change detection task. At the same time, is used to further optimize the training of the model. Specifically, a pre-trained VGG16 model on ImageNet is frozen and used as the backbone network of this loss function. By comparing the prediction results with the ground truth at multiple scales, the model's extraction of real features is improved. and respectively represent the features extracted by VGG16 from the predicted true value and the ground truth of each layer. The root mean square error loss function is used to calculate the difference between the predicted value and the true value of each layer, and they are accumulated as the final depth loss.

[0029] Preferably, in the third step, the model is implemented using the PyTorch framework, and the batch size, number of epochs, initial learning rate, and weight decay are set to 8, 200, 0.02, and 0.0003 respectively. The optimization algorithm used is the Nesterov momentum stochastic gradient descent method.

[0030] Preferably, in the fourth step, the performance of the model is evaluated using the IoU, Precision, Recall, F1-score, and OverallAccuuracy metrics, and the specific formulas are as follows:

[0031]

[0032] In the formula, TP, FP, FN, and TN represent the numbers of true positive predictions, false positive predictions, false negative predictions, and true negative predictions respectively.

[0033] Beneficial effects

[0034] The method of the present invention solves the problems of the interaction ability limitation and computational cost in global-local feature extraction of CNN and Transformer. In application, a remote sensing target detection technology with practical value is obtained, in order to effectively promote the application and development of the monitoring of target ground objects in remote sensing images. Description of the drawings

[0035] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0036] Figure 1 It is the model architecture diagram of the present invention;

[0037] Figure 2 It is the model change detection result diagram based on the CDD dataset (note: in the figure, the white part represents TP, the black part represents TN, the dark gray part represents FP, and the light gray part represents FN). Detailed implementation manners

[0038] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the protection scope of the present invention.

[0039] Please refer to Figure 1-2, the present invention provides a technical solution: a lightweight feature smoothing group attention network for remote sensing image change detection, including the following steps:

[0040] Step 1: Collect remote sensing image datasets

[0041] Download the CDD, SECOND, LEVIR-CD, LEVIR-CD+ and WHU-CD datasets from the Internet respectively, and uniformly crop the image size to 256×256 pixels. In addition, operations such as random flipping, rotation, and scaling are used to enhance the images, and all patches are randomly divided into a training set, a validation set, and a test set according to a 7:1:2 division ratio.

[0042] Step 2: Model construction

[0043] LAGANet is a double-encoder and double-decoder structure. The encoding part consists of a CGSA module and a downsampling module. At the same time, depthwise separable convolutions are used instead of standard convolution blocks. The downsampling module consists of multiple groups of depthwise separable convolutional layers, batch normalization layers, and ReLU activation function layers, which are used to extract features of different scales while minimizing computational complexity and information loss. In the decoding stage, the decoding part is formed by integrating two groups of depthwise separable convolutional layers, batch normalization layers, and ReLU activation function layers;

[0044] In the feature fusion stage of the decoder, first, the features from two time periods are fused. Then, through the filtering and reconstruction process, the differences between the convolutional kernels are smoothed and merged with the features of the previous stage.

[0045] Specifically, this model includes a CGSA module, a dual feature fusion module, and a loss function module.

[0046] CGSA module: To solve the computational redundancy problem in attention calculation, the multi-head attention mechanism (MHSA) is applied to couple lightweight, inter-channel features. The CGSA module evenly divides the input features into four attention heads in the channel dimension to ensure that each attention head can learn different features. The specific formula is as follows:

[0047] Q, K, V = Conv Q,K,V (w(x i ))

[0048]

[0049] In the formula, x i represents the full input feature of the i-th block, and w(x i ) represents window segmentation processing. Among them, Q, K, and V are obtained by 3 1×1 convolutions, T represents the transpose operation, and d is for dimensionality reduction. The channel, i.e., MHSA, divides the input window features into n groups according to the number of fronts. Then, self-attention calculations are performed on these n channel features respectively. Finally, forward propagation τ is used. f Perceive and analyze the local features of the cumulative attention feature map.

[0050] Dual-feature fusion module: is the feature of the previous time stage, is the feature of the later time stage. The specific formula is as follows:

[0051] P 1 = |F 1 - F 2 |

[0052] P 2 = Conv(SE(Cat(F 1 , F 2 )))

[0053] FF = Cat(σ(P 2 ) × P 1 ) + P 2

[0054] In the formula, P 1 represents the feature enhancement branch, P 2 represents the feature retention branch, Cat represents the feature concatenation of F 1 and F 2 along the channel dimension. The SE module is used to strengthen the channel features of the input feature map, and 1×1 convolution is used to align the channels. To enhance the interaction between P 1 and P 2 , a sigmoid activation function is applied to normalize P 2 . Then, the normalized P 2 is multiplied by the enhanced P 1 features. Finally, the features of the two paths are fused;

[0055]

[0056] In the formula, in the group normalization (GN) operation, the dynamically updated parameter γ is used to record the weight of each feature channel. By comparing the preset empirical coefficient τ, different channels are filtered out. Then, these channels are multiplied by the original input feature F i and the weighted attention feature to obtain the final high-weight and low-weight features. Finally, inter-kernel smoothing is performed through channel concatenation.

[0057] Loss function module: Design LossNet to provide comprehensive supervision across different feature levels to better perceive the inherent features of the changing regions. The specific formula is as follows:

[0058]

[0059] In the formula, represents the binary cross-entropy (BCE) loss, which is a commonly used loss function in change detection tasks. At the same time, is used to further optimize the training of the model. Specifically, a pre-trained VGG16 model on ImageNet is frozen and used as the backbone network of this loss function. By comparing the prediction results with the ground truth at multiple scales, the model's extraction of real features is improved. and respectively represent the features extracted by VGG16 from the predicted true values and ground truth values of each layer. The root mean square error loss function is used to calculate the difference between the predicted values and the true values of each layer, and they are accumulated as the final depth loss.

[0060] Step 3: Model implementation and training

[0061] The model is implemented using the PyTorch framework. The batch size, number of epochs, initial learning rate, and weight decay are set to 8, 200, 0.02, and 0.0003 respectively. The optimization algorithm used is Nesterov momentum stochastic gradient descent (SGD).

[0062] Step 4: Model accuracy evaluation

[0063] The performance of the model is evaluated using metrics such as IoU, Precision, Recall, F1-score, and Overall Accuuracy. The specific formulas are as follows:

[0064]

[0065] In the formula, TP, FP, FN, and TN represent the numbers of true positive predictions, false positive predictions, false negative predictions, and true negative predictions respectively;

[0066] Specifically, using the CDD dataset according to the above steps, the present invention provides an implementation case. The experiment is carried out on the Windows10 operating system, using an Intel Core i7-13700F CPU and an NVIDIA GeForce GTX 4060Ti GPU (16GB of memory). At the same time, models such as FC-EF, FC-EF-Cat, FC-EF-diff, SNUNet, AFCF3DNet, Changeformer, BiT, and ICIFNet are selected for comparative analysis to verify the effectiveness and superiority of this method. As shown in Table 1 and Figure 2As shown, it can be seen that the F1-score and IoU of LAGANet are increased by 3.25% and 5.97% respectively compared with the latest AFCF3DNet model. This indicates that LAGANet performs very well in extracting the boundaries of changed regions, detecting changes in changed regions, and avoiding false detections in unchanged regions, and is able to proficiently extract low-frequency features from images.

[0067] Table 1 Comparison of results of different model methods

[0068]

[0069] In the description of this specification, the descriptions referring to terms such as "one embodiment", "example", "specific example", etc. mean that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0070] The preferred embodiments of the present invention disclosed above are only used to help illustrate the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made according to the content of this specification. These embodiments are selected and specifically described in this specification in order to better explain the principles and practical applications of the present invention, so that those skilled in the art can well understand and utilize the present invention. The present invention is only limited by the claims and their full scope and equivalents.

Claims

1. A lightweight feature smoothing group attention network for remote sensing image change detection, characterized by: The following steps are involved: Step 1: Collect remote sensing image datasets; Step 2: Model construction; Step 3: Model implementation and training; Step 4: Model accuracy evaluation.

2. The lightweight feature smoothing group attention network for remote sensing image change detection according to claim 1 is characterized in that: In step 1, CDD, SECOND, LEVIR-CD, LEVIR-CD+, and WHU-CD datasets were downloaded from the Internet, and the images were uniformly cropped to 256×256 pixels. Random flipping, rotation, and scaling operations were used to enhance the images, and all patches were randomly divided into training, validation, and test sets in a 7:1:2 ratio.

3. The lightweight feature smoothing group attention network for remote sensing image change detection according to claim 2 is characterized in that: The LAGANet in step 2 is composed of a dual encoder and a dual decoder. The encoding part consists of a CGSA module and a downsampling module. Depth-separable convolution is used instead of a standard convolution block. The downsampling module consists of multiple groups of depth-separable convolution layers, batch normalization layers, and ReLU activation function layers to extract features of different scales while minimizing computational complexity and information loss. In the decoding stage, the decoding part is formed by integrating two sets of depth-wise separable convolutional layers, batch normalization layers, and ReLU activation function layers. In the feature fusion stage of the decoder, the features from the two time periods are first fused, and then the differences between the convolution kernels are smoothed through the filtering and reconstruction process, and then merged with the features of the previous stage.

4. The lightweight feature smoothing group attention network for remote sensing image change detection according to claim 3 is characterized in that: The model in step 2 includes a CGSA module, a dual feature fusion module and a loss function module.

5. The lightweight feature smoothing group attention network for remote sensing image change detection according to claim 4 is characterized in that: In step 2, the multi-head attention mechanism is used to couple lightweight and inter-channel features. The CGSA module evenly divides the input features into four attention heads in the channel dimension to ensure that each attention head can learn different features. The specific formula is as follows: Q,K,V=Conv Q,K,V (w(x i )) In the formula, x i represents the full input feature of the i-th block, w(x i ) represents the window segmentation process; where Q, K, and V are obtained by three 1×1 convolutions; T represents the transposition operation, and d represents the dimension reduction; is a channel, i.e., MHSA. The input window features are divided into n groups according to the number of positive faces. Then, the self-attention calculation is performed on these n channel features respectively. Finally, the forward propagation τ is used f Perceiving and analyzing local features of cumulative attention feature maps.

6. The lightweight feature smoothing group attention network for remote sensing image change detection according to claim 5, characterized in that: In the step 2, Characteristic of the previous time stage, is the post-time stage feature, and the specific formula is as follows: P1=|F1-F2| P2=Conv(SE(Cat(F1,F2))) FF=Cat(σ(P2)×P1)+P2 In the formula, P1 represents the feature enhancement branch, P2 represents the feature retention branch, Cat represents the feature concatenation of F1 and F2 along the channel dimension, the SE module is used to enhance the channel features of the input feature map, and 1×1 convolution is used to align the channels. In order to enhance the interaction between P1 and P2, a sigmoid activation function is applied to normalize P2, and then the normalized P2 is multiplied by the enhanced P1 feature. Finally, the features of the two paths are fused; In the group normalization (GN) operation, the dynamically updated parameter γ is used to record the weight of each feature channel. By comparing the preset empirical coefficient τ, different channels are filtered out. Then, these channels are compared with the original input feature F i and weighted attention features Multiply them together to get the final high-weight and low-weight features, and finally, perform inter-kernel smoothing through channel cascading.

7. The lightweight feature smoothing group attention network for remote sensing image change detection according to claim 6, characterized in that: In step 2, LossNet is designed to provide comprehensive supervision across different feature levels. The specific formula is as follows: In the formula, Represents binary cross entropy loss, which is a commonly used loss function in change detection tasks. At the same time, use Further optimize the model training, specifically, freeze a VGG16 model pre-trained on ImageNet and use it as the backbone network for this loss function to improve the model's extraction of real features by comparing the prediction results with the ground truth at multiple scales. and They represent the features extracted by VGG16 from the predicted true value and the ground truth of each layer, respectively. The root mean square error loss function is used to calculate the difference between the predicted value and the true value of each layer, and they are accumulated as the final depth loss.

8. The lightweight feature smoothing group attention network for remote sensing image change detection according to claim 7, characterized in that: In the step 3, the model is implemented using the PyTorch framework, and the batch size, number of epochs, initial learning rate, and weight decay are set to 8, 200, 0.02, and 0.0003, respectively. The optimization algorithm used is the Nesterov momentum stochastic gradient descent method.

9. The lightweight feature smoothing group attention network for remote sensing image change detection according to claim 8, characterized in that: In step 4, the performance of the model is evaluated using the IoU, Precision, Recall, F1-score, and OverallAccuuracy indicators. The specific formula is as follows: Where TP, FP, FN and TN represent the number of true positive predictions, false positive predictions, false negative predictions and true negative predictions, respectively.

Citation Information

Patent Citations

  • High-resolution remote sensing image target detection method of M-F-Y type lightweight convolutional neural network

    CN111666836A

  • Visual target tracking method based on twin residual attention convergence network

    CN116934796A

  • Strong supervision change detection method based on convolutional neural network and visual attention model

    CN119206487A

  • Remote sensing image change detection method based on adaptive Transform and deformable convolution

    CN119418204A

  • Convolutional-neutral-network based filter for video coding

    US20210329286A1