Forest grass change detection method based on multi-scale convolution attention

By using the multi-scale convolutional attention method in forest and grass change detection, the Siamese encoder and multi-scale convolutional attention decoder are used for feature extraction and segmentation output, combined with the feature learning of deep supervision optimization model, the problem of poor forest and grass change detection effect in the existing technology is solved, and a higher quality change detection effect is achieved.

CN120125848APending Publication Date: 2025-06-10AEROSPACE INFORMATION RES INST CAS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510244532.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The existing deep learning methods have poor results in forest and grass change detection scenarios, making it difficult to effectively detect targets of different shapes and scales, and there are problems such as incomplete segmentation of changing areas and difficult to distinguish changing areas.

Method used

The forest and grass change detection method based on multi-scale convolutional attention is adopted. By obtaining the two-time phase images for preprocessing and inputting the trained forest and grass change detection model, the feature extraction and segmentation output is performed using the Siamese encoder and the multi-scale convolutional attention decoder, combined with the feature learning of the deep supervision optimization model.

Benefits of technology

The quality of forest and grass change detection is improved, the multi-scale feature extraction capability of the model is enhanced, the targets of different shapes and scales can be accurately detected, missed and missed detection, and the change areas can be refined.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125848A_ABST
    Figure CN120125848A_ABST
Patent Text Reader

Abstract

The invention provides a forest grass change detection method based on multi-scale convolution attention, and the method comprises the steps: obtaining a to-be-detected dual-time-phase image which is two remote sensing images with different time phases; preprocessing a to-be-detected dual-time-phase image, and inputting the to-be-detected dual-time-phase image into the forest grass change detection model to obtain a change detection binary image; wherein the forest grass change detection model is obtained through training, and the training process comprises the following steps: encoding a to-be-trained dual-time-phase image through an encoder to obtain dual-time-phase features; after the dual-time-phase features are integrated through a feature fusion module, the dual-time-phase features are input into a decoder to be decoded, and two target features of different scales are obtained; and performing segmentation output on the target features by using a dynamic up-sampling segmentation head to obtain a first prediction feature map and a second prediction feature map, the first prediction feature map being used for representing a change detection binary result, and the second prediction feature map being used for depth supervision to optimize feature learning of the forest and grass change detection model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of forest and grass change detection, and specifically relates to a forest and grass change detection method based on multi-scale convolutional attention. Background Art

[0002] Forests and grasslands are important components of terrestrial natural ecosystems, playing a huge role in carbon sequestration, climate regulation, wind prevention and sand fixation, and biodiversity protection, which are crucial for human well-being and sustainable development. With the intensification of human activities, the destruction of forests and grasslands has gradually become serious, and the timely monitoring of forests and grasslands has become a necessary means to protect forest and grass resources. Timely understanding of the destruction and degradation of forests and grasslands is helpful for management work and is of great significance for vegetation restoration and ecological environment construction.

[0003] With the booming development of remote sensing technology, the method of using remote sensing images for identification has replaced methods such as manual field inspections and has become an important way for dynamic supervision of forests and grasslands. By using high-resolution remote sensing images to monitor the changes in forests and grasslands and update the quarterly and monthly change information of forest and grass resources, the timeliness of forest and grass resource monitoring has been significantly improved. Traditional change detection methods usually have a large amount of salt-and-pepper noise and pseudo-changes on high-resolution images. With the breakthrough of deep learning technology in the field of computer vision, high-resolution change detection has become a research hotspot, but the existing deep learning methods cannot well meet the application environment of the forest and grass change detection scenario, and the forest and grass change detection effect is poor. Summary of the Invention

[0004] The present invention provides a forest and grass change detection method based on multi-scale convolutional attention, including: obtaining a pair of temporal images to be detected, where the pair of temporal images are two remote sensing images at different times; preprocessing the pair of temporal images to be detected and inputting them into a forest and grass change detection model to obtain a change detection binary map; wherein, the forest and grass change detection model is obtained through training on a forest and grass change detection training data set, and the training process of the model includes: encoding the pair of temporal images to be trained through an encoder to obtain pair of temporal features; after integrating the pair of temporal features through a feature fusion module, inputting them into a decoder for decoding to obtain target features at two different scales; using a dynamic upsampling segmentation head to separately segment and output the target features to obtain a first prediction feature map and a second prediction feature map, where the first prediction feature map is used to represent the change detection binary result, and the second prediction feature map is used for deep supervision to optimize the feature learning of the forest and grass change detection model.

[0005] In the above solution, the encoder is a Siamese encoder with a VGG-16 neural network structure as the backbone, including four convolutional stages, and each convolutional stage is a convolutional block after batch normalization.

[0006] In the above solution, the dual-temporal images to be trained are encoded by an encoder to obtain dual-temporal features, including: obtaining the dual-temporal images to be trained; preprocessing the dual-temporal images to be trained and inputting them into a Siamese encoder with a VGG-16 neural network structure as the backbone to obtain the corresponding dual-temporal features encoded by each convolutional block.

[0007] In the above solution, a feature fusion module is provided after each convolutional stage corresponding to the Siamese encoder. Within each feature fusion module, the dual-temporal features are integrated, including: after splicing each pair of dual-temporal features, generating an intermediate feature map through the first convolutional operation; performing a splitting operation on the intermediate feature map to obtain a first sub-feature map and a second sub-feature map after splitting, where the second sub-feature map is used for input into the Bottleneck for processing; splicing the second sub-feature map processed by the Bottleneck with the first sub-feature map and completing channel transformation and feature fusion through the second convolutional operation to obtain fused features.

[0008] In the above solution, the decoder is a multi-scale convolutional attention decoder, including: a large kernel grouped attention gate, an upsampling convolutional block, and a multi-scale convolutional attention module; and the decoder includes four decoding stages corresponding to the feature fusion modules respectively.

[0009] In the above solution, the two target features are the first target feature and the second target feature respectively. Multiple fused features are input into the multi-scale convolutional attention decoder for decoding to obtain two target features of different scales, including: in the first decoding stage of the encoder, based on the multi-scale convolutional attention module, outputting the second target feature; in the fourth decoding stage of the encoder, based on the large kernel grouped attention gate, the upsampling convolutional block, and the multi-scale convolutional attention module, outputting the first target feature.

[0010] In the above solution, the dynamic upsampling segmentation head is used to separately segment and output the target features to obtain a first predicted feature map and a second predicted feature map, including: using the dynamic upsampling segmentation head to perform channel conversion and dynamic upsampling operations on the second target feature and segmenting and outputting the second predicted feature map.

[0011] In the above solution, the dynamic upsampling segmentation head is used to separately segment and output the target features to obtain a first predicted feature map and a second predicted feature map, further including: using the dynamic upsampling segmentation head to perform channel conversion and dynamic upsampling operations on the first target feature and segmenting and outputting the first predicted feature map.

[0012] In the above solution, the method further includes: based on the second predicted feature map, creating an auxiliary loss function for deep supervision.

[0013] In the above solution, the method further includes: after the first predicted feature map is processed by an activation function, obtaining a binary change detection result.

[0014] The technical solution of the embodiment of the present invention has at least the following beneficial effects:

[0015] (1) In the scenario of forest and grass change detection, the forest and grass change detection model needs to have strong multi-scale feature extraction capabilities to accurately detect targets of different shapes and scales.

[0016] (2) This method improves the quality of forest and grass change detection, solves the problems of incomplete segmentation of change regions in double-temporal images and difficult discrimination of difficult change regions. The forest and grass change detection model in this method has strong discrimination capabilities, reducing missed detections and false detections.

[0017] (3) For targets with small scales, the forest and grass change detection model in this method can fully extract spatial detail features to achieve refined identification and differentiation of change regions in double-temporal images. Description of the Drawings

[0018] Figure 1 Schematically shows a flowchart of a method for forest and grass change detection based on multi-scale convolutional attention according to an embodiment of the present invention;

[0019] Figure 2 Schematically shows an overall framework diagram of a forest and grass change detection model according to an embodiment of the present invention;

[0020] Figure 3 Schematically shows a flowchart of the training process of a forest and grass change detection model according to an embodiment of the present invention;

[0021] Figure 4A Schematically shows a framework diagram of a feature fusion module according to an embodiment of the present invention;

[0022] Figure 4B Schematically shows a framework diagram of a Bottleneck layer according to an embodiment of the present invention;

[0023] Figure 5 Schematically shows a framework diagram of a decoder according to an embodiment of the present invention;

[0024] Figure 6 Schematically shows a process diagram of dynamic upsampling according to an embodiment of the present invention;

[0025] Figure 7 Schematically shows a comparison diagram of change detection visualization results on a dataset to be measured by different methods according to an embodiment of the present invention; and

[0026] Figure 8 Schematically shows a comparison diagram of change detection visualization results on another dataset to be measured by different methods according to an embodiment of the present invention. Detailed implementation manners

[0027] To make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to specific embodiments and the accompanying drawings.

[0028] Figure 1 The flowchart of the forest and grass change detection method based on multi-scale convolutional attention according to an embodiment of the present invention is schematically shown. Figure 2 The overall framework diagram of the forest and grass change detection model according to an embodiment of the present invention is schematically shown.

[0029] Please specifically refer to Figure 1 , the specific process of the forest and grass change detection method based on multi-scale convolutional attention according to the embodiment of the present invention includes operations S110 to S120.

[0030] In operation S110, dual-temporal images to be detected are obtained, and the dual-temporal images are two remote sensing images with different time phases.

[0031] In operation S120, the dual-temporal images to be detected are preprocessed and input into the forest and grass change detection model to obtain a change detection binary map.

[0032] Exemplarily, the original dual-temporal images to be detected are obtained, where the dual-temporal images are two remote sensing images with different time phases. Image registration, radiometric correction, image fusion and other processes are performed on the original dual-temporal images. In order to facilitate input and output of the forest and grass change detection model, slicing operations also need to be performed on the original dual-temporal images.

[0033] Further, as Figure 2 , after the original dual-temporal images after the slicing operation are input into the forest and grass change detection model, after an encoding stage, feature fusion, a decoding stage and a segmentation stage, a change detection binary map of the corresponding region is generated.

[0034] It should be noted that the above forest and grass change detection model is obtained by training on a forest and grass change detection training data set, where the forest and grass change detection training data set includes multiple groups of dual-temporal images to be trained and corresponding labels.

[0035] The training process of the forest and grass change detection model will be described in detail below.

[0036] Figure 3 The flowchart of the training process of the forest and grass change detection model according to an embodiment of the present invention is schematically shown.

[0037] Please specifically refer to Figure 3 , the specific process of the training process of the forest and grass change detection model according to the embodiment of the present invention includes operations S310 to S330.

[0038] At operation S310, the encoder encodes the dual-temporal image to be trained to obtain dual-temporal features.

[0039] In an embodiment of the present invention, encoding the dual-temporal image to be trained by an encoder to obtain dual-temporal features includes: acquiring the dual-temporal image to be trained; preprocessing the dual-temporal image to be trained and inputting it into a Siamese encoder with a VGG-16 neural network structure as the backbone to obtain the corresponding dual-temporal features after encoding in each convolutional stage.

[0040] Specifically, the embodiments of the present invention select high-resolution satellite data of the Gaofen series (GF1 and GF6) as high-resolution remote sensing image data for the RGB band combination (i.e., the forest and grassland change detection training dataset), and this forest and grassland change detection training dataset includes multiple dual-temporal high-resolution satellite images. Further, batch preprocessing is performed on these dual-temporal high-resolution satellite images, including operations such as image registration, radiometric correction, and image fusion. For the convenience of model input and output, these dual-temporal high-resolution satellite images are sliced and then input into the forest and grassland change detection model for training.

[0041] Further, as Figure 2 shown, the encoder of this forest and grassland change detection model is a Siamese encoder with a VGG-16 neural network structure as the backbone, including four convolutional stages (stages), and each convolutional stage is a convolutional block after batch normalization. When the dual-temporal image to be trained is input into the model, during the encoding process, the encoder generates the corresponding dual-temporal features after encoding in each convolutional stage, that is, the VGG-16 backbone gradually generates multi-level features.

[0042] It should be noted that in VGG-16, the pooling layer is used for downsampling, and the continuous downsampling process often loses accurate spatial position information, which may lead to inaccurate edge detection of the change area of the dual-temporal image and omission of small targets. Based on this, in this embodiment, the pooling layer in the fourth convolutional stage of VGG-16 is discarded so that this step will not be performed during feature extraction.

[0043] At operation S320, after integrating the dual-temporal features through the feature fusion module, it is input into the decoder for decoding to obtain two target features of different scales.

[0044] Exemplarily, first, the integration process of the dual-temporal features by the feature fusion module will be described in detail.

[0045] Figure 4A Schematically shows a framework diagram of the feature fusion module according to an embodiment of the present invention. Figure 4B Schematically shows a framework diagram of the Bottleneck layer according to an embodiment of the present invention.

[0046] It should be understood that remote sensing image segmentation only needs to extract features in a single-temporal image for segmentation, while change detection must combine the features of double-temporal images to segment the changes in target categories. The feature fusion module can integrate the feature maps of the front and back temporal phases, so it is an important part of the change detection task. Most Siamese network models use 3×3 convolutions for feature fusion, which may lead to insufficient integration of double-temporal features and make it difficult to segment difficult samples. To solve this problem, this implementation proposes a Cross-stage Double Convolution Feature Fusion (CDF) module to achieve efficient information flow and integration between the channels of the double-temporal feature maps.

[0047] As Figure 2 shown, a Cross-stage Double Convolution Feature Fusion (CDF) module is provided at each convolution stage corresponding to the Siamese encoder. Within each Cross-stage Double Convolution Feature Fusion (CDF) module, double-temporal features are integrated, including: after splicing each double-temporal feature, through the first convolution operation, an intermediate feature map is generated; the intermediate feature map is split to obtain the first sub-feature map and the second sub-feature map after splitting, where the second sub-feature map is used for input to the Bottleneck layer for processing; the second sub-feature map processed by the Bottleneck is spliced with the first sub-feature map, and through the second convolution operation, channel transformation and feature fusion are completed to obtain the fused feature.

[0048] Specifically, as Figure 4A shown, in each Cross-stage Double Convolution Feature Fusion (CDF) module, first, after splicing the double-temporal features, a 1×1 convolution is used to perform feature transformation on the input data, and the convolution layer expands the number of channels to generate an intermediate feature map. Then, the intermediate feature map is split by into y1 (the first sub-feature map) and y2 (the second sub-feature map), where the branch y2 passes through processing.

[0049] Furthermore, as Figure 4B shown, the Bottleneck layer uses 3×3 grouped convolutions and 1×1 pointwise convolutions to simultaneously process the same input feature map, that is, the branch y2, and adds the outputs, fusing different types of features to optimize information processing and feature extraction. Then, the processed branch y2 is spliced with y1 . Finally, after The convolutional layer performs channel transformation and feature fusion to obtain the fused features output by each Cross-stage Dual Convolution Feature Fusion (CDF) module. In this Cross-stage Dual Convolution Feature Fusion (CDF) module, the information from different branches enriches the feature expression ability. Such a design helps to increase the non-linearity ability and representation ability of the network, thereby improving the spatio-temporal modeling ability of the network for dual-temporal complex data.

[0050] Based on the above, the fusion module CDF can be summarized as the following formula:

[0051]

[0052] Based on the above, the branch structure and efficient convolutional operations in the Cross-stage Dual Convolution Feature Fusion (CDF) module help to integrate features at different levels and abstraction degrees in the dual-temporal image data, realizing the effective fusion of dual-temporal features.

[0053] It should be understood that in the embodiments of the present invention, in dual-temporal image change detection, factors such as vegetation phenological changes and light intensity result in significant feature differences between some image pairs. During the feature fusion process, the Cross-stage Dual Convolution Feature Fusion (CDF) module has a strong ability to represent change features, thereby modeling the position and time relationships in dual-temporal images and enhancing the discriminative ability of the model for difficult samples. Change detection requires comparing the feature differences between two time phases. Traditional methods usually directly splice or perform simple convolutional operations, which easily ignore the deep interactions between the features of the two time phases.

[0054] Furthermore, the Bottleneck structure in the Cross-stage Dual Convolution Feature Fusion (CDF) module retains the original features through residual connections, avoiding gradient disappearance and making the network pay more attention to the changing regions. The Bottleneck layer combines Group Convolution and Pointwise Convolution, which can capture local and global features simultaneously.

[0055] Through the embodiments of the present invention, the Cross-stage Dual Convolution Feature Fusion (CDF) module realizes dual-path feature fusion. This dual-path design can enhance the feature expression ability and improve the discriminative ability for difficult samples (such as changing regions with high similarity to the background), enabling the model to better handle complex dual-temporal feature fusion.

[0056] Exemplarily, the process of the decoder decoding the fused features to obtain target features at two different scales will be described in detail below.

[0057] Figure 5 Schematically shows the framework diagram of the decoder according to an embodiment of the present invention.

[0058] It should be understood that in the forest and grassland change detection task, in order to segment the change regions of different sizes and shapes in the forest and grassland images, the ability of the model to capture refined features at different scales is extremely important. Since change detection can be regarded as a binary image segmentation problem, a suitable advanced semantic segmentation architecture is introduced in this embodiment to solve the change detection problem.

[0059] Specifically, the decoder in this embodiment is a new multi-scale convolutional attention decoder. As Figure 5 shown, the decoder includes key components such as multiple large kernel grouped attention gates (LGAG), up-convolution blocks (EUCB), and multi-scale convolutional attention modules (MSCAM).

[0060] Among them, the large kernel grouped attention gate (LGAG) uses the gating signal from the advanced stage to control the features at different stages of the network, thereby activating the sum of relevant features and suppressing irrelevant features. The up-convolution block (EUCB) is used to upsample the feature map of the current stage to achieve alignment with the feature map of the skip connection in both spatial and channel dimensions. The multi-scale convolutional attention module (MSCAM) consists of three components, namely, a channel attention block, a spatial attention block, and a multi-scale convolutional block. First, the channel attention block can assign different weights to each channel, dynamically adjusting the importance of the channels, thereby emphasizing more relevant feature channels. Second, the spatial attention block is used to focus on specific regions of the input image, helping the model to identify and respond to task-related spatial local features. Finally, the efficient multi-scale convolutional block is used to retain and enhance the context relationship in the multi-scale features. Traditional convolutional operations lack the ability to dynamically adjust the importance of channels and spatial positions and have a fixed receptive field, which limits the flexible processing ability of the network for features. However, the multi-scale convolutional attention module (MSCAM) of the decoder in this embodiment enhances the significant features through channel and spatial attention and captures target features at different scales through parallel multi-scale separable convolutions, improving the adaptability and recognition ability of the model to various scale features in the image.

[0061] Please refer to Figure 5 again. The multi-scale convolutional attention decoder in this embodiment includes four decoding stages corresponding to the cross-stage double convolutional feature fusion (CDF) module respectively. In the embodiment of the present invention, multiple fused features are input into the multi-scale convolutional attention decoder for decoding to obtain two target features of different scales, including: in the first decoding stage of the encoder, based on the multi-scale convolutional attention module, the second target feature is output; in the fourth decoding stage of the encoder, based on the large kernel grouped attention gate, the up-convolution block, and the multi-scale convolutional attention module, the first target feature is output.

[0062] Specifically, as Figure 5As shown, considering that VGG-16 does not have residual connections, in the first decoding stage of the encoder, only based on the multi-scale convolutional attention module (MSCAM), the second target feature is output, which is used for subsequent deep supervision to alleviate the gradient vanishing problem and improve the feature learning and generalization ability of the model. In addition, in the fourth decoding stage of the encoder, based on the upsampled data of the large kernel grouped attention gate, the multi-scale convolutional attention module, and the transposed convolutional block, the first target feature is output, which is used for subsequent generation of the change detection binary result.

[0063] Through the embodiments of the present invention, the multi-scale convolutional attention decoder captures features through efficient multi-scale convolutions, and at the same time integrates complex spatial relationships and local attention by using channel, spatial, and gated attention mechanisms. The model can focus on the target features that have a decisive impact on task completion, thereby improving its accuracy in identifying change detection.

[0064] Furthermore, through the embodiments of the present invention, based on the setting of the multi-scale convolutional attention decoder, the forest and grass change detection model needs to have a strong multi-scale feature extraction ability to accurately detect targets of different shapes and scales.

[0065] In operation S330, the dynamic upsampling segmentation head is used to separately segment and output the target features to obtain the first predicted feature map and the second predicted feature map. The first predicted feature map is used to represent the change detection binary result, and the second predicted feature map is used for deep supervision to optimize the feature learning of the forest and grass change detection model.

[0066] Figure 6 Schematically shows a process diagram of dynamic upsampling according to an embodiment of the present invention.

[0067] It should be understood that the decoder outputs two target features of different scales. In order to better obtain the change detection binary segmentation map, these two low-resolution target features need to be channel-converted and upsampled. To better capture the details and features of the image, this embodiment proposes to generate a segmentation output from the target feature map of the decoder through a dynamic upsampling segmentation head (DySH).

[0068] In the embodiments of the present invention, the dynamic upsampling segmentation head is used to perform channel conversion and dynamic upsampling operations on the second target feature, and the second predicted feature map is segmented and output; and the dynamic upsampling segmentation head is used to perform channel conversion and dynamic upsampling operations on the first target feature, and the first predicted feature map is segmented and output.

[0069] Specifically, DySH first converts the multi-channel feature map into a single channel using a 1×1 convolution, and then applies a dynamic upsampler (DySample) to restore the binary segmentation single channel to the original size. As a lightweight and efficient upsampler, DySample adopts a point-based sampling method, which improves the sampling accuracy by dynamically adjusting the offset, while the computational burden is very small.

[0070] As Figure 6 shown, for example, for the first target feature map X with an input size of C×H×W and an upsampling factor of s, DySample first generates an initial offset, then uses a 1×1 convolution to calculate the dynamic offset. By combining the initial offset and the dynamic offset, the final offset O is obtained. After the offset O is rearranged, it is summed with the original sampling grid G to obtain the new sampling coordinates , as shown in the following formula:

[0071]

[0072] After normalizing the sampling coordinates , bilinear interpolation is used to upsample the input feature map and , and the first predicted feature map with an output size of is obtained, which can be expressed by the following formula:

[0073]

[0074] Exemplarily, the same upsampling operation as the above first target feature Figure 1 is performed on the second target feature map, and the second predicted feature map is output.

[0075] Furthermore, in the embodiments of the present invention, based on the second predicted feature map, an auxiliary loss function is created for depth (DS) supervision. After the first predicted feature map is processed by an activation function, a binary change detection result is obtained.

[0076] Specifically, based on the above, at different stages of the decoder, a dynamic upsampling segmentation head is used to generate two predicted feature maps, namely the first predicted feature map and the second predicted feature map . The first predicted feature map is used to generate the final segmentation result after passing through the function, and a binary change detection result is obtained. The second predicted feature map is used to create the overall loss function for depth supervision to alleviate the problems of gradient disappearance and slow convergence speed, where the loss function is shown in the following formula:

[0077]

[0078] is the cross - entropy loss function, which can be expressed as:

[0079]

[0080] where is the predicted probability of the model, obtained by inputting the output features into the activation function, and y is the true label.

[0081] Through the embodiments of the present invention, based on the settings of the VGG - 16 encoder and the dynamic up - sampling segmentation head, for small - scale targets (such as small roads and tiny changes in forest and grass), the forest and grass change detection model in this method can fully extract spatial detail features to achieve refined recognition and differentiation of the changed areas.

[0082] Based on the above - mentioned forest and grass change detection method based on multi - scale convolutional attention, in the embodiments of the present invention, ablation experiments are also carried out to verify the improvement of the model detection ability by the deep supervision (DS) strategy, CDF, and DySH.

[0083] Exemplarily, for the baseline model (BASE), the deep supervision strategy is not used. The dual - temporal feature fusion is performed using 3×3 convolution, batch normalization, and ReLU activation function. The segmentation head uses 1×1 convolution and bilinear interpolation for channel conversion and up - sampling to obtain the output. When the deep supervision strategy is used for this baseline model, the IoU increases by 1.39% and the F1 increases by 0.99%. When the feature fusion module in the baseline model is replaced by CDF, while the number of parameters of the model decreases and the computational efficiency increases, the IoU increases by 0.49% and the F1 increases by 0.35%. This is because CDF optimizes the integration of dual - temporal features of the model, enhances the spatio - temporal modeling ability of features for difficult samples, and effectively reduces missed detections. When DySH is replaced with almost the same number of parameters, the IoU of the model increases by 1.01% and the F1 increases by 0.7%. This is because DySH can dynamically adjust the sampling offset according to different features, making the mapping relationship between low - resolution and high - resolution features more flexible and accurate, and well identifying local detail information. For example, for subtle road changes in the image, after introducing the dynamic up - sampling head, the model can better locate the boundaries of the changed areas in the image and improve the recognition accuracy.

[0084] Through the embodiments of the present invention, in the scenario of forest and grass change detection, the forest and grass change detection model needs to have a powerful multi-scale feature extraction ability to accurately detect targets of different shapes and scales. Secondly, for difficult samples (such as change regions with high similarity to the background), the forest and grass change detection model in this method has strong discriminative ability to reduce missed detections and false detections. Finally, for targets with small scales (such as small roads and tiny forest and grass changes), the forest and grass change detection model in this method can fully extract spatial detail features to achieve refined recognition and differentiation of the change regions.

[0085] Figure 7 Schematically shows a comparison chart of the visualization results of change detection on a dataset to be measured by different methods according to embodiments of the present invention. Figure 8 Schematically shows a comparison chart of the visualization results of change detection on another dataset to be measured by different methods according to embodiments of the present invention.

[0086] Based on the above forest and grass change detection method based on multi-scale convolutional attention, the detection method of this embodiment and other detection methods based on existing models are compared and verified according to two sets of samples to be detected below.

[0087] Exemplarily, other models include FC-EF, FC-Siam-Conc, FC-Siam-Diff, SNUNet, HANet, CGNet, BIT, Change Former, and Changer. The two sets of samples to be detected are SHENMU-CD and GUIYANG-CD respectively. SHENMU-CD and GUIYANG-CD are forest and grass change detection datasets constructed based on GF-1 and GF-6 satellite remote sensing images and manually interpreted change patches.

[0088] Among them, SHENMU-CD contains a total of 11,559 groups of non-overlapping samples. The study area is located at the junction of the Mu Us Desert and the Loess Plateau, which is an agro-pastoral ecotone and belongs to a temperate semi-arid continental climate. The average annual temperature is 9.2 °C, and the average annual rainfall is about 466.6 mm. The terrain is relatively flat, and the vegetation is mainly shrub forest. GUIYANG-CD includes a total of 4,042 non-overlapping image pairs of size 256×256. The study area is located in the subtropical monsoon climate zone with superior hydrothermal conditions, lush vegetation growth, and a high surface cover rate. The landform types in this area are mainly mountains and hills. The complex terrain undulation leads to an increase in image registration errors and intensifies the ground object shadow effect at the same time.

[0089] Further, input various sample data in SHENMU-CD into the above-mentioned other models and the forest and grass change detection model of this embodiment to obtain an evaluation index table of the change detection results output by each model (as shown in Table 1) and a comparison chart of the change detection visualization results of different methods on SHENMU-CD (such as Figure 7 ).

[0090] Table 1 One of the evaluation index tables of the change detection results output by each model

[0091]

[0092] In the embodiment of the present invention, the detection performance is evaluated according to 4 indicators, namely Precision (Pre), Recall (Rec), Intersection over Union (IoU), and F1-score (F1). Higher IoU and F1 values indicate better overall change detection performance. The definitions of these indicators are as follows: Pre represents the precision rate, that is, the proportion of samples that are actually positive among all samples predicted as positive by the model. Rec represents the recall rate, that is, the ratio of the samples correctly judged as positive by the model to all samples that are actually positive. The higher the recall rate, the fewer positive samples are missed. The recall rate is a key indicator reflecting the model's ability to identify positive samples. IoU is the ratio of the intersection to the union of the predicted region and the true region, and it is a key indicator for measuring the overlap degree between the detection result and the ground truth data. When the IoU value is high, the similarity between the two is large. F1 is the harmonic mean of precision and recall, which is used to quantify the balance between them.

[0093] According to Table 1 above, it can be seen that the change detection results output by the forest and grass change detection model of this embodiment are better than the output results of other models in two comprehensive evaluation indicators such as IoU and F1, indicating that the overall change detection performance of the forest and grass change detection model is more excellent.

[0094] Further, such as Figure 7, After the dual-temporal images in the SHENMU-CD test sample set are input into each model, the change detection visualization results of each model are obtained. It should be noted that, in order to view different change regions, in this embodiment, TP: True Positive, the model prediction result is a change region, and it is actually a change region, that is, the number of correctly identified change regions. FP: False Positive, the model prediction result is a change region, and it is actually an unchanged region, that is, the number of misdetected unchanged regions. TN: True Negative, the model prediction result is an unchanged region, and it is actually an unchanged region, that is, the number of correctly identified unchanged regions. FN: False Negative, the model prediction result is an unchanged region, and it is actually a change region, that is, the number of missed change regions. Further, for convenience of viewing, different colors are used to represent TP (green), TN (black), FP (yellow), and FN (red).

[0095] It can be observed that the forest and grass change detection model in this embodiment has obtained more accurate results than other models, and the red and yellow parts are relatively less on the visualization graph (that is, relatively fewer missed and misdetected parts). In Figure 7 (1), the change regions detected by the model in this embodiment are closest to the ground truth, the generated boundaries are more accurate, and the missed detection rate is lower. As Figure 7 (3), 7(4), the model in this embodiment has good recognition ability for small regions and irregular regions. Figure 7 (5), there are a total of three smaller change regions. Other models can only identify one or two of them, while the model in this embodiment successfully and accurately identifies all change regions. This is because the forest and grass change detection model in this embodiment has a strong ability to extract detailed features, enhancing the detection effect of small targets.

[0096] In another embodiment, various sample data in GUIYANG-CD are input into the above-mentioned other models and the forest and grass change detection model in this embodiment, and an evaluation index table of the change detection results output by each model (as shown in Table 2) and a comparison graph of the change detection visualization results of different methods on GUIYANG-CD (such as Figure 8 ) are obtained.

[0097] Table 2 The second evaluation index table of the change detection results output by each model

[0098]

[0099] According to the above table, it can be seen that the change detection results output by the forest and grass change detection model in this embodiment are superior to the output results of other models in two comprehensive evaluation indexes such as IoU and F1, indicating that the overall performance of the change detection of the forest and grass change detection model is more excellent.

[0100] Further, as Figure 8 , after the dual-temporal images in the GUIYANG-CD test sample set are input into each model, the change detection visualization results of each model are obtained. It can be observed that the forest and grass change detection model of this embodiment has better recognition results for various types of forest and grassland changes, such as newly built roads, forest land logging, grassland damage, and pests and diseases compared with other models. For Figure 8 (1) the newly built road, the road detected by the forest and grass change detection model of this embodiment has better continuity. For Figure 8 (2) and Figure 8 (3) the forest land logging and grassland damage, the forest and grass change detection model of this embodiment has stronger detection ability for small change areas, and the detected edges are more complete. For Figure 8 (4) the pest and disease area, the forest and grass change detection model of this embodiment clearly identifies the forest land change area through the spectral feature difference.

[0101] In the above specific embodiments, the purpose, technical solution, and beneficial effects of the present invention are further described in detail. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A forest and grass change detection method based on multi-scale convolutional attention, characterized in that: The method comprises: Acquire a dual-phase image to be detected, wherein the dual-phase image is two remote sensing images of different phases; Preprocessing the dual-phase image to be detected and inputting it into a forest and grass change detection model to obtain a change detection binary image; The forest and grass change detection model is obtained by training with a forest and grass change detection training data set, and the training process of the model includes: The encoder is used to encode the dual-phase image to be trained to obtain the dual-phase feature; After the bi-phase features are integrated through the feature fusion module, they are input into the decoder for decoding to obtain target features of two different scales; The target features are segmented and outputted respectively using a dynamic upsampling segmentation head to obtain a first prediction feature map and a second prediction feature map. The first prediction feature map is used to characterize the binary result of change detection, and the second prediction feature map is used for deep supervision to optimize the feature learning of the forest and grass change detection model.

2. The forest and grass change detection method according to claim 1, characterized in that: in, The encoder is a Siamese encoder with a VGG-16 neural network structure as the backbone, including four convolution stages, each of which is a batch-normalized convolution block.

3. The forest and grass change detection method according to claim 2 is characterized in that: The step of encoding the dual-phase image to be trained by the encoder to obtain the dual-phase features includes: Obtaining a dual-phase image to be trained; The dual-phase image to be trained is preprocessed and input into a Siamese encoder with a VGG-16 neural network structure as the backbone to obtain the corresponding dual-phase features after being encoded by each convolutional block.

4. The forest and grass change detection method according to claim 3 is characterized in that: A feature fusion module is provided after each convolution stage corresponding to the Siamese encoder, and in each of the feature fusion modules, the bi-phase features are integrated, including: After each of the two-phase features is concatenated, an intermediate feature map is generated through a first layer of convolution operation; Performing a splitting operation on the intermediate feature map to obtain a first sub-feature map and a second sub-feature map after splitting, wherein the second sub-feature map is input into Bottleneck for processing; The second sub-feature map after the Bottleneck processing is concatenated with the first sub-feature map, and channel transformation and feature fusion are completed through a second layer of convolution operation to obtain a fused feature.

5. The forest and grass change detection method according to claim 1 is characterized in that: The decoder is a multi-scale convolutional attention decoder including: a large core grouping attention gate, an up-convolution block and a multi-scale convolutional attention module, and; The decoder includes four decoding stages corresponding to the feature fusion modules respectively.

6. The forest and grass change detection method according to claim 4 or 5, characterized in that: The two target features are respectively the first target feature and the second target feature. The plurality of fused features are input into a multi-scale convolutional attention decoder for decoding to obtain target features of two different scales, including: In a first decoding stage of the encoder, outputting a second target feature based on a multi-scale convolutional attention module; In the fourth decoding stage of the encoder, a first target feature is output based on a large kernel grouping attention gate, an up-convolution block, and a multi-scale convolution attention module.

7. The forest and grass change detection method according to claim 6, characterized in that: The method of using the dynamic upsampling segmentation head to segment and output the target features respectively to obtain a first prediction feature map and a second prediction feature map includes: The dynamic upsampling segmentation head is used to perform channel conversion and dynamic upsampling operations on the second target feature, and the second prediction feature map is segmented and output.

8. The forest and grass change detection method according to claim 6, characterized in that: The method further comprises: using a dynamic upsampling segmentation head to segment and output the target features respectively to obtain a first prediction feature map and a second prediction feature map; A dynamic upsampling segmentation head is used to perform channel conversion and dynamic upsampling operations on the first target feature, and a first prediction feature map is segmented and output.

9. The forest and grass change detection method according to claim 1 or 7, characterized in that: The method further comprises: Based on the second predicted feature map, an auxiliary loss function is created to perform deep supervision.

10. The forest and grass change detection method according to claim 1 or 8, characterized in that: The method further comprises: After the first prediction feature map is processed by the activation function, the change detection binary result is obtained.