Picea forest high-resolution remote sensing identification method and system based on MMA-U-Net model
By employing a high-resolution remote sensing identification method for spruce forests based on the MMA-U-Net model, and utilizing various image data and an improved attention mechanism, the method solves the problems of low data acquisition efficiency and inaccurate identification of spruce forests in existing technologies, and achieves efficient and accurate spruce forest identification.
Patent Information
- Application Number
- CN202510981438.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-16
- Publication Date
- 2025-10-31
AI Technical Summary
Existing methods for acquiring spruce forest data are inefficient and costly, and the U-Net model has insufficient generalization ability in complex image scenarios, resulting in inaccurate recognition results.
A high-resolution remote sensing identification method for spruce forests based on the MMA-U-Net model was adopted. By acquiring various different image data, an improved algorithm that integrates the CBAM attention mechanism and the DCA attention mechanism was used to train a semantic segmentation model for identifying spruce forests.
It improves the accuracy and generalization ability of spruce forest identification, enabling accurate extraction of spruce forest features and boundaries in complex environments, and reduces identification costs.
Smart Images

Figure CN120877102A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of remote sensing technology, and in particular to a high-resolution remote sensing identification method and system for spruce forests based on the MMA-U-Net model. Background Technology
[0002] The spruce forests of the Tianshan Mountains are an important ecological barrier, playing a crucial role in water source protection and preventing soil erosion. However, due to human activities such as over-logging, overgrazing, and tourism development, these spruce forests are suffering severe damage. Therefore, obtaining accurate information on the distribution of spruce forests is essential for effective resource management.
[0003] However, current methods for acquiring spruce forest data mainly rely on literature and manual measurement, which are inefficient and costly. Although the rise of deep learning technology has brought improvements, and the U-Net model has performed well in image segmentation, its generalization ability is limited, its ability to handle complex image scenes is insufficient, and it is easily affected by the quality of input data, leading to inaccurate recognition results. Summary of the Invention
[0004] The purpose of this application is to provide a high-resolution remote sensing identification method and system for spruce forests based on the MMA-U-Net model, which can accurately identify spruce forests.
[0005] To achieve the above objectives, this application provides the following solution:
[0006] Firstly, this application provides a high-resolution remote sensing identification method for spruce forests based on the MMA-U-Net model, including:
[0007] Acquire historical remote sensing image data of the target area; the historical remote sensing image data includes remote sensing image data obtained by at least two different image data acquisition methods; the investigation time of each of the remote sensing image data is the same;
[0008] The historical remote sensing image data of the target area is divided into a training set and a validation set according to a set ratio;
[0009] The MMA-U-Net semantic segmentation model is trained based on the training set and validation set to obtain a trained MMA-U-Net semantic segmentation model. The MMA-U-Net semantic segmentation model is a model based on the MMA-U-Net algorithm. The MMA-U-Net algorithm is an improved algorithm that incorporates the CBAM attention mechanism and the DCA attention mechanism on the U-Net model architecture.
[0010] Based on the trained MMA-U-Net semantic segmentation model, spruce forests are identified from real-time remote sensing image data of the target area.
[0011] Optionally, the remote sensing image data of the target area is preprocessed, specifically including:
[0012] ENVI 5.3 software was used to perform radiometric calibration and atmospheric correction on the remote sensing image data of the target area;
[0013] Orthorectification and image registration are performed on the radiometrically calibrated and atmospherically corrected remote sensing image data of the target area to obtain data after distortion elimination.
[0014] The data after distortion removal is cropped to complete data preprocessing.
[0015] Optionally, the MMA-U-Net semantic segmentation model includes a CBAM module and a DCA module.
[0016] Optionally, the CBAM module includes a channel attention submodule and a spatial attention submodule.
[0017] Optionally, the formula expression for the channel attention submodule is:
[0018] M C (F)=σ(MLP(Avgpool(F))+MLP(MaxPool(F)));
[0019] The formula expression for the spatial attention submodule is:
[0020] M s (F)=σ(f 7*7 ([Avgpool(F);MaxPool(F)]));
[0021] In the formula, M C (F) represents the channel attention module, M s (F) represents the spatial attention module, f 7*7 A 7×7 convolution kernel is used for the convolution operation, and F is the feature map.
[0022] Optionally, the DCA module includes a channel cross-attention submodule and a spatial cross-attention submodule; the channel cross-attention submodule utilizes cross-channel token cross-attention of multi-scale encoder features to extract global channel dependencies; the spatial cross-attention submodule performs cross-attention to capture spatial dependencies across spatial tokens to capture long-range dependencies.
[0023] Optionally, the formula expression for the channel cross-attention submodule is:
[0024]
[0025] The formula expression for the spatial cross-attention submodule is:
[0026]
[0027] In the formula, Q represents the query, K represents the key, V represents the value, and the subscript i indicates the i-th encoder. and is the scaling factor, CCA is channel cross attention, and SCA is spatial cross attention.
[0028] Optionally, after obtaining the trained MMA-U-Net semantic segmentation model, the following steps are also included:
[0029] The accuracy of the trained MMA-U-Net semantic segmentation model is verified based on the accuracy, recall, precision, F1 score, and mIOU metrics.
[0030] Optionally, the formulas for the Accuracy, Recall, Precision, F1Score, and mIOU metrics are as follows:
[0031]
[0032] Wherein, TP represents the correct classification of the target category by the trained MMA-U-Net semantic segmentation model, TN represents the correct classification of the non-target category by the trained MMA-U-Net semantic segmentation model, FN represents the incorrect classification of the non-target category by the trained MMA-U-Net semantic segmentation model, and FP represents the incorrect classification of the target category by the trained MMA-U-Net semantic segmentation model.
[0033] Secondly, this application provides a high-resolution remote sensing identification system for spruce forests based on the MMA-U-Net model, the high-resolution remote sensing identification system for spruce forests comprising:
[0034] The data acquisition module is used to acquire historical remote sensing image data of the target area; the historical remote sensing image data includes remote sensing image data obtained by at least two different image data acquisition methods; the survey time for each of the remote sensing image data is the same;
[0035] The dataset partitioning module is used to divide the historical remote sensing image data of the target area into training set and validation set according to a set ratio;
[0036] The model training module is used to train the MMA-U-Net semantic segmentation model based on the training set and validation set to obtain the trained MMA-U-Net semantic segmentation model; the MMA-U-Net semantic segmentation model is a model based on the MMA-U-Net algorithm; the MMA-U-Net algorithm is an improved algorithm that incorporates the CBAM attention mechanism and the DCA attention mechanism on the U-Net model architecture;
[0037] The identification module is used to identify spruce forests in the acquired real-time remote sensing image data of the target area based on the trained MMA-U-Net semantic segmentation model.
[0038] According to the specific embodiments provided in this application, the following technical effects are disclosed:
[0039] This application provides a high-resolution remote sensing identification method and system for spruce forests based on the MMA-U-Net model. First, historical remote sensing image data of the target area is acquired. This data contains rich ground feature information and forms the basis for training the model. Using remote sensing image data obtained through at least two different image acquisition methods can improve data diversity and accuracy, thereby enhancing the model's generalization ability. Second, the historical remote sensing image data of the target area is divided into a training set and a validation set according to a set ratio. The training set is used for model training, and the validation set is used to evaluate the model's performance. This division method ensures that the model can perform well on unseen data, thereby improving identification accuracy. Then, the MMA-U-Net semantic segmentation model is trained using the training set and the validation set. The MMA-U-Net algorithm is an improved algorithm that incorporates CBAM attention mechanisms and DCA attention mechanisms on the U-Net model architecture. This allows the model to more accurately capture key information in remote sensing images, such as the features of spruce forests. Through training, the model can learn how to extract the features of spruce forests from remote sensing images and achieve accurate identification of spruce forests. Finally, based on the trained MMA-U-Net semantic segmentation model, spruce forests were identified from the acquired real-time remote sensing image data of the target area. Since the model has learned the characteristics of spruce forests, it can accurately identify spruce forests in the real-time remote sensing images. Attached Figure Description
[0040] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0041] Figure 1A flowchart illustrating a high-resolution remote sensing identification method for spruce forests based on the MMA-U-Net model, provided as an embodiment of this application;
[0042] Figure 2 This is a schematic diagram of a dataset image provided in one embodiment of this application;
[0043] Figure 3 A structural diagram of the CBAM model provided in an embodiment of this application;
[0044] Figure 4 This is a structural diagram of a DCA module provided in an embodiment of this application;
[0045] Figure 5 This is a schematic diagram of a U-Net model provided in an embodiment of this application;
[0046] Figure 6 This is a schematic diagram of the MMA-U-Net model provided in an embodiment of this application;
[0047] Figure 7 A schematic diagram of a loss curve provided for an embodiment of this application;
[0048] Figure 8 This is a schematic diagram of the ablation experiment prediction results provided in an embodiment of this application;
[0049] Figure 9 This is a schematic diagram of the comparative experimental prediction results provided in an embodiment of this application;
[0050] Figure 10 A predicted distribution map of spruce forests provided in an embodiment of this application;
[0051] Figure 11 This is a schematic diagram showing the area and changes of a spruce forest according to an embodiment of this application;
[0052] Figure 12 This is a schematic diagram of the functional modules of a high-resolution remote sensing identification system for spruce forests based on the MMA-U-Net model, provided as an embodiment of this application. Detailed Implementation
[0053] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0054] The Tianshan spruce forests are a vital ecological barrier for the region, playing a crucial role in protecting water resources, preventing soil erosion, improving soil quality, and mitigating climate change. These forests are primarily distributed in cold temperate or tundra zones at altitudes of 1500-2800 meters, typically in remote and topographically complex areas. With the continuous expansion of human activities, the Tianshan spruce forests face serious ecological threats. Excessive deforestation, grazing, and tourism development have caused varying degrees of damage. Therefore, timely, effective, and accurate information on the area and spatial distribution of spruce forests is essential for forestry departments to formulate scientific and rational resource management plans and ensure the sustainable use of these resources.
[0055] In the initial stages of forestry surveys, spruce forest land data could only be obtained by consulting relevant literature and manual measurement, a method that was inefficient and required significant manpower and resources. Before the rise of deep learning, semantic segmentation mainly relied on traditional image processing techniques. These methods typically classify pixels based on low-level features such as color and texture. However, because these features are sensitive to changes in lighting and angle, the performance of traditional methods is often significantly limited. With the rise of deep learning technology, the U-Net model, due to its encoder-decoder structure, fuses feature information from different levels through skip connections, making it excellent at capturing detailed information and achieving good results in image segmentation. Compared to other deep learning models, the U-Net model has fewer parameters and is easy to extend and improve, allowing it to be trained in a shorter time and achieve good performance on limited training data. However, its generalization ability is limited for complex and varied image scenes, failing to effectively extract and utilize image features. Furthermore, traditional skip connections increase the dependence on input data; when the input data is of poor quality or contains noise, this negative information will be passed to the decoding stage, affecting the final segmentation result and leading to a decrease in segmentation accuracy. Therefore, scholars have conducted extensive research on improving U-Net.
[0056] With the introduction of attention mechanisms, numerous experiments have demonstrated that these mechanisms can focus on specific parts of the target object, reducing interference from noisy regions and effectively addressing the problem of inefficiently extracting and utilizing image features, thus improving the accuracy of semantic segmentation. Among them, the CBAM module can simultaneously model the channel and spatial information of an image, thereby better capturing important information and enhancing the model's ability to focus on targets and capture details. The DCA module improves upon traditional skip connections by sequentially capturing the channel and spatial dependencies between multi-scale encoder features to resolve the semantic gap between encoder and decoder features. However, the expressive power of a single attention mechanism is limited by its singular focus, failing to fully express the complexity and diversity of input data. It can only focus on a specific aspect or feature of the input data, and therefore may not capture all key information when handling complex tasks. Therefore, this study explores how to address the issue of partial information loss during downsampling in the U-Net network, affecting the integrity of the original image information, and how to solve the problem that traditional skip connections directly pass encoder feature information to the decoder, easily propagating noise or outliers. By combining attention mechanisms, we can better extract edge information of spruce forests and improve recognition accuracy. This is of great significance for promoting the application of deep learning models with multi-attention mechanisms in forest remote sensing identification.
[0057] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0058] Example 1
[0059] like Figure 1 As shown, this embodiment provides a high-resolution remote sensing identification method for spruce forests based on the MMA-U-Net model, including:
[0060] Step 101: Acquire historical remote sensing image data of the target area; the historical remote sensing image data includes remote sensing image data obtained by at least two different image data acquisition methods; the survey time for each of the remote sensing image data is the same;
[0061] Step 102: Divide the historical remote sensing image data of the target area into a training set and a validation set according to a set ratio;
[0062] Step 103: Train the MMA-U-Net semantic segmentation model based on the training set and validation set to obtain the trained MMA-U-Net semantic segmentation model; the MMA-U-Net semantic segmentation model is a model based on the MMA-U-Net algorithm; the MMA-U-Net algorithm is an improved algorithm that incorporates the CBAM attention mechanism and the DCA attention mechanism on the U-Net model architecture;
[0063] Step 104: Based on the trained MMA-U-Net semantic segmentation model, identify spruce forests in the acquired real-time remote sensing image data of the target area.
[0064] In some embodiments, when performing step 101, the specific steps may be as follows:
[0065] The target area in this embodiment is located in the western region of the Tianshan Mountains, along both sides of the Tianshan mountain range, with an altitude of 174m-4544m, situated at 45°26′N-42°19′N and 79°52′E-85°00′E.
[0066] Specifically, the remote sensing image data used are GF-1PMS data from 2015 and 2022. The original GF-1PMS data has multispectral images with a spatial resolution of 8 meters. The images have four bands: red, green, blue, and near-infrared. Considering the impact of clouds and snow caused by high latitude and high altitude, as well as the area of the study area, images from June to September were selected. The DEM data was jointly measured by NASA and NIMA and has been publicly released since 2003. In this embodiment, the remote sensing images obtained through another acquisition method are a combination of forest land Class II survey data provided by relevant departments and visual interpretation. The time of the forest land Class II survey data corresponds one-to-one with the year of the remote sensing images.
[0067] In some embodiments, after performing step 101, the specific steps may be as follows:
[0068] ENVI 5.3 software was used to perform radiometric calibration and atmospheric correction on the remote sensing image data of the target area;
[0069] Orthorectification and image registration are performed on the radiometrically calibrated and atmospherically corrected remote sensing image data of the target area to obtain data after distortion elimination.
[0070] The data after distortion removal is cropped to complete data preprocessing.
[0071] Specifically, the Gaofen-1 PMS data used in this embodiment is Level 1 product data, requiring preprocessing. ENVI 5.3 software is used for radiometric calibration and atmospheric correction of the images to eliminate errors caused by the atmosphere and sensors, improving image quality. Orthorectification and image registration are then performed to eliminate geometric distortions caused by atmospheric refraction and Earth's curvature, which affect the true coordinate information of objects. Finally, the images are cropped to obtain the desired area. Considering the altitude of spruce forest growth and the effects of cloud and snow cover and occlusion in the images, sample areas are selected through visual interpretation and Class II survey data.
[0072] In some embodiments, when performing step 102, the specific steps may be as follows:
[0073] Using 2015 and 2022 imagery as base maps, mixed samples were created in three different regions. Regional labels were generated using visual interpretation in ArcGIS 10.8 software, and spruce forests and other land types were manually classified. After manual classification, the vector data was converted to raster data with the same resolution as the GF-1PMS imagery; spruce forests were assigned pixel values of 0, and other types were assigned pixel values of 1. Python was used to crop the images and corresponding labels to 256*256 pixels, and the labels were matched with the images. Data augmentation techniques were employed, including horizontal image flipping and rotation at different angles, to construct the dataset. Images with NODATA pixel values were removed to prevent interference. The dataset was split into training and validation sets in a 4:1 ratio, resulting in 5715 images. 4572 images were used for the training set, and 1143 images were used for the validation set (e.g., ...). Figure 2 (As shown).
[0074] In addition, radar data (such as Sentinel-1) can be fused to solve the cloud and fog obstruction problem. Specifically, this can be achieved through the following steps:
[0075] First, radar data, such as Sentinel-1, is used to obtain surface information, as it is unaffected by clouds and fog. Radar data can penetrate clouds and fog, providing information on surface structure, soil moisture, vegetation, and other data, which is used to understand the type and state of land cover.
[0076] Then, the radar data is fused with optical remote sensing imagery. While optical remote sensing imagery, such as GF-1PMS, has limited information acquisition under cloud and fog conditions, it can provide high-resolution land cover information in clear weather. By matching and fusing radar data with optical remote sensing imagery using algorithms, a high-resolution land cover information map can be generated that is unaffected by cloud and fog obstruction.
[0077] The specific steps mentioned in the embodiments, such as creating mixed samples, visual interpretation using ArcGIS software, manual classification, data augmentation, and dataset partitioning, are all aimed at building high-quality training and validation datasets for training deep learning models or performing other types of analysis. Ultimately, by fusing radar data and optical remote sensing imagery, accurate land cover information can be generated even under cloud and fog conditions.
[0078] In some embodiments, when performing step 103, the specific steps may be as follows:
[0079] This embodiment uses the U-Net architecture as its base network. By adding Hybrid Attention (CBAM), it focuses on important channels and spatial locations in the input feature map of the network model, thereby improving the feature extraction capability of spruce forests. It also replaces the original skip connections with DCA modules to resolve the semantic gap between encoder and decoder features and improve the model segmentation accuracy.
[0080] Specifically, the CBAM model framework is as follows: Figure 3 As shown, CBAM is a lightweight attention mechanism in deep learning designed to enhance the ability of convolutional neural networks to model and represent image features. It performs attention operations in both spatial and channel dimensions, enabling the model to dynamically adjust the weights of feature maps to adapt to different tasks and scenarios. CBAM comprises two sub-modules: Channel Attention (CAM) and Spatial Attention (SAM), which help the network focus more on the feature regions of objects from both channel and spatial perspectives, improving classification accuracy and better adapting to different image features. Therefore, this embodiment adds a CBAM module between the downsampling convolution and the activation function, directly affecting the weight allocation of features extracted after the convolution operation. This allows the module to directly affect the convolution output, directly influencing the feature representation ability, more accurately focusing on important regions or feature channels in the input, and thus optimizing the subsequent activation function processing. In this binary classification example, it can focus more on the target itself, thereby achieving better performance in capturing the boundaries and features of the spruce forest.
[0081] M C (F)=σ(MLP(Avgpool(F))+MLP(MaxPool(F))).
[0082] M s (F)=σ(f 7*7 ([Avgpool(F);MaxPool(F)])).
[0083] In the formula, M C (F) is the CAM module, M s (F) represents the SAM module, f 7*7 A 7×7 convolution kernel is used for the convolution operation, and F is the feature map.
[0084] Specifically, the DCA module, such as Figure 4As shown, Dual Cross-Attention (DCA) is a simple yet effective attention module proposed by a research team at the University of Miami in 2023. It aims to enhance skip connections in the U-Net architecture simply and effectively with only a slight increase in parameters and complexity. The DCA module mainly consists of Channel Cross-Attention (CCA) and Spatial Cross-Attention (SCA). First, a multi-scale block embedding module is used to obtain encoder tokens. The CCA module extracts global channel dependencies by utilizing cross-channel token cross-attention of multi-scale encoder features. The SCA module performs cross-attention to capture spatial dependencies across spatial tokens to capture long-range dependencies. These fine-grained encoder features are upsampled and connected to their corresponding decoder parts to form a skip connection scheme. By sequentially capturing the channel and spatial dependencies between multi-scale encoder features, the semantic gap between encoder and decoder features is resolved. This connection method helps optimize the feature fusion process, allowing the network to fully utilize feature information from different levels, thereby improving model performance while maintaining a simple network structure.
[0085] The formula expression for the channel cross-attention submodule is:
[0086]
[0087] The formula expression for the spatial cross-attention submodule is:
[0088]
[0089] In the formula, Q represents the query, K represents the key, V represents the value, and the subscript i indicates the i-th encoder. and is the scaling factor, CCA is channel cross attention, and SCA is spatial cross attention.
[0090] Among them, such as Figure 5 As shown, the U-Net model is an improved FCN (Fully Convolutional Network) structure. This model consists of a compressed channel on the left and an expanded channel on the right. It utilizes convolutional and pooling layers for feature extraction, and then uses deconvolutional layers to restore the image size. Employing an encoder-decoder structure, it effectively extracts feature information from the image through downsampling and upsampling operations, and maps this feature information back to the original image's spatial dimensions during the decoding stage. Furthermore, the U-Net model introduces skip connections to connect features between the encoder and decoder, which helps improve the model's segmentation accuracy. Compared to other deep learning models, the U-Net model has fewer parameters, allowing it to be trained in a shorter time and achieve good performance on limited training data.
[0091] Among them, such as Figure 6 As shown in this embodiment, an MMA-U-Net algorithm is proposed, which is built on the U-Net model as the basic model framework and incorporates CBAM and DCA attention mechanisms. By adding a CBAM module during downsampling, attention weights are applied to the input feature map in both channel and spatial dimensions, dynamically adjusting the feature map weights to make the network focus more on important features and reduce interference from redundant information, thus improving performance without increasing network complexity. A DCA module is used to replace skip connections. The Channel Cross-Attention (CCA) part in the module extracts global channel dependencies through cross-attention of cross-channel tokens of multi-scale encoder features. Then, the Spatial Cross-Attention (SCA) part in the module performs cross-attention to capture spatial dependencies of cross-spatial tokens. This solves the semantic gap caused by skip connections when connecting encoder and decoder features, which prevents the locality of convolution from capturing long-distance dependencies between different features. This allows the network to focus on key regions in the input image and allocate more computational resources to them, helping the network understand the interactions and dependencies between different channels, thereby better representing the features of the input image.
[0092] In the decoding stage, a reverse attention module can be designed to enhance boundary features by suppressing background noise. First, the input and output of the decoding stage are defined. The input to the decoding stage typically comes from the encoder's feature maps, while the output is the feature maps after upsampling or transposed convolution, which are used to restore the resolution of the original image. Next, the reverse attention module is designed. The core idea of the reverse attention module is to suppress background noise while enhancing boundary features. This can be achieved through the following steps:
[0093] I. Calculating the attention map from the feature map. The attention map is obtained by calculating the similarity at each position in the feature map. Methods such as cosine similarity and dot product similarity can be used. The size of the attention map is usually the same as that of the feature map.
[0094] 2. Reverse processing of the attention map. To suppress background noise, high-value regions (usually representing background regions) in the attention map can be set to low values, while low-value regions (usually representing boundaries or target regions) can be set to high values. This can be achieved by thresholding the attention map, inverting it, or applying a non-linear function.
[0095] Third, multiply the inversely processed attention map with the original feature map. This suppresses background noise while enhancing boundary features. The result of the multiplication is used as the output of the inverse attention module.
[0096] Finally, the inverse attention module is integrated into the decoding stage. The output of the inverse attention module can be used as input to a convolutional or upsampling layer in the decoder, or directly added to or multiplied by the decoder's output. In this way, the decoder is influenced by the inverse attention module when generating high-resolution feature maps, thereby suppressing background noise and enhancing boundary features.
[0097] Specifically, during the training and validation sets of the MMA-U-Net semantic segmentation model, the decoding stage is designed to incorporate a reverse attention module.
[0098] 1) Define the input and output of the decoding stage; the input of the decoding stage comes from the feature map of the encoder, and the output is the feature map after upsampling or transposed convolution, which is used to restore the resolution of the original image;
[0099] 2) Design the reverse attention module; calculate the attention map of the feature map using methods such as cosine similarity and dot product similarity; reverse process the attention map, setting high-value regions to low values and low-value regions to high values to suppress background noise and enhance boundary features; multiply the reverse-processed attention map with the original feature map to obtain the output of the reverse attention module;
[0100] 3) Integrate the inverse attention module into the decoding stage; use the output of the inverse attention module as the input of a convolutional layer or upsampling layer in the decoder, or directly add or multiply it with the output of the decoder;
[0101] 4) Based on the trained MMA-U-Net semantic segmentation model, combined with the reverse attention module in the decoding stage, the real-time remote sensing image data of the acquired target area is used to identify spruce forests.
[0102] In some embodiments, after obtaining the trained MMA-U-Net semantic segmentation model, the following is also included:
[0103] The accuracy of the trained MMA-U-Net semantic segmentation model is verified based on the accuracy, recall, precision, F1 score, and mIOU metrics.
[0104] Specifically, to demonstrate the performance of the improved model proposed in this embodiment in semantic segmentation, the same benchmark dataset was used, and the SE-U-Net, ResUNet, and ECA-U-Net models were compared and validated with the improved model proposed in this embodiment. To quantitatively evaluate the effectiveness of the proposed method, five metrics were selected: Accuracy, Recall, Precision, F1 Score, and mIOU (mean Intersection over Union) to provide a comprehensive evaluation of recognition accuracy. In analyzing the results, visual interpretation was used as the ground truth, the model's prediction results were used as the predicted values, and the accuracy based on pixel count was finally calculated.
[0105] The formulas for the Accuracy, Recall, Precision, F1Score, and mIOU metrics are as follows:
[0106]
[0107] Wherein, TP represents the correct classification of the target category by the trained MMA-U-Net semantic segmentation model, TN represents the correct classification of the non-target category by the trained MMA-U-Net semantic segmentation model, FN represents the incorrect classification of the non-target category by the trained MMA-U-Net semantic segmentation model, and FP represents the incorrect classification of the target category by the trained MMA-U-Net semantic segmentation model.
[0108] In some embodiments, the identification result obtained when performing step 104 may be as follows:
[0109] Spruce, a dominant tree species in the Tianshan Mountains, forms spectacular mountain coniferous forests along the mountain range. It mainly grows in the low-to-mid-mountain forest-steppe zone to the subalpine sparse forest zone at altitudes of 1500 to 2800 meters. The altitude range of this embodiment is 174 to 4544 meters. Based on the vertical zonation of mountain ranges in parts of the Tianshan Mountains by Zhang Xinshi et al., and appropriately adjusted according to the actual conditions of this embodiment, appropriate adjustments were made. Using Arc MAP 10.8 software, this embodiment divides the study area into six vertical zones: low-to-mid-mountain forest-steppe zone (below 150 to 1500 meters), low-to-mid-mountain forest-steppe zone (1500 to 1700 meters), mid-mountain forest-meadow zone (1700 to 2250 meters), upper mid-mountain forest-meadow zone (2250 to 2550 meters), subalpine sparse forest zone (2550 to 2700 meters), and the area above the subalpine sparse forest zone (2700 to 4600 meters).
[0110] In setting the experimental parameters, this embodiment used a desktop computer equipped with a Windows 11 operating system. The experiment used Python 3.7 and the PyTorch 1.13 framework for programming. The computer's processor was an Intel(R) Xeon(R) Silver4210R CPU @ 2.40GHz, and the graphics processing unit (GPU) was an NVIDIA RTX A5000 with 24GB of video memory. During model training, this embodiment selected the SGD algorithm as the optimizer, with an initial learning rate of 0.01, a momentum parameter of 0.9, a batch size of 8, and 250 epochs.
[0111] To reflect the dynamic trends of network training and understand whether the model has converged or overfitted, loss curves are used to evaluate model performance, such as... Figure 7 As shown. If the model is overfitted, the validation loss curve will become increasingly higher as the training loss curve becomes lower. Overfitting can exhibit unstable performance fluctuations during training; the model's performance on the validation or test set may fluctuate wildly as it tries to fit every detail of the training data. In the first round of training, the loss values on the training and validation sets differ significantly, indicating that the model has not yet completed its learning. After 100 rounds of training, the two loss curves drop to the same level and remain stable, indicating that the model training is complete.
[0112] To verify that the proposed method in this embodiment improves accuracy, an ablation experiment was designed for comparative analysis. The CBAM and DCA modules were added sequentially to the U-Net model using the same configuration parameters and training environment. The test results are shown in Table 1, and the prediction results are as follows: Figure 8 As shown.
[0113] Table 1 Ablation Experiment Results
[0114] CBAM DCA Accuracy / % Recall / % Precision / % F1Score / % mIOU / % χ χ 81.61 79.80 77.80 77.69 70.8 √ χ 87.98 88.44 88.56 88.57 72.6 χ √ 88.18 88.18 88.11 88.06 72.4 √ √ 93.06 92.22 93.26 93.89 81.8
[0115] Table 1 shows the segmentation accuracy results, indicating that the initial U-Net model performs worse than the model with the added attention mechanism module across all evaluation metrics. In contrast, using the DCA and CBAM modules respectively shows relative improvements in the evaluation metrics. Adding the CBAM module improves Accuracy, Recall, Precision, and F1Score by 6.37%, 8.64%, 10.76%, and 10.88%, respectively, and mIOU by 1.8%. The DCA module improves Accuracy, Recall, Precision, F1Score, and mIOU accuracy by 6.57%, 8.38%, 10.31%, 10.37%, and 1.6%, respectively. Comparing the two modules, the DCA module shows higher Accuracy than the CBAM module, while the DCA module performs worse in other evaluation metrics, with only minor differences in accuracy (all within 1%). With the addition of both DCA and CBAM modules, the segmentation performance and accuracy of the network model are significantly improved compared to the original U-Net model. Accuracy, Recall, Precision, and F1Score are improved by 11.45%, 16.20%, 15.46%, and 12.42%, respectively, and mIOU is improved by 10.0%. This indicates that the DCA and CBAM modules adopted in this embodiment can extract image features more fully, retain more complete semantic information, and better extract the area and boundary of spruce trees.
[0116] To verify the model's generalization ability, this embodiment selected three validation areas: urban areas and mountainous regions with dense spruce forests. The prediction results showed that the initial U-Net model suffered significant feature loss in spruce forest extraction, resulting in instances where spruce forests were not identified. Adding the CBAM module alone improved spruce forest extraction, but some misclassifications occurred. Adding the DCA module alone resulted in no misclassifications, but incomplete extraction of fine boundaries. Adding both the DCA and CBAM modules simultaneously resulted in predictions that were nearly identical to the labels, demonstrating excellent performance in extracting image feature information from each module and effectively identifying spruce forest boundaries and features.
[0117] This embodiment also includes a comparative analysis of different methods, as detailed below:
[0118] The model proposed in this embodiment shows a significant improvement in extraction capability compared to other models (Table 2), achieving the highest performance across all evaluation metrics. ECA-U-Net is the second best performing model, while ResUNet performs the worst, ranking last in all evaluation metrics. Compared to models SE-U-Net, ResUNet, and ECA-U-Net that incorporate additional attention mechanisms, the method in this embodiment achieves higher accuracy (10.29%, 19.84%, and 5.42%), higher recall (10.72%, 20.08%, and 11.73%), higher precision (12.29%, 22.65%, and 11.97%), higher F1 score (11.97%, 20.35%, and 10.13%), and higher average interaction ratio (8.9%, 9.1%, and 9.00%). In terms of single-image prediction time, U-Net is the fastest, while MMA-U-Net is the slowest due to its complex structure and numerous modules.
[0119] Table 2 Comparison of experimental results
[0120] Model Accuracy / % Recall / % Precision / % F1Score / % mIOU / % speed / (it.s-1) U-Net 81.61 79.80 77.80 77.69 71.8 7.01 SE-U-Net 82.77 81.50 80.97 83.76 72.9 6.86 ResUNet 73.22 72.14 70.61 73.54 72.7 6.98 ECA-U-Net 87.64 80.49 83.29 81.92 72.8 4.98 MMA-U-Net 93.06 92.22 93.26 93.89 81.8 1.28
[0121] from Figure 9 It can be seen that the SE-U-Net model in the validation area extracts a rather mottled spruce forest, with only the basic shape of the spruce forest and poor integrity in terms of extracted boundary details and spruce forest area. The ResUNet model extracts a spruce forest area that differs significantly from the label, and it is basically unable to identify spruce forests, resulting in a large number of misidentifications. The ECA-U-Net model performs well in extracting edge details and can roughly extract spruce forests, but it still has misclassifications, classifying other forest areas as spruce forests. In comparison, the method proposed in this embodiment has the best performance, and the extracted spruce forests have higher segmentation accuracy and edge information than other models.
[0122] in, Figure 10 This is a map showing the predicted distribution of spruce forests in the study area in 2015 and 2022 using the MMA-U-Net model proposed in this embodiment. Figure 10 (a) shows the distribution of spruce forests in 2015, and (b) shows the distribution in 2022. Visual interpretation confirmed the predicted results, which largely match the visual interpretation results. The figures show the distribution of spruce forests across different vertical zones. The distribution patterns are largely similar between the two years, primarily concentrated between altitudes of 1700-2250m and 2250-2550m, which are also the main growth areas for spruce forests. Within the study area, spruce forests are relatively concentrated in the upper left corner, with most of the spruce forests at altitudes of 100-1500m located in this area. Spruce forests in other areas are mainly distributed along the sides of the central mountain range.
[0123] Figure 11The distribution of spruce forest area at different altitudes is shown, with (a) representing the spruce forest area in 2015, (b) representing the spruce forest area in 2022, and (c) showing the change in spruce forest area from 2015 to 2022. The spruce forest area below the mid-to-low mountain forest-steppe zone in 2015 was 112.54 km². 2 In 2022, the area of spruce forest in this region was 75.73 km². 2 The area decreased by 36.81 km² between 2015 and 2022. 2 In areas below the low-to-mid-mountain forest-steppe zone, mainly affected by tourism and animal husbandry, the overall trend is downward; in 2015, the area of spruce forest in the low-to-mid-mountain forest-steppe zone was 141.00 km². 2 In 2022, the area of spruce forest in this region was 162.20 km². 2 The area increased by 21.2 km². 2 The area of the spruce forests did not change significantly compared to the overall area. The mid-mountain forest-meadow zone to the subalpine sparse forest zone is the main growing area of spruce forests and also the area with the most significant area changes, accounting for more than 80% of the total spruce forest area in the study area. In 2015, the area of spruce forests in the mid-mountain forest-meadow zone was 1398.62 km². 2 In 2022, the area of spruce forest in this region was 1515.20 km². 2 The area of spruce forest in the upper-middle mountain forest-meadow zone in 2015 was 1282.49 km². 2 The area in 2022 was 1363.90 km². 2 The subalpine sparse forest zone covered an area of 378.03 km² in 2015. 2 The area in 2022 was 392.70 km². 2 Due to the influence of the spruce forest's growing environment, this region accounts for the majority of the spruce forest distribution area. Thanks to the vigorous development of forest resources by the Xitianshan Forest Farm, the original spruce forests in the region are growing normally, and the area of spruce forests in all three vertical zones is steadily increasing; above the subalpine sparse forest zone, the area of spruce forests in the low-to-mid-mountain forest-steppe zone was 195.51 km² in 2015. 2 In 2022, the area of spruce forest in this region was 179.56 km². 2 The area was reduced by 15.95km. 2 The overall area showed an increasing trend from 2015 to 2022, with an increase of 181.10 km². 2 .
[0124] Example 2
[0125] like Figure 12 As shown, this embodiment provides a high-resolution remote sensing identification system for spruce forests based on the MMA-U-Net model. The high-resolution remote sensing identification system for spruce forests includes:
[0126] The data acquisition module 1201 is used to acquire historical remote sensing image data of the target area; the historical remote sensing image data includes remote sensing image data obtained by at least two different image data acquisition methods; the survey time of each of the remote sensing image data is the same;
[0127] The dataset partitioning module 1202 is used to divide the historical remote sensing image data of the target area into a training set and a validation set according to a set ratio.
[0128] The model training module 1203 is used to train the MMA-U-Net semantic segmentation model according to the training set and validation set to obtain the trained MMA-U-Net semantic segmentation model; the MMA-U-Net semantic segmentation model is a model based on the MMA-U-Net algorithm; the MMA-U-Net algorithm is an improved algorithm that incorporates the CBAM attention mechanism and the DCA attention mechanism on the U-Net model architecture;
[0129] The identification module 1204 is used to identify spruce forests in the acquired real-time remote sensing image data of the target area based on the trained MMA-U-Net semantic segmentation model.
[0130] In summary, this application has the following technical effects:
[0131] This application proposes a multi-hybrid attention semantic segmentation model, MMA-U-Net, using Gaofen-1 PMS data and employing U-Net as the backbone network. This method addresses the common problem of traditional U-Net models neglecting channel and spatial information and semantic gaps in images. During downsampling, a CBAM module is added to mitigate the loss of spruce forest boundary information. Furthermore, the original skip connections are replaced with dual cross-attention (DCA), enhancing the model's ability to represent spruce forest features within the network layers and reducing semantic gaps. Through the coordinated use of multiple attention mechanisms, the accuracy of spruce forest identification is improved.
[0132] This application also utilizes DEM data to conduct a comprehensive analysis of the spatial distribution and changes in spruce forest area. This study expands its application in forest management and provides technical support for refined management.
[0133] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0134] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A high-resolution remote sensing identification method for spruce forests based on the MMA-U-Net model, characterized in that, The high-resolution remote sensing identification method for spruce forests includes: Acquire historical remote sensing image data of the target area; the historical remote sensing image data includes remote sensing image data obtained by at least two different image data acquisition methods; the investigation time of each of the remote sensing image data is the same; The historical remote sensing image data of the target area is divided into a training set and a validation set according to a set ratio; The MMA-U-Net semantic segmentation model is trained based on the training set and validation set to obtain a trained MMA-U-Net semantic segmentation model. The MMA-U-Net semantic segmentation model is a model based on the MMA-U-Net algorithm. The MMA-U-Net algorithm is an improved algorithm that incorporates the CBAM attention mechanism and the DCA attention mechanism on the U-Net model architecture. Based on the trained MMA-U-Net semantic segmentation model, spruce forests are identified from real-time remote sensing image data of the target area.
2. The high-resolution remote sensing identification method for spruce forests based on the MMA-U-Net model according to claim 1, characterized in that, Data preprocessing of the remote sensing image data of the target area specifically includes: ENVI 5.3 software was used to perform radiometric calibration and atmospheric correction on the remote sensing image data of the target area; Orthorectification and image registration are performed on the radiometrically calibrated and atmospherically corrected remote sensing image data of the target area to obtain data after distortion elimination. The data after distortion removal is cropped to complete data preprocessing.
3. The high-resolution remote sensing identification method for spruce forests based on the MMA-U-Net model according to claim 1, characterized in that, The MMA-U-Net semantic segmentation model includes a CBAM module and a DCA module.
4. The high-resolution remote sensing identification method for spruce forests based on the MMA-U-Net model according to claim 3, characterized in that, The CBAM module includes a channel attention submodule and a spatial attention submodule.
5. The high-resolution remote sensing identification method for spruce forests based on the MMA-U-Net model according to claim 4, characterized in that, The formula expression for the channel attention submodule is: M C (F)<σ(MLP(Avgpool(F))+MLP(MaxPool(F))); The formula expression for the spatial attention submodule is: M s (F)=σ(f 7*7 ([Avgpool(F);MaxPool(F)])); In the formula, M C (F) represents the channel attention module, M s (F) represents the spatial attention module, f 7*7 A 7×7 convolution kernel is used for the convolution operation, and F is the feature map.
6. The high-resolution remote sensing identification method for spruce forests based on the MMA-U-Net model according to claim 3, characterized in that, The DCA module includes a channel cross-attention submodule and a spatial cross-attention submodule; the channel cross-attention submodule uses cross-channel tagging cross-attention of multi-scale encoder features to extract global channel dependencies; The spatial cross-attention submodule performs cross-attention to capture spatial dependencies across spatial tokens in order to capture remote dependencies.
7. The high-resolution remote sensing identification method for spruce forests based on the MMA-U-Net model according to claim 6, characterized in that, The formula expression for the channel cross-attention submodule is: The formula expression for the spatial cross-attention submodule is: In the formula, Q represents the query, K represents the key, V represents the value, and the subscript i indicates the i-th encoder. and is the scaling factor, CCA is channel cross attention, and SCA is spatial cross attention.
8. The high-resolution remote sensing identification method for spruce forests based on the MMA-U-Net model according to claim 1, characterized in that, After obtaining the trained MMA-U-Net semantic segmentation model, the following is also included: The accuracy of the trained MMA-U-Net semantic segmentation model is verified based on the accuracy, recall, precision, F1 score, and mIOU metrics.
9. A high-resolution remote sensing identification method for spruce forests based on the MMA-U-Net model according to claim 8, characterized in that, The formulas for the Accuracy, Recall, Precision, F1Score, and mIOU metrics are as follows: Wherein, TP represents the correct classification of the target category by the trained MMA-U-Net semantic segmentation model, TN represents the correct classification of the non-target category by the trained MMA-U-Net semantic segmentation model, FN represents the incorrect classification of the non-target category by the trained MMA-U-Net semantic segmentation model, and FP represents the incorrect classification of the target category by the trained MMA-U-Net semantic segmentation model.
10. A high-resolution remote sensing identification system for spruce forests based on the MMA-U-Net model, characterized in that, The high-resolution remote sensing identification system for the spruce forest includes: The data acquisition module is used to acquire historical remote sensing image data of the target area; the historical remote sensing image data includes remote sensing image data obtained by at least two different image data acquisition methods; the survey time for each of the remote sensing image data is the same; The dataset partitioning module is used to divide the historical remote sensing image data of the target area into training set and validation set according to a set ratio; The model training module is used to train the MMA-U-Net semantic segmentation model based on the training set and validation set to obtain the trained MMA-U-Net semantic segmentation model; the MMA-U-Net semantic segmentation model is a model based on the MMA-U-Net algorithm; the MMA-U-Net algorithm is an improved algorithm that incorporates the CBAM attention mechanism and the DCA attention mechanism on the U-Net model architecture; The identification module is used to identify spruce forests in the acquired real-time remote sensing image data of the target area based on the trained MMA-U-Net semantic segmentation model.