Tree species AI identification method based on multi-modal data fusion
By adopting multimodal data fusion AI method in tree species recognition, combined with GF-2, LiDAR and RGB data, the problem of insufficient recognition accuracy in complex forest environments is solved, and higher recognition accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510243887.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-06-27
AI Technical Summary
Traditional tree species recognition methods rely on a single data source, making it difficult to achieve high-precision recognition in complex and changeable forest environments, especially in subtropical or tropical forests.
A tree species AI recognition method based on multimodal data fusion is used to identify tree species through a fusion network (MTSCFNet) using high-resolution multispectral data (GF-2), lidar point cloud feature (LiDAR) and ultra-high resolution RGB images.
It significantly improves the accuracy and robustness of tree species classification, and realizes large-area, automated and high-precision forest resource survey, management and ecological protection.
Smart Images

Figure CN120219950A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the cross - technical field of forestry science and computer science, and relates to a tree species AI recognition method based on multi - modal data fusion. Background Art
[0002] Accurate tree species recognition is of great significance for forest resource management, biodiversity conservation, and ecological research. However, traditional tree species recognition methods mainly rely on a single data source (such as spectral images, LiDAR data, or RGB images). These methods face many challenges in complex and variable forest environments, such as insufficient information, noise interference, and insufficient feature representation, resulting in limited classification accuracy. Especially in subtropical or tropical forests, due to the diverse tree species, complex canopy structures, and variable background environments, single - data - source methods often fail to achieve ideal recognition results. In addition, with the continuous development of high - resolution remote sensing technology, although it provides more detailed spatial information, it also increases the intra - class variability, making pixels within the same tree species likely to be misclassified into different categories, further reducing the overall classification accuracy. Therefore, how to effectively fuse multi - modal data to improve the accuracy and robustness of tree species recognition has become a hot and difficult issue in current research. Summary of the Invention
[0003] The present invention solves the technical problems existing in the prior art, and thus provides a tree species AI recognition method based on multi - modal data fusion, which involves using artificial intelligence technology for tree species recognition and is a deep - learning recognition method for tree species based on multi - modal data of high - resolution image GF - 2, LiDAR point - cloud features, and ultra - high - resolution RGB.
[0004] The existing technical problems are solved through the following technical solutions: A tree species AI recognition method based on multi - modal data fusion, comprising the following steps: A tree species recognition method based on a fusion network (MTSCFNet) of GF - 2, LiDAR point - cloud features, and ultra - high - resolution RGB multi - modal data, which makes up for the limitations of single - data - source or dual - data - source in current tree species recognition, improves the tree species recognition accuracy in multi - canopy, multi - tree - species subtropical forests, and aims to achieve large - area, automated, and high - precision forest resource surveys, management, and ecological protection.
[0005] The advantages of the present invention are as follows: The proposed tree species AI recognition method based on multi - modal data fusion (MTSCFNet) aims to break through the limitations of single - modal data in complex multi - canopy, multi - tree - species subtropical forest environments, such as insufficient information, noise interference, and insufficient feature representation, which lead to problems such as limited recognition accuracy and insufficient robustness. At the same time, the recognition accuracy of MTSCFNet has fully met the actual application requirements in traditional forest resource inventory, management, and protection and other fields.
[0006] By fusing multi-modal data such as GF-2 (high-resolution multi-spectral data), LiDAR (LiDAR point cloud feature data), and RGB (high-resolution color images), MTSCFNet significantly improves the accuracy of tree species classification. The experimental results show that the classification effect using multi-modal data (R+L+S) is the best, with the average F1-score and MCC reaching 0.928 and 0.924 respectively, significantly superior to single-modal data (such as using only RGB, LiDAR, or GF-2 data).
[0007] Based on the UNet architecture, MTSCFNet introduces residual blocks and a gated attention mechanism, enhancing the model's feature extraction ability. Residual blocks can extract higher-level abstract features by deepening the network hierarchy, while the gated attention mechanism can dynamically integrate features at different scales, suppress noise, and highlight important regions, thus improving the classification accuracy. The experimental results show that the average F1-score and MCC of MTSCFNet in multiple test areas are 0.90 and 0.82 respectively, significantly superior to other comparison models (such as DeepLabV3+, ResUNet, AUNet, etc.), especially showing stronger robustness when dealing with complex forest environments.
[0008] MTSCFNet performs excellently in cross-region transfer experiments under different forest density conditions. In the three test areas with different densities (high density, medium density, low density) in Huangfengqiao Forest Farm, the average F1-score and MCC of MTSCFNet are 0.85 and 0.80 respectively, superior to other comparison models. This shows that MTSCFNet has strong generalization ability and can adapt to complex environments with different forest densities and structures. Especially in high-density and medium-density forests, the classification accuracy of MTSCFNet remains at a high level. In low-density forests, although the accuracy decreases slightly, it is still superior to other models, showing its robustness in complex backgrounds.
[0009] By visualizing the feature attention weights of MTSCFNet at different levels, it is found that the model can effectively integrate multi-level feature information. High-level features (Level3 and Level4) capture global context information, while low-level features (Level1 and Level2) retain local spatial details (such as edges and textures). This dynamic integration of multi-level features enables MTSCFNet to achieve more accurate tree species classification in complex forest environments. The visualization of feature attention weights also shows that MTSCFNet can effectively suppress noise and highlight regions that contribute more to classification, thus improving the classification accuracy.
[0010] Visual analysis of the features extracted by MTSCFNet using t-SNE technology reveals that the features extracted by MTSCFNet exhibit high inter-class separability in the low-dimensional space. Especially in the case of multi-modal data fusion (R+L+S), the feature distributions of different tree species are clearer and there is less inter-class overlap. This indicates that MTSCFNet can effectively extract discriminative features, thereby improving the classification accuracy. In contrast, the feature separability of single-modal data (such as only using RGB or LiDAR) is poor and there is more inter-class overlap, further demonstrating the necessity of multi-modal data fusion.
[0011] Compared with traditional pixel-based random forest (PBRF) and object-based random forest (OBRF), MTSCFNet has a significant advantage in classification accuracy. The average F1-scores of PBRF and OBRF are 0.74 and 0.73 respectively, while the average F1-score of MTSCFNet reaches 0.90. In addition, MTSCFNet can effectively reduce the "salt-and-pepper noise" in classification and generate smoother and more accurate classification results. Brief Description of the Drawings
[0012] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. As shown in the following figures:
[0013] Figure 1a It is the main structure diagram of MTSCFNet of the present invention.
[0014] Figure 1b It is the diagram of the gated attention mechanism, upsampling module and basic residual module of the present invention.
[0015] Figure 2a It is the diagram of the change of the loss value in the training stage of the MTSCFNet model of the present invention.
[0016] Figure 2b It is the diagram of the change of the F1 score in the training stage of the MTSCFNet model of the present invention.
[0017] Figure 3a It is the diagram of the change of the precision rate of the tree species classification results of MTSCFNet of the present invention under different modal data.
[0018] Figure 3b It is the diagram of the change of the recall rate of the tree species classification results of MTSCFNet of the present invention under different modal data.
[0019] Figure 3c It is a graph showing the change of F1 score of the tree species classification results of the MTSCFNet of the present invention under different modality data.
[0020] Figure 3d It is a graph showing the change of Matthews correlation coefficient MCC of the tree species classification results of the MTSCFNet of the present invention under different modality data.
[0021] Figure 4 It is a graph showing the tree species classification results of MTSCFNet, DeepLabv3+, ResUNet, AUNet, UNet, OBRF and PBRF in the S1-S3 regions of the Yamashita Forest Farm of the present invention.
[0022] Figure 5 It is a graph showing the tree species classification results of MTSCFNet, DeepLabV3+, ResUNet, AUNet and UNet of the present invention under different forest density scenarios in 3 target migration areas of the Huangfengqiao State-owned Forest Farm.
[0023] Figure 6a It is a graph showing the tree species classification accuracy of MTSCFNet, DeepLabV3+, ResUNet, AUNet and UNet of the present invention in the migration area T1.
[0024] Figure 6b It is a graph showing the tree species classification accuracy of MTSCFNet, DeepLabV3+, ResUNet, AUNet and UNet of the present invention in the migration area T2.
[0025] Figure 6c It is a graph showing the tree species classification accuracy of MTSCFNet, DeepLabV3+, ResUNet, AUNet and UNet of the present invention in the migration area T3.
[0026] Figure 7 It is the graph of the attention weight of each scale feature of the present invention.
[0027] Figure 8 It is a two-dimensional representation graph of the tree species pixel feature space in the S1-S3 regions of the Yamashita Forest Farm of the present invention.
[0028] Figure 9 It is the feature separability graph of tree species classification using different combinations of multi-modal data in the Yamashita Forest Farm in Jiangxi Province of the present invention. Detailed implementation manners
[0029] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0030] Embodiment 1: As Figure 1a , Figure 1b , Figure 2a , Figure 2b , Figure 3a , Figure 3b , Figure 3c , Figure 4 , Figure 5 , Figure 6a , Figure 6b , Figure 6c , Figure 7 , Figure 8 and Figure 9 shown, a tree species AI recognition method based on multi-modal data fusion, a tree species recognition method based on a fusion network (MTSCFNet) of GF-2, LiDAR and ultra-high resolution RGB multi-modal data, to make up for the limitations of single data source or dual data sources in current tree species recognition, and improve the tree species recognition accuracy in multi-canopy and multi-tree species subtropical forests. By fusing data features of different modalities, including multi-dimensional features such as spectrum, structure and texture, it aims to achieve large-area, automated and high-precision forest resource survey, management and ecological protection.
[0031] The tree species AI recognition method based on multi-modal data fusion (MTSCFNet), its overall architecture is as Figure 1a and Figure 1b shown.
[0032] Figure 1a is the main structure of MTSCFNet, where Res1 to Res4 are composed of multiple stacked basic residual blocks for extracting hierarchical features.
[0033] Figure 1b is the architecture of the gated attention module, upsampling module and basic residual block.
[0034] Xf and Xc respectively represent the input features from the fine scale and the coarse scale.
[0035] A tree species AI recognition method based on multi-modal data fusion includes the following steps: MTSCFNet adopts an encoder-decoder architecture, where the encoder part of MTSCFNet is used to extract multi-level features from these different data sources (including GF-2, LiDAR, and RGB), and transform them into abstract representations to capture key information at different levels of detail.
[0036] Subsequently, by integrating these multi-level abstract representations, the decoder enhances the model's ability to make accurate predictions at the pixel level, ensuring that the fine details required for high-resolution classification are retained and effectively utilized.
[0037] Given that tree species exhibit spectral variations as well as spatial and structural characteristics, efficiently capturing and representing these unique features is crucial for accurate analysis and classification.
[0038] To enhance the extraction of high-level semantic features of the target species, the encoder is improved by integrating a residual network (ResNet), enabling it to obtain more distinctive and deeper features.
[0039] In MTSCFNet, the final pooling layer and fully connected layer of ResNet are omitted, and ResNet-50 is selected as the encoder to achieve the best balance between classification accuracy and computational efficiency.
[0040] In addition, although high-level features provide extensive context information about the trees, such as tree shape, canopy structure, and spectral variations, the downsampling process may lead to a reduction in spatial information that is crucial for pixel-level precise fine-grained predictions.
[0041] In UNet, traditional skip connections are designed to connect the encoder and decoder layers to retain spatial details and enhance feature reconstruction. However, these connections are often insufficient in effectively fusing low-level and high-level features, which may lead to the introduction of redundant noise or the loss of key context information.
[0042] A tree species AI recognition method based on multi-modal data fusion includes the following steps: MTSCFNet introduces a gated attention mechanism in its decoder to dynamically fuse coarse-scale and fine-scale features by generating soft spatial weights, addressing the challenge of feature fusion at different levels.
[0043] Therefore, MTSCFNet effectively reduces irrelevant noise and emphasizes important regions, thereby improving its ability to accurately define the boundaries of tree species.
[0044] In addition, the gated attention module of MTSCFNet enhances feature fusion by focusing on pixels that make significant contributions to prediction refinement.
[0045] By effectively integrating high-level and low-level features, this mechanism not only emphasizes key details but also provides insights into the types of features extracted by MTSCFNet at different levels.
[0046] In terms of model training, cross-entropy loss has been widely used to train deep neural networks and is very effective in various classification problems due to its excellent performance in penalizing incorrect predictions based on the class probability distribution. However, due to class imbalance, cross-entropy loss faces challenges in tree species classification, where the dataset usually contains a skewed distribution, that is, the occurrence frequency of some classes (such as some tree species) is much lower than that of other species, which may cause the model to overly favor the majority class.
[0047] To address these problems, a tree species AI recognition method based on multi-modal data fusion includes the following steps: The MTSCFNet model is trained using Dice loss. The Dice coefficient is used to measure the overlap between the prediction result and the actual tree species classification. By focusing on the accuracy of the prediction overlap rather than just the class distribution accuracy, it is particularly effective in dealing with unbalanced class distributions.
[0048] The expression of Dice loss L is shown in Equation (1).
[0049]
[0050] Where y i and represent the predicted output and the actual ground truth of pixel i, respectively. N is the number of pixels in the input image. During the calculation, to avoid numerical instability, especially to prevent division by zero, a very small constant ε is added.
[0051] Example 2: Taking the Chinese fir mixed forest in the Shanxia Experimental Forest Farm in Jiangxi Province as the specific research area, this area shows extremely high species diversity. A tree species AI recognition method based on multi-modal data fusion includes the following steps:
[0052] Among them, Chinese fir (Cunninghamia lanceolata), Machilus pauhoi, and Schima superba dominate the tree population in this area, accounting for more than 95%, forming a complex Chinese fir mixed forest with multiple canopy layers and coexisting tree species.
[0053] 1. Experimental design
[0054] 1.1 Comparative experiments on different combinations of modal data
[0055] To comprehensively evaluate the effectiveness of the tree species AI recognition method based on multi-modal data fusion (MTSCFNet) in tree species classification using different combinations of multi-modal remote sensing data, the present invention tested various combinations, including (1) GF-2, LiDAR, and RGB data; (2) GF-2 and LiDAR data; (3) GF-2 and RGB data; (4) LiDAR and RGB data; (5) LiDAR data; (6) RGB data; (7) GF-2 data.
[0056] The present invention adopts an early fusion strategy, combines these two or three data types into a unified multi-band image for classification, and thus systematically evaluates the classification accuracy of all combinations and the impact of each data source on the tree species classification performance.
[0057] It is worth noting that before merging the datasets, the GF-2 and LiDAR point cloud feature maps were resampled to align with the resolution (0.0623m) of the RGB data, ensuring the spatial consistency of the input data.
[0058] 1.2 Ablation experiments of MTSCFNet
[0059] To evaluate the performance of the tree species AI recognition method based on multi-modal data fusion (MTSCFNet) in recognizing tree species with residual blocks and gated attention mechanisms, the present invention developed ResUNet and AUNet. Both of these networks are derived from the standard UNet architecture and were compared and evaluated with MTSCFNet for tree species classification performance.
[0060] Specifically, ResUNet deepens the network hierarchy by incorporating residual blocks, thereby enhancing the performance of UNet; while AUNet integrates an attention mechanism to improve the fusion effect between the hierarchical features of the UNet encoder.
[0061] In addition, the present invention also introduced DeepLabV3+ as a comparison benchmark for MTSCFNet. DeepLabV3+ uses ResNet-50 as the backbone network and is a leading convolutional neural network (CNN) model in semantic segmentation applications.
[0062] To further compare the differences between deep learning and traditional machine learning methods, the present invention also adopted two common methods, namely pixel-based random forest (PBRF) and object-based random forest (OBRF).
[0063] 1.3 Transfer experiments of different forest features
[0064] The size and shape of the tree crown are largely influenced by forest characteristics such as stand density, which significantly affects the crown characteristics in mixed forests.
[0065] To evaluate the transferability and generalization ability of the tree species AI recognition method based on multi-modal data fusion (MTSCFNet) in forests with different densities, the present invention selected the Yamashita Forest Farm as the original area for model training and the Huangfengqiao Forest Farm as the target area for transfer testing.
[0066] According to the percentage of adjacent tree crown overlap in the selected study area, the Huangfengqiao Forest Farm was further divided into three stand density levels: dense (T1), medium (T2), and sparse (T3). In the transfer experiment, MTSCFNet was initialized with random weights and trained using limited data in each stand density level (T1, T2, and T3).
[0067] 1.4 Experimental process design
[0068] All models of the present invention underwent 1000 training cycles to achieve a stable convergence state in ablation experiments and comparative experiments. To accelerate the training process, the present invention adopted the Adam optimizer, in which β1 and β2 were set to 0.9 and 0.999.
[0069] To further improve the generalization ability and optimization effect of the model, the present invention implemented a multi-step learning rate decay strategy, with the initial learning rate set to 1e-2.
[0070] In addition, all experiments were based on the TensorFlow (2.8.0) architecture, using an NVIDIA GeForce RTX 3070 graphics card with 8GB VRAM, CUDA version 11.4, an Intel Core i7-10700K processor (main frequency 3.80GHz), 128GB of memory, and were conducted in the Ubuntu 18.04.6 LTS operating system environment.
[0071] 1.5 Performance evaluation
[0072] Precision, Recall, F1-score, and Matthews Correlation Coefficient (MCC) were used to evaluate the performance of tree species classification. Precision measures the proportion of false positives, while Recall measures the proportion of false negatives. The F1-score combines them by calculating the harmonic mean of Precision and Recall, providing a measure of overall accuracy for identifying tree species. At the same time, MCC provides a reliable metric for quantifying the agreement between predicted and true results, especially useful when dealing with imbalanced datasets (Waldner and Diakogiannis, 2020). The four evaluation metrics are calculated as shown in equations (2)-(5):
[0073]
[0074]
[0075]
[0076]
[0077] Among them, TP, TN, FP, and FN represent the number of pixels correctly classified as a specific tree species, pixels classified as non-specific tree species, pixels misclassified as a specific tree species, and pixels misclassified as non-specific tree species, respectively. Precision corresponds to the commission error, while Recall corresponds to the omission error. The F1-score is a measure of the overall accuracy of tree species classification. The Matthews Correlation Coefficient (MCC) measures the agreement between predicted and true values.
[0078] 2. Experimental Results
[0079] 2.1 Model Performance
[0080] Figure 2a 、 Figure 2b Shows the changes in loss values and F1-scores during the training of MTSCFNet.
[0081] Overall, the gap between the training set and the validation set is small, and there are no obvious signs of overfitting during the training process.
[0082] Specifically, the loss value drops rapidly in the initial stage of training, while the F1-score rises sharply, indicating that MTSCFNet can quickly learn key features in the early stage of training.
[0083] From approximately the 200th epoch to the 400th epoch, the loss value decreased at a slower rate after a brief stabilization, and the F1 score of the validation set showed minor fluctuations, which reflected that the model was adapting to more complex patterns. By the 400th epoch, the loss values of both the training set and the validation set reached a stable state, and the final loss value of the validation set stabilized at around 0.2( Figure 2a ).
[0084] Meanwhile, the F1 score continued to increase gradually. By the end of training, the F1 score of the training set approached the optimal value of 0.9, and the F1 score of the validation set was approximately 0.8( Figure 2b ). These results indicate that MTSCFNet has strong tree species classification accuracy even in subtropical forest environments.
[0085] 2.2 Results of comparative experiments with different multimodal data combinations
[0086] Figure 3a , Figure 3b , Figure 3c and Figure 3d show the performance of MTSCFNet in tree species identification when using different combinations of multimodal image datasets (i.e., R+L+S, R+S, R+L, L+S, R, L, S). Among them, the R+L+S combination showed the best performance, achieving the highest precision, recall, F1 score, and Matthews correlation coefficient (MCC) of 0.928, 0.928, 0.928, and 0.924 respectively. In addition, the R+L+S had the smallest standard deviation in all metrics, highlighting its robustness and stability in accurately classifying tree species by leveraging complementary data sources.
[0087] In contrast, the performance of MTSCFNet relying solely on a single data source (i.e., R, L, S) was lower but still acceptable, with its precision, recall, and F1 score ranging from 0.884 to 0.912, and the MCC being approximately 0.88 to 0.89.
[0088] However, the standard deviations of these models were higher, indicating an increase in their variability and unreliability when dealing with complex classification tasks. The performance of the bimodal datasets (i.e., R+S, R+L, L+S) was at a medium level, with their precision, recall, and F1 score ranging from 0.922 to 0.927, and the MCC approaching 0.92. Among the bimodal combinations, R+S and R+L performed better than L+S, indicating that combining RGB data with spectral data or LiDAR point cloud feature maps can provide stronger complementary features.
[0089] Nonetheless, the performance of the R+L+S dataset was still better than all other combinations, indicating that integrating more diverse data sources enables the model to capture a wider range of information, thereby improving the accuracy and consistency of classification.
[0090] 2.3 Ablation Experiment Results of MTSCFNet
[0091] Figure 4 Shows the tree species classification results of five deep learning models (i.e., MTSCFNet, DeepLabV3+, ResUNet, AUNet, and UNet) and two machine learning methods (i.e., RBRF and OBRF) in three test areas (S1 - S3) of the Yamashita Forest Farm. The classification accuracies evaluated using precision, recall, F1-score, and Matthews correlation coefficient (MCC) are summarized in Table 1.
[0092] Among all the methods, MTSCFNet performed the best, achieving the highest average MCC (0.82) and F1-score (0.90), indicating its robustness and effectiveness in accurately distinguishing tree species.
[0093] Other deep learning models, such as ResUNet and AUNet, achieved comparable average MCC values (0.79), slightly higher than UNet (0.78) but lower than MTSCFNet (0.82), which reflects the contribution of residual modules and attention mechanisms to improving classification accuracy.
[0094] In contrast, the machine learning-based methods performed relatively poorly. The average MCC and F1-score of PBRF were below 0.60 and 0.70 respectively, with the weakest classification ability. As Figure 4 shown, the maps generated by it had obvious "salt and pepper" noise. Although OBRF classified based on objects rather than pixels, significantly reducing this noise, its accuracy metrics were still mediocre, with average F1-score, precision, and recall of about 0.73, showing only a slight improvement compared to PBRF.
[0095] However, deep learning methods significantly reduced commission errors and omission errors, resulting in more accurate and reliable classification results.
[0096] Notably, MTSCFNet outperformed DeepLabV3+ in all test areas, thanks to its adaptive hierarchical feature fusion strategy driven by the attention mechanism.
[0097] In contrast, the decoder in DeepLabV3+ was difficult to fully utilize multi-level features, leading to classification errors, especially for tree species with similar texture, spectral features, and vertical structures. These limitations resulted in a higher commission error rate for DeepLabV3+.
[0098] These results together highlight the effectiveness of advanced architectures like MTSCFNet in handling complex forest classification tasks, especially in heterogeneous environments.
[0099] Table 1 Accuracy of tree species classification using MTSCFNet, DeepLabV3+, ResUNet, AUNet, UNet, OBRF, and PBRF in the S1 - S3 areas of the Shanxia Forest Farm.
[0100]
[0101]
[0102] Figure 4 Tree species classification results of MTSCFNet, DeepLabv3+, ResUNet, AUNet, UNet, OBRF, and PBRF in the S1 - S3 areas of the Shanxia Forest Farm. By comparing the model outputs with the ground truth data, the correctness, commission, and omission situations of tree species classification were determined.
[0103] 2.4 Migration performance of MTSCFNet among different forest densities
[0104] Figure 5 The tree species classification results under different forest density conditions in three model migration areas (T1 - T3) are shown, while Figure 6 presents the corresponding classification accuracies.
[0105] Among the tested models, MTSCFNet, ResUnet, and AUNet demonstrated the best robustness, with the average MCC and F1 scores stably remaining above 0.80 and 0.85 respectively in all migration areas. MTSCFNet showed excellent performance, with its average MCC and F1 scores increasing by 0.07 and 0.02 respectively compared to DeepLabV3+ with the lowest migratability. DeepLabV3+ performed particularly poorly under heterogeneous forest density conditions.
[0106] MTSCFNet performed particularly well in high - density (T1) and medium - density (T2) forest areas, maintaining a high degree of consistency with the reference data due to its advanced feature extraction and fusion mechanisms.
[0107] However, under sparse forest conditions with complex understory vegetation and backgrounds (such as T3), its performance decreased slightly, but it was still the most robust compared to other models.
[0108] In contrast, DeepLabV3+ failed to effectively capture multi - level hierarchical features, resulting in higher commission errors and omission errors as well as greater performance variability. ResUNet and AUNet incorporated residual blocks and attention mechanisms respectively, and their accuracies were higher than that of UNet, especially in the T1 area, which highlighted the contributions of these enhanced architectures.
[0109] However, their accuracy in all transferred regions is consistently lower than that of MTSCFNet, indicating that the combined use of residual learning and attention mechanism in MTSCFNet provides significant advantages in tree species classification. MTSCFNet integrates RGB, GF2, and LiDAR point cloud feature data, enabling it to effectively adapt to different forest characteristics such as stand density and canopy structure, thus ensuring high transferability.
[0110] Overall, these results confirm that MTSCFNet is the most robust and transferable model for large-scale tree species identification, demonstrating its ability to identify with superior accuracy and reliability while handling different forest environments and density conditions.
[0111] Figure 6a , Figure 6b and Figure 6c Tree species classification accuracy of MTSCFNet, DeepLabV3+, ResUNet, AUNet and UNet in the target migration areas of the three models.
[0112] 2.5 Analysis of MTSCFNet’s feature extraction capabilities
[0113] 2.5.1 Feature Attention Weights Across Different Scales
[0114] Figure 7 The feature attention weights of the three test regions at different levels are shown. Level 1 to Level 4 represent the attention weights extracted from the shallower to deeper layers of MTSCFNet, where brighter areas correspond to higher attention weights, indicating that these features contribute more to the classification, while darker areas represent lower attention weights.
[0115] Despite the differences in forest characteristics among the three testing areas, the spatial distribution of attention weights presents a consistent pattern, which highlights the robustness of MTSCFNet in adapting to different forest environments.
[0116] The high-level features produced by deeper layers of the network (i.e., Level3 and Level4) are highly consistent with the ground-truth, highlighting their importance in capturing global contextual and semantic information that is crucial for tree species classification.
[0117] These features with large receptive fields effectively incorporate spatial relationships to identify category changes. However, they are limited in their effectiveness for fine-grained prediction at the pixel level due to the loss of details caused by downsampling.
[0118] Conversely, the low-level features extracted from the shallower layers of the network (i.e., Level1 and Level2) focus on capturing local spatial details, and higher attention weights are concentrated at the edges and boundaries of tree species. These features are good at identifying spatial details such as textures, edges, and corners, which are crucial for fine classification tasks. This hierarchical feature representation - from fine details at the low level to global context information at the high level - enables MTSCFNet to comprehensively analyze and classify tree species with different characteristics.
[0119] Overall, the combination of hierarchical feature representation and the attention mechanism makes the performance of MTSCFNet superior to traditional machine learning methods such as OBRF and PBRF, which lack the ability to effectively balance local and global feature extraction.
[0120] Figure 7 Feature attention weights at different scales. Level1 to Level4 represent the attention weights extracted from the shallower to deeper layers of MTSCFNet, where brighter / darker regions correspond to higher / lower attention weights respectively.
[0121] 2.5.2 Feature Separability between Different Test Regions
[0122] From the test regions S1, S2, and S3, the present invention randomly selected 2000 pixel points of each tree species (i.e., Schima superba, Michelia figo, Osmanthus fragrans, Vernicia fordii, Cunninghamia lanceolata, Machilus pauhoi), and 300 pixel points of Liquidambar formosana for t-SNE analysis. Using the t-SNE technique, the present invention converts the high-dimensional features extracted from GF-2, LiDAR point cloud features, RGB data, and the final encoding layer of the deep neural network into a two-dimensional space to visually display the feature separability between tree species.
[0123] Figure 8 Shows the feature distribution in the three test regions.
[0124] The separability of the original features is the lowest, with significant overlap between different categories in the two-dimensional space, indicating that relying solely on the original spectral, structural, and texture features cannot effectively distinguish these seven tree species.
[0125] Conversely, the features extracted using advanced deep learning models (such as UNet and MTSCFNet) have shown a significant improvement in separability, thanks to hierarchical modeling and multi-scale feature fusion. These features provide a clearer distinction for tree species categories and demonstrate superiority in classification performance compared to traditional machine learning methods (such as OBRF and PBRF) that rely only on the original features.
[0126] Among the evaluated models, MTSCFNet outperformed UNet, producing a feature space with less overlap and higher inter-class separability in all three test regions (S1, S2, S3). The higher separability achieved by MTSCFNet can be attributed to its integration of residual connections and attention mechanisms, which enhance the abstraction ability of representative and discriminative features for each tree species.
[0127] These results indicate that MTSCFNet is more effective than UNet in exploring the complex, high-dimensional relationships in the data and transforming them into unique low-dimensional representations.
[0128] Figure 8 Two-dimensional representation of the tree species pixel feature space in the S1 - S3 regions. This feature space demonstrates the ability of specific features to distinguish tree species pixels. The greater the separability between tree species in the feature space, the higher the accuracy of tree species classification.
[0129] 2.5.3 Feature Separability under Different Combinations of Multimodal Datasets
[0130] In this study, the present invention randomly selected 2000 pixel points for each tree species (i.e., Schima superba, Michelia figo, Osmanthus fragrans, Vernicia fordii, Cunninghamia lanceolata, Machilus pauhoi) and 300 pixel points of Liquidambar formosana for t-SNE visualization analysis to evaluate the feature separability of tree species using different combinations of multimodal datasets.
[0131] Figure 9 Shows the two-dimensional feature separability under different combinations of RGB, GF2, and LiDAR point cloud feature data, including single-modal datasets (R0, L0, S0), combinations of two datasets (R1+S1, L1+R1, S1+L1), and combinations of three datasets (S1+L1+R1).
[0132] The results show that the R1+L1+S1 combination, which combines the abstract features of GF-2, LiDAR point cloud features, and RGB data, exhibits the highest feature separability among the seven tree species with the smallest inter-class overlap.
[0133] This indicates that multimodal datasets provide richer and more comprehensive information, which helps in abstractly representing features and enhancing the distinguishability between tree species. In contrast, the single-modal datasets (R0, L0, S0) have the worst separability with significant inter-class overlap, indicating that a single data source is insufficient to represent the diverse features of different tree species.
[0134] The combinations of two datasets (R1+S1, R1+L1, L1+S1) perform between single-modal and the combination of three datasets. Although the combinations of two datasets provide more information than single-modal, they still cannot achieve the high separability shown by the combination of three datasets.
[0135] Compared with single-modal or two data sets, the multi-modal feature space constructed by MTSCFNet shows less overlap and higher inter-class separability among seven tree species, indicating its excellent ability to abstract and fuse features from multiple data sources in the complex masson pine mixed forest environment.
[0136] Figure 9 Feature separability for tree species classification using different combinations of multi-modal data in the mountainous area of Jiangxi Province. L0 and R0 represent the original features from LiDAR data and RGB data respectively. S1, L1, and R1 represent the abstract features from GF-2, LiDAR, and RGB data respectively. S1+L1, L1+R1, R1+S1, and S1+L1+R1 represent the abstract features of the combination of GF-2 and LiDAR data, LiDAR and RGB data, RGB and GF-2 data, and the combination of GF-2, LiDAR, and RGB data respectively.
[0137] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A tree species AI identification method based on multimodal data fusion, characterized in that , contains the following steps: A tree species identification method based on the fusion network (MTSCFNet) of GF-2, LiDAR and ultra-high resolution RGB multimodal data makes up for the limitations of single data source or dual data sources in current tree species identification, improves the accuracy of tree species identification in multi-canopy and multi-species subtropical forests, and aims to achieve large-scale, automated, and high-precision forest resource survey, management, and ecological protection.
2. The tree species AI identification method based on multimodal data fusion according to claim 1 is characterized in that MTSCFNet adopts an encoder-decoder architecture, where the encoder part of MTSCFNet is used to extract multi-level features from these different data sources (including GF-2, LiDAR, and RGB) and transform them into abstract representations to capture key information at different levels of detail. In MTSCFNet, the final pooling layer and fully connected layer of ResNet are omitted, and ResNet-50 is selected as the encoder to achieve the best balance between classification accuracy and computational efficiency. MTSCFNet introduces a gated attention mechanism in its decoder to dynamically fuse coarse-scale and fine-scale features by generating soft spatial weights. The MTSCFNet model is trained using Dice loss, which uses the Dice coefficient to measure the overlap between the predicted results and the actual tree species classification. By focusing on the accuracy of the predicted overlap rather than just the accuracy of the category distribution, The expression of Dice loss L is shown in formula (1): where y i and They represent the predicted output and actual true value of pixel i respectively, N is the number of pixels in the input image. During the calculation process, a small constant ε is added to avoid numerical instability, especially to prevent division by zero.
3. The tree species AI identification method based on multimodal data fusion according to claim 1 is characterized in that: It also contains the following steps: an early fusion strategy is used to merge two or three data types into a unified multi-band image for classification, the classification accuracy of all combinations and the impact of each data source on the tree species classification performance are evaluated, and before merging the datasets, the GF-2 and LiDAR point cloud feature maps are resampled to align with the resolution of the RGB data (0.0623m) to ensure the spatial consistency of the input data. ResUNet and AUNet are used. Both networks are derived from the standard UNet architecture and are compared with MTSCFNet for tree species classification performance evaluation. ResUNet deepens the network hierarchy by incorporating residual blocks to enhance the performance of UNet; AUNet integrates the attention mechanism to improve the fusion effect between the hierarchical features of the UNet encoder. DeepLabV3+ is used as the comparison benchmark for MTSCFNet. DeepLabV3+ uses ResNet-50 as the backbone network and adopts pixel-based random forest (PBRF) and object-based random forest (OBRF) steps.