Fine tree species recognition model construction method and system, recognition method and system

By constructing a high-dimensional multi-source feature dataset and introducing a cross-scale deformable attention fusion module to optimize the ConvNeXt architecture, the problems of insufficient examination of temporal feature contributions and difficulty in capturing detailed features in remote sensing tree species classification are solved, thereby improving the accuracy and generalization ability of the remote sensing tree species classification model.

CN120673265BActive Publication Date: 2025-11-04NORTHEAST INST OF GEOGRAPHY & AGRIECOLOGY C A S
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511188371.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-25
Publication Date
2025-11-04
Estimated Expiration
2045-08-25

AI Technical Summary

Technical Problem

Existing technologies fail to fully examine the overall contribution of the entire set of temporal features to model performance, and deep learning models struggle to capture detailed features in Sentinel-2 images, resulting in a significant decline in the performance of remote sensing tree species classification models.

Method used

A high-dimensional multi-source feature dataset is constructed, and the optimal classification feature set is obtained through feature optimization. A deep learning sample library is constructed by combining the measured sample data, and a cross-scale deformable attention fusion module is introduced to perform feature alignment and fusion. The ConvNeXt architecture is optimized to improve classification accuracy.

Benefits of technology

It improves the ability to extract spatial structural details and express spatiotemporal features in medium-resolution remote sensing images, enhances the accuracy and generalization ability of remote sensing tree species classification, and is suitable for fine-grained tree species identification and image segmentation tasks under complex forest stand structures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673265B_ABST
    Figure CN120673265B_ABST
Patent Text Reader

Abstract

The fine tree species recognition model construction method, construction system, recognition method and system belong to the field of forest management and ecological protection, and solve the technical problems that the existing technology cannot comprehensively examine the overall contribution of the whole time sequence characteristics to the model performance, and due to the resolution limitation of remote sensing images and the randomness of forest tree species distribution space and quantity, the performance of the remote sensing tree species classification model is obviously reduced. Step 1, a high-dimensional multi-source feature data set is constructed; step 2, feature optimization is carried out based on the high-dimensional multi-source feature data set to obtain an optimal classification feature set; step 3, real measurement sample data is collected, and a deep learning sample library is constructed in combination with the optimal classification feature set; step 4, a preset fine tree species recognition model is input into the deep learning sample library to train the preset fine tree species recognition model, and a fine tree species recognition model is obtained. The present application is used for obtaining accurate and latest tree species distribution map.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of forest management and ecological protection, specifically to a method and system for constructing a refined tree species identification model, and an identification method and system thereof. Background Technology

[0002] Forest ecosystems constitute the core of global terrestrial ecosystems. Accurate understanding of the spatial distribution of forest types and tree species is crucial for forestry management, monitoring of forest pests and diseases and fires, inversion of forest physiological parameters, and optimization of vegetation ecological models. Furthermore, a comprehensive understanding of existing forest vegetation types and their distribution is fundamental for achieving global carbon monitoring, forest ecosystem service assessment, and biodiversity conservation. For ecologists and policymakers, obtaining accurate and up-to-date tree species distribution maps is a key prerequisite for the scientific development of relevant strategies and management measures.

[0003] Compared to time-consuming and expensive field sampling, remote sensing can efficiently and economically acquire large-area data with high temporal resolution, and has therefore been widely used in tree species classification. Currently, remote sensing-based tree species classification mainly relies on two categories of methods: Machine Learning and Deep Learning. Traditional machine learning techniques excel at modeling nonlinear relationships in high-dimensional remote sensing data. However, these methods typically rely on manually constructed feature sets and lack the ability to automatically extract hierarchical spatial and spectral representations. Furthermore, their versatility and transferability are limited when facing different data distributions and task settings. To overcome these shortcomings, deep learning methods have been widely applied in tree species identification in recent years. Compared to traditional methods, deep learning has stronger capabilities in spatial structure modeling and automatic extraction of multi-level features. Among them, CNN (Convolutional Neural Network) is one of the most commonly used deep learning models. Its classic architectures, such as VGG (Visual Geometry Group), ResNet (Residual Network), DenseNet (Densely Connected Convolutional Neural Network), and EfficientNet (Efficient Neural Network), possess more complex deep nonlinear structures, achieving higher classification accuracy and stronger robustness in diverse scenarios. However, these classic architectures still face challenges in accurately capturing fine-grained spatial features, such as canopy boundaries and morphological details.

[0004] In remote sensing tree species classification tasks, temporal features (such as multi-temporal spectral reflectance and vegetation indices) can effectively reflect the changes in spectral response of vegetation at different growth stages, serving as an important source of information for improving classification accuracy. However, directly extracting a large number of temporal features often presents two challenges: first, the sheer number and redundancy of features significantly increase the computational cost and complexity of the model; second, most existing feature selection methods evaluate individual variables independently, neglecting the synergistic effects of the same type of feature at different time points, making it difficult to reflect the overall contribution of the entire set of temporal features to classification performance. This may not only obscure dynamic change patterns crucial to the discrimination results but also reduce the model's stability and generalization ability when dealing with seasonal changes and phenological differences. Therefore, there is an urgent need for a method that can evaluate the contribution of a feature group as a whole and effectively select high-value temporal features, balancing model accuracy and computational efficiency.

[0005] In the prior art, Chinese patent document CN118228005A discloses a "method and device for monitoring forest tree species diversity," which involves collecting forest remote sensing images; obtaining spatiotemporal heterogeneity indices of the forest based on the forest remote sensing images during the plant growth season, using these indices as an image feature set; collecting measured tree species diversity data; obtaining a spatial matching relationship between remote sensing image features and diversity indices based on the image feature set and tree species diversity data; obtaining input variables based on the spatial matching relationship between the remote sensing image features and diversity indices; and training a preset deep learning model based on the input variables. The constructed model is then used to process forest images of the target area to obtain diversity results. However, this technical solution fails to comprehensively examine the overall contribution of the entire set of temporal features to the model performance, easily overlooking the synergistic effect of features at different time points, thus causing information loss during feature selection. Furthermore, due to the limited spatial resolution of Sentinel-2 imagery, fine structures in forest stands (such as crown morphology and leaf texture) are difficult to distinguish accurately. In addition, the high heterogeneity of forest tree species in spatial distribution and the random fluctuation in their numbers make it difficult for deep learning models to effectively capture high-precision detailed features. This leads to a decrease in the model's ability to distinguish complex forest stands and significantly restricts its overall classification performance.

[0006] In summary, existing technologies have several limitations. They fail to comprehensively examine the overall contribution of the entire set of temporal features to the model's performance, resulting in information loss. Furthermore, due to the spatial resolution of Sentinel-2 imagery and the randomness of the spatial and quantitative distribution of forest tree species, deep learning models struggle to capture high-precision detailed features, leading to a significant decline in the performance of remote sensing tree species classification models. Summary of the Invention

[0007] This invention solves the technical problem that existing technologies fail to fully examine the overall contribution of the entire set of temporal features to the model performance, and that the performance of remote sensing tree species classification models is significantly reduced because deep learning models have difficulty capturing detailed features in Sentinel-2 images.

[0008] The method for constructing a refined tree species identification model according to the present invention includes the following steps:

[0009] Step 1: Construct a high-dimensional multi-source feature dataset;

[0010] Step 2: Perform feature optimization based on the high-dimensional multi-source feature dataset to obtain the optimal classification feature set;

[0011] Step 3: Collect actual sample data and combine it with the optimal classification feature set to construct a deep learning sample library;

[0012] Step 4: Pre-set a fine tree species identification model, input the deep learning sample library to train the pre-set fine tree species identification model, and obtain the fine tree species identification model.

[0013] Furthermore, in this embodiment of the invention, the high-dimensional multi-source feature dataset in step 1 includes principal component features, texture features, terrain features, spectral features, and vegetation indices.

[0014] Furthermore, in this embodiment of the invention, step 2 involves feature optimization based on a high-dimensional multi-source feature dataset to obtain the optimal classification feature set, specifically as follows:

[0015] The observation values ​​of each type of feature in the high-dimensional multi-source feature dataset at different time phases are divided into multiple feature groups. The classification model is then used to train the multiple feature groups to obtain the optimal classification feature set.

[0016] Furthermore, in this embodiment of the invention, the step of training multiple feature sets using a classification model to obtain the optimal classification feature set specifically involves:

[0017] The classification model is trained on multiple feature groups to obtain multiple single-channel feature importance indices. All single-channel feature importance indices within each feature group are aggregated to obtain the overall contribution of each feature group. The feature group with the largest overall contribution is taken as the optimal classification feature set.

[0018] Furthermore, in this embodiment of the invention, the deep learning sample library in step 3 is divided into a training set and a validation set in an 8:2 ratio.

[0019] Furthermore, in this embodiment of the invention, the preset refined tree species identification model in step 4 specifically refers to:

[0020] The input data is extracted using a data extraction module, and the multi-scale features are aligned using a cross-scale deformable attention fusion module. The aligned multi-scale features are then fused to output the recognition result.

[0021] Furthermore, in this embodiment of the invention, the cross-scale deformable attention fusion module specifically comprises:

[0022] The multi-scale feature map is spatially offset, and deformable convolution is used to align the features of the spatially offset multi-scale feature map. The features of the feature-aligned multi-scale feature map are then concatenated. Lightweight self-attention is used to obtain the self-attention result. Dynamic channel gating is used to fuse the self-attention result with the concatenated multi-scale feature map to obtain the feature-fused multi-scale feature map.

[0023] The refined tree species identification method of this invention is implemented based on any of the refined tree species identification model construction methods described above, specifically as follows:

[0024] Forest images of the area to be tested are used to identify detailed tree species.

[0025] The refined tree species identification model construction system of the present invention is based on the above-mentioned refined tree species identification model construction method and includes the following modules:

[0026] The data acquisition module constructs a high-dimensional, multi-source feature dataset.

[0027] The feature optimization module performs feature optimization based on a high-dimensional multi-source feature dataset to obtain the optimal classification feature set.

[0028] The sample library construction module collects actual test sample data and combines it with the optimal classification feature set to construct a deep learning sample library;

[0029] The training module uses a pre-set refined tree species recognition model. The model is trained by inputting a deep learning sample library to obtain the refined tree species recognition model.

[0030] The refined tree species identification system of the present invention is implemented based on the above-mentioned refined tree species identification model construction system, and the identification system includes:

[0031] The identification module identifies forest images of the area under test and obtains detailed tree species identification results.

[0032] This invention addresses the technical problems of existing technologies, such as failing to comprehensively examine the overall contribution of the entire set of temporal features to model performance, and the significant performance degradation of remote sensing tree species classification models due to the difficulty of deep learning models in capturing detailed features in Sentinel-2 imagery. Specific beneficial effects include:

[0033] 1. The present invention proposes a refined tree species identification method, such as... Figure 1 As shown, in the task of tree species identification in medium-resolution remote sensing images, a fusion-optimized ConvNeXt architecture was constructed to address core issues such as improving the ability to extract spatial structural details and enhancing the ability to express spatiotemporal features. By fusing multi-scale feature maps, the classification accuracy of ConvNeXt is improved, and the cross-scale high-resolution classification capability is enhanced while maintaining low computational overhead. In addition, by combining multi-source feature fusion and feature optimization mechanisms and introducing temporal phenological information, the system optimizes the classification performance of remote sensing images under complex forest stand conditions. This provides a lightweight, high-performance, and highly generalizable fine tree species identification method for remote sensing tree species classification tasks, which is particularly suitable for fine-grained tree species identification and image segmentation tasks under complex forest stand structures.

[0034] 2. The present invention proposes a method for constructing a fine tree species identification model, which constructs an end-to-end framework for the fusion of spatial-spectral-temporal multidimensional features for tree species classification. By introducing joint training of temporal-spatial-spectral dimensional features, a fine tree species identification model with broad adaptability is constructed, thereby achieving a balance between classification accuracy and generalization ability.

[0035] 3. The method for constructing a refined tree species identification model proposed in this invention can overcome the key technical bottlenecks of traditional convolutional neural networks, such as insufficient spatial detail representation ability in forest type classification. It can also significantly improve the classification accuracy and inference efficiency of forest tree species identification in large-scale areas, enhance the deep perception ability of forest stand spatial structure and pattern, and provide high-value parameter support for forest ecosystem monitoring and modeling. This solves the problem of insufficient spatial detail representation of deep learning models in medium-resolution remote sensing images. Attached Figure Description

[0036] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:

[0037] Figure 1 It is the tree species identification and mapping process based on the ConvNeXt architecture described in the invention.

[0038] Figure 2 It is the original framework of ConvNeXt described in Implementation Method 5;

[0039] Figure 3This is a schematic diagram of the improved ConvNeXt described in Implementation Method 5;

[0040] Figure 4 This is an example result of forest tree species classification as described in Implementation Method Seven. Detailed Implementation

[0041] Various embodiments of the present invention will now be clearly and completely described with reference to the accompanying drawings. The embodiments described with reference to the drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.

[0042] Implementation Method 1. The method for constructing a refined tree species identification model as described in this implementation method includes the following steps:

[0043] Step 1: Construct a high-dimensional multi-source feature dataset;

[0044] Step 2: Perform feature optimization based on the high-dimensional multi-source feature dataset to obtain the optimal classification feature set;

[0045] Step 3: Collect actual sample data and combine it with the optimal classification feature set to construct a deep learning sample library;

[0046] Step 4: Pre-set a fine tree species identification model, input the deep learning sample library to train the pre-set fine tree species identification model, and obtain the fine tree species identification model.

[0047] Existing technologies have technical problems such as failing to fully examine the overall contribution of the entire set of temporal features to the model performance, and the difficulty of deep learning models in capturing detailed features in Sentinel-2 images, leading to a significant decrease in the performance of remote sensing tree species classification models.

[0048] To address the aforementioned technical problems, this embodiment provides a method for constructing a refined tree species identification model, specifically including the following steps:

[0049] Step 1: Construct a high-dimensional multi-source feature dataset;

[0050] This implementation method constructs a high-dimensional multi-source feature dataset based on Sentinel-2 multi-temporal images to improve the accuracy and stability of forest tree species classification.

[0051] Step 2: Perform feature optimization based on the high-dimensional multi-source feature dataset to obtain the optimal classification feature set;

[0052] Step 3: Collect actual sample data and combine it with the optimal classification feature set to construct a deep learning sample library;

[0053] To construct high-quality, representative training samples of pure forest land, this implementation method integrates vector data of forest sub-plots from the 2018 Class II forest survey provided by the Tarim Forestry Bureau with high-resolution remote sensing imagery, selecting six major dominant tree species as classification targets. Given that the survey data does not record the spatial distribution of different tree species in mixed forests, making pixel-level precise labeling difficult, only pure forest land was selected as the source of training and validation samples. Within each class of pure forest land, a random sampling strategy was employed, extracting samples from a 15-meter buffer zone within the polygon to avoid noise introduced by spatial positioning errors and inaccurate boundaries. The collected measured sample data, combined with the optimal feature set obtained in step 2, constructed a deep learning sample library for model training. Considering the limited size of the deep learning sample library and its potential for overfitting, the samples in the deep learning sample library were further cropped into 256×256 pixel image blocks, and data augmentation methods such as vertical flipping, horizontal flipping, random rotation, and channel permutation were used to expand the deep learning sample library to 100,000 images.

[0054] Step 4: Pre-set a fine tree species identification model, input the deep learning sample library to train the pre-set fine tree species identification model, and obtain the fine tree species identification model.

[0055] Therefore, this implementation method constructs a high-dimensional multi-source feature dataset, optimizes its features, combines it with measured sample data to form deep learning samples, and uses the deep learning samples to train a preset fine tree species identification model to obtain a fine tree species identification model. This solves the technical problems of existing technologies that fail to comprehensively examine the overall contribution of the entire set of time-series features to the model performance, and that the performance of remote sensing tree species classification models is significantly reduced due to the randomness of the spatial and quantitative distribution of forest tree species.

[0056] Implementation Method 2. This implementation method further defines the fine tree species identification model construction method described in Implementation Method 1. In step 1, the high-dimensional multi-source feature dataset includes principal component features, texture features, terrain features, spectral features, and vegetation indices.

[0057] Because tree species identification is highly sensitive to seasonal spectral differences, incorporating phenological information helps improve the performance of multispectral remote sensing data in tree species classification. Sentinel-2, in particular, possesses high temporal resolution and multi-band imaging capabilities, effectively capturing the dynamic changes in vegetation reflectance characteristics at different phenological stages. To maximize the application potential of Sentinel-2 time series data in tree species classification, this implementation selects image data covering seven key time points throughout the growing season (April to October), and incorporates observation images from the same period two years prior and subsequent years to construct a time series dataset with low cloud and fog interference and high data integrity. To systematically evaluate the contribution of different time series data to classification performance, this implementation proposes an importance aggregation method based on a gradient boosting tree model for the importance assessment and selection of remote sensing multi-temporal feature groups, ultimately obtaining the most discriminative time series features. By fusing comprehensive features from temporal, spatial, and spectral dimensions, an optimal classification feature set is constructed. This not only enhances the model's ability to perceive phenological changes but also improves the accuracy of characterizing fine-grained forest structural information. Ultimately, the tree species identification framework based on the spatial continuity, temporal density, and feature fusion of Sentinel-2 time series data demonstrated good adaptability and application potential in addressing challenges such as complex species mixing and strong heterogeneity in temperate mixed forests.

[0058] Furthermore, to fully utilize temporal information, the Sentinel-2 vegetation index time series features from April 1st to October 31st were introduced to enhance the ability to identify phenological changes. To improve the classification accuracy of the refined tree species identification system in evaluating tree species with different feature types, based on the feature selection of the RE-RFE (Recursive Feature Elimination in Random Forest) algorithm, the most discriminative feature combinations were obtained. The four feature combination schemes are as follows:

[0059] (1) Basic scheme: integrating original spectra, independent principal components and vegetation indices;

[0060] (2) Texture enhancement scheme: Introduce texture features into the basic scheme;

[0061] (3) Terrain expansion scheme: Based on the above, add terrain variables;

[0062] (4) Time series fusion scheme: further integrate time series features.

[0063] The above four feature combination schemes help to systematically analyze the differences and complementarities in the performance of different feature dimensions in tree species identification, and provide a theoretical basis for constructing the optimal classification feature set.

[0064] First, the imagery at seven key time points was dimensionality-reduced using ICA (Independent Principal Component Analysis) to obtain three independent non-Gaussian components. Forty-eight texture features were extracted from the ICA components, covering eight categories: mean, variance, contrast, entropy, correlation, homogeneity, second moment of angle, and difference. Window sizes were set to 3×3 and 5×5, and the gray quantization level was 64. To enhance the discriminative power of spectral differences and reduce environmental variation interference, 41 vegetation indices were calculated, covering key bands such as red-edge, near-infrared, and shortwave infrared. Simultaneously, three topographic factors—elevation, slope, and aspect—were calculated based on SRTM 30m DEM (Digital Elevation Model) data to characterize the ecological gradient effect of forest spatial distribution. Finally, the principal component features, texture features, vegetation indices, spectral features, and topographic features were integrated to construct a unified multidimensional feature space, resulting in a high-dimensional multi-source feature dataset. Topographic features include elevation, slope, and aspect. The table below shows the vegetation indices calculated based on Sentinel-2 imagery.

[0065] Table 1

[0066]

[0067]

[0068] in, , , , , , These are the blue, green, red, red edge 1, red edge 2, and near-infrared bands in the Sentinel-2 image, respectively.

[0069] Implementation Method 3. This implementation method further defines the refined tree species identification model construction method described in Implementation Method 1. In step 2, feature optimization is performed based on a high-dimensional multi-source feature dataset to obtain the optimal classification feature set, specifically as follows:

[0070] The observations of each feature class in the high-dimensional multi-source feature dataset at different time phases are divided into groups to obtain multiple feature groups. The classification model is used to train multiple feature groups to obtain multiple single-channel feature importance indices. All single-channel feature importance indices in each feature group are aggregated to obtain the overall contribution of each feature group. The feature group with the largest overall contribution is taken as the optimal classification feature set.

[0071] Implementation method two extracts multiple sets of temporal features to enhance the model's expressive power. However, large feature sets often suffer from severe information redundancy, increasing model complexity and potentially weakening classification performance. Traditional feature selection methods typically construct a temporal feature set first, then select features from all temporal features (each feature being an independent variable). This approach fails to comprehensively consider the overall contribution of the entire set of temporal features to model performance, easily leading to the neglect of synergistic effects among temporal features and the inability to effectively capture key dynamic temporal patterns. Consequently, it affects the model's stable recognition ability for different growth stages and seasonal changes, resulting in decreased generalization ability, blurred inter-class distinctions, and a tendency to misclassify and predict unstable results in classification tasks involving complex ecological environments and highly similar tree species. To address these issues, this implementation method proposes an importance aggregation method based on a gradient boosting tree model for the importance assessment and selection of multi-temporal feature sets from remote sensing.

[0072] This implementation method uses physically meaningful feature categories of vegetation indices as units, dividing the observations of each feature across multiple time phases into a group. All groups are then fed into XGBoost (Gradient Boosting Classification Model) for training. After training, the single-channel feature importance index is used to aggregate the importance indices of all channels within each feature group, thus obtaining the overall contribution of each feature group. Subsequently, all feature groups are sorted, identifying the time-series feature groups with the greatest impact on classification performance. Based on these, an optimal classification feature set is constructed to improve model performance and simplify feature dimensions. Table 2 shows the 16 most discriminative features ultimately selected for subsequent modeling.

[0073] Table 2

[0074]

[0075] in, This refers to the Sentinel-2 red-edge position index. For excessive greenness index, To convert the chlorophyll absorption ratio index, The terrestrial chlorophyll index, and This is a gray-level index, representing the row and column in the gray-level co-occurrence matrix, corresponding to the two gray-level values ​​of a pixel in the image. The expected value or Mean is the weighted average gray value of the gray-level co-occurrence matrix.

[0076] Implementation Method 4. This implementation method further defines the fine tree species identification model construction method described in Implementation Method 1. In step 3, the deep learning sample library is divided into a training set and a validation set in an 8:2 ratio.

[0077] Implementation Method 5. This implementation method further defines the refined tree species identification model construction method described in Implementation Method 1. The preset refined tree species identification model in step 4 is specifically as follows:

[0078] The input data is extracted using a data extraction module, and the multi-scale features are aligned using a cross-scale deformable attention fusion module. The aligned multi-scale features are then fused to output the recognition result.

[0079] Due to limitations in the ability of medium spatial resolution remote sensing imagery to represent fine-grained information such as tree canopy outlines, boundaries, and textures, CNNs face significant challenges in classification accuracy and class separability when dealing with temperate forest regions characterized by highly mixed species and complex species distribution. To address these issues, this implementation introduces a structurally optimized ConvNeXt as the model backbone. The optimized ConvNeXt represents a systematic reconstruction based on traditional CNNs.

[0080] Inspired by the Swing Transformer (a hierarchical vision system using shift windows), ConvNeXt systematically optimizes the ResNet-50 (Residual Network-50) architecture, combining powerful local detail capture and global context modeling capabilities, significantly enhancing the model's ability to capture both local details and global contextual features. Figure 2As shown, multi-scale feature extraction of input data is performed based on ConvNeXt (convolutional neural network). First, initial image encoding is performed using 7×7 convolutional kernels to effectively expand the receptive field and enhance the perception of global context. The parameter scale is controlled by compressing the number of channels in the underlying convolutions. In the feature extraction stage, depthwise separable convolution and pointwise convolution are introduced to improve computational efficiency while maintaining high expressive power. Subsequently, a multi-scale fusion module, borrowing the multi-resolution parallel structure and multi-scale fusion idea from HRNet (high-resolution network), is introduced to form a fused feature tensor. This fused feature tensor has multi-scale and multi-level contextual feature expression capabilities. However, the fusion process cannot automatically align semantics and details between different scales and lacks a globally adaptive channel weight allocation mechanism. To address these technical issues, this implementation proposes a CSDAF module (cross-scale deformable attention fusion module) based on deformable convolution and lightweight multi-head self-attention. This module aligns multi-scale features, fuses the aligned multi-scale features, and outputs the recognition result. This implementation borrows the multi-scale fusion idea of ​​the multi-resolution parallel structure in HRNet, but no longer adopts the multi-branch stacking method. Instead, it is based on the existing multi-stage features of ConvNeXt, and uses deformable convolution to align spatial positions, lightweight multi-head attention to achieve cross-scale interaction, and adaptive channel gating for dynamic fusion. Thus, without introducing additional branches, it achieves a feature enhancement effect similar to HRNet, while significantly reducing model parameters and computational costs.

[0081] In summary, this implementation uses the deep learning sample library constructed in step 3 as input and, based on the PyTorch (deep learning framework) architecture within the Python programming language, proposes an improved model that integrates the ConvNeXt backbone network and the CSDAF module to enhance classification accuracy and spatial localization capabilities. Figure 3As shown, the CSDAF module is integrated to achieve efficient interaction and dynamic alignment of multi-resolution features, adaptively capturing key information flows across different scales and enhancing the collaborative expression of spatial details and semantic information. Simultaneously, an SE channel attention mechanism is introduced in the key feature extraction stage to adaptively allocate feature weights for each channel, enhancing discriminative feature expression and suppressing redundant information. This structure achieves accurate fusion of cross-scale information and fine-tuning of channel features while maintaining low computational overhead, effectively improving the accuracy of pixel-level classification and the model's robustness in complex scenarios. Applying the improved ConvNeXt to Sentinel-2 data fully leverages its unique remote sensing capabilities, including rich spectral bands, high spatial resolution, and short access cycles. The improved ConvNeXt enables high-precision forest species mapping and dynamic monitoring over long-term sequences and large scales. It more effectively combines Sentinel-2 imagery with a deep learning architecture to achieve long-term, large-scale forest species mapping and monitoring. This improved ConvNeXt significantly enhances the spatial alignment and semantic interaction capabilities of multi-scale features without significantly increasing the model size. S1, S2, S3 and S4 are the feature maps extracted by the ConvNeXt model from the first to the fourth stage, respectively, and X1, X2, X3 and X4 are the feature maps after upsampling of the feature maps in each stage.

[0082] Implementation Method Six. This implementation method further defines the refined tree species identification model construction method described in Implementation Method Five. The cross-scale deformable attention fusion module is specifically as follows:

[0083] The multi-scale feature map is spatially offset, and deformable convolution is used to align the features of the spatially offset multi-scale feature map. The features of the feature-aligned multi-scale feature map are then concatenated. Lightweight self-attention is used to obtain the self-attention result. Dynamic channel gating is used to fuse the self-attention result with the concatenated multi-scale feature map to obtain the feature-fused multi-scale feature map.

[0084] First, spatial offset is performed on the multi-scale features output by each stage of ConvNeXt through a small convolutional subnet. Deformable convolution is then used to align local details at different resolutions. Subsequently, the aligned features are stitched together, and lightweight self-attention with channel compression is used to capture long-range dependencies across scales. Finally, the self-attention results are fused with the original stitched features through dynamic channel gating.

[0085] Implementation Method Seven. The refined tree species identification method described in this implementation method is implemented based on the refined tree species identification model construction method described in any one of Implementation Methods One to Six, specifically as follows:

[0086] Forest images of the area to be tested are used to identify detailed tree species.

[0087] like Figure 4 As shown, the refined tree species identification results of the output area are identified by the refined tree species identification model, and then further transferred to the tree species identification of different areas.

[0088] To evaluate the effectiveness of the proposed refined tree species identification model in temperate forest tree species classification, four metrics were used to measure the classification performance: F1 score, OA (overall accuracy), PA (producer accuracy), and UA (user accuracy). These four evaluation metrics together provide a systematic quantitative basis for evaluating the model's predictive performance and accuracy, ultimately leading to a more robust method for identifying forest tree species.

[0089] In summary, this invention not only has significant economic value in optimizing forest restoration projects and estimating ecological carbon sinks, but also has important application prospects in the fields of forest management and ecological protection.

[0090] Implementation Method 8. The refined tree species identification model construction system described in this implementation method is based on the refined tree species identification model construction method described in Implementation Method 1, and includes the following modules:

[0091] The data acquisition module constructs a high-dimensional, multi-source feature dataset.

[0092] The feature optimization module performs feature optimization based on a high-dimensional multi-source feature dataset to obtain the optimal classification feature set.

[0093] The sample library construction module collects actual test sample data and combines it with the optimal classification feature set to construct a deep learning sample library;

[0094] The training module uses a pre-set refined tree species recognition model. The model is trained by inputting a deep learning sample library to obtain the refined tree species recognition model.

[0095] Implementation Method Nine. The refined tree species identification system described in this implementation method is based on the refined tree species identification model construction system described in Implementation Method Eight. The identification system includes:

[0096] The identification module identifies forest images of the area under test and obtains detailed tree species identification results.

[0097] The above provides a detailed description of the refined tree species identification model construction method and system, as well as the identification method and system proposed in this invention. Specific examples have been used to illustrate the principles and implementation methods of this invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this invention. Therefore, the content of this specification should not be construed as a limitation of this invention.

Claims

1. A method for constructing a refined tree species identification model, characterized in that, Includes the following steps: Step 1: Construct a high-dimensional multi-source feature dataset; Step 2: Perform feature optimization based on the high-dimensional multi-source feature dataset to obtain the optimal classification feature set; Step 3: Collect actual sample data and combine it with the optimal classification feature set to construct a deep learning sample library; Step 4: Pre-set a fine tree species identification model, input the deep learning sample library to train the pre-set fine tree species identification model, and obtain the fine tree species identification model; The preset fine tree species identification model in step 4 is specifically as follows: The input data is extracted using a data extraction module, and the multi-scale features are aligned using a cross-scale deformable attention fusion module. The aligned multi-scale features are then fused to output the recognition result. The aforementioned cross-scale deformable attention fusion module specifically includes: The multi-scale feature map is spatially offset, and deformable convolution is used to align the features of the spatially offset multi-scale feature map. The features of the feature-aligned multi-scale feature map are then concatenated. Lightweight self-attention is used to obtain the self-attention result. Dynamic channel gating is used to fuse the self-attention result with the concatenated multi-scale feature map to obtain the feature-fused multi-scale feature map.

2. The method for constructing a refined tree species identification model according to claim 1, characterized in that, The high-dimensional multi-source feature dataset in step 1 includes principal component features, texture features, terrain features, spectral features, and vegetation indices.

3. The method for constructing a refined tree species identification model according to claim 1, characterized in that, In step 2, feature optimization is performed based on a high-dimensional multi-source feature dataset to obtain the optimal classification feature set, specifically as follows: The observation values ​​of each type of feature in the high-dimensional multi-source feature dataset at different time phases are divided into multiple feature groups. The classification model is then used to train the multiple feature groups to obtain the optimal classification feature set.

4. The method for constructing a refined tree species identification model according to claim 3, characterized in that, The method of training multiple feature sets using a classification model to obtain the optimal classification feature set is as follows: The classification model is trained on multiple feature groups to obtain multiple single-channel feature importance indices. All single-channel feature importance indices within each feature group are aggregated to obtain the overall contribution of each feature group. The feature group with the largest overall contribution is taken as the optimal classification feature set.

5. The method for constructing a refined tree species identification model according to claim 1, characterized in that, In step 3, the deep learning sample library is divided into a training set and a validation set in an 8:2 ratio.

6. A system for constructing a refined tree species identification model, wherein the system is implemented based on the method for constructing a refined tree species identification model as described in claim 1, characterized in that, Includes the following modules: The data acquisition module constructs a high-dimensional, multi-source feature dataset. The feature optimization module performs feature optimization based on a high-dimensional multi-source feature dataset to obtain the optimal classification feature set. The sample library construction module collects actual test sample data and combines it with the optimal classification feature set to construct a deep learning sample library; The training module has a preset high-precision tree species recognition model. The high-precision tree species recognition model is trained by inputting a deep learning sample library to obtain the high-precision tree species recognition model. The preset fine-grained tree species recognition model in the training module is specifically as follows: The input data is extracted using a data extraction module, and the multi-scale features are aligned using a cross-scale deformable attention fusion module. The aligned multi-scale features are then fused to output the recognition result. The aforementioned cross-scale deformable attention fusion module specifically includes: The multi-scale feature map is spatially offset, and deformable convolution is used to align the features of the spatially offset multi-scale feature map. The features of the feature-aligned multi-scale feature map are then concatenated. Lightweight self-attention is used to obtain the self-attention result. Dynamic channel gating is used to fuse the self-attention result with the concatenated multi-scale feature map to obtain the feature-fused multi-scale feature map.

Citation Information

Patent Citations

  • Forest tree species diversity monitoring method and device

    CN118228005A

  • Tree species refined classification method based on deep learning algorithm and time sequence sentinel image

    CN113869370A