Method for identifying dominant tree species based on spatio-temporal-spectral multi-domain perception and guidance
Through the dual-stream network architecture and feature fusion module based on time-space-spectral multi-domain perception and guidance, the problem that traditional models are difficult to explore seasonal changes in tree species recognition in large-scale regional tree species recognition is solved, and more efficient space-time and spatial spectrum feature fusion and forest tree species recognition accuracy are achieved.
Patent Information
- Application Number
- CN202510398729.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-04-01
AI Technical Summary
Traditional multispectral tree species classification model is difficult to fully explore seasonal changes characteristics and construct tree species space-time spectrum constraint relationships, resulting in poor results in large-scale forest tree species identification.
Using a method based on time-space-spectral multi-domain perception and guidance, parallel feature extraction through dual-stream network architecture, combined with a hierarchical time-sequence spectrum integration mechanism and dynamic spatiotemporal excitation optimization mechanism, multi-time phase dynamic features are fused across domains to generate spatiotemporal spectral features.
It effectively captures the seasonal changes in the spectral characteristics of forest tree species, enhances the model's adaptability in complex space-time scenarios, and significantly improves the accuracy and robustness of large-scale forest tree species recognition.
Smart Images

Figure CN119919817B_ABST
Abstract
Description
Technical Field
[0001] The technical field to which the present invention belongs is multi - spectral image processing and analysis in remote sensing technology, specifically for forest tree species identification at the regional scale using the space - time - spectrum multi - domain perception and guidance method of space - borne remote sensing. Background Art
[0002] Currently, multi - spectral tree species classification models mostly combine the spatial resolution advantages of airborne multi - spectral data, emphasizing the fine expression of local spectral and spatial features. However, due to the differences in observation scale and sensor characteristics, there are significant differences in data characteristics and modeling requirements between space - borne and airborne multi - spectral data, and these methods are difficult to be directly applied to the large - scale space - borne multi - spectral tree species classification task. Specifically, the coverage range of space - borne multi - spectral images is relatively wide, and there are large differences in environmental conditions, tree species composition, and growth status in different regions, resulting in strong spatial heterogeneity in the spectral characteristics of tree species.
[0003] In addition, due to the relatively large map size and relatively low spatial resolution of space - borne multi - spectral data, it is more vulnerable to background interference and has a higher spectral mixing effect. This mixing effect weakens the spectral feature independence of a single tree species, which poses a challenge to the current multi - spectral tree species classification models that lack dynamic learning ability. In addition, the current multi - spectral tree species classification research focuses on single - temporal data, and improves the classification performance through the coupling of spatial and spectral features. However, there are significant seasonal changes in the spectral characteristics of tree species in the forest. Multi - temporal multi - spectral data can capture these dynamic features and provide more comprehensive information. However, existing multi - temporal multi - spectral models mostly simply stack or average the data at each time step, fail to fully mine the seasonal change features, and are also difficult to effectively construct the spatio - temporal - spectral constraint relationship of tree species, which limits their application effect in large - scale regional forest tree species identification. Summary of the Invention
[0004] Aiming at the deficiencies of the existing technology, the present invention provides a dominant tree species identification method based on space - time - spectrum multi - domain perception and guidance, which solves the technical problems that traditional tree species classification models cannot fully mine the seasonal change features and effectively construct the spatio - temporal - spectral constraint relationship of tree species.
[0005] To solve the above - mentioned technical problems, the present invention provides the following technical solution: A dominant tree species identification method based on space - time - spectrum multi - domain perception and guidance, the method comprising the following steps:
[0006] S1. Obtain multi - spectral image data of space - borne multi - temporal covering the target area;
[0007] S2. Perform parallel feature extraction on the multi - spectral image data through a two - stream network architecture to obtain two - stream features. The two - stream network architecture includes a first branch and a second branch, wherein:
[0008] The first branch is a lightweight residual Transformer branch for capturing long-range spectral dependencies across bands during spectral information modeling;
[0009] The second branch is a hybrid-scale spectral-spatial fusion branch for dynamically fusing local and global spectral and spatial features for joint representation;
[0010] S3. Input the extracted two-stream features into the spatio-temporal-spectral feature fusion module, and through the hierarchical temporal-spectral integration mechanism and the dynamic spatio-temporal excitation optimization mechanism, perform cross-domain fusion on multi-temporal dynamic features to obtain spatio-temporal-spectral features;
[0011] S4. Based on the fused spatio-temporal-spectral features, output the recognition results of the dominant tree species in the forest through a classifier.
[0012] Furthermore, the lightweight residual Transformer branch includes a multi-head self-attention mechanism that captures the non-linear dependence between spectral bands by parallel computing multiple groups of query, key, and value matrices, and an asymmetric convolution block that enhances the spectral feature expression ability through horizontal 1×k convolution and vertical k×1 convolution operations along the spectral dimension. The asymmetric convolution block The expression is:
[0013] ;
[0014] In the formula, and are the horizontal and vertical convolution kernels respectively; and are both bias terms; is the ReLU activation function; is the input semantic information;
[0015] The lightweight residual Transformer branch also includes a residual connection module that dynamically adjusts the residual weights based on the hierarchical scaling residual connection mechanism through a scaling factor ;
[0016] Furthermore, the scaling factor is generated by global average pooling and a fully connected network.
[0017] Furthermore, the specific process of dynamically adjusting the residual weights based on the hierarchical scaling residual connection mechanism through a scaling factor includes:
[0018] Perform global average pooling on the current layer feature to extract global context information , and the expression is:
[0019] ;
[0020] In the formula, are the number of channels, height, and width of the feature map, respectively; is the feature vector at the position; represents a feature tensor with the shape of ;
[0021] Generate a scaling factor through a two-layer fully connected network, and the expression is:
[0022] ;
[0023] In the formula, is the ReLU activation function, which is used to increase the non-linear expression ability; is the Sigmoid loss function, which restricts the scaling factor within the range of [0, 1]; and are both learnable weight matrices, and are real number domain matrices with shapes and respectively, represents the number of feature channels, represents the ratio of channel compression and expansion; is the bias term;
[0024] Use the generated scaling factor to adjust the weight of the contribution of the residual connection, and the expression is:
[0025] ;
[0026] In the formula, is the input feature of the current layer; is the feature representation after transformation of the layer; ⊙ represents the element-wise multiplication operation.
[0027] Furthermore, the hybrid-scale spectral-spatial fusion branch includes a multi-attention mechanism for dynamically enhancing key features and suppressing redundant information in multi-spectral image data, and a hybrid-scale convolutional block that uses multi-level cascaded convolutional kernels to extract spectral features and spatial features from macro to micro layer by layer;
[0028] The multi-attention mechanism includes a spectral attention module that weights spectral channel features through 1×1 convolution and multi-head self-attention, and a spatial attention module that dynamically generates spatial weights through depth convolution and pointwise convolution;
[0029] The expression of the mixed-scale convolution block is as follows:
[0030] ;
[0031] In the formula, is the output feature map of the current branch; is the feature map obtained after being transformed by the activation function.
[0032] Furthermore, the spectral attention module first performs inter-channel information fusion using a 1×1 convolution kernel, then introduces non-linear features through batch normalization and the ReLU activation function, and then calculates the correlation of spectral channels through the multi-head self-attention mechanism to generate spectral attention weights, thereby enhancing the important spectral information in the feature map. The expression is as follows:
[0033] ;
[0034] In the formula, Q, K, and V are the query, key, and value matrices respectively; is the regularization term to prevent overfitting; is the dimension of the key vector; is the Softmax normalization function; is the spectral semantic information enhanced by spectral attention;
[0035] Finally, the enhanced spectral features are refined into detailed representations with a higher semantic level through attention weighting and linear transformation. The expression is as follows:
[0036] ;
[0037] In the formula, is the final output of the spectral attention feature; is the linear transformation matrix represents the index set of spectral channels;
[0038] The spatial attention module uses depth convolution for spatial feature extraction to maintain channel independence; then it fuses channel information through pointwise convolution and generates spatial attention weights using the Sigmoid activation function. The generated spatial attention weights are multiplied by the original feature map to achieve adaptive enhancement and suppression of spatial features. The expression is as follows:
[0039] ;
[0040] In the formula, is the adaptive learning parameter for dynamically weighting spectral and spatial information; W is the suppression weight obtained through adaptive learning; is the Softmax normalization function; is the Sigmoid activation function; ⊙ represents the per-pixel multiplication operation.
[0041] Furthermore, the hierarchical temporal spectral integration mechanism uses a multi-level GRU network to recursively model the dynamic dependencies of time steps, and uses a dynamically denoised optimized temporal attention mechanism to perform weighted fusion on multi-temporal spectral features;
[0042] The dynamic spatio-temporal excitation optimization mechanism optimizes the spatio-temporal spectral feature expression by adjusting the channel weights of temporal perception and combining the weight information of the previous time step.
[0043] Furthermore, the use of a multi-level GRU network to recursively model the dynamic dependencies of time steps is specifically as follows:
[0044] A multi-level GRU network is used to capture the dynamic dependencies between time steps, and the hidden state features of each time step are extracted through recursive calculation , and the expression of the GRU network recursive calculation process is:
[0045] ;
[0046] In the formula, is the update gate; is the reset gate, which controls the fusion and reset of the previous time step state and the current input; is the time step 's final hidden state feature; ⊙ represents element-wise multiplication; is the Sigmoid function; is the hyperbolic tangent function; represents the candidate hidden state; is the temporal spectral information of the current time step ; respectively represent the weight calculation matrices of the update gate, reset gate, and candidate hidden state; respectively represent to the weight calculation matrices of the update gate, reset gate, and candidate hidden state; respectively represent the biases of the update gate, reset gate, and candidate hidden state;
[0047] The calculation process of performing weighted fusion on multi-temporal spectral features using the temporal attention mechanism is expressed as:
[0048] ;
[0049] In the formula, is the attention score of the time step ; are the candidate hidden state features of the time steps respectively and Cosine similarity; is the denoising coefficient used to balance the influence of feature similarity on the attention score; is the total number of time steps; is the attention weight at time step t; is calculated by the exponential function; is the weight matrix of the linear transformation of the time attention score; is the bias term in the calculation of the attention score.
[0050] Furthermore, the process of optimizing the spatio-temporal spectral feature expression is as follows:
[0051] After extracting the spatio-temporal spectral features of each time step First, apply the time-optimized squeeze-and-excitation module to adaptively adjust the spatial and spectral dependencies between multi-temporal channels to obtain a feature sequence The expression of the processing process of the squeeze-and-excitation module is:
[0052] ;
[0053] In the formula, is the channel weight; is the global average pooling; is the ReLU non-linear activation function; is the Sigmoid function; ⊙ represents the element-wise multiplication operation; is the mapping matrix between time steps, used to fuse the channel weight information of the previous time step; is the linear transformation matrix; are the bias terms of the channel attention and the channel weight between time steps respectively;
[0054] For the adjusted feature sequence , calculate the fusion weight of each time step through the fully connected layer, and perform multi-temporal spatio-temporal spectral feature fusion. The expression is:
[0055] ;
[0056] In the formula, is the scoring function, used to calculate the attention weight at time step t.
[0057] Furthermore, the dominant tree species identification method is applicable to the multi-temporal dynamic monitoring of large-scale forest areas and supports the adaptive modeling of cross-seasonal spectral changes.
[0058] With the above technical solutions, the present invention provides a dominant tree species identification method based on time-space-spectrum multi-domain perception and guidance, which has at least the following beneficial effects:
[0059] 1. The dual-stream network architecture and the spatio-temporal-spectral feature fusion module proposed by the present invention focus on computational efficiency in design. Through effective feature extraction and integration strategies, it reduces the consumption of computing resources, improves the efficiency of processing large-scale spaceborne multispectral image data, and meets the high-efficiency requirements in practical applications.
[0060] 2. By constructing a dual-stream architecture and a spatio-temporal-spectral feature fusion module specifically for spaceborne multispectral characteristics, the present invention not only performs excellently in dealing with large ranges and diverse environmental conditions, but also can effectively capture seasonal changes, showing excellent performance in both static and dynamic spatio-temporal feature modeling. At the same time, it has strong scalability, providing efficient and reliable technical support for mapping the dominant tree species in the study area.
[0061] 3. The present invention combines the spatial-spectral information characteristics of spaceborne multispectral data, focuses on global information collaborative modeling and spectral dynamic information capture for the first time, and proposes a dual-stream network architecture specifically designed for spaceborne multispectral classification. This architecture effectively overcomes the performance bottlenecks of existing multispectral classification models under the challenges of spectral mixing interference and spatial heterogeneity, and significantly improves the accuracy and robustness of tree species recognition.
[0062] 4. The present invention innovatively designs a plug-and-play spatio-temporal-spectral feature fusion module. Based on multi-domain perception and guidance strategies, it can fully explore and integrate the dynamic semantic information in multi-temporal spaceborne multispectral data. It effectively captures the seasonal changes in the spectral characteristics of forest tree species, enhances the adaptability of the model in complex spatio-temporal scenarios, and improves the accuracy of large-scale forest tree species recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:
[0064] Figure 1 is a flowchart of the method for identifying dominant tree species in the present invention;
[0065] Figure 2 is the TS in the present invention 2 network structure diagram of the PGNet model;
[0066] Figure 3 is the network structure diagram of the spatio-temporal-spectral feature fusion module in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0067] To make the above objects, features, and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. Thereby, a full understanding of the implementation process of how this application uses technical means to solve technical problems and achieve technical effects can be obtained and implemented accordingly.
[0068] In order to improve the efficiency of processing large-scale spaceborne multi-spectral image data in this embodiment, to meet the high-efficiency requirements in practical applications, and to perform excellently under a wide range of diverse environmental conditions, and to effectively capture seasonal changes, it shows excellent performance in both static and dynamic spatio-temporal feature modeling. Please refer to Figures 1-3 , this embodiment is based on TS 2 PGNet framework (Temporal-spatial-spectral multi-domain perception and guidance network) to propose a method for identifying dominant tree species based on spatio-temporal-spectral multi-domain perception and guidance, as Figure 2 shown, the TS 2 PGNet framework consists of a parallel two-stream architecture and a plug-and-play spatio-temporal-spectral feature fusion module (TS 2 FM). Please refer to Figure 1 , this method includes the following steps:
[0069] S1. Obtain multi-spectral image data of spaceborne multi-temporal covering the target area. In this embodiment, multi-temporal spaceborne multi-spectral image data is obtained through a satellite platform, and then radiation correction, geometric registration, normalization and other processing are performed to improve the accuracy of the obtained data.
[0070] S2. Parallel feature extraction is performed on the multi-spectral image data through a two-stream network architecture to obtain two-stream features, where the two-stream architecture includes two key branches: a lightweight residual Transformer branch (Lightweight Residual Transformer Branch, LRTM) and a mixed-scale spectral-spatial fusion branch (Mixed-Scale Spectral-Spatial Fusion Module, MS 2 F). The lightweight residual Transformer branch can effectively capture long-range spectral dependencies across bands through global spectral dependence modeling, thereby improving the modeling ability of spectral information. The mixed-scale spectral-spatial fusion branch adopts a local feature decomposition and global dynamic fusion strategy to realize the joint representation of spatial features and spectral features, enhancing the model's ability in the fusion of spatial and spectral information.
[0071] To solve the spectral mixing effect of forest canopies in multispectral image processing, this embodiment proposes a lightweight residual Transformer branch combined with a dynamic hierarchical scaling mechanism, which can capture long-range spectral dependencies across bands when modeling the input spectral information. The overall architecture of the lightweight residual Transformer branch is as shown in Figure 2 shown. Its core idea is to enhance the model's ability to model complex spectral relationships by finely extracting and adaptively scaling spectral features. In the lightweight residual Transformer branch, each Transformer encoder layer first extracts features through the multi-head self-attention mechanism (MHSA) and the asymmetric convolution block (ACB). Subsequently, the dynamic hierarchical scaling residual connection mechanism (LSRC) is used to adaptively scale the input features to adapt to feature information at different scales. Finally, residual connection and layer normalization are used to ensure the stable transmission and effective flow of information. The specific process is as follows.
[0072] Specifically, the lightweight residual Transformer branch adopts an improved lightweight Transformer module, named L-Transformer, for pixel-level spectral information extraction. L-Transformer first processes multiple attention heads in parallel through the multi-head self-attention mechanism (MHSA) to capture complex correlations between input features, especially the highly non-linear relationships between spectral bands. The multi-head self-attention mechanism can effectively mine the correlations between different bands by adaptively adjusting the weights of different attention heads, thereby improving the model's ability to capture spectral features. Its calculation process is as follows:
[0073] ;
[0074] In the formula, represents the input feature the output obtained after applying the multi-head self-attention mechanism; represents the concatenation process; represents the th attention head; is the output weight matrix; represents the linear transformation; are the learnable weight matrices of query, key, and value, respectively.
[0075] The above formula represents the calculation process of the multi-head self-attention mechanism (MHSA). It divides the input features into multiple attention heads, calculates the attention weights of each attention head separately, and splices and linearly transforms the outputs of these attention heads to obtain the final output. This mechanism can capture different aspects of information in the input features and improve the performance of the model.
[0076] On this basis, in order to further improve the expression ability and computational efficiency of the model, this module adopts an asymmetric convolution block (ACB) to replace the traditional feed-forward neural network in the Transformer architecture. The asymmetric convolution block (ACB) enhances the expression ability of spectral features through horizontal 1×k convolution and vertical k×1 convolution operations along the spectral dimension. Specifically, the input features first undergo asymmetric convolution operations of 1×k and k×1 to capture the dependencies and enhance the information in the spectral domain respectively. Subsequently, these features are fused with the features extracted by the multi-head self-attention mechanism in the Transformer, and finally, the activation function is used to further improve the model's expression ability for spectral information to enhance the effectiveness and robustness of the model in processing multi-spectral images. Asymmetric convolution block The expression is as follows:
[0077] ;
[0078] In the formula, and are the horizontal and vertical convolution kernels respectively; and are both bias terms; is the ReLU activation function; is the input semantic information.
[0079] The asymmetric convolution block enhances the sensitivity and expression accuracy to spectral information while maintaining a lightweight design by refining the feature modeling in the spectral direction.
[0080] To address the challenges of spectral mixing and high redundancy in complex forest backgrounds, this embodiment designs a residual connection module with a layer-wise scaling residual connection mechanism (LSRC) to further enhance the learning ability of the model. The layer-wise scaling residual connection mechanism adaptively scales each layer of features, allowing the network to automatically adjust the importance of information at different levels, thereby optimizing the transmission process of features in the network. The residual connection module proposed in this embodiment dynamically adjusts the residual weights based on the layer-wise scaling residual connection mechanism through the scaling factor The specific process is as follows:
[0081] First, perform global average pooling on the features of the current layer to extract global context information , and the expression is:
[0082] ;
[0083] In the formula, are the number of channels, height, and width of the feature map respectively; is the position at the feature vector; is a feature tensor with the shape of ; is a real number field matrix with n rows and c columns;
[0084] Then, generate a scaling factor through a two-layer fully connected network , and the expression is:
[0085] ;
[0086] In the formula, is the ReLU activation function, which is used to increase the non-linear expression ability; is the Sigmoid loss function, which restricts the scaling factor within the range of [0,1]; and are both learnable weight matrices, and are real number field matrices with shapes and respectively, represents the number of feature channels, represents the ratio of channel compression and expansion; is the bias term;
[0087] Finally, use the generated scaling factor to adjust the weights of the contribution of the residual connection, and the expression is:
[0088] ;
[0089] In the formula, is the input feature of the current layer, is the feature representation after transformation of the layer; ⊙ represents the element-wise multiplication operation.
[0090] The residual connection module adjusts the weights through the scaling factor Dynamically adjust the weights of residual connections, enabling the model to adaptively adjust the intensity of information transmission according to the importance of current features, enhancing the expression of key stand features. In the case of severe spectral mixing phenomena, this can effectively help the model distinguish the spectral features of different tree species.
[0091] To overcome the problem of inconsistent global information expression caused by spatial heterogeneity in multi-temporal spaceborne multi-spectral image data and the difficulty of spectral-spatial fusion in high-dimensional feature spaces, this embodiment constructs a mixed-scale spectral-spatial fusion branch (MS 2 F) for dynamically fusing local and global spectral features and spatial features for joint representation. This branch aims to achieve dynamic fusion of local and global spectral-spatial features based on spaceborne multi-spectral by combining a multi-attention mechanism (Multi-Attention Mechanism, MAM) with a mixed-scale convolution design, improving the reliability of the model in large-scale forest monitoring.
[0092] During the feature extraction process in the spectral domain and spatial domain, the mixed-scale spectral-spatial fusion branch (MS 2 F) introduces a multi-attention mechanism and an adaptive weight fusion with suppression strategy (Adaptive Weight Fusion with Suppression, AWFS) to achieve efficient mining of key features and effective suppression of redundant information in multi-spectral data. Different from traditional models that only rely on a single spectral or spatial attention mechanism, the mixed-scale spectral-spatial fusion branch (MS 2 F) significantly improves the richness and discrimination ability of feature expression by simultaneously integrating a spectral attention module (Spectral Attention Block) and a spatial attention module (MultiGroup Spatial Attention), achieving dual attention between spectral channels and within spatial regions.
[0093] Specifically, the spectral attention module captures the global dependence relationship between spectral channels through 1×1 convolution and a multi-head self-attention mechanism, generating spectral attention weights to highlight key band features. This module first performs inter-channel information fusion using a 1×1 convolution kernel, then introduces non-linear features through batch normalization and the ReLU activation function, and then calculates the correlation of spectral channels through the multi-head self-attention mechanism to generate spectral attention weights, thereby enhancing the important spectral information in the feature map. The expression is:
[0094] ;
[0095] In the formula, Q, K, and V are the query, key, and value matrices respectively; is the regularization term to prevent overfitting; is the dimension of the key vector; is the Softmax normalization function; is the spectral semantic information enhanced by spectral attention;
[0096] Finally, the enhanced spectral features are refined into detailed representations with a higher semantic level through attention weighting and linear transformation. The expression is:
[0097] ;
[0098] In the formula, is the output of the final spectral attention feature; is the linear transformation matrix represents the index set of spectral channels.
[0099] Meanwhile, the spatial attention module dynamically generates spatial attention weights through depth convolution and pointwise convolution operations to capture important feature regions in the spatial domain and suppress background noise and irrelevant information. Specifically, this module uses depth convolution (Depthwise Convolution) for spatial feature extraction to maintain channel independence; then it fuses channel information through pointwise convolution (1×1 Convolution) and generates spatial attention weights using the Sigmoid activation function. The generated spatial attention weights are multiplied by the original feature map to achieve adaptive enhancement and suppression of spatial features. The expression is:
[0100] ;
[0101] In the formula, is the adaptive learning parameter for dynamically weighting spectral and spatial information; W is the suppression weight obtained through adaptive learning; is the Softmax normalization function; is the Sigmoid activation function; ⊙ represents the per-pixel multiplication operation.
[0102] After multi-slice learning, to address the highly heterogeneous spectral-spatial feature problem in hyperspectral image data due to environmental conditions and tree species composition, this embodiment adopts a hybrid-scale convolution design (Hybrid-Scale Convolution, HSC). This method uses three cascaded hybrid-scale convolution blocks (7×7, 5×5, and 3×3) to extract spectral and spatial features from macro to micro layer by layer and gradually refine the feature representation. The process of hybrid-scale convolution can be expressed as:
[0103] ;
[0104] In the formula, is the output feature map of the current branch; is the feature map obtained after being transformed by the activation function.
[0105] Meanwhile, batch normalization and ReLU activation operations are added between each layer to enhance the training stability and non-linear expression ability. Finally, the feature information is aggregated through the global pooling layer (Global Pooling) to achieve the consistent expression of local and global features. Compared with traditional multi-scale convolutions, this design can effectively alleviate the interference caused by global heterogeneity in multi-spectral data processing through the feature extraction strategy of layer-by-layer refinement.
[0106] In summary, the hybrid-scale spectral-spatial fusion branch can optimize the extraction and fusion of complex spectral-spatial features in spaceborne multi-spectral image data through the synergistic effect of multiple attention mechanisms and slice reinforcement learning. The hybrid-scale convolution design can further enhance the global consistency of feature extraction, enabling the model to have higher classification performance and stability in large-scale forest monitoring tasks.
[0107] S3. Input the extracted dual-stream features into the spatio-temporal-spectral feature fusion module (TS 2 FM), and through the hierarchical temporal-spectral integration mechanism (HTSI) and the dynamic spatio-temporal excitation optimization mechanism (DSEO), cross-domain fusion of multi-temporal dynamic features is performed to obtain spatio-temporal-spectral features. In the multi-spectral image data processing task, the fusion of features at different time steps contained in multi-temporal multi-spectral image data lacks pertinence, which may lead to information redundancy and feature conflicts. Therefore, this embodiment innovatively proposes a spatio-temporal-spectral feature fusion module (TS 2 FM) based on multi-domain perception and guidance, as Figure 3 shown. The core idea of this module is to use the multi-domain perception strategy to perform specific temporal modeling and fusion on different feature domains, and fully consider the inherent attributes of features in each domain and the dependencies across time steps, guiding the model to focus on key temporal dynamic changes and avoiding information redundancy and feature conflicts.
[0108] Among them, the hierarchical temporal-spectral integration mechanism (Hierarchical time-spectral integration, HTSI) uses a multi-level GRU network to recursively model the dynamic dependencies of time steps, and uses a dynamic denoising-optimized temporal attention mechanism to perform weighted fusion on multi-temporal spectral features, that is, to capture the dynamic changes of multi-temporal spectra through the GRU network and the dynamic denoising attention mechanism.
[0109] In multi-temporal multi-spectral image data, spectral features can change significantly in the time dimension due to vegetation phenological characteristics or environmental fluctuations. This dynamic change not only affects the importance of spectral features at a single time step but may also lead to the introduction of redundant information, resulting in the dilution or conflict of time information in dynamic spectral features. Therefore, in this embodiment, a hierarchical temporal spectral integration mechanism is constructed to enhance the precise screening ability of time spectral dynamic information.
[0110] Specifically, first, this embodiment uses a multi-level GRU network to capture the dynamic dependencies between time steps and extracts the hidden state features of each time step through recursive calculation. , is a real-number domain matrix; the expression for the recursive calculation process of the GRU network is:
[0111] ;
[0112] In the formula, is the update gate; is the reset gate, which controls the fusion and reset of the previous time step state and the current input; is the final hidden state feature at time step ; ⊙ represents element-wise multiplication; is the Sigmoid function; is the hyperbolic tangent function; represents the candidate hidden state; is the temporal spectral information at the current time step ; respectively represent the weight calculation matrices of the update gate, reset gate, and candidate hidden state; respectively represent to the weight calculation matrices of the update gate, reset gate, and candidate hidden state; respectively represent the biases of the update gate, reset gate, and candidate hidden state.
[0113] Subsequently, to further avoid the conflict of features between time steps and the interference of redundant information, this embodiment introduces a dynamic denoising optimized time attention mechanism to perform weighted fusion on multi-temporal spectral information, and the expression for its calculation process is:
[0114] ;
[0115] In the formula, is the attention score at time step ; is the cosine similarity between the candidate hidden state features at time steps and and ; is the denoising coefficient used to balance the influence of feature similarity on the attention score. is the total number of time steps; is the attention weight at time step t; is the exponential function calculation; is the weight matrix of the linear transformation of the time attention score; is the bias term in the attention score calculation.
[0116] In this embodiment, the gating mechanism of the GRU network can effectively screen out important features with high temporal correlation. At the same time, combined with the dynamic focusing of the attention mechanism and the robust design of the denoising strategy, redundant information can be effectively removed to ensure that the model only focuses on the key temporal dynamics. At the same time, its lightweight structure can also balance the contradiction between complex time relationship modeling and computational efficiency.
[0117] The Dynamic Spatiotemporal Excitation Optimization (DSEO) mechanism optimizes the spatiotemporal spectral feature representation by adjusting the channel weights based on temporal perception and combining the weight information of the previous time step, that is, optimizing the spatiotemporal spectral feature fusion by adjusting the channel weights based on temporal perception.
[0118] There are dynamic dependencies in the spatial-spectral features of different time steps, and simple local modeling cannot capture these characteristic changes across time steps. Therefore, this embodiment designs a Dynamic Spatiotemporal Excitation Optimization (DSEO) mechanism to enhance the model's adaptability to the evolution of spatiotemporal features by capturing the dynamic changes and cross-time step dependencies of spatiotemporal spectral features.
[0119] Specifically, after extracting the spatiotemporal spectral features of each time step first, apply the time-optimized squeeze-and-excitation module (D-SE) to adaptively adjust the spatial and spectral dependencies between multi-temporal channels to obtain a feature sequence . Different from the traditional SE module, the squeeze-and-excitation module introduces a time perception layer, that is, this module not only considers the global features of the current time step but also introduces the channel weight information of the previous time step to achieve dynamic adjustment in time series and more precisely capture the changes in spatiotemporal features. The expression of the processing process of the squeeze-and-excitation module is:
[0120] ;
[0121] In the formula, is the channel weight; is the global average pooling; is the ReLU non-linear activation function; is the Sigmoid function; ⊙ represents the element-wise multiplication operation; is the time-step mapping matrix, which is used to fuse the channel weight information of the previous time step; is the linear transformation matrix; are the bias terms of channel attention and time-step channel weight respectively; in is the previous time step.
[0122] Through the squeeze-and-excitation module, the channel weights are dynamically adjusted with the change of time steps, so as to more accurately capture the dynamic dependencies and change patterns of spatio-temporal spectral features. For the adjusted feature sequence , the fusion weights of each time step are calculated through a fully connected layer, and multi-temporal spatio-temporal spectral feature fusion is carried out. The expression is:
[0123] ;
[0124] In the formula, is the scoring function, which is used to calculate the attention weight at time step t.
[0125] In this embodiment, the spatio-temporal-spectral feature fusion module (TS 2 FM) deeply mines the semantic information of spatio-temporal spectrum in the tree species classification task of multi-temporal multi-spectral images. The spatio-temporal-spectral feature fusion module can dynamically integrate the changing features in different time dimensions, improve the expression ability of multi-dimensional data features, and further enhance the model's identification ability for dynamic changes and complex spectral features.
[0126] S4. Based on the fused spatio-temporal spectral features, the recognition results of the dominant tree species in the forest are output through a classifier (MLP). According to the recognition results, it can be seen that the two-stream network architecture and spatio-temporal-spectral feature fusion module proposed in this embodiment focus on computational efficiency in design. Through effective feature extraction and integration strategies, the consumption of computational resources is reduced, and the efficiency of processing large-scale spaceborne multi-spectral image data is improved, meeting the high-efficiency requirements in practical applications.
[0127] To sum up, the TS 2 PGNet model proposed by the present invention makes full use of the spatial, spectral and time information of multi-spectral data through its innovative two-stream architecture and spatio-temporal-spectral feature fusion module, significantly improving the accuracy and robustness of forest tree species classification, overcoming many challenges in the large-scale application of multi-spectral data, and having broad application prospects and promotion value.
[0128] This embodiment verifies the proposed TS 2 PGNet model, and the specific content is as follows:
[0129] 1. Sample labels
[0130] Field forest tree species surveys and experiments were carried out in the study area from 2023 to 2024. At the same time, referring to the forest resource inventory data of Huoshan County (2023), the spatial positions of more than 4000 typical forest dominant tree species and forest stand sample plots (30 m × 30 m) were obtained. Then, typical and representative areas were selected in the study area to produce a high-quality labeled dataset with a resolution of 10 m. The dataset sizes are 345×321 and 320×310 pixels respectively, covering 8 common dominant tree species and forest stand types in the study area, with a total of 166934 tree species pixels for classification, as shown in Table 1. Aiming at the problem of unbalanced label numbers for different tree species, data augmentation was performed on the spaceborne multi-spectral dataset, including random flipping, radiation enhancement, and mixed noise enhancement. Through these methods, the spatial variability can be increased, different lighting conditions can be simulated, and the within-class feature distribution can be enriched, which helps to improve the robustness and generalization ability of the model, so as to show higher stability in the forest tree species classification task in real scenarios.
[0131] Table 1 Distribution of label numbers for different tree species
[0132]
[0133] 2. Implementation details
[0134] This experiment was implemented in the deep learning framework PyTorch 2.1.0 and Python 3.11 environment. Relying on the NVIDIA GeForce RTX 4090 graphics card (24GB video memory) and CUDA 11.8 acceleration unit, model construction and optimization were carried out to make full use of hardware resources to improve the training speed and efficiency. During model training, the drop rate was set to 0.1 to reduce the risk of overfitting. The loss function was selected as CrossEntropyLoss, the initial learning rate was 1e-4, and the cosine annealing learning rate decay strategy was combined to gradually reduce the learning rate, so as to better adapt to the dynamic changes during training. In this embodiment, the ADAM optimizer was used to improve the convergence speed and stability of the model. The total number of training batches was 500 (epochs), and both the batch size and the patch size were set to 16 to achieve high training efficiency and memory management. To ensure the reproducibility of the results, the random seed was set to 14 in all experiments.
[0135] In addition, in order to comprehensively measure the constructed TS 2Regarding the classification accuracy of the PGNet model, this embodiment uses multiple evaluation metrics, including Overall Accuracy (OA), Kappa coefficient, and F1 Score for each category. OA can reflect the overall performance of the model across all categories, the Kappa coefficient is used to quantify the classification consistency of the model, and F1 Score can provide independent accuracy analysis for each category, thus presenting the classification performance in more detail. The combined use of these metrics helps to verify the effectiveness and stability of the model from multiple perspectives.
[0136] 3. Introduction to Comparative Models
[0137] To test the 2 performance of the PGNet model, this embodiment conducts a comparative evaluation of the state-of-the-art methods covering CNN and Transformer architectures. The specific description is as follows.
[0138] DBCTNet: A dual-branch convolutional Transformer network that extracts spatial and global spectral features using a multi-scale spectral feature extraction module and an improved Transformer encoder ConvTE, and constructs a dual-branch module based on 3-D CNN and ConvTE for accurate classification of remote sensing images.
[0139] GMA-Net: A CNN network based on multi-attention mechanisms and group guidance that extracts mixed information using spatial and spectral attention mechanisms and 3D CNN, and simultaneously learns pixel-level and patch-level spatial spectral features, effectively improving the model classification performance.
[0140] CLMA-Net: A dual-branch CNN network that constructs the backbone network with dual-branch convolutional neural layers, optimizes spectral features through a cross-layer multi-attention module that fuses multiple convolutional attention information, and completes the effective segmentation of spectral remote sensing images.
[0141] SpectralFormer: A highly flexible Transformer network that uses the Transformer model as the backbone network and designs grouped spectral embedding and cross-layer adaptive fusion modules for learning local detailed spectral representations respectively, and its accuracy is significantly better than that of the classical Transformer model.
[0142] 3D-CNN: An improved three-dimensional convolutional neural network (3D-CNN) tree species classification model that captures high-level semantic information of joint spatial spectral features and combines the earlystop method to prevent overfitting, achieving the highest tree species classification accuracy using multi-spectral remote sensing images among multiple baseline models.
[0143] GSC-ViT: A lightweight ViT network based on grouped separable convolution, which captures local spectral-spatial information in multi-spectral images with grouped pointwise convolution and group convolution, and designs a grouped separable multi-head self-attention module to provide global spatial feature extraction, showing excellent classification performance with relatively few training samples.
[0144] 4. Experimental Results
[0145] In the classification tasks of multiple multi-spectral images, TS²PGNet shows significant advantages in terms of F1-score, overall accuracy (OA), and Kappa coefficient. Compared with other state-of-the-art models, it demonstrates strong spatio-temporal adaptability and stability. The comparison results of the tree species mapping accuracy of the same model are shown in Table 2.
[0146] Specifically, the OA (Kappa) in spring, summer, and autumn are 81.09% (75.92%), 80.51% (75.49%), and 80.77% (75.7%) respectively, while the OA (Kappa) under multi-temporal conditions reaches 83.23% (78.89%). At the same time, by comparing the performance of other models, it can be found that there are certain limitations in other models in this classification task. For example, although GMA-Net and CLMA-Net perform well in some seasons, the accuracy fluctuates greatly between different seasons, indicating insufficient stability; the classification accuracy of 3-DCNN and DBCTNet under single-temporal conditions is significantly lower than that of TS²PGNet, while the accuracy improves under multi-temporal conditions, which shows that their architectures have limited ability to model the features of multi-spectral data under single-temporal conditions and need to rely on multi-temporal data to make up for the deficiencies. In addition, although the classification performance of SpectralFormer and GSC-ViT is relatively stable for each task, the overall accuracy is still lower than that of TS²PGNet, and they fail to fully exploit the spatio-temporal-spectral features in multi-spectral data.
[0147] Table 2 Comparison of tree species mapping accuracy of different models
[0148]
[0149] The above results indicate that TS²PGNet can maintain stable and superior classification performance in different seasons and multi-temporal classification tasks through spatio-temporal-spectral dynamic modeling, demonstrating strong generalization ability.
[0150] Those of ordinary skill in the art at the intersection of remote sensing and artificial intelligence can understand that all or part of the steps in implementing the methods of the above embodiments can be completed by instructing relevant hardware through a program. Therefore, this application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, this application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.
[0151] The above embodiments have introduced the present invention in detail. Specific examples are used in this article to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A method for identifying dominant tree species based on time-space-spectrum multi-domain perception and guidance, characterized in that: The method comprises the following steps: S1. Acquire multi-spectral image data from satellites covering multiple time periods of the target area; S2. Performing parallel feature extraction on the multispectral image data through a dual-stream network architecture to obtain dual-stream features, wherein the dual-stream network architecture includes a first branch and a second branch, wherein: The first branch is a lightweight residual Transformer branch for capturing long-range spectral dependencies across bands when modeling spectral information; The second branch is a hybrid scale spectral-spatial fusion branch for dynamically fusing local and global spectral features with spatial features for joint representation; S3, input the extracted dual-stream features into the time-space-spectrum feature fusion module, and perform cross-domain fusion of multi-phase dynamic features to obtain time-space spectrum features through the hierarchical time-series spectrum integration mechanism and dynamic time-space excitation optimization mechanism; S4. Based on the fused spatiotemporal spectral features, the recognition results of the dominant tree species in the forest are output through the classifier.
2. The method for identifying dominant tree species according to claim 1, characterized in that: The lightweight residual Transformer branch includes a multi-head self-attention mechanism that captures the nonlinear dependencies between spectral bands by computing multiple sets of query, key, and value matrices in parallel, and an asymmetric convolution block that enhances the expression of spectral features by performing horizontal 1×k convolutions and vertical k×1 convolutions along the spectral dimension. The expression is: ; In the formula, and They are horizontal and vertical convolution kernels respectively; and All are bias terms; is the ReLU activation function; is the semantic information of the input; The lightweight residual Transformer branch also includes a layer-based scaling residual connection mechanism through a scaling factor Residual connection module with dynamically adjusted residual weights.
3. The potential tree species identification method according to claim 2, characterized in that: The scaling factor Generated by global average pooling and a fully connected network.
4. The method for identifying dominant tree species according to claim 2 or 3, characterized in that: The layer-based scaling residual connection mechanism uses a scaling factor The specific process of dynamically adjusting the residual weight includes: For the current layer features Perform global average pooling to extract global context information , the expression is: ; In the formula, are the number of channels, height, and width of the feature map, respectively; For location The eigenvector at ; Represents the shape The feature tensor of Generate scaling factors through a two-layer fully connected network , the expression is: ; In the formula, It is the ReLU activation function, which is used to increase the nonlinear expression ability; For the Sigmoid loss function, the scaling factor Restricted to the range [0,1]; and are all learnable weight matrices, and The shapes and The real field matrix of represents the number of feature channels, Indicates the ratio of channel compression and expansion; is the bias term; Use the resulting scaling factor The contribution of the residual connection is weighted and expressed as: ; In the formula, For the current The input features of the layer; For the The feature representation after layer transformation; ⊙ represents the element-by-element multiplication operation.
5. The method for identifying dominant tree species according to claim 1, characterized in that: The mixed-scale spectral-spatial fusion branch includes a multiple attention mechanism for dynamically enhancing key features in multispectral image data and suppressing redundant information, and a mixed-scale convolution block that uses multi-layer cascade convolution kernels to extract spectral features and spatial features from macroscopic to microscopic layer by layer; The multi-attention mechanism includes a spectral attention module that weights spectral channel features through 1×1 convolution and multi-head self-attention, and a spatial attention module that dynamically generates spatial weights through deep convolution and point-by-point convolution; The expression of the mixed-scale convolution block is: ; In the formula, is the output feature map of the current branch; The feature map is obtained after the activation function transformation.
6. The method for identifying dominant tree species according to claim 5, characterized in that: The spectral attention module first uses a 1×1 convolution kernel to fuse information between channels, then introduces nonlinear features through batch normalization and ReLU activation function, and then calculates the correlation of spectral channels through a multi-head self-attention mechanism to generate spectral attention weights, thereby enhancing important spectral information in the feature map. The expression is: ; Where Q, K, and V are query, key, and value matrices, respectively; It is a regularization term to prevent overfitting; is the dimension of the key vector; is the Softmax normalization function; is the spectral semantic information after spectral attention enhancement; Finally, the enhanced spectral features are refined into a detailed representation with a higher semantic level after attention weighting and linear transformation, expressed as: ; In the formula, The final spectral attention feature output; is the linear transformation matrix A collection of indices representing spectral channels; The spatial attention module uses deep convolution to extract spatial features and maintain channel independence. It then fuses channel information through point-by-point convolution and uses the Sigmoid activation function to generate spatial attention weights. The generated spatial attention weights are multiplied with the original feature map to achieve adaptive enhancement and suppression of spatial features. The expression is: ; In the formula, is an adaptive learning parameter used to dynamically weight spectral and spatial information; W is the inhibition weight obtained through adaptive learning; is the Softmax normalization function; is the Sigmoid activation function; ⊙ represents the pixel-by-pixel multiplication operation.
7. The method for identifying dominant tree species according to claim 1, characterized in that: The hierarchical temporal spectral integration mechanism adopts a multi-level GRU network to recursively model the dynamic dependency of time steps, and uses a dynamic denoising optimized temporal attention mechanism to perform weighted fusion of multi-phase spectral features; The dynamic spatiotemporal excitation optimization mechanism optimizes the spatiotemporal spectral feature expression by adjusting the channel weights based on time series awareness and combining the weight information of the previous time step.
8. The method for identifying dominant tree species according to claim 7, characterized in that: The multi-level GRU network is used to recursively model the dynamic dependency of time steps, specifically: A multi-level GRU network is used to capture the dynamic dependencies between time steps, and the hidden state features of each time step are extracted through recursive calculation. , the expression of the recursive calculation process of the GRU network is: ; In the formula, To update the door; To reset the gate, control the fusion and reset of the previous time step state and the current input; is the time step The final hidden state features of ⊙ represents element-by-element multiplication; is the Sigmoid function; is the hyperbolic tangent function; represents the candidate hidden state; is the current time step Time series spectrum information; Represent the weight calculation matrices of the update gate, reset gate, and candidate hidden state respectively; Respectively To the update gate, reset gate, and weight calculation matrix of the candidate hidden state; Represent the update gate, reset gate, and bias of candidate hidden states respectively; The time attention mechanism is used to perform weighted fusion of multi-phase spectral features. The calculation process is expressed as follows: ; In the formula, is the time step Attention score; The time steps are Candidate hidden state features and The cosine similarity of is the denoising coefficient used to balance the effect of feature similarity on the attention score; is the total number of time steps; is the attention weight at time step t; Calculate for exponential function; is the linear transformation weight matrix of the temporal attention score; is the bias term in the attention score calculation.
9. The method for identifying dominant tree species according to claim 7, characterized in that: The process of optimizing the spatiotemporal spectral feature expression is as follows: After extracting the spatiotemporal features of each time step Then, a time-optimized squeezing-excitation module is applied to adaptively adjust the spatial and spectral dependencies between multi-temporal channels to obtain the feature sequence. , the expression of the squeeze-excitation module processing process is: ; In the formula, is the channel weight; is global average pooling; is the ReLU nonlinear activation function; is the Sigmoid function; ⊙ represents the element-by-element multiplication operation; It is the mapping matrix between time steps, which is used to fuse the channel weight information of the previous time step; is a linear change matrix; They are the bias terms of channel attention and channel weight between time steps respectively; For the adjusted feature sequence , the fusion weight of each time step is calculated through the fully connected layer, and multi-phase spatiotemporal feature fusion is performed. The expression is: ; In the formula, is a scoring function used to calculate the attention weight at time step t.
10. The method for identifying dominant tree species according to any one of claims 1 to 9, characterized in that: The dominant tree species identification method is suitable for multi-temporal dynamic monitoring of large-scale forest areas and supports adaptive modeling of cross-seasonal spectral changes.
Citation Information
Patent Citations
Hyperspectral image classification method based on Transform and non-local neural network double-branch architecture
CN117218537A
Double-flow time high-sensitivity space-time fusion method for remote sensing image mutation
CN118072137A