A dual-alignment gate-driven multi-modal satellite image forest tree species classification method

By using the TDAGNet model in conjunction with multi-temporal multispectral imagery, panchromatic imagery, and lidar imagery, the problem of insufficient accuracy and stability in remote sensing tree species identification in complex forest scenes was solved, achieving higher accuracy and stability in tree species classification and identification.

CN122116363AActive Publication Date: 2026-05-29HEFEI UNIV OF TECH
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HEFEI UNIV OF TECH
Filing Date
2026-04-28
Publication Date
2026-05-29

Smart Images

  • Figure CN122116363A_ABST
    Figure CN122116363A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of spaceborne remote sensing image processing and analysis, solves the technical problem of insufficient precision and stability of existing technology in the face of complex forest scene remote sensing tree species identification, especially relates to a kind of double alignment gate driven multi-modal satellite image forest tree species classification method, TDAGNet model is used to extract multi-temporal multi-spectral spatio-temporal feature, panchromatic spatial texture feature and laser radar structure feature respectively, unified fusion representation is formed through different scale modal alignment and cross-modal gate collaborative mechanism;Cross-modal alignment and joint discrimination are realized on the unified fusion network to obtain tree species prediction classification result map.The present application fully excavates the complementary information and its correlation of multi-temporal multi-spectral image, panchromatic image and laser radar image through double-branch panchromatic modeling, different scale grid attention aggregation and double alignment gate mechanism, significantly improves the expression ability and fusion effect of multi-source remote sensing classification in complex forest scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of spaceborne remote sensing image processing and analysis technology, and in particular to a method for classifying forest tree species in multimodal satellite imagery driven by dual-alignment gating. Background Technology

[0002] Forest tree species identification is a crucial aspect of forest resource monitoring and management. Spaceborne remote sensing, with its advantages of wide coverage, short revisit cycles, and high acquisition efficiency, has been widely applied to forest tree species identification at regional scales. With the development of multi-temporal and multispectral remote sensing, high-resolution optical remote sensing, spaceborne lidar, and deep learning technologies, multimodal remote sensing fusion is gradually becoming an important means to improve tree species classification performance. Furthermore, existing research indicates that mid-level feature fusion methods are generally superior to early data stitching and late-stage decision fusion methods.

[0003] However, existing technologies still have the following shortcomings: First, there is a mismatch between the spatial resolution of spaceborne remote sensing observations and the scale of forest canopy structure. In complex forest stands, pixel mixing, canopy edge overlap, and shadow interference are prone to occur, leading to blurred interclass boundaries, increased intraclass differences, and insufficient cross-regional generalization ability of the model. Second, although multi-temporal multispectral images can reflect phenological change information, they are difficult to accurately describe fine-grained spatial features such as individual tree canopy width, canopy porosity, and boundary areas. Third, structural data such as spaceborne lidar are limited by footprint scale and spatial registration accuracy, making it difficult to form a unified discriminative feature with optical images. Fourth, although panchromatic images have high spatial resolution, existing methods mostly use them for panchromatic sharpening or super-resolution preprocessing, failing to fully utilize their independent spatial discriminative characterization role. Moreover, the detail injection process is prone to spectral distortion, making it difficult to fully reflect their complementary advantages with multispectral images.

[0004] Therefore, there is an urgent need for a multimodal satellite remote sensing tree species classification method that can synergistically utilize panchromatic imagery, multi-temporal multispectral imagery, and structural information to improve the accuracy and stability of remote sensing tree species identification in complex forest scenarios. Summary of the Invention

[0005] To address the shortcomings of existing technologies, this invention provides a dual-alignment gating-driven multimodal satellite image forest tree species classification method, which solves the technical problems of insufficient accuracy and stability in remote sensing tree species identification under complex forest scenarios.

[0006] To address the aforementioned technical problems, this invention provides a dual-alignment gating-driven multimodal satellite imagery forest tree species classification method, which includes the following steps: Acquire and preprocess multi-source, multimodal data, including multi-temporal, multispectral, panchromatic, and lidar images, of complex forest scenes within the study area; Multi-source, multi-modal data are input into the TDAGNet model to extract multi-temporal, multi-spectral spatiotemporal features, panchromatic space texture features, and lidar structural features, and a unified fusion representation is formed through heteroscale modal alignment and cross-modal gating collaborative mechanisms; The TDAGNet model includes a dual-branch panchromatic feature extractor for mining high-resolution spatial information in panchromatic images to obtain panchromatic spatial texture features, which include panchromatic detail features and panchromatic semantic features. Furthermore, by using pixel-level importance scoring and attention aggregation within grid cells, high-resolution pixel representations are adaptively mapped to grid-level full-color alignment features corresponding to the structural features of the lidar, thereby realizing a heteroscale modal alignment module that explicitly aligns from pixel space to structural grid space. In addition, a multimodal collaborative alignment and joint discrimination module is used to explicitly model the cross-modal consistency relationship between full-color space texture features, multi-temporal multispectral spatiotemporal features and lidar structural features on a unified fusion grid, so as to realize multimodal joint discrimination. Cross-modal alignment and joint discrimination are achieved on a unified and fused grid to obtain a tree species prediction classification result map including tree species category probabilities, so as to realize tree species classification mapping of the study area.

[0007] Furthermore, the dual-branch panchromatic feature extractor adopts a dual-path structure design that includes a high-frequency detail branch and a semantic backbone branch running in parallel; The high-frequency detail branch is used to extract local texture, crown boundary and fine-grained spatial variation information from panchromatic images to obtain panchromatic detail features, so as to enhance the TDAGNet model's ability to represent high-frequency structural details in complex forest scenes. The semantic backbone branch is used to extract high-level spatial semantic features from panchromatic images to obtain panchromatic semantic features, so as to make up for the limitation of the high-frequency detail branch, which focuses on local texture and edge information.

[0008] Furthermore, the TDAGNet model also includes a spatiotemporal feature extractor for extracting temporal-spatial-spectral features from multi-temporal multispectral images and for highlighting key phenological phases for tree species identification through temporal attention aggregation to obtain multi-temporal multispectral spatiotemporal features; And a lidar feature extractor for extracting stand vertical structure information from lidar images to obtain lidar structural features.

[0009] Furthermore, the process by which the spatiotemporal feature extractor extracts multi-temporal and multispectral spatiotemporal features includes: Let the input multi-temporal multispectral image be ,in B For batch size, T For the time phase number, CFor single-phase bands, H mul and W mul These are the height and width of the space, respectively. pass N 3D convolutional blocks extract spatiotemporal spectral joint features with spatiotemporal spectral coupling characteristics for deep characterization. ; The response at each time phase is obtained by global averaging in the spatial dimension, and in the time dimension... Attention weights during generation time Polymerization characteristics were characterized by stable time-series spectra. To adaptively highlight key phenological phases that are more discriminative for tree species identification; Introducing SEGate for aggregated features Channel recalibration was performed to obtain multi-temporal and multispectral spatiotemporal characteristics. This is used to enhance the representation of representative time-series spectra; The expression for the spatiotemporal feature extractor is as follows:

[0010] in, This represents an N-layer spatiotemporal convolutional block consisting of 3D convolutions, normalization, and activation functions; Spatiotemporal spectral joint features In the Time-dimensional index, spatial location The characteristic response corresponding to the location; For the first The spatiotemporal spectrum joint features corresponding to each time dimension index; This indicates element-wise weighting; This indicates a channel recalibration operation; Indexed by time dimension; For spatial location index in the height direction; This is the spatial location index in the width direction.

[0011] Furthermore, the lidar feature extractor employs two layers of depth-separable convolutional blocks for lightweight encoding, enhancing the ability to express structural patterns while controlling the number of parameters and computational load. The expression for the lidar feature extractor is:

[0012] in, The input is a lidar image; and They represent the first N Depth convolution kernels and point convolution kernels for layers; This represents the convolution operation; For the first N Layer depth can separate the input features of convolutional blocks; Indicates group normalization; Represents a nonlinear activation function; For the first N Output features of layered coded blocks; Encode features for the initial structure; For the first N -1 layer structure encoding features; For the first N Layered structure encoding features; The final output of the lidar structure features.

[0013] Furthermore, the high-frequency detail branch adopts N Depthwise separable convolutions, passed through concatenated layers, are used for feature extraction, outputting panchromatic detail features that maintain spatial resolution. The expression is:

[0014] in, Indicates the first N One deep convolutional kernel; This represents the corresponding point convolution kernel; This represents the convolution operation; Indicates group normalization; This represents the SiLU nonlinear activation function; For the first N The input features of a depthwise separable convolutional block on the high-frequency detail branch, when hour, For input panchromatic image ; The extraction of full-color semantic features by the semantic backbone branch includes: First, panchromatic images are processed using shallow convolutional blocks. Initial semantic features are obtained through preliminary encoding. ; Subsequently, the initial semantic features were processed using the anti-aliasing downsampling module. Smooth downsampling is performed to extract smooth semantic features. And for smooth semantic features Extracting mid-level semantic features through smooth downsampling To reduce aliasing distortion of high-frequency textures during scale compression; Next, the receptive field is gradually expanded using lightweight hollow residual blocks to obtain enhanced deep semantic features. Furthermore, enhanced contextual semantic features are obtained by improving the ability to characterize canopy spatial organization, regional structure, and global semantic patterns through multi-scale context modules. ; Finally, through channel compression and convolutional mapping, full-color semantic features are obtained for subsequent cross-modal fusion. ; The expression for the semantic backbone branch is:

[0015] in, This represents the initial encoding operation consisting of two convolutional layers, normalization, and activation functions; Indicates anti-aliasing downsampling; This represents a semantic extraction process consisting of convolutions and lightweight dilated residual blocks. This represents a semantic enhancement operation for further stacked holed residual blocks; and These represent downsampling and upsampling, respectively. Indicates a lightweight void residual block; , and These represent depthwise separable convolution, pointwise convolution, and channel convolution operations, respectively.

[0016] Furthermore, the process by which the heteroscale modal alignment module obtains grid-level panchromatic alignment features includes: Suppose the full-color detail features output by the high-frequency detail branch. ,in, B For batch size, C h The number of feature channels, H P and W P These represent the height and width of the full-color detail features, respectively. Full color detail features Divided into grids according to the target resolution G H × G W There are 1 grid cells, each grid cell being 1. H × W ,satisfy H P = G H × H , W P = G W × W , H , WThese are the height and width of the grid cell, respectively; A pixel-level importance score map is obtained through a pointwise convolutional mapping layer. S And within each grid cell, pixel-level importance scores are generated. S Perform normalization to generate normalized attention weights; full-color detail features of the corresponding area After performing weighted summation, the grid-level pancolor alignment feature is finally obtained. P Its core calculation process is expressed as follows:

[0017] in, This represents the pointwise convolutional kernel used to generate a single-channel score map; Indicates the ( G H , G W Position within each grid cell ( U , V Normalized attention weights; Indicates the ( G H , G W Traversing positions within each grid cell Unnormalized pixel-level importance score map; U and V These represent the spatial position indices along the height and width directions within the current grid cell, respectively. and These represent the summation indices for traversing all positions within the grid cell in the height and width directions, respectively.

[0018] Furthermore, the multimodal collaborative alignment and joint discrimination module takes as input the panchromatic semantic features output from the semantic backbone branch, the multi-temporal and multispectral spatiotemporal features output from the spatiotemporal feature extractor, the grid-level panchromatic alignment features after grid alignment of the high-frequency detail branch, and the lidar structural features, respectively constructs panchromatic-spectral gating and panchromatic-laser gating to constrain the intermodal relationships from two levels: spatial semantic-temporal spectrum and spatial detail-vertical structure. Finally, multimodal joint discrimination is achieved through feature fusion and tree species classification head.

[0019] Furthermore, the panchromatic-spectral gating mechanism is used to model the consistency relationship between panchromatic semantic features and multi-temporal multispectral spatiotemporal features, wherein: Full-color semantic features are and multi-temporal and multi-spectral spatiotemporal characteristics Each element is mapped to a unified common embedding space via 1×1 convolution and normalized along the channel dimension. Then, position-wise cosine similarity is calculated, and a panchromatic-spectral gated map is generated using a learnable scaling parameter and a sigmoid function. The expression for this calculation process is as follows:

[0020] in, This represents the normalized representation of full-color semantic features in the public embedding space; Normalized representation of multi-temporal and multispectral spatiotemporal features in a common embedding space; This is a cosine similarity graph; A panchromatic-spectral gated graph; and These represent the projected convolution kernels for panchromatic semantic features and multi-temporal multispectral spatiotemporal features, respectively; D For public embedded dimensions; This represents L2 normalization along the channel dimension; This represents element-wise multiplication; This represents the Sigmoid activation function; Softplus learnable scaling parameters; They represent and In the c Characteristic responses on each channel; The panchromatic-laser gating mechanism is used to model the consistency relationship between grid-level panchromatic alignment features and lidar structural features, wherein: When grid-level full-color alignment feature P and lidar structural features When spatial dimensions are inconsistent, the structural features of the lidar should be considered first. Interpolation to grid-level full-color alignment features P Consistent space size; They are then mapped to a unified common embedding space, and panchromatic-laser gated maps are generated through normalization and cosine similarity calculations. The expression for the calculation process is as follows:

[0021] in, O This represents the structural features of a lidar after alignment with the grid-level full-color alignment feature space. This represents the normalized representation of grid-level full-color alignment features in the common embedding space. This represents the normalized representation of the aligned lidar structural features in the common embedding space. A panchromatic laser similarity map; A full-color laser-gated image; This represents the bilinear interpolation operation; and These are the height and width of the grid-level full-color alignment feature, respectively. and These represent the projected convolution kernels for grid-level full-color alignment features and lidar structural features, respectively. For learnable scaling parameters; and They are respectively and In the c Characteristic responses on each channel.

[0022] Furthermore, cross-modal alignment and joint discrimination are achieved on the unified and fused grid, including: First, the full-color semantic features LiDAR structural features Multi-temporal and multi-spectral spatiotemporal characteristics Align to a unified, blended grid; Subsequently, the panchromatic-spectral gating map And full-color laser-gated chart As an auxiliary consistency channel, it is used in conjunction with full-color semantic features. LiDAR structural features Multi-temporal and multi-spectral spatiotemporal characteristics The components are spliced ​​together in the channel dimension to form a fusion feature. Input, i.e.:

[0023] Subsequently, fusion features The results are fused through convolutional mapping and channel attention, and finally output through a tree-type classification head. The expression is:

[0024] in, Indicates channel splicing; This represents a fusion operation consisting of convolutional mapping and channel attention; Represents a classification mapping; Y This is the final classification output.

[0025] By employing the above technical solution, the present invention provides a method for classifying forest tree species using dual-alignment gating driven multimodal satellite imagery, which has at least the following beneficial effects: 1. This invention employs a dual-path structure in the panchromatic image branch: a high-frequency detail branch preserves local texture, edge, and fine-grained spatial variation information, while a semantic backbone branch extracts stable and global semantic representations. Compared to existing technologies that use a single path to extract panchromatic image features, this invention balances detail preservation and semantic modeling requirements. This allows the panchromatic image to provide fine-grained support for local spatial alignment and stable support for subsequent fusion and classification, thereby improving the utilization efficiency and representational capability of panchromatic modalities in multimodal fusion processes.

[0026] 2. This invention establishes a panchromatic image gridding attention aggregation mechanism for cross-scale modal alignment. High-resolution panchromatic image features are divided into grid units according to low-resolution LiDAR data, and pixel features are adaptively aggregated within each grid unit using an attention mechanism. Compared to traditional coarse compression methods such as mean pooling or direct downsampling, this invention can highlight more discriminative local textures, edges, and structural responses during scale matching, reducing the ineffective loss of high-resolution information during downscaling. This improves the accuracy of mapping panchromatic image features to the low-resolution grid space, providing more reliable spatial priors and detailed support for subsequent cross-modal fusion.

[0027] 3. This invention constructs a three-modal collaborative fusion mechanism based on dual alignment gating. It explicitly models the feature relationships between panchromatic and lidar images, and between panchromatic and multi-temporal multispectral images, respectively, and generates corresponding alignment gating information based on intermodal similarity. Simultaneously, it combines a hierarchical fusion framework with different resolutions and representational forms to perform targeted feature encoding on panchromatic, lidar, and multi-temporal multispectral images, achieving cross-modal alignment and joint discrimination on a unified fusion grid. This improves the matching and synergy between data from different sources, with different spatial resolutions, and with different representational forms, fully leveraging the complementary advantages of each modality, reducing interference from modal differences, and thus enhancing the robustness, reliability, and comprehensive discrimination capability of tree species classification and identification results in complex forest remote sensing scenarios. Attached Figure Description

[0028] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 This is a network architecture diagram of the TDAGNet model in this invention; Figure 2 This is a schematic diagram of the principle architecture of the dual-branch panchromatic feature extractor in this invention; Figure 3 This is a schematic diagram of the principle architecture of the heteroscale modal alignment module in this invention; Figure 4 This is a schematic diagram of the principle architecture of the multimodal collaborative alignment and joint discrimination module in this invention; Figure 5 This is a tree species truth label diagram created in this invention; Figure 6 This is a diagram showing the tree species prediction and classification results generated by the TDAGNet model in this invention. Figure 7 To compare the tree species prediction classification results generated by LF-DLM in the model; Figure 8 To compare the tree species prediction classification results generated by ExViT in the model; Figure 9 To compare the tree species prediction classification results generated by AMIANet in the model; Figure 10 To compare the tree species prediction classification results generated by MAESTRO in the model; Figure 11 This is a comparison image of the tree species prediction classification results generated by MultiSenseSeg in the model. Detailed Implementation

[0029] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. This will allow for a full understanding of how the present application uses technical means to solve technical problems and achieve technical effects, and to facilitate its implementation.

[0030] This embodiment provides a multimodal satellite image forest tree species classification method driven by dual alignment gating. Through bi-branch panchromatic modeling, heteroscale grid attention aggregation, and dual alignment gating mechanism, it fully explores the complementary information and correlation between multi-temporal multispectral images, panchromatic images, and lidar images, significantly improving the expressive power and fusion effect of multi-source remote sensing classification in complex forest scenes.

[0031] First, multi-temporal multispectral images, panchromatic images, lidar images, and corresponding tree species sample labels of the study area are acquired. The acquired multi-source multimodal data are then preprocessed, and the tree species sample labels are tree species information that has been collected and recorded in detail beforehand.

[0032] Multi-temporal multispectral imagery can be obtained free of charge or for a fee through platforms such as Landsat, Sentinel-2, and the WorldView series, or by acquiring high-resolution images using a multispectral camera (such as MicaSense RedEdge), suitable for fine-grained monitoring of small areas. Panchromatic imagery is usually acquired simultaneously with multispectral imagery (such as the 0.31m panchromatic band of WorldView-3), requiring matching of the resolution and time series of the multispectral imagery. LiDAR imagery can be obtained through global airborne LiDAR data provided by Opentopography and the Illinois Height Modernization Project, or by using the Velodyne 16-line LiDAR, parsing the raw point cloud data via UDP protocol, converting it to PCD format for post-processing.

[0033] Preprocessing includes geometric correction and registration, radiometric correction, lidar point cloud processing, image mosaicking and cropping, cloud removal and shadow processing, etc. After preprocessing, it provides a high-quality data foundation for subsequent tree species classification, change detection or ecological parameter inversion.

[0034] Then, multi-temporal multispectral images, panchromatic images, and lidar images are input into the TDAGNet model to extract multi-temporal multispectral temporal features, panchromatic spatial texture features, and lidar structural features, respectively. A unified fusion representation is formed through heteroscale modal alignment and cross-modal gating collaborative mechanisms.

[0035] This invention constructs a TDAGNet (Tri-modal Dual-Alignment Gating Network) model for multimodal feature extraction and forms a unified fusion representation through heteroscale modal alignment and cross-modal gating collaborative mechanisms, thereby achieving cross-modal alignment and joint discrimination on a unified fusion grid.

[0036] like Figure 1 As shown, the TDAGNet model consists of three modality-specific feature extraction branches and a multimodal collaborative alignment and joint discrimination module. The overall design aims to fully integrate forest "structure-spectrum" information to improve the performance and robustness of dominant tree species identification in complex forest scenarios. Overall, the TDAGNet model comprises three key parts: a dual-branch panchromatic feature extractor consisting of a semantic backbone branch and a high-frequency detail branch; and heteroscale modality alignment and multimodal collaborative alignment and joint discrimination modules.

[0037] Simultaneously, lightweight feature extraction is required for both multi-temporal multispectral images and lidar images. For multi-temporal multispectral images, spatiotemporal joint modeling is primarily performed using a spatiotemporal feature extractor. This involves extracting coupled temporal-spatial-spectral features through 3D convolutional blocks and utilizing temporal attention to highlight key phenological phases, forming a stable temporal spectral representation. Lidar images, on the other hand, employ a lightweight convolutional encoder to extract structural features, providing vertical structure information for subsequent fusion.

[0038] In this invention, the dual-branch panchromatic feature extractor is one of the core innovations of the TDAGNet model. This branch adopts a dual-path design: the semantic backbone branch extracts high-level spatial semantic features of panchromatic images through anti-aliasing downsampling, lightweight holed residual blocks, and multi-scale contextual modeling; while the high-frequency detail branch focuses on preserving edge, texture, and local structural change information, so that the panchromatic image can provide global semantic support and characterize complex forest boundaries and fine-grained heterogeneity.

[0039] The TDAGNet model introduces a gridded attention aggregation mechanism, which divides high-frequency detail features into blocks according to the LiDAR grid and performs adaptive aggregation within each grid cell, thereby enhancing the accuracy of heteroscale modal alignment.

[0040] In the multimodal collaborative alignment and joint discrimination module, the TDAGNet model constructs panchromatic-spectral gating and panchromatic-laser gating mechanisms respectively, generating double-alignment gating information based on similarity. Finally, the three main features and the double-alignment gating information are jointly input into the feature fusion and tree classification head on a unified fusion grid to achieve multimodal joint discrimination.

[0041] In summary, the TDAGNet model, through bi-branch panchromatic modeling, heteroscale grid attention aggregation, and dual-alignment gating mechanism, fully exploits the complementary information and correlations of the three modalities, significantly improving the expressive power and fusion effect of multi-source remote sensing classification in complex forest scenes.

[0042] The spatiotemporal feature extractor jointly extracts temporal, spatial, and spectral features from multi-temporal multispectral images, and highlights key phenological phases for tree species identification through temporal attention aggregation. Let the input multi-temporal multispectral images be... ,in B For batch size, T For the time phase number, C For single-phase bands, H mul and W mul These represent the spatial height and width, respectively. The spatiotemporal feature extractor first uses... N 3D convolutional blocks extract spatiotemporal spectral joint features with spatiotemporal spectral coupling characteristics for deep characterization. Based on this, the response of each temporal phase is obtained by global averaging across the spatial dimension, and then further processed in the time dimension. Attention weights during generation time Polymerization characteristics were characterized by stable time-series spectra. This approach adaptively highlights key phenological phases that are more discriminative for tree species identification; finally, it introduces SEGate to aggregate features. Channel recalibration was performed to obtain multi-temporal and multispectral spatiotemporal characteristics. This further enhances the representative temporal spectral representation. The core calculation process of the spatiotemporal feature extractor is expressed as follows:

[0043] in, This represents an N-layer spatiotemporal convolutional block consisting of 3D convolutions, normalization, and activation functions; Spatiotemporal spectral joint features In the Time-dimensional index, spatial location The characteristic response corresponding to the location; For the first The spatiotemporal spectrum joint features corresponding to each time dimension index; This indicates element-wise weighting; This indicates a channel recalibration operation; Indexed by time dimension; For spatial location index in the height direction; This is the spatial location index in the width direction.

[0044] like Figure 2 As shown, the dual-branch panchromatic feature extractor is mainly used to mine high-resolution spatial information in panchromatic images. It adopts a dual-path structure design with high-frequency detail branches and semantic backbone branches running in parallel, while taking into account both local detail preservation and high-level semantic modeling capabilities.

[0045] The high-frequency detail branch is used to extract local texture, crown boundary, and fine-grained spatial variation information from the panchromatic image to obtain panchromatic detail features, thereby enhancing the TDAGNet model's ability to represent high-frequency structural details in complex forest scenes. Let the input panchromatic image be... High-frequency detail branches are adopted N Depthwise separable convolutions, passed through concatenated layers, are used for feature extraction, outputting panchromatic detail features that maintain spatial resolution. Compared to standard convolution, depthwise separable convolution first captures local edge and texture responses through channel-wise spatial convolution, and then uses pointwise convolution to complete channel blending, thereby improving the ability to model high-dimensional spatial patterns with a lower number of parameters. Its core computational process can be represented as:

[0046] in, Indicates the first N One deep convolutional kernel; This represents the corresponding point convolution kernel; This represents the convolution operation; Indicates group normalization; This represents the SiLU nonlinear activation function; For the first N The input features of a depthwise separable convolutional block, when hour, For input panchromatic image , and These represent the height and width of the input feature, respectively.

[0047] The semantic backbone branch is used to extract high-level spatial semantic features from the panchromatic image to obtain panchromatic semantic features. Panchromatic semantic features have stronger stability and global discriminative ability, compensating for the limitations of the high-frequency detail branch, which focuses on local texture and edge information. Let the input panchromatic image be... The semantic backbone branch first processes the panchromatic image through shallow convolutional blocks. Initial semantic features are obtained through preliminary encoding. Subsequently, the initial semantic features were processed using an anti-aliasing downsampling module. Smooth downsampling is performed to extract smooth semantic features. And for smooth semantic features Extracting mid-level semantic features through smooth downsampling To reduce aliasing distortion of high-frequency textures during scale compression, a lightweight, hollow residual block is used to progressively expand the receptive field and obtain enhanced deep semantic features. Furthermore, enhanced contextual semantic features are obtained by improving the ability to characterize canopy spatial organization, regional structure, and global semantic patterns through multi-scale context modules. Finally, through channel compression and convolutional mapping, full-color semantic features are obtained for subsequent cross-modal fusion. Its core calculation process can be represented as follows:

[0048] in, This represents the initial encoding operation consisting of two convolutional layers, normalization, and activation functions; Indicates anti-aliasing downsampling; This represents a semantic extraction process consisting of convolutions and lightweight dilated residual blocks. This represents a semantic enhancement operation for further stacked holed residual blocks; and These represent downsampling and upsampling, respectively. Indicates a lightweight void residual block; , and These represent depthwise separable convolution, pointwise convolution, and channel convolution operations, respectively.

[0049] A lidar feature extractor is used to extract forest stand vertical structure information from lidar imagery. Assume the input lidar image is... Due to lidar imagery The structure has been organized into a low-resolution regular grid, primarily carrying information such as canopy height, vertical distribution, and structural differences. Therefore, the lidar feature extractor no longer employs a complex downsampling-decoding structure, but instead uses two depthwise separable convolutional blocks for lightweight encoding, enhancing the expressive power of structural patterns while controlling the number of parameters and computational cost. Its core computational process can be represented as follows:

[0050] in, and They represent the first N Depth convolution kernels and point convolution kernels for layers; This represents the convolution operation; For the first N Layer depth can separate the input features of convolutional blocks; Indicates group normalization; Represents a nonlinear activation function; For the first N Output features of layered coded blocks; Encode features for the initial structure; For the first N -1 layer structure encoding features; For the first N Layered structure encoding features; The final output of the lidar structure features.

[0051] The heteroscale modal alignment module is used to divide high-resolution panchromatic detail features into grid cells according to the low-resolution LiDAR structural features, and adaptively aggregate pixel features within each grid cell using an attention-based approach. This module addresses the spatial scale mismatch between high-resolution panchromatic detail features and the low-resolution LiDAR grid by adaptively selecting more discriminative edges, textures, and local structural responses within each low-resolution structural cell.

[0052] The heteroscale modal alignment module takes the panchromatic detail features output from the high-frequency detail branch as input. Through pixel-level importance scoring and in-grid attention aggregation, it adaptively maps high-resolution pixel representations to grid-level panchromatic alignment features corresponding to the LiDAR structural features, thereby achieving explicit alignment from pixel space to structural grid space, such as... Figure 3 As shown.

[0053] Suppose the full-color detail features output by the high-frequency detail branch. ,in, B For batch size, C h The number of feature channels, H P and W P These represent the height and width of the panchromatic detail features, respectively. The grid is then divided according to the target resolution. G H × G W There are 1 grid cells, each grid cell being 1. H × W ,satisfy H P = G H × H , W P = G W × W , H , W These represent the height and width of the grid cell, respectively.

[0054] First, the heteroscale modality alignment module generates a single-channel importance score map for each pixel location. Specifically, a pixel-level importance score map is obtained by first performing a pointwise convolutional mapping. S Subsequently, pixel-level importance scores were generated within each grid cell. S Perform normalization to generate normalized attention weights, and apply them to the panchromatic detail features of the corresponding region. After performing weighted summation, the grid-level pancolor alignment feature is finally obtained. P Its core calculation process can be represented as follows:

[0055] in, This represents the pointwise convolutional kernel used to generate a single-channel score map; Indicates the ( G H , G W Position within each grid cell ( U ,V Normalized attention weights; Indicates the ( G H , G W Traversing positions within each grid cell Unnormalized pixel-level importance score map; U and V These represent the spatial position indices along the height and width directions within the current grid cell, respectively. and These represent the summation indices for traversing all positions within the grid cell in the height and width directions, respectively.

[0056] The multimodal collaborative alignment and joint discrimination module is used to explicitly model the cross-modal consistency relationships between panchromatic semantic features, grid-level panchromatic alignment features, multi-temporal multispectral spatiotemporal features, and lidar structural features on a unified fusion grid, and to perform joint discrimination based on this. This module takes the panchromatic semantic features output from the semantic backbone branch, the multi-temporal multispectral spatiotemporal features output from the spatiotemporal feature extractor, the grid-level panchromatic alignment features after grid alignment of the high-frequency detail branch, and lidar structural features as inputs. It constructs panchromatic-spectral gating and panchromatic-liquid gating respectively to constrain the intermodal relationships from two levels: "spatial semantics-temporal spectrum" and "spatial detail-vertical structure." Finally, it achieves trimodal joint discrimination through feature fusion and a tree-type classification head. Figure 4 As shown.

[0057] The panchromatic-spectral gating mechanism is used to model the consistency relationship between panchromatic semantic features and multi-temporal, multispectral spatiotemporal features. Let the panchromatic semantic features output by the semantic backbone branch be... The spatiotemporal feature extractor outputs multi-temporal and multispectral spatiotemporal features as follows: First, the full-color semantic features are: and multi-temporal and multi-spectral spatiotemporal characteristics Each image is mapped to a unified common embedding space via 1×1 convolution and normalized along the channel dimension. Then, position-wise cosine similarity is calculated, and a panchromatic-spectral gated map is generated using a learnable scaling parameter and a sigmoid function. The expression for the calculation process is as follows:

[0058] in, This represents the normalized representation of full-color semantic features in the public embedding space; Normalized representation of multi-temporal and multispectral spatiotemporal features in a common embedding space; This is a cosine similarity graph; A panchromatic-spectral gated graph; and These represent the projected convolution kernels for panchromatic semantic features and multi-temporal multispectral spatiotemporal features, respectively; D For public embedded dimensions; This represents L2 normalization along the channel dimension; This represents element-wise multiplication; This represents the Sigmoid activation function; Softplus learnable scaling parameters; They represent and In the c Characteristic responses on each channel.

[0059] The panchromatic-laser gating mechanism is used to model the consistency relationship between grid-level panchromatic alignment features and lidar structural features. Let the grid-level panchromatic alignment feature obtained after heteroscale grid alignment of the high-frequency detail branch be... P The lidar structural features output by the lidar feature extractor are: .

[0060] When grid-level full-color alignment feature P and lidar structural features When spatial dimensions are inconsistent, the structural features of the lidar should be considered first. Interpolation to grid-level full-color alignment features P A consistent spatial size is used; these are then mapped to a unified common embedding space, and panchromatic-laser gated maps are generated through normalization and cosine similarity calculations. The expression for the calculation process is as follows:

[0061] in, O This represents the structural features of a lidar after alignment with the grid-level full-color alignment feature space. This represents the normalized representation of grid-level full-color alignment features in the common embedding space. This represents the normalized representation of the aligned lidar structural features in the common embedding space. A panchromatic laser similarity map; A full-color laser-gated image; This represents the bilinear interpolation operation; and These are the height and width of the grid-level full-color alignment feature, respectively. and These represent the projected convolution kernels for grid-level full-color alignment features and lidar structural features, respectively. For learnable scaling parameters; and They are respectively and In thec Characteristic responses on each channel.

[0062] Finally, the tree species classification head is used to output the predicted classification results of each sample, thereby realizing the tree species classification mapping of the study area.

[0063] After completing panchromatic-spectral gating and panchromatic-laser gating modeling, the three main features are fed into the feature fusion and tree classification head. First, the panchromatic semantic features are... LiDAR structural features Multi-temporal and multi-spectral spatiotemporal characteristics Align to a unified, blended grid; then, gate the panchromatic-spectral map. And full-color laser-gated chart As an auxiliary consistency channel, it is used in conjunction with full-color semantic features. LiDAR structural features Multi-temporal and multi-spectral spatiotemporal characteristics The components are spliced ​​together in the channel dimension to form a fusion feature. Input, i.e.:

[0064] Subsequently, fusion features The results are fused through convolutional mapping and channel attention, and finally output through a tree-type classification head. The expression is:

[0065] in, Indicates channel splicing; This represents a fusion operation consisting of convolutional mapping and channel attention; Represents a classification mapping; Y This is the final classification output.

[0066] Tree species sample labels are used to supervise the training of the TDAGNet model. The classification performance is obtained by comparing the tree species prediction classification result map generated by the TDAGNet model with the sample label map. Its classification performance is visually measured by the tree species prediction classification result map generated by the TDAGNet model and quantitatively measured by OA, Kappa, and F1-score values.

[0067] Verification Example: This embodiment evaluates the classification performance of the TDAGNet model by collecting data and combining it with the constructed tree species ground truth label map. The verification process is as follows: 1. Acquisition and preprocessing of multi-source, multi-modal data.

[0068] Tree Species True Value Label Map: The study area is located in the temperate continental forest-covered region of the ecological transition zone between the Greater Khingan Mountains and the Yanshan Mountains in Kalaqin Banner, Inner Mongolia Autonomous Region, with a forest coverage rate of 83%. Based on the national forest resource inventory data, samples of major forest stand types were labeled and validated, obtaining the spatial locations of over 4000 typical dominant forest tree species and forest stand sample areas (30m×30m). Then, within the study area, typical and representative areas were selected to create high-quality tree species true value label maps with a 10m resolution, such as... Figure 5 As shown in Table 1, the tree species ground truth label map uses different colors to represent the distribution areas of different tree species, and is divided into training and test sets. Specifically, pink (color value code: #E93BC3) represents non-forest land, grass green (color value code: #A9C837) represents aspen, bright green (color value code: #49EC6A) represents birch, brown (color value code: #CB7D67) represents Mongolian oak, purple (color value code: #806ACF) represents larch, and cyan (color value code: #43B9D9) represents Chinese pine. Non-forest land is not considered part of forests and is used for illustrative purposes only.

[0069] The datasets, with dimensions of 300×300 and 400×200 pixels, cover five common dominant tree species and forest stands in the study area, totaling 131,304 tree species pixels for classification. These methods increase spatial variability, simulate different lighting conditions, and enrich intra-class feature distributions, helping to improve the robustness and generalization ability of the TDAGNet model, thus demonstrating higher stability in real-world forest tree species classification tasks. The main tree species include white birch (Betula platyphylla), larch (Larix gmelinii), Chinese pine (Pinus tabulaeformis), aspen (Populus davidiana), and Mongolian oak (Quercus mongolica).

[0070] Table 1. Distribution of the number of truth labels for each tree species

[0071] Sentinel-2: Sentinel-2 is a high-resolution multispectral imaging satellite, equipped with a multispectral imager (MSI). It consists of two polar-orbiting satellites, Sentinel 2A and Sentinel 2B, operating in the same orbit with a 180° phase difference. It can revisit the Earth's equatorial region every five days. To achieve long-term time-series observations of ground features, this invention integrates and overlays multiple Sentinel-2 images from 2020 to 2023, with each image having less than 20% cloud cover. The data product level is L2A, which has undergone radiometric and orthorectification correction. Due to the different resolutions of different bands in Sentinel-2, a bicubic convolution resampling method was used to resample all bands to 10m. The experiment was conducted using 10 bands: B2, B3, B4, B5, B6, B7, B8, B8A, B11, and B12.

[0072] GEDI (Global Ecosystem Dynamics Investigation) is a full-waveform lidar mission aboard the International Space Station (ISS) that provides information on forest vertical structure and canopy characteristics. This invention selects structural indices from GEDI, including RH98, PAI, COVER, PAVD, PGAP_THETA_MERGED, and FHD_NORMAL_MERGED, to characterize key structural attributes such as canopy height, leaf area, and porosity. Considering the sparse distribution and discontinuous spatial cover of GEDI observations along the track, to obtain rasterized input consistent with Sentinel-2, this invention first performs quality control and outlier removal on the GEDI indices. Then, discrete footprint information is interpolated into a continuous raster, and the results are unified to a 10m spatial resolution to achieve spatial alignment with multi-temporal optical data and support subsequent fusion mapping experiments.

[0073] Jilin-1: The Jilin-1 satellite possesses high spatial resolution optical imaging capabilities, providing panchromatic (PAN) and other products to support the extraction of detailed ground feature information. This invention selects Jilin-1 panchromatic data as a supplement to high-resolution spatial details, filtering data from image dates with cloud cover below 20%. After radiometric calibration and orthorectification, to align with multi-source data on a spatial scale, this invention employs a resampling strategy to uniformly process the panchromatic images to a 1m resolution, and completes cropping and registration with the study area. Finally, it serves as the high-resolution spatial prior input for subsequent experiments.

[0074] 2. Implementation details.

[0075] This embodiment is implemented in the deep learning framework PyTorch 2.1.0 and Python 3.10 environment, relying on an NVIDIA GeForce RTX 4090 graphics card (24GB VRAM) and CUDA 11.8 acceleration units for model building and optimization, so as to make full use of hardware resources to improve training speed and efficiency. In model training, cross-entropy loss is used as the loss function, AdamW is used as the optimizer, and the initial learning rate is set to 4×10⁻⁶. -4 The weight decay coefficient is set to 5×10. -6 The learning rate scheduler employs a multi-step learning rate decay strategy, with a total of 60 training epochs and a batch size of 8. After data preparation and model framework construction are completed, the training script is run. By continuously adjusting hyperparameters, the TDAGNet model is supervised for forward propagation and backward optimization using the ground truth label graphs of tree species corresponding to the training set, thereby generating the tree species prediction classification result graph. That is: Based on the sample annotation information, a ground truth tree label map of the study area is created for model training supervision and classification performance evaluation. The TDAGNet model proposed in this invention, along with various comparative models, generates predicted tree species classification results for the study area, which are used for qualitative comparison with the ground truth tree label map. This compares the classification effectiveness of different methods; the closer the predicted tree species classification results are to the ground truth tree label map, the better the model's classification performance. To help pinpoint the classification effectiveness of different models, areas with significantly different classification results are locally selected using red boxes.

[0076] Furthermore, to comprehensively evaluate the classification performance of the TDAGNet model, the tree species prediction classification results generated by the TDAGNet model were compared pixel-by-pixel with the ground truth tree species labels in the corresponding test areas. The correspondence between the prediction results and the ground truth labels for each category was statistically analyzed, and the overall accuracy (OA), Kappa coefficient, and F1-score for each tree species were calculated to quantitatively evaluate the classification mapping results. The experiment used overall accuracy (OA), Kappa coefficient, and F1-score for each category as evaluation metrics. OA measures the overall classification accuracy, Kappa coefficient reflects the consistency difference between the classification results and random classification, and F1-score comprehensively considers precision and recall, more effectively characterizing the recognition performance of each category. Among these metrics, higher values ​​for OA, Kappa, and F1-score indicate better classification performance of the TDAGNet model.

[0077] 3. Introduction to the comparative model.

[0078] To verify the technical effectiveness of the TDAGNet model proposed in this invention, LF-DLM, ExViT, AMIANet, MAESTRO, and MultiSenseSeg were selected as comparative models for experimental comparison under the same dataset, the same training / testing partition, and the same evaluation metrics.

[0079] After training and testing the TDAGNet model, to further evaluate the classification performance and technical effectiveness of this invention, a comparative experiment was conducted. The same input data was fed into comparative models such as LF-DLM, ExViT, AMIANet, MAESTRO, and MultiSenseSeg, generating tree species prediction classification result maps output by each comparative model. The differences between these maps and the ground truth tree species labels were visualized, and the classification performance was quantitatively quantified using OA, Kappa, and F1-Score metrics. The specific description is as follows: LF-DLM: A deep multimodal late-stage fusion architecture for remote sensing semantic segmentation. It adopts a dual-branch independent modeling approach and performs fusion at the decision layer: one branch extracts high-resolution spatial texture using UNetFormer, and the other branch captures temporal dynamics using U-TAE. Finally, the probability outputs of the two branches are fused by weighted geometric mean to simultaneously utilize spatial details and temporal information to improve segmentation accuracy and robustness.

[0080] ExViT: An extended ViT framework for multimodal remote sensing classification. It processes heterogeneous modal supplementary information with parallel "location sharing" branches, introduces separable convolutions at the input to jointly model spatial and channel features, and promotes pixel-level correlation-driven information interaction through cross-modal attention. Finally, it fuses at the decision level to further improve classification and discrimination capabilities and overall performance.

[0081] AMIANet is an asymmetric multimodal interactive enhancement network for remote sensing semantic segmentation. It directly processes images and point clouds together, alleviates the imbalance in feature extraction between modalities by pre-extracting point cloud features, and introduces a collaborative multimodal interaction (SMI) module in the encoding stage to achieve bidirectional complementary enhancement through "image-driven point cloud fusion (IDPF) + point cloud-driven image fusion (PDIF)". At the same time, it designs multi-scale convolutional projection (MCP) to improve the integrity of the projection from point cloud to image space, thereby making fuller use of complementary structural and textural information and improving segmentation accuracy.

[0082] MAESTRO: A self-supervised framework for multimodal, multitemporal, and multispectral Earth observation data. By designing cross-modal fusion strategies and introducing normalized priors grouped by relevant bands in multispectral reconstruction, it learns more robust spatiotemporal-spectral representations for downstream mapping tasks.

[0083] MultiSenseSeg is a low-cost unified framework for remote sensing multimodal semantic segmentation. It introduces modality-specific lightweight experts (MSEs) to mitigate modality distribution differences and leverages adaptive multimodal matching (AMM) to achieve cross-modal coarse alignment and complementary information retrieval. Furthermore, it performs multi-scale semantic modeling on a shared backbone, thereby learning more consistent and complementary fusion representations for downstream remote sensing segmentation tasks with less additional computational overhead.

[0084] 4. Experimental results.

[0085] This embodiment systematically compares the classification results of the TDAGNet model and other comparative models on the test set. The classification performance of the models is visualized through tree species prediction classification result graphs generated by different models. Figure 6 - Figure 11 As shown, different indicators are used to quantitatively reflect classification performance. From Figure 6 As shown in Table 2, the TDAGNet model achieved the best overall classification performance, demonstrating strong class discrimination ability and cross-class stability. Specifically, the OA and Kappa of the TDAGNet model reached 90.90% and 87.33%, respectively, both outperforming all the comparison models; among them, compared with the second-best MultiSenseSeg, the improvements were 2.79% and 3.82%, respectively.

[0086] Table 2 Comparison of tree species classification performance of different models

[0087] Meanwhile, the TDAGNet model achieved the highest F1-Score across multiple tree species categories. Specifically, it improved by 2.86% compared to the second-best model, AMIANet, in the Mongolian oak category, and by 3.27% compared to the second-best model, MultiSenseSeg, in the Chinese pine category, demonstrating the proposed method's stronger discriminative advantage in distinguishing complex categories and identifying easily confused tree species. The performance improvement of the TDAGNet model is particularly significant for categories like Mongolian oak, larch, and Chinese pine, which rely more heavily on multimodal detail representation, indicating that the model can more effectively integrate complementary information between different modalities, thereby enhancing its ability to extract key discriminative features. Compared to some methods that excel in individual categories but lack overall category balance, the TDAGNet model demonstrates superior global generalization ability and overall classification robustness.

[0088] Furthermore, compared to the LF-DLM and MAESTRO models, the TDAGNet model shows a more significant improvement. This indicates that traditional or weak feature modeling methods are insufficient to fully characterize the spatial structure, spectral response, and class differences in multimodal remote sensing data, while the TDAGNet model, through a more effective feature interaction and discrimination mechanism, significantly enhances the model's ability to identify multiple tree species in complex forest scenes.

[0089] In summary, the TDAGNet model demonstrates superior overall accuracy, class balance, and F1 scores for most tree species, proving the effectiveness of the proposed method in multimodal remote sensing forest tree species classification tasks. In particular, its comprehensive lead in F1 scores for OA, Kappa, and the four key categories indicates that the model can more fully exploit complementary information between different modalities, improving the stability and reliability of tree species identification in complex scenarios, and possesses significant application potential.

[0090] Those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing related hardware. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0091] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. Since the above embodiments are substantially similar to the method embodiments, their descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0092] The above embodiments provide a detailed description of the present invention. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A method for classifying forest tree species using dual-alignment gating driven multimodal satellite imagery, characterized in that, The method includes the following steps: Acquire and preprocess multi-source, multimodal data, including multi-temporal, multispectral, panchromatic, and lidar images, of complex forest scenes within the study area; Multi-source, multi-modal data are input into the TDAGNet model to extract multi-temporal, multi-spectral spatiotemporal features, panchromatic space texture features, and lidar structural features, and a unified fusion representation is formed through heteroscale modal alignment and cross-modal gating collaborative mechanisms; The TDAGNet model includes a dual-branch panchromatic feature extractor for mining high-resolution spatial information in panchromatic images to obtain panchromatic spatial texture features, which include panchromatic detail features and panchromatic semantic features. Furthermore, by using pixel-level importance scoring and attention aggregation within grid cells, high-resolution pixel representations are adaptively mapped to grid-level full-color alignment features corresponding to the structural features of the lidar, thereby realizing a heteroscale modal alignment module that explicitly aligns from pixel space to structural grid space. In addition, a multimodal collaborative alignment and joint discrimination module is used to explicitly model the cross-modal consistency relationship between full-color space texture features, multi-temporal multispectral spatiotemporal features and lidar structural features on a unified fusion grid, so as to realize multimodal joint discrimination. Cross-modal alignment and joint discrimination are achieved on a unified and fused grid to obtain a tree species prediction classification result map including tree species category probabilities, so as to realize tree species classification mapping of the study area.

2. The multimodal satellite image forest tree species classification method according to claim 1, characterized in that, The dual-branch panchromatic feature extractor adopts a dual-path structure design that includes a high-frequency detail branch and a semantic trunk branch running in parallel. The high-frequency detail branch is used to extract local texture, crown boundary and fine-grained spatial variation information from panchromatic images to obtain panchromatic detail features, so as to enhance the TDAGNet model's ability to represent high-frequency structural details in complex forest scenes. The semantic backbone branch is used to extract high-level spatial semantic features from panchromatic images to obtain panchromatic semantic features, so as to make up for the limitation of the high-frequency detail branch, which focuses on local texture and edge information.

3. The multimodal satellite imagery forest tree species classification method according to claim 1, characterized in that, The TDAGNet model also includes a spatiotemporal feature extractor for extracting time-space-spectral features from multi-temporal multispectral images and for highlighting key phenological phases for tree species identification through temporal attention aggregation to obtain multi-temporal multispectral spatiotemporal features; And a lidar feature extractor for extracting stand vertical structure information from lidar images to obtain lidar structural features.

4. The multimodal satellite image forest tree species classification method according to claim 3, characterized in that, The process of extracting multi-temporal and multispectral spatiotemporal features by the spatiotemporal feature extractor includes: Let the input multi-temporal multispectral image be ,in B For batch size, T For the time phase number, C For single-phase bands, H mul and W mul These are the height and width of the space, respectively. pass N 3D convolutional blocks extract spatiotemporal spectral joint features with spatiotemporal spectral coupling characteristics for deep characterization. ; The response at each time phase is obtained by global averaging in the spatial dimension, and in the time dimension... Attention weights during generation time Polymerization characteristics were characterized by stable time-series spectra. To adaptively highlight key phenological phases that are more discriminative for tree species identification; Introducing SEGate for aggregated features Channel recalibration was performed to obtain multi-temporal and multispectral spatiotemporal characteristics. This is used to enhance the representation of representative time-series spectra; The expression for the spatiotemporal feature extractor is as follows: ; in, This represents an N-layer spatiotemporal convolutional block consisting of 3D convolutions, normalization, and activation functions; Spatiotemporal spectral joint features In the Time-dimensional index, spatial location The characteristic response corresponding to the location; For the first The spatiotemporal spectrum joint features corresponding to each time dimension index; This indicates element-wise weighting; This indicates a channel recalibration operation; Indexed by time dimension; For spatial location index in the height direction; This is the spatial location index in the width direction.

5. The multimodal satellite imagery forest tree species classification method according to claim 3, characterized in that, The lidar feature extractor uses two layers of depth-separable convolutional blocks for lightweight encoding, which enhances the ability to express structural patterns while controlling the number of parameters and computational load. The expression for the lidar feature extractor is: ; in, The input is a lidar image; and They represent the first N Depth convolution kernels and point convolution kernels for layers; This represents the convolution operation; For the first N Layer depth can separate the input features of convolutional blocks; Indicates group normalization; Represents a nonlinear activation function; For the first N Output features of layered coded blocks; Encode features for the initial structure; For the first N -1 layer structure encoding features; For the first N Layered structure encoding features; The final output of the lidar structure features.

6. The multimodal satellite image forest tree species classification method according to claim 2, characterized in that, The high-frequency detail branch adopts N Depthwise separable convolutions, passed through concatenated layers, are used for feature extraction, outputting panchromatic detail features that maintain spatial resolution. The expression is: ; in, Indicates the first N One deep convolutional kernel; This represents the corresponding point convolution kernel; This represents the convolution operation; Indicates group normalization; This represents the SiLU nonlinear activation function; For the first N The input features of a depthwise separable convolutional block on the high-frequency detail branch, when hour, For input panchromatic image ; The extraction of full-color semantic features by the semantic backbone branch includes: First, panchromatic images are processed using shallow convolutional blocks. Initial semantic features are obtained through preliminary encoding. ; Subsequently, the initial semantic features were processed using the anti-aliasing downsampling module. Smooth downsampling is performed to extract smooth semantic features. And for smooth semantic features Extracting mid-level semantic features through smooth downsampling To reduce aliasing distortion of high-frequency textures during scale compression; Next, the receptive field is gradually expanded using lightweight hollow residual blocks to obtain enhanced deep semantic features. Furthermore, enhanced contextual semantic features are obtained by improving the ability to characterize canopy spatial organization, regional structure, and global semantic patterns through multi-scale context modules. ; Finally, through channel compression and convolutional mapping, full-color semantic features are obtained for subsequent cross-modal fusion. ; The expression for the semantic backbone branch is: ; in, This represents the initial encoding operation consisting of two convolutional layers, normalization, and activation functions; Indicates anti-aliasing downsampling; This represents a semantic extraction process consisting of convolutions and lightweight dilated residual blocks. This represents a semantic enhancement operation for further stacked holed residual blocks; and These represent downsampling and upsampling, respectively. Indicates a lightweight void residual block; , and These represent depthwise separable convolution, pointwise convolution, and channel convolution operations, respectively.

7. The multimodal satellite imagery forest tree species classification method according to claim 1, characterized in that, The process by which the heteroscale modal alignment module obtains grid-level panchromatic alignment features includes: Suppose the full-color detail features output by the high-frequency detail branch. ,in, B For batch size, C h The number of feature channels, H P and W P These represent the height and width of the full-color detail features, respectively. Full color detail features Divided into grids according to the target resolution G H × G W There are 1 grid cells, each grid cell being 1. H × W ,satisfy H P = G H × H , W P = G W × W , H , W These are the height and width of the grid cell, respectively; A pixel-level importance score map is obtained through a pointwise convolutional mapping layer. S And within each grid cell, pixel-level importance scores are generated. S Perform normalization to generate normalized attention weights; full-color detail features of the corresponding area After performing weighted summation, the grid-level pancolor alignment feature is finally obtained. P Its core calculation process is expressed as follows: ; in, This represents the pointwise convolutional kernel used to generate a single-channel score map; Indicates the ( G H , G W Position within each grid cell ( U , V Normalized attention weights; Indicates the ( G H , G W Traversing positions within each grid cell Unnormalized pixel-level importance score map; U and V These represent the spatial position indices along the height and width directions within the current grid cell, respectively. and These represent the summation indices for traversing all positions within the grid cell in the height and width directions, respectively.

8. The multimodal satellite image forest tree species classification method according to claim 1, characterized in that, The multimodal collaborative alignment and joint discrimination module takes as input the panchromatic semantic features output from the semantic backbone branch, the multi-temporal and multispectral spatiotemporal features output from the spatiotemporal feature extractor, the grid-level panchromatic alignment features after grid alignment of the high-frequency detail branch, and the lidar structural features. It constructs panchromatic-spectral gating and panchromatic-laser gating respectively to constrain the intermodal relationships from two levels: spatial semantic-temporal spectrum and spatial detail-vertical structure. Finally, it achieves multimodal joint discrimination through feature fusion and tree species classification head.

9. The multimodal satellite imagery forest tree species classification method according to claim 8, characterized in that, The panchromatic-spectral gating mechanism is used to model the consistency relationship between panchromatic semantic features and multi-temporal multispectral spatiotemporal features, wherein: Full-color semantic features are and multi-temporal and multi-spectral spatiotemporal characteristics Each element is mapped to a unified common embedding space via 1×1 convolution and normalized along the channel dimension. Then, position-wise cosine similarity is calculated, and a panchromatic-spectral gated map is generated using a learnable scaling parameter and a sigmoid function. The expression for this calculation process is as follows: ; in, This represents the normalized representation of full-color semantic features in the public embedding space; Normalized representation of multi-temporal and multispectral spatiotemporal features in a common embedding space; This is a cosine similarity graph; A panchromatic-spectral gated graph; and These represent the projected convolution kernels for panchromatic semantic features and multi-temporal multispectral spatiotemporal features, respectively. D For public embedded dimensions; This represents L2 normalization along the channel dimension; This represents element-wise multiplication; This represents the Sigmoid activation function; Softplus learnable scaling parameters; They represent and In the c Characteristic responses on each channel; The panchromatic-laser gating mechanism is used to model the consistency relationship between grid-level panchromatic alignment features and lidar structural features, wherein: When grid-level full-color alignment feature P and lidar structural features When spatial dimensions are inconsistent, the structural features of the lidar should be considered first. Interpolation to grid-level full-color alignment features P Consistent space size; They are then mapped to a unified public embedding space, and panchromatic-laser gated maps are generated through normalization and cosine similarity calculations. The expression for the calculation process is as follows: ; in, O This represents the structural features of a lidar after alignment with the grid-level full-color alignment feature space; This represents the normalized representation of grid-level full-color alignment features in the common embedding space. This represents the normalized representation of the aligned lidar structural features in the common embedding space. A panchromatic laser similarity map; A full-color laser-gated image; This represents the bilinear interpolation operation; and These are the height and width of the grid-level full-color alignment feature, respectively. and These represent the projected convolution kernels for grid-level full-color alignment features and lidar structural features, respectively. For learnable scaling parameters; and They are respectively and In the c Characteristic responses on each channel.

10. The multimodal satellite image forest tree species classification method according to claim 8, characterized in that, Achieving cross-modal alignment and joint discrimination on a unified and integrated grid includes: First, the full-color semantic features LiDAR structural features Multi-temporal and multi-spectral spatiotemporal characteristics Align to a unified, blended grid; Subsequently, the panchromatic-spectral gating map And full-color laser-gated map As an auxiliary consistency channel, it is used in conjunction with full-color semantic features. LiDAR structural features Multi-temporal and multi-spectral spatiotemporal characteristics The components are spliced ​​together in the channel dimension to form a fusion feature. Input, i.e.: ; Subsequently, fusion features The results are fused through convolutional mapping and channel attention, and finally output through a tree-type classification head. The expression is as follows: ; in, Indicates channel splicing; This represents a fusion operation consisting of convolutional mapping and channel attention; Represents a classification mapping; Y This is the final classification output.