A Vegetation Classification Method Based on Spatiotemporal Multimodal Deep Learning

By integrating optical, radar, and terrain data using a spatiotemporal multimodal deep learning approach, the problem of unstable vegetation classification in existing technologies has been solved. This approach enables efficient expression and fine classification of vegetation characteristics in complex areas, improving the reliability and adaptability of classification results.

CN121600411BActive Publication Date: 2026-04-21XIAN UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
XIAN UNIV OF POSTS & TELECOMM
Filing Date
2026-01-28
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing vegetation classification methods are unstable in mountainous areas, areas with frequent cloud cover, or areas with significant topographic relief. They also have limited ways to fuse multimodal remote sensing data, lack a unified spatiotemporal modeling mechanism, make it difficult to fully utilize the complementary information from multi-source remote sensing data, and lack adaptability to multidimensional features.

Method used

A spatiotemporal multimodal deep learning-based approach is adopted. By acquiring multi-temporal optical images, radar images, and digital elevation model data, spectral, microwave, topographic, and texture features are extracted using a derived feature extraction network and a spatiotemporal multimodal encoder. Cross-modal self-attention fusion is performed, and feature fusion and classification are carried out using a gated spatiotemporal graph convolution module and a dual-path classification decoder. Finally, a refined vegetation classification result is generated.

Benefits of technology

It achieves deep fusion of multimodal remote sensing data, enhances the model's ability to express vegetation characteristics in complex areas, improves the reliability and accuracy of classification results, enhances adaptability, and maintains robustness under complex terrain and lighting changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121600411B_ABST
    Figure CN121600411B_ABST
Patent Text Reader

Abstract

This application discloses a vegetation classification method based on spatiotemporal multimodal deep learning, relating to the field of image processing technology. The method includes: acquiring multi-temporal optical images, radar images, and digital elevation model data of the region to be classified, forming multimodal data; extracting spectral features, microwave features, topographic features, and texture features to form fused features; calculating graph node features and updating the graph node features to form graph features; fusing the fused features and graph features to obtain pixel-level coarse classification logits; extracting and fusing the fused features to form region-level coarse classification logits; fusing to obtain coarse classification probabilities; and performing fine classification on the coarse class basic features to obtain the final classification result. The method of this application can adapt to mountainous environments with frequent cloud cover, shadows, and large topographic relief, improving the stability and classification precision of vegetation type identification, and is suitable for large-scale vegetation monitoring and ecological assessment scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to a method for vegetation classification using optical, radar and topographic data. Background Technology

[0002] Vegetation classification is a crucial foundational task in fields such as ecological monitoring, natural resource surveys, land use assessment, and environmental change analysis. With the development of remote sensing technology, multi-source data, including multispectral imagery, synthetic aperture radar imagery, and digital elevation models, can reflect the spectral characteristics, structural features, and topographic information of surface vegetation from different dimensions, providing important data support for the refined classification of regional vegetation.

[0003] Existing vegetation classification methods primarily rely on single optical images or spectral indices for feature extraction and classification. However, in mountainous areas, regions frequently obscured by clouds, or areas with significant topographic relief, optical images are easily affected by cloud shadows, illumination angles, and topographic effects, leading to unstable classification results. To compensate for the shortcomings of optical data, some studies have incorporated radar data. However, optical, radar, and topographic data differ significantly in resolution, noise structure, and observation mechanisms, making direct fusion prone to information redundancy or feature conflicts.

[0004] Furthermore, vegetation exhibits distinct seasonal and temporal variations, making it difficult for single-temporal data to fully reflect its temporal patterns. While existing technologies attempt to classify vegetation based on multi-temporal imagery, a unified spatiotemporal modeling mechanism is lacking, failing to effectively characterize the relationships between multimodal and multi-temporal data. Simultaneously, current deep learning models often focus on single-modal or local spatial features, lacking a unified framework capable of simultaneously processing multi-dimensional features such as optical, radar, topographic, and textural data.

[0005] In summary, existing technologies still have shortcomings in the following aspects:

[0006] (1) There are limited ways to fuse multimodal remote sensing data, making it difficult to simultaneously utilize spectral, scattering, and topographic features;

[0007] (2) There is a lack of spatiotemporal modeling mechanisms that take into account both spatial structural characteristics and temporal evolution features;

[0008] (3) The vegetation classification model is not adaptable to multi-source features, and the classification precision and stability are still limited.

[0009] Therefore, it is necessary to propose a new multimodal spatiotemporal remote sensing data-driven vegetation classification method to fully utilize the complementary information of multi-source remote sensing data, enhance the model's ability to express vegetation characteristics in complex areas, and improve the reliability of classification results. Summary of the Invention

[0010] This application provides a vegetation classification method based on spatiotemporal multimodal deep learning to solve the above-mentioned problems in the prior art.

[0011] This application provides a vegetation classification method based on spatiotemporal multimodal deep learning, including:

[0012] Acquire multi-temporal optical images, radar images, and digital elevation model data of the area to be classified to form multimodal data;

[0013] Multimodal data is input into a derived feature extraction network, which extracts spectral data, microwave data, terrain data, and texture data from the multimodal data.

[0014] A spatiotemporal multimodal encoder is used to extract spectral features, microwave features, topographic features, and texture features from spectral data, microwave data, topographic data, and texture data, respectively. The spectral features, microwave features, topographic features, and texture features are then fused together through cross-modal self-attention to form fused features.

[0015] The fused features are divided into multiple superpixel blocks using a gated spatiotemporal graph convolution module. The graph node features of the superpixel blocks are calculated, and the graph node features are updated using graph convolution operations. The updated graph node features are then reshaped into spatial features, and the spatial features are upsampled to form graph features.

[0016] A dual-path classifier decoder is used to fuse fusion features and graph features at the pixel level to obtain pixel-level coarse classification logits. At the same time, multi-scale feature extraction and cross-scale fusion are performed on the fusion features at the region level to form region-level coarse classification logits. The pixel-level coarse classification logits and the region-level coarse classification logits are fused to obtain the fused coarse classification probability. A fine classifier is then used to perform fine classification on the coarse basic features extracted from the fusion features and graph features according to the coarse classification probability to obtain the final classification result.

[0017] The spatiotemporal multimodal segmentation model consists of a derived feature extraction network, a spatiotemporal multimodal encoder, a gated spatiotemporal graph convolutional module, and a dual-path classification decoder. During the training of the spatiotemporal multimodal segmentation model, sample data is acquired and combined into spatiotemporal cube samples, which are then input into the spatiotemporal multimodal segmentation model.

[0018] The vegetation classification method based on spatiotemporal multimodal deep learning in this application has the following advantages:

[0019] 1. Achieving deep fusion of multimodal remote sensing data: This application achieves comprehensive characterization of the spectral-scattering-topographic integrated features of vegetation by jointly modeling multiple remote sensing features such as optical, radar, topography and texture, thereby improving the feature expression capability of the model.

[0020] 2. Construct a unified spatiotemporal feature representation method: By using a spatiotemporal cube structure, features of different times and different modalities are organized in a unified framework, which helps to enhance the model's ability to understand the laws of temporal change and the correlation between multiple modalities.

[0021] 3. Enhance the model's adaptability to complex regions: This application adopts cross-modal fusion and spatiotemporal graph convolution structure to achieve deep modeling of spatial neighborhood features and time series features, thereby improving the model's robustness to complex terrain, lighting changes and noisy environments.

[0022] 4. Improve the consistency and precision of classification results: Based on the multi-path decoding structure and joint loss mechanism, the ability to recover local details can be enhanced while maintaining semantic consistency, making the classification results more continuous and reasonable.

[0023] In summary, the multimodal spatiotemporal remote sensing data-driven vegetation classification method proposed in this application can be widely used in fields such as ecological monitoring, natural resource surveys, vegetation mapping and environmental analysis, and has good application prospects. Attached Figure Description

[0024] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0025] Figure 1 A flowchart illustrating a vegetation classification method based on spatiotemporal multimodal deep learning, provided for an embodiment of this application.

[0026] Figure 2 This is an architecture diagram of a spatiotemporal multimodal encoder provided in an embodiment of this application.

[0027] Figure 3 This is an architecture diagram of the gated spatiotemporal graph convolution module provided in an embodiment of this application.

[0028] Figure 4 This is an architecture diagram of the dual-path classification decoder provided in an embodiment of this application.

[0029] Figure 5 A comparison chart of the overall accuracy, Kappa coefficient, and average crossover ratio (CROR) of the model during training, provided in the embodiments of this application.

[0030] Figure 6 Five sets of vegetation classification results are visualized and compared in the embodiments of this application. Detailed Implementation

[0031] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0032] Figure 1 This application provides a flowchart of a vegetation classification method based on spatiotemporal multimodal deep learning, which includes the following steps:

[0033] S100 acquires multi-temporal optical images, radar images, and digital elevation model data of the area to be classified, forming multimodal data.

[0034] For example, the data obtained from the region to be classified is called multimodal data, and the data obtained from the sample region is called sample data. In addition to multi-temporal optical imagery, radar imagery, and digital elevation model data of the sample region, the sample data also includes ground truth data.

[0035] Specifically, the ground truth data collection process is as follows: combining field survey data with multiple publicly available high-precision land use products such as GlobeLand30, FROM-GLC (Global Land Cover Dataset), and GLC_FCS30 (Global 30-meter resolution fine-classification land cover product), accurate vegetation classification labels are generated through cross-validation, and the vegetation classification labels combined with the corresponding vegetation regions constitute the ground truth data.

[0036] The optical image acquisition process is as follows: Collect multispectral data from the Sentinel-2 satellite released by the European Space Agency's Copernicus Data Center. Four representative time phases are collected annually: spring (March-May), summer (June-August), autumn (September-November), and winter (December-February) to capture seasonal changes in vegetation. L2A-level atmospheric correction products are selected, with cloud cover below 5%, and four core bands are extracted: B2 (490nm blue light), B3 (560nm green light), B4 (665nm red light), and B8 (842nm near-infrared).

[0037] The radar image acquisition process is as follows: Sentinel-1 SAR (Satellite Synthetic Aperture Radar) data released by the European Space Agency's Copernicus Data Center is collected, maintaining the same time series as the optical imagery. IW (Interferometric Wide Swath) mode L1A-level products are collected, including VH (Vertical Transmit-Horizontal Receive) and VV (Vertical Transmit-Vertical Receive) dual-polarization data, to obtain radar backscattering information of the target area.

[0038] The data acquisition process for digital elevation model (DEM) is as follows: 30-meter resolution DEM data is used, which is derived from the SRTM (Space Shuttle Radar Topography Mission) global DEM product to provide terrain feature information and terrain correction reference.

[0039] Furthermore, after acquiring multimodal data and sample data, both multimodal data and sample data are preprocessed.

[0040] Specifically, the preprocessing of multimodal data for the region to be classified includes: mosaicking and cropping optical images and unifying coordinates / resolution; SAR preprocessing and postprocessing radar images; and mosaicking and cropping digital elevation model data and unifying coordinates / resolution.

[0041] The sample data includes ground truth data, optical imagery, radar imagery, and digital elevation model data of the sample area. The preprocessing of the ground truth data includes format conversion, coordinate unification, and mode voting fusion.

[0042] The preprocessing workflow for ground truth data is as follows: land use products are standardized to the WGS1984 (1984 World Geodetic System) coordinate system and resampled to a spatial resolution of 10 meters. Using a mode voting fusion algorithm, multiple label products are cross-validated through a 3×3 sliding window. Following the principle of "based on high-confidence field survey data, and determined by voting with other products," the final eight vegetation classification labels are generated.

[0043] The preprocessing workflow for optical images is as follows: Sentinel-2 data is mosaicked and cropped to extract regional data. All bands are resampled to 10-meter resolution and converted to the WGS1984 coordinate system to ensure spatial alignment with the label data. During processing, four bands, B2, B3, B4, and B8, are retained to form a multi-temporal spectral dataset.

[0044] The radar image preprocessing workflow is as follows: Establish a Sentinel-1 SAR data preprocessing workflow and configure parameters according to the following process:

[0045] (1) Data reading: Load L1A level raw SAR data;

[0046] (2) Application of track files: Use Sentinel precise track files;

[0047] (3) Thermal noise removal: handling VH and VV polarization;

[0048] (4) Radiometric calibration: output gamma0 band; speckle noise filtering: use Refined Lee filter;

[0049] (5) Terrain correction: Use external DEM data with a pixel spacing of 10 meters and map projection WGS84; Linear value to decibel: Convert Gamma0_VH and Gamma0_VV to dB values;

[0050] (6) Output results: Save as GeoTIFF (geotagged image file format). The preprocessed VH and VV band data are also resampled to 10-meter resolution, then mosaicked and cropped to obtain microwave data within the range.

[0051] The preprocessing workflow for digital elevation model (DEM) data is as follows: DEM data is mosaicked and cropped to extract regional topographic data. It is then resampled to a 10-meter resolution and converted to the WGS1984 coordinate system.

[0052] All data preprocessing steps ensured pixel-level spatial alignment across different data sources, namely: a unified coordinate system of WGS1984, a unified spatial resolution of 10 meters, and a unified spatial extent of the same bounding rectangle coordinate range. This stringent alignment standard laid a solid foundation for subsequent multimodal feature fusion.

[0053] S110 inputs multimodal data into a derived feature extraction network, which extracts spectral data, microwave data, terrain data, and texture data from the multimodal data.

[0054] For example, the optical image contains four temporal phases. The derived feature extraction network adds six indices to the optical image: normalized difference vegetation index, enhanced vegetation index, normalized difference water index, modified soil-regulated vegetation index, green leaf index, and near-infrared to red band ratio. It also calculates texture in the B8 band. After standardizing the 13 channels consisting of the four temporal phases, the six indices, and the texture, the spectral data and texture data are obtained.

[0055] Specifically, in the process of extracting spectral features, the input spectral data in the original format (TIME_STEPS=4, 4, H, W) contains 4 time phases and four bands B2 / B3 / B4 / B8, with values ​​ranging from 0 to 10000.

[0056] The spectral data is then preprocessed again as follows:

[0057] (1) Reflectance restoration: Divide by the scaling factor of 10000 to convert to the actual reflectance value;

[0058] (2) Handling invalid values: Replace 0 values ​​with the mean effective value of the corresponding band at that time to improve data integrity;

[0059] (3) Calculation of vegetation indices: Six new indices with clear ecological and physical significance have been added: NDVI (Normalized Difference Vegetation Index), which reflects vegetation cover; EVI (Enhanced Vegetation Index), which is more sensitive to areas with high vegetation cover; NDWI (Normalized Difference Water Index), which identifies water body information; MSAVI (Modified Soil-Regulated Vegetation Index), which reduces the influence of soil background; GLI (Green Leaf Index), which highlights the characteristics of green vegetation; and NIR / R (Near Infrared to Red Band Ratio), which reflects the vegetation growth status.

[0060] (4) Texture feature extraction: Three texture features are calculated from the B8 band using the gray-level co-occurrence matrix (GLCM): contrast, entropy, and correlation, which reflect the local spatial structure information of the image;

[0061] (5) Standardization and extreme value clipping: The 13 channels (4 original + 6 exponential + 3 texture) are standardized, and the mean / standard deviation is calculated using the training set statistics. The values ​​are clipped to the range of [-3, 3] to eliminate the influence of outliers.

[0062] The output spectral and texture features are in the format (TIME_STEPS=4, 13, H, W), with the number of channels increasing from 4 to 13, and include both spectral and texture features.

[0063] In the process of microwave feature extraction, the input radar data in the original format (TIME_STEPS=4, 2, H, W) contains the dB values ​​of four time phases and two polarization bands VH(0) / VV(1).

[0064] The radar data is then preprocessed again as follows:

[0065] (1) Invalid value handling: Replace the 0.000 value with the mean of the effective value corresponding to the polarization at that time;

[0066] (2) Feature engineering: Four new radar vegetation indices have been added: RVI (Radar Vegetation Index), which reflects the scattering characteristics of vegetation; PR (Polarization Ratio), which identifies land cover types; NRVI (Normalized Radar Vegetation Index), which enhances the distinction between vegetation and non-vegetation; and DPVI (Dual Polarization Composite Index), which optimizes vegetation monitoring capabilities.

[0067] (3) Standardization: The polarization bands (VV / VH) are mapped to the range of [-1,1]. The newly added exponents are standardized using the statistical information of the training set and reordered as VV(0) and VH(1).

[0068] (4) Extreme value clipping: Clip all channel values ​​to the range of [-3, 3].

[0069] The output microwave features are in the format (TIME_STEPS=4, 6, H, W), with the number of channels increasing from 2 to 6, and include polarization information and radar vegetation index.

[0070] During the extraction of terrain features, input digital elevation model data in the original format (H, W), in centimeters. Invalid values ​​are -2147483647.

[0071] The digital elevation model data is then preprocessed again as follows:

[0072] (1) Invalid value handling: Replace -2147483647 with the mean of the valid values ​​within the 3×3 window, and use mirror expansion to process the edge area;

[0073] (2) Unit conversion: from centimeters to meters, which conforms to the actual terrain analysis habits;

[0074] (3) DEM standardization: normalize to the range of [0, 1] using the minimum / maximum values ​​of the training set statistics;

[0075] (4) Calculation of terrain-derived features: The slope is calculated using the Sobel operator (normalized to [0,1], corresponding to 0-90°). After calculating the aspect angle, the periodicity is eliminated by cosine transformation to obtain the range of values ​​in [-1, 1].

[0076] The output terrain features are in the format (3, H, W), expanding from a single-channel DEM to a 3-channel terrain feature that includes DEM, slope, and aspect.

[0077] As can be seen from the above process:

[0078] The data dimensions have increased: spectral features have increased from 4 channels to 13 channels; microwave features have increased from 2 channels to 6 channels; and topographic features have increased from 1 channel to 3 channels. The total input feature dimensions have increased significantly, providing richer information on ground features.

[0079] Data quality improvement: invalid values ​​are effectively filled, outliers are pruned, and multimodal data is standardized to ensure the stability of model training.

[0080] Enhanced feature richness: Expanded from raw sensor data to include physically meaningful indices, adding spatial texture and terrain-derived features, with multimodal data (optical, microwave, terrain) collaboratively providing complementary information.

[0081] S120 uses a spatiotemporal multimodal encoder (ST-MME) to extract spectral features, microwave features, topographic features, and texture features from spectral data, microwave data, terrain data, and texture data, respectively. It then performs cross-modal self-attention fusion of the spectral features, microwave features, topographic features, and texture features to form fused features.

[0082] For example, a spatiotemporal multimodal encoder contains four parallel branches, such as Figure 2 As shown, the four branches extract spectral features, microwave features, terrain features, and texture features respectively. The branches for extracting spectral features and microwave features contain dynamically deformable convolution, spatial reconstruction units, and spatiotemporal reconstruction units connected in sequence. The branch for extracting terrain features contains dynamically deformable convolution and spatial reconstruction units. The branch for extracting texture features contains spatial reconstruction units and spatiotemporal reconstruction units.

[0083] Specifically, lightweight, dynamically deformable convolution uses dynamic offsets to adjust the kernel shape, adaptively adjusting the receptive field to adapt to complex terrain features. The specific process is as follows:

[0084] 1. Lightweight design. A channel-based dimensionality reduction strategy is adopted. First, the number of input feature channels is reduced by 1×1 convolution. The reduction ratio can be adjusted by the reduction parameter to reduce the number of parameters and computational complexity of offset calculation. At the same time, a 1×1 convolution branch is introduced to preserve fine-grained features and avoid information loss caused by dimensionality reduction.

[0085] 2. Dynamic offset adjustment of convolution kernel shape. An offset learning network consisting of two convolutional layers (1×1 dimensionality reduction convolution + 3×3 convolution) generates two offsets (horizontal and vertical) for each convolution kernel position, realizing the dynamic adjustment of the convolution kernel sampling position, thereby changing the shape of the effective convolution kernel.

[0086] 3. Adaptive adjustment of receptive field. The network automatically learns offsets based on the input terrain features, enabling the convolutional kernels to adaptively adjust the size and shape of the receptive field according to the complexity of local features—using a smaller effective receptive field for flat areas to preserve details, and automatically expanding the effective receptive field for complex terrain areas to capture more contextual information, thereby better adapting to complex terrain features.

[0087] The Spatial Reconstruction Unit (SC-RU) comprises the Adaptive SRU (Adaptive Spatial Reconstruction Unit) and the Adaptive CRU (Adaptive Channel Reconstruction Unit). It optimizes feature quality through grouping normalization and dynamic thresholding, thereby improving spatial structure and channel dependencies. The specific process is as follows:

[0088] 1. Grouping Normalization Optimizes Feature Quality. In AdaptiveSRU, grouping normalization (the number of groups can be adjusted via the group_norm_num parameter) is used to normalize the input features, effectively reducing redundancy between features and improving feature expressiveness. The spatial importance score of the features after grouping normalization is calculated to evaluate the amount of feature information at different spatial locations.

[0089] 2. Dynamic Thresholding Enables Adaptive Feature Filtering. A threshold estimation network consisting of adaptive average pooling and two 1×1 convolutions is designed to dynamically generate an adaptive threshold between 0 and 1 based on the global statistical information of the input features. This threshold is then used to binarize the features, generating information-rich region masks and redundant region masks, thereby enabling the differentiation and filtering of features of different importance.

[0090] 3. Spatial structure optimization. Based on the generated mask, the original features of information-rich regions are preserved, and redundant regions are filled and reconstructed using the average of the grouping of information-rich regions, effectively optimizing the spatial feature structure and enhancing the spatial consistency and information density of features.

[0091] 4. Channel Dependency Optimization. In AdaptiveCRU, the feature channels are divided into two groups and processed separately by dynamically calculating the channel splitting ratio (based on feature global statistical adaptive learning); cross-enhancement between channels is achieved by using a broadcast mechanism and 1×1 convolution, thereby optimizing the channel dependency relationship and improving the channel expressive power of the features.

[0092] The Spatiotemporal Reconstruction Unit (ST-RU) combines spatial attention and temporal gating (GRU) mechanisms to enhance key spectral channels and model temporal dependencies between multi-temporal data. The specific process is as follows:

[0093] 1. Implementation of Spatial Attention Mechanism. A lightweight SE-Net (Squeeze-Activated Network) structure is adopted to process the spatiotemporal features of the input. First, channel statistics are extracted at each time step through adaptive average pooling. Then, the dependencies between channels are learned through two 1×1 convolutional layers (with a ReLU activation function introduced in between). Finally, the sigmoid activation function is used to generate channel weights between 0 and 1, thereby achieving adaptive enhancement of key spectral channels.

[0094] 2. Temporal Gating (GRU) Mechanism Implementation. A gated recurrent unit (GRU) is used to capture temporal dependencies between multi-temporal data. First, the original features and SE-weighted features are concatenated as GRU input. Then, the spatial dimension is merged into the batch dimension. The GRU network processes the temporal sequence of each spatial location independently, learning the dynamic dependencies between temporal phases.

[0095] 3. Lightweight optimization design. 2D convolution is used instead of 3D convolution to reduce computational overhead. At the same time, the number of feature channels of the GRU output is adjusted back to the number of channels of the original input through residual connection adjustment layer (1×1 convolution). Finally, spatiotemporal refined features are output through residual summation and ReLU (corrected linear unit) activation function, which effectively balances model performance and computational efficiency.

[0096] In the spectral branch, the processing proceeds sequentially through Dynamically Deformable Convolution → SC-RU → ST-RU; the processing flow in the microwave branch is the same as that in the spectral branch; the terrain branch focuses on the spatial reconstruction of terrain features, using Dynamically Deformable Convolution → SC-RU; and the texture branch employs a unique processing order, first performing temporal feature extraction (ST-RU) and then spatial feature extraction (SC-RU).

[0097] Furthermore, the features output from the four branches are concatenated and then the 4×128 channels are reduced to 256 channels through a cross-modal projection layer. The result is then fed into the CrossModalTransformer module, which uses a sliding window attention mechanism to perform cross-modal self-attention fusion.

[0098] Specifically, the output features of the four branches (128 channels each) are first concatenated, and then the dimensionality of the 4×128 channels is reduced to 256 channels through a cross-modal projection layer. Next, the input is fed into the CrossModalTransformer module. This module implements the sliding window attention mechanism as follows: first, query, key, and value are generated through convolution; then, grouped convolutions are used to simulate local window attention, and attention weights are generated through sigmoid activation; finally, the weights are applied to the value features to achieve cross-modal self-attention fusion. Simultaneously, the location encoding of DEM information is combined to enhance the feature representation. The DEM data is first converted to 256-channel features through 1×1 convolution, and then added to the learnable fixed location encoding and input features to achieve enhanced fusion of terrain and location information. Finally, the processed features are upsampled to 128×128 resolution and fused with the input features through global residual connections, ultimately outputting a 256-channel fused feature.

[0099] S130 uses a gated spatiotemporal graph convolution module (ST-GGCM) to divide the fused features into multiple superpixel blocks, calculates the graph node features of the superpixel blocks, updates the graph node features using graph convolution operations, reshapes the updated graph node features into spatial features, and upsamples the spatial features to form graph features.

[0100] For example, such as Figure 3 As shown, the gated spatiotemporal graph convolution module contains two gated graph convolution layers, namely gated graph convolution layer one and gated graph convolution layer two. The gated spatiotemporal graph convolution module integrates the spatial proximity, temporal similarity and terrain consistency of graph node features to generate a dynamic adjacency matrix. Gated graph convolution layer one performs graph convolution operation on graph node features according to the dynamic adjacency matrix to form initial graph features. Gated graph convolution layer two updates the initial graph features according to the dynamic adjacency matrix to form updated graph features, and reshapes the updated graph features into spatial features.

[0101] Specifically, the workflow of the gated spatiotemporal graph convolution module is as follows:

[0102] (1) Graph node construction: The fused features output by ST-MME are divided into multiple superpixel blocks through superpixel segmentation. The mean value of the features in each superpixel block is calculated as the graph node feature, reducing the dimensionality of the 512×512 features to about 50 superpixel nodes, which significantly reduces the amount of computation.

[0103] (2) Dynamic Adjacency Matrix Generation: Adaptive generation is achieved by calculating the adjacency matrices of three modalities and performing weighted fusion. When calculating the spatial adjacency matrix, spectral similarity is first obtained by matrix multiplication of node features, and then spatial proximity is calculated based on Gaussian decay of node coordinate differences. The two are multiplied to obtain the spatial adjacency matrix. The temporal adjacency matrix is ​​based on node feature similarity and only retains the connection between adjacent temporal phases. The terrain adjacency matrix is ​​calculated based on Gaussian decay of node elevation differences. The three adjacency matrices are weighted and fused using learnable weight parameters (normalized by softmax and range restriction). After fusion, the adjacency matrix is ​​subjected to Top-K sparsification (retaining the connection with the largest weight), symmetry processing, addition of self-loops, and row normalization. The resulting dynamic adjacency matrix can adaptively adjust the adjacency weights according to the input features to adapt to different sample characteristics.

[0104] (3) Gated graph convolution operation: Two gated graph convolution layers, GatedGraphConv, are used to perform graph convolution operation. The gating mechanism controls the flow of information between nodes through GRU. Each layer first performs symmetric normalization on the adjacency matrix and selects the neighbor features with the largest weight using Top-K aggregation. The first layer inputs the aggregated features into GRU gating after linear transformation (mapping to 256 channels), layer normalization, and LeakyReLU activation, and outputs 256-channel graph features. The second layer takes the output of the first layer as input, performs neighbor aggregation and GRU gating processing in the same way, and adds the output features to the original node features through residual connection to realize further updates of node features and enhance the global context modeling capability.

[0105] (4) Feature reshaping and upsampling: The updated graph node features and updated graph features are reshaped into spatial features, and the resolution is gradually restored from 16×16 to 256×256 through progressive upsampling to preserve spatial details.

[0106] S140: The dual-path classifier decoder is used to fuse the fused features and graph features in the pixel-level path to obtain pixel-level coarse classification logits. At the same time, the fused features are also subjected to multi-scale feature extraction and cross-scale fusion in the region-level path to form region-level coarse classification logits. The pixel-level coarse classification logits and the region-level coarse classification logits are fused to obtain the fused coarse classification probability. The fine classifier is then used to perform fine classification on the coarse basic features extracted from the fused features and graph features according to the coarse classification probability to obtain the final classification result.

[0107] For example, such as Figure 4 As shown, the dual-path classification decoder in this embodiment adopts a semi-cascaded classification strategy, which includes two parallel paths: pixel-level and region-level. Finally, the classification result is generated through gating fusion.

[0108] In the pixel-level path, the input consists of the fused features output by ST-MME and the graph features output by ST-GGCM. The computational cost of these features is reduced by using inverse residual blocks. A lightweight spatial attention mechanism is used to refine the pixel-level features, and the output is pixel-level coarse classification logits (6 categories). The computational cost reduction by inverse residual blocks includes expanding channels, extracting features using depthwise separable convolutions, and projecting back to low dimensions. The lightweight spatial attention mechanism for refining pixel-level features consists of a spatial attention module and a channel attention module. Spatial attention generates weights through 3x3 convolutions, and channel attention generates weights through adaptive average pooling and 1x1 convolutions.

[0109] In the regional path, the input is the fused feature output by ST-MME. The feature is extracted using three pooling layers of different scales (2×2, 4×4, 8×8, 16×16). These features are then upsampled to a uniform scale (8×8) and concatenated. Cross-scale fusion is performed through 1x1 convolution, and global feature processing is performed through a fully connected layer (including linear transformation and Dropout). The output is a regional coarse classification logits (6 categories).

[0110] Gated fusion uses two mechanisms: global gating and spatial adaptive gating. Global gating generates weights by concatenating pixel-level global features and region-level features and then passing them through a fully connected layer. Spatial adaptive gating generates pixel-level weights through 3x3 convolution. The two weights are combined to perform weighted fusion of pixel-level coarse classification logits and region-level coarse classification logits to generate the fused coarse classification probability.

[0111] Semi-cascaded classification is based on the fused coarse classification probabilities (6 classes). It uses 8 independent coarse-class specific fine classifiers to generate the final fine classification logits (8 classes). Each classifier is a 1x1 convolution and is created only when a corresponding fine class exists. Each fine classifier is dedicated to the subdivision task within its corresponding coarse class. Finally, the semantic segmentation probability map of 8 vegetation classes is output.

[0112] In the embodiments of this application, the dual-path classification decoder simultaneously outputs edge detection results and superpixel segmentation results. These auxiliary tasks improve the boundary recognition accuracy and internal consistency of the main task through multi-task learning. The edge detection result is a single-channel edge probability map obtained by processing pixel features through two layers of convolution and finally using Sigmoid. During training, the Sobel operator is used to generate the edge map from the labels. The superpixel segmentation result is obtained by processing pixel features through multiple layers of convolution and attention mechanism, and the output 16-channel embedding is used for the superpixel segmentation result.

[0113] As can be seen, the model architecture adopted in this application has the following advantages:

[0114] Multimodal fusion: Fully utilizes complementary information from optical, microwave, and terrain data to improve segmentation accuracy in complex terrains;

[0115] Adaptive mechanism: SC-RU and ST-RU use adaptive parameters that are dynamically adjusted according to the input, enhancing the model's flexibility;

[0116] Spatiotemporal joint modeling: It considers both spatial and temporal dimensions, making it suitable for processing multi-temporal remote sensing data;

[0117] Introduction of Graph Neural Networks: Modeling long-range spatial dependencies through ST-GGCM to capture global context;

[0118] Lightweight design: Employs techniques such as depthwise separable convolution and inverse residual blocks to balance performance and efficiency;

[0119] Semi-cascaded classification: Generates fine classifications based on coarse classification results, improving segmentation accuracy and efficiency.

[0120] The S150, consisting of a derived feature extraction network, a spatiotemporal multimodal encoder, a gated spatiotemporal graph convolutional module, and a dual-path classification decoder, forms a spatiotemporal multimodal segmentation model. During the training of the spatiotemporal multimodal segmentation model, sample data is acquired and combined into spatiotemporal cube samples, which are then input into the spatiotemporal multimodal segmentation model.

[0121] For example, the preprocessed sample data is combined into spatiotemporal cube samples and converted to NPZ (NumPy compressed archive format) format for direct and efficient reading by deep learning models. The specific format of each sample is as follows:

[0122] Spectral data: The shape is (4, 4, 256, 256), representing 4 time phases (spring, summer, autumn, winter) × 4 bands (B2 / B3 / B4 / B8) × 256 × 256 pixels;

[0123] Radar data: Shape (4, 4, 256, 256), representing 4 time phases × 2 polarizations (VH+VV) × 256 × 256 pixels;

[0124] Digital Elevation Model (DEM) data: Shape (256, 256) represents DEM elevation data;

[0125] Labels: Shape (256, 256) represent ground truth labels (8 categories, pixel values ​​1-8, 0 is an invalid value).

[0126] The dataset, consisting of the above data, is divided into training, validation, and test sets to ensure a balanced distribution of samples across all categories. Data augmentation is applied dynamically during the training phase without pre-expanding the dataset to avoid storage overhead.

[0127] Specifically, before training, the following data augmentation processing is performed:

[0128] Space enhancement: Horizontal flip (probability 0.5), vertical flip (probability 0.3), random cropping (fill from 480×480 to 512×512, probability 0.8), finite angle random rotation (-15° to +15°, probability 0.4).

[0129] Spectral enhancement: Gaussian noise (probability 0.4, standard deviation 0.01-0.03) and brightness shift (visible channel only, probability 0.3, shift range ±0.05) are applied only to the data of the spectral channels.

[0130] Category-aware enhancement: Category balanced cropping (prioritizes the preservation of rare category regions), Category-aware copying (targeted copying of severely under-resourced category regions), Pixel-level Mixup (used in high-intensity enhancement mode).

[0131] Enhanced intensity control: Supports high, medium, and low intensity modes, achieved by adjusting the probability and cropping size.

[0132] Furthermore, the loss functions used for training the spatiotemporal multimodal segmentation model include fine classification loss, coarse classification loss, and graph correlation loss.

[0133] Specifically, the detailed classification loss includes:

[0134] Cross-entropy loss (weight 1.0): Basic classification loss, supports class weighting;

[0135] Dice loss (weight 1.2): solves the class imbalance problem and smooths the boundary prediction;

[0136] Terrain constraint loss (weight 0.2): Constrains the distribution of a specific vegetation type based on the DEM height range.

[0137] Coarse classification loss includes:

[0138] Coarse-layer cross-entropy loss (weight 0.6): Improves region-level classification consistency;

[0139] Coarse-to-fine consistency regularity (weight 0.3): Ensures consistency of classification results across different levels.

[0140] The graph correlation loss uses graph regularization loss (weight 0.5): it constrains the node feature distribution based on the graph adjacency matrix.

[0141] Furthermore, the loss function also includes auxiliary task loss, which includes edge detection loss and superpixel segmentation loss. The edge detection loss and superpixel segmentation loss are based on the output of the dual-path classification decoder. The edge detection loss is obtained by calculating the binary cross-entropy between the edge probability map output by the decoder and the automatically generated edge map. The superpixel segmentation loss is obtained by calculating the distance between the pixel embedding and the corresponding class center embedding.

[0142] Specifically, the losses in auxiliary tasks include:

[0143] Edge detection loss (weight 0.2): Automatically generates edge maps from labels, improving boundary prediction accuracy;

[0144] Superpixel segmentation loss (weight 0.3): encourages similar pixel embeddings within the same category.

[0145] In the loss function optimization process, Focal Loss variants are used to handle scarce classes. Label smoothing (coefficient 0.01-0.1, setting the probability of the correct class to 1-smoothing, and the probability of the other classes to smoothing / (C-1)) is applied to reduce the risk of overfitting. A loss clipping threshold (0.5×loss clipping value×log(number of classes)) is set to prevent the influence of outliers. The Focal Loss variant is used to handle scarce classes by dynamically adjusting the gamma value. The basic gamma for ordinary classes is 2.0, and the basic gamma for scarce classes is 8.0. The focal factor (1-p_t)^gamma is calculated.

[0146] During training, the optimizer is configured as follows:

[0147] (1) Optimizer type:

[0148] AdamW optimizer, with a base learning rate of 1e-4, weight decay of 1e-5, momentum parameters (0.9, 0.999), and numerical stability parameter of 1e-8;

[0149] (2) Layered learning rate strategy:

[0150] ST-MME encoder: The weights, biases, and normalization layers all have a base learning rate of 1.0x;

[0151] ST-GGCM graph convolution: weights and biases are 1.2 times the base learning rate, and the normalization layer is 1.0 times.

[0152] Decoder: Weights and biases are 1.5 times the base learning rate, and the normalization layer is 1.0 times.

[0153] Gradient accumulation: 2 steps, reducing memory usage.

[0154] The following learning rate scheduling mechanism is used during the training iteration process:

[0155] (1) Three-stage scheduling:

[0156] Warm-up phase: Flexible warm-up scheduler (supports three warm-up methods: linear, cosine, and exponential, and dynamically calculates the learning rate adjustment factor according to the mathematical formula of the warm-up type within the warm-up rounds), 5 rounds of warm-up, with an initial learning rate of 1e-6, which is gradually increased to the base learning rate;

[0157] Main scheduling phase: Adaptive cosine annealing scheduler, initial cycle length T_0=20, cycle multiplication factor T_mult=2, minimum learning rate eta_min=1e-6, dynamically adjusts the cycle length based on validation performance. This scheduler receives validation performance parameters in the step method, calculates the performance improvement compared to a threshold, and maintains or increases the cycle length when the performance improvement exceeds the threshold (0.01). When performance decreases, T_0 is reduced by a cooling factor (0.8) and the current cycle count T_cur is reset to 0, thus achieving dynamic adjustment of the cycle length.

[0158] Gradient-aware adjustment: When the gradient norm continues to increase, the learning rate is scaled by 0.9 times to prevent training oscillations.

[0159] The regularization techniques used during training are as follows:

[0160] Weight regularization: L2 regularization is implemented through the AdamW optimizer, with a weight decay rate of 1e-4;

[0161] Normalization techniques: SC-RU and ST-RU use group normalization to adapt to mini-batch training, and batch normalization is used after convolutional layers to accelerate convergence;

[0162] Training process regularization: Dropout (scale 0.1) is used in CrossModalTransformer and graph convolution modules, early stopping strategy patience value of 50 rounds monitoring validation set mIoU, mixed precision training uses AMP (Automatic Mixed Precision Training) to reduce memory usage and speed up training, gradient pruning thresholds are set by module (ST-MME=5.0, ST-GGCM=8.0, decoder=10.0).

[0163] The model initialization and training process is as follows:

[0164] Initialization: Convolutional layers use Kaiming normal initialization to adapt the ReLU activation function, linear layers use Xavier initialization to ensure consistent input and output variance, and normalized layers initialize γ to 1.0 and β to 0.0.

[0165] Training monitoring: Training logs are recorded every 10 batches, and TensorBoard visualization is recorded every 50 batches. Monitoring metrics include total loss and loss of each component, gradient norm, learning rate, validation set mIoU, class precision, recall, and F1 score.

[0166] Special training techniques: Dynamic feature management (resetting intermediate features of ST-MME and ST-GGCM before each training batch), skipping abnormal batches for NaN / Inf (non-numeric / infinity) detection, prioritizing sampling regions containing rare classes for category-aware sampling, and module-differentiated training by setting different learning rates and gradient clipping thresholds for different modules.

[0167] After completing model training, classify the vegetation in the region to be classified according to the following procedure:

[0168] 1. Prediction process:

[0169] Input preprocessing: Preprocessing and feature extraction are performed on the optical imagery, radar imagery, and digital elevation model data of the region to be classified;

[0170] Model inference: The preprocessed multimodal data is input into the trained model for forward propagation calculation;

[0171] Post-processing: The argmax operation is performed on the probability map output by the model to obtain the final class label of each pixel. Small region filtering is applied to remove isolated patches with too small area, and edge smoothing is performed to optimize boundary details.

[0172] 2. Model performance validation:

[0173] At the 151st epoch of training, the model demonstrated excellent performance on the independent validation set, with the specific evaluation results as follows:

[0174] Overall performance indicators:

[0175] The average intersection-union ratio (Val mIoU) of the validation set is 0.6648.

[0176] Overall accuracy of the validation set (Val OA): 0.8342;

[0177] The average crossover and union ratio (Val coarse mIoU) for coarse classification on the validation set is 0.7251.

[0178] Overall accuracy of coarse classification on the validation set (Val coarse OA): 0.9847.

[0179] Detailed intersection and union ratios for each category:

[0180] Category 1 IoU: 0.7586; Category 2 IoU: 0.6306; Category 3 IoU: 0.8407; Category 4 IoU: 0.4658; Category 5 IoU: 0.4988; Category 6 IoU: 0.5950; Category 7 IoU: 0.8336; Category 8 IoU: 0.7833.

[0181] Coarse classification intersection-union ratio:

[0182] Crude category 1 IoU: 0.7652; Crude category 2 IoU: 0.8776; Crude category 3 IoU: 0.5514; Crude category 4 IoU: 0.5982; Crude category 5 IoU: 0.8479; Crude category 6 IoU: 0.7102.

[0183] Training loss analysis:

[0184] Fine-class cross-entropy loss (ce_loss): 0.076303;

[0185] Dice loss (dice_loss): 0.464607;

[0186] Coarse classification cross-entropy loss (coarse_ce): 0.106256;

[0187] Consistency of coarse-grained regularization: 0.004507;

[0188] Topological constraint loss (topo_loss): 0.000124;

[0189] Graph regularization loss: 0.000005;

[0190] Total training loss: 0.6990, validation loss: 0.6980;

[0191] Gradient norm: 0.6244, current learning rate: 0.000005.

[0192] The above results demonstrate that the proposed method achieves a high level of accuracy in vegetation classification tasks. The overall accuracy (OA) exceeds 83%, and the coarse classification accuracy approaches 99%, fully validating the effectiveness of the semi-cascaded classification strategy. The IoU values ​​for each category show that the model has stable recognition capabilities for different vegetation types, the training loss converges smoothly, and the gradient norm is moderate, indicating that the training process is stable and effective.

[0193] Figure 5 The graph shows the performance curves of the model during training, including overall accuracy (OA), Kappa coefficient, and mean intersection-over-union ratio (mIoU). As can be seen from the graph, all three curves steadily increase and eventually stabilize as training progresses. The validation set curve is largely consistent with the training set curve, indicating no significant overfitting, thus demonstrating that the model has good generalization ability and convergence.

[0194] The model ultimately outputs eight vegetation classification categories, each corresponding to a specific vegetation or land cover type. Based on the classification system of this application and the regional vegetation distribution characteristics, these eight categories and their correspondences are as follows:

[0195] Category 1: Cultivated land; Category 2: Deciduous broad-leaved forest; Category 3: Evergreen broad-leaved forest; Category 4: Deciduous coniferous forest; Category 5: Evergreen coniferous forest; Category 6: Shrubland and grassland; Category 7: Water bodies and wetlands; Category 8: Artificial surface.

[0196] The classification results are output as raster images, with each pixel value (1-8) corresponding to a vegetation type. A classification accuracy report is also output, including quantitative evaluation indicators such as overall accuracy, Kappa coefficient, producer accuracy for each category, and user accuracy. These results can be directly used to create regional vegetation type distribution maps, supporting applications such as ecological monitoring, resource management, and environmental protection.

[0197] Figure 6 This paper presents a visual comparison of the vegetation classification performance of the proposed method in five different sub-regions. Each set of illustrations includes three sub-images: the left side is a true-color composite image of the original remote sensing image taken by the Sentinel-2 satellite; the middle side is the manually corrected image of the actual vegetation category labels for the corresponding region; and the right side is the vegetation classification prediction result output by the proposed method. Through comparison of five regions with different terrains and vegetation types, it is evident that the prediction results visually match the actual labels highly, accurately distinguishing various land types such as woodland, water bodies, cultivated land, and shrubs. Especially in complex mountainous terrain, it maintains high boundary accuracy and category consistency. The five sets of comparison images intuitively demonstrate the practicality and effectiveness of the proposed method, particularly its stable classification ability in complex terrain and mixed vegetation scenarios.

[0198] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.

[0199] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A vegetation classification method based on spatiotemporal multimodal deep learning, characterized in that, include: Acquire multi-temporal optical images, radar images, and digital elevation model data of the area to be classified to form multimodal data; The multimodal data is input into a derived feature extraction network, which extracts spectral data, microwave data, terrain data, and texture data from the multimodal data. A spatiotemporal multimodal encoder is used to extract spectral features, microwave features, terrain features, and texture features from the spectral data, microwave data, terrain data, and texture data, respectively. The spectral features, microwave features, terrain features, and texture features are then fused together using cross-modal self-attention to form fused features. The fused features are divided into multiple superpixel blocks using a gated spatiotemporal graph convolution module. The graph node features of the superpixel blocks are calculated. The graph node features are updated using graph convolution operations. The updated graph node features are reshaped into spatial features. The spatial features are upsampled to form graph features. The fusion feature and the graph feature are fused at the pixel level using a dual-path classifier decoder to obtain pixel-level coarse classification logits. At the same time, the fusion feature is subjected to multi-scale feature extraction and cross-scale fusion at the region level to form region-level coarse classification logits. The pixel-level coarse classification logits and the region-level coarse classification logits are fused to obtain the fused coarse classification probability. A fine classifier is then used to perform fine classification on the coarse basic features extracted from the fusion feature and the graph feature according to the coarse classification probability to obtain the final classification result. The derived feature extraction network, the spatiotemporal multimodal encoder, the gated spatiotemporal graph convolutional module, and the dual-path classifier decoder constitute a spatiotemporal multimodal segmentation model. During the training of the spatiotemporal multimodal segmentation model, sample data is acquired and combined into spatiotemporal cube samples, which are then input into the spatiotemporal multimodal segmentation model.

2. The vegetation classification method based on spatiotemporal multimodal deep learning according to claim 1, characterized in that, After acquiring the multimodal data and the sample data, both the multimodal data and the sample data are preprocessed.

3. The vegetation classification method based on spatiotemporal multimodal deep learning according to claim 2, characterized in that, The multimodal data preprocessing for the region to be classified includes: mosaicking and cropping the optical image and unifying the coordinates / resolution; performing SAR preprocessing and postprocessing on the radar image; and mosaicking and cropping the digital elevation model data and unifying the coordinates / resolution. The sample data includes ground truth data of the sample area, the optical image, the radar image, and the digital elevation model data. The preprocessing of the ground truth data includes format conversion, coordinate unification, and mode voting fusion.

4. The vegetation classification method based on spatiotemporal multimodal deep learning according to claim 1, characterized in that, The optical image contains four temporal phases. The derived feature extraction network adds six indices to the optical image: normalized difference vegetation index, enhanced vegetation index, normalized difference water index, modified soil-regulated vegetation index, green leaf index, and near-infrared to red band ratio. It also calculates texture in the B8 band. After standardizing the 13 channels composed of the four temporal phases, the six indices, and the texture, the spectral data and the texture data are obtained.

5. The vegetation classification method based on spatiotemporal multimodal deep learning according to claim 1, characterized in that, The spatiotemporal multimodal encoder comprises four parallel branches, which respectively extract the spectral features, microwave features, terrain features, and texture features. The branches for extracting the spectral features and microwave features contain a dynamically deformable convolution, a spatial reconstruction unit, and a spatiotemporal reconstruction unit connected in sequence. The branch for extracting the terrain features contains a dynamically deformable convolution and a spatial reconstruction unit. The branch for extracting the texture features contains a spatial reconstruction unit and a spatiotemporal reconstruction unit.

6. The vegetation classification method based on spatiotemporal multimodal deep learning according to claim 5, characterized in that, The features output from the four branches are concatenated and then the 4×128 channels are reduced to 256 channels through a cross-modal projection layer. The result is then input into the CrossModalTransformer module, which uses a sliding window attention mechanism for cross-modal self-attention fusion.

7. The vegetation classification method based on spatiotemporal multimodal deep learning according to claim 1, characterized in that, The gated spatiotemporal graph convolution module includes two gated graph convolutional layers, namely gated graph convolutional layer one and gated graph convolutional layer two. The gated spatiotemporal graph convolution module fuses the spatial proximity, temporal similarity and terrain consistency of the graph node features to generate a dynamic adjacency matrix. The gated graph convolutional layer one performs graph convolution operations on the graph node features according to the dynamic adjacency matrix to form initial graph features. The gated graph convolutional layer two updates the initial graph features according to the dynamic adjacency matrix to form updated graph features, and reshapes the updated graph features into the spatial features.

8. The vegetation classification method based on spatiotemporal multimodal deep learning according to claim 1, characterized in that, The loss functions used to train the spatiotemporal multimodal segmentation model include fine classification loss, coarse classification loss, and graph correlation loss.

9. A vegetation classification method based on spatiotemporal multimodal deep learning according to claim 8, characterized in that, The loss function also includes auxiliary task loss, which includes edge detection loss and superpixel segmentation loss. The edge detection loss and the superpixel segmentation loss are based on the output of the dual-path classification decoder. The edge detection loss is obtained by calculating the binary cross-entropy between the edge probability map and the edge map output by the decoder. The superpixel segmentation loss is obtained by calculating the distance between the pixel embedding and the corresponding class center embedding.

Citation Information

Patent Citations

  • Binary arithmetic subtraction circuit

    CN106569775A

  • Land utilization monitoring method and system based on remote sensing and big data

    CN120599485A