Ground feature classification method and device for multivariate time sequence remote sensing data
Through the MSSTAN algorithm, combined with the AM-GTU module, coding layer and Residual-Gate KAN module, the inaccuracy problem in the geographic classification of multi-time sequence remote sensing data is solved, and high-precision land coverage classification is achieved.
Patent Information
- Application Number
- CN202510243662.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-03-03
AI Technical Summary
The prior art has problems with inaccuracy in the classification of multi-time sequence remote sensing data, especially when processing high-dimensional time sequence data, it is easy to lose detailed information and it is difficult to capture complex space-time patterns.
The MSSTAN algorithm is adopted, which includes the AM-GTU module, the encoding layer, the Residual-Gate KAN module and the classifier. The AM-GTU module extracts spectral and timing information through the multi-scale 1D convolution layer and the gating unit layer, the encoding layer captures long-term dependence through the multi-head attention mechanism, and the Residual-Gate KAN module extracts local and global features through the KAN module and the linear network layer, and adjusts the feature extraction ratio through the gating unit.
The model's attention to details and global dependence has been improved, and the geographic classification accuracy of multi-time sequence remote sensing data has been effectively improved, and the inaccuracy problem of existing models in multi-time sequence geographic classification has been solved.
Smart Images

Figure CN120180264A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing data processing, and particularly to a method and device for classifying ground objects of multi-source time-series remote sensing data. Background Art
[0002] Due to the rich time-dimensional information contained in time-series data, it has been widely used in many fields. In the field of remote sensing, time-series data is a continuous observation sequence composed of ground object spectral information and phase information, which can effectively reflect the spectral change characteristics of ground objects in the time dimension. In recent years, with the increasing abundance of open-source remote sensing satellite data, the accessibility of time-series data has been significantly improved, which provides an important basis for the application of time series classification (TSC) in the field of remote sensing data mining.
[0003] To achieve accurate classification of ground objects, researchers have proposed a variety of TSC algorithms based on machine learning. TSC aims to build a machine learning model for predicting the class labels of a continuous ordered sequence of real-valued observations. Since multi-source time-series remote sensing data belongs to high-dimensional data, which contains features of multiple bands, it is difficult to achieve high-precision results using existing methods for multi-source time-series ground object classification. Therefore, applying multi-source time-series remote sensing data for ground object classification is a complex process, and there are many reasons affecting the classification accuracy, such as the feature selection of multi-source data, the selection of different ground object feature classification methods, etc.
[0004] Traditional machine learning methods (such as DTW, RF, SVM) still have limitations in the discrimination and mining ability of sequence features, and have limited ability to extract the temporal and spectral features of time series, making it difficult to capture complex spatio-temporal patterns. Existing deep learning models (such as CNN, Bi-LSTM) have problems such as insufficient local context capture and easy loss of detail information when processing high-dimensional time-series remote sensing data. Therefore, existing methods and models all have inaccurate problems in multi-source time-series ground object classification, especially for the land cover classification task with unbalanced time-series data samples in remote sensing, the above problems are more prominent and difficult to meet the actual needs. Summary of the Invention
[0005] In view of this, the embodiments of the present invention provide a method and device for classifying ground objects of multi-source time-series remote sensing data, based on MSSTAN, which improves the model's attention to details and global dependencies, solves the inaccurate problems of existing models in multi-source time-series ground object classification, and realizes accurate land cover classification of multi-source time-series remote sensing data.
[0006] In a first aspect, the embodiments of the present invention provide a method for classifying ground objects of multi-source time-series remote sensing data, including the following steps:
[0007] S1: Obtain multi-temporal remote sensing data, where the multi-temporal remote sensing data includes multi-band spectral features and time series;
[0008] S2: Input the multi-temporal remote sensing data into the MSSTAN algorithm (Multi-Scale Spectral-Temporal Attention Network) to output land object classification and probability;
[0009] Among them, the MSSTAN algorithm includes an AM-GTU module, an encoding layer, a Residual-Gate KAN module, and a classifier; the AM-GTU module includes a multi-scale 1D convolutional layer and a gated unit layer. The multi-scale 1D convolutional layer is used to extract spectral and temporal information at different scales, and the gated unit layer controls the incoming proportion of information; the encoding layer focuses on the relationships between different time steps in the sequence through a multi-head attention mechanism to capture long-term dependencies within the sequence; the Residual-Gate KAN module includes a KAN module and a linear network layer. The KAN module is used to extract local features. The splines in the KAN module are accurate for low-dimensional functions and local adjustments. The linear network layer can well retain and transmit global feature information to extract global features. The global features retained by the linear network layer are input into the gated unit together with the local adjustment results obtained by the KAN module through a residual connection, and the gated unit adjusts the extraction proportion of local and global features; the classifier is used to output land object classification and probability.
[0010] The technical effects of the above embodiments are as follows: For the existing Transformer model that directly inputs data into the encoder, resulting in sparse local detail expressions of spectra and time series. This land object classification method fuses the important spectral and temporal information at different scales of the input data through an improved multi-scale gated Tanh unit (AM-GTU), and after fusion, it pays more attention to the local context information of spectra and time series in the samples, which can effectively improve the classification accuracy. For the problem that the existing Transformer linear layer has limited ability to learn subtle patterns from high-dimensional datasets, this land object classification method proposes a Residual-Gate KAN module that combines KAN and a residual gated network, replacing the linear mapping process with a learnable non-linear mapping process, and at the same time effectively reducing the problem of high-dimensional information loss. In summary, the MSSTAN method proposed by this land object classification method realizes the dual extraction of local detail features and pixel global relationships while maintaining the original advantages of the Transformer, achieving accurate land cover classification of multi-temporal remote sensing data.
[0011] According to a specific implementation manner of the embodiment of the present invention, the multi-scale 1D convolutional layer includes 1D causal convolutional kernels of 5 scales, namely 1×3, 1×5, 1×7, 1×9, and 1×12.
[0012] According to a specific implementation manner of an embodiment of the present invention, the gating unit layer of the AM-GTU module controls the information input ratio by using the activation functions Sigmoid and Tanh, and then performs a residual connection with the original input. The Sigmoid and Tanh gating units are used to screen important local features; the residual connection is used to alleviate the problem of feature loss in deep networks.
[0013] According to a specific implementation manner of an embodiment of the present invention, the encoding layer includes an input embedding layer, a position encoding layer, a multi-head attention layer, a cross-layer and normalization layer, a feed-forward network layer, and a cross-layer and normalization layer. The input embedding layer and the position encoding layer perform encoding operations on the input features, and then, through the multi-head attention mechanism in the Transformer, train the input features to continuously optimize the matching degree between the classification result and the true value, so as to output a more accurate classification result graph.
[0014] According to a specific implementation manner of an embodiment of the present invention, the linear network layer includes a linear layer, a ReLU activation function, and a BN normalization layer.
[0015] According to a specific implementation manner of an embodiment of the present invention, the gating unit in the Residual-Gate KAN module uses the sigmoid function to dynamically adjust the fusion ratio of local and global features.
[0016] According to a specific implementation manner of an embodiment of the present invention, the classifier is a linear layer that maps the feature dimension from d_model / / 2 to the number of categories.
[0017] In a second aspect, an embodiment of the present invention provides a ground object classification device for multi-source time-series remote sensing data, including:
[0018] An acquisition module for acquiring multi-source time-series remote sensing data, where the multi-source time-series remote sensing data includes multi-band spectral features and time series;
[0019] A detection module that inputs the multi-source time-series remote sensing data into the MSSTAN algorithm and outputs the classification and probability of ground object targets;
[0020] Among them, the MSSTAN algorithm includes an AM-GTU module, an encoding layer, a Residual-Gate KAN module, and a classifier; the AM-GTU module includes a multi-scale 1D convolutional layer and a gated unit layer. The multi-scale 1D convolutional layer is used to extract spectral and temporal information at different scales, and the gated unit layer controls the incoming ratio of information; the encoding layer focuses on the relationship between different time steps in the sequence through a multi-head attention mechanism to capture long-term dependencies within the sequence; the Residual-Gate KAN module is used to aggregate multi-level context features, including a KAN module and a linear network layer. The KAN module is used to extract local features, and the linear network layer is used to extract global features, and the gated unit adjusts the extraction ratio of local and global features; the classifier is used to output the classification and probability of ground object targets. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. In all the drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, the elements or parts are not necessarily drawn to scale.
[0022] Figure 1 FIG. shows a flowchart of a method for classifying ground objects in multi-source temporal remote sensing data provided by an embodiment of the present invention;
[0023] Figure 2 FIG. shows an overall framework diagram of the MSSTAN algorithm in an embodiment of the present invention;
[0024] Figure 3 FIG. shows a schematic structural diagram of the AM-GTU module in an embodiment of the present invention;
[0025] Figure 4 FIG. shows a schematic structural diagram of the Residual-Gate KAN module in an embodiment of the present invention;
[0026] Figure 5 FIG. shows a schematic diagram of local visualization classification results of different methods in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0027] The following will describe in detail the embodiments of the technical solutions of the present invention with reference to the drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention, so they are only examples and cannot be used to limit the protection scope of the present invention.
[0028] It should be noted that unless otherwise specified, the technical terms or scientific terms used in this application should have the ordinary meaning understood by those skilled in the art to which the present invention belongs.
[0029] Figure 1 The flowchart of the steps of a method for classifying ground objects of multi-source time-series remote sensing data provided by an embodiment of the present invention is shown in Figure 1 , and the method includes the following steps:
[0030] S1: Obtain multi-source time-series remote sensing data, where the data includes multi-band spectral features and time series;
[0031] In this embodiment, the publicly available land use dataset in the TiSeLaC challenge is used for illustration. The TiSeLaC dataset is collected from the 2A-level Landsat 8 satellite image of Reunion Island with a size of 2866*2633 pixels and a spatial resolution of 30m taken in 2014. There are 23 scenes throughout the year and a total of 10 bands. Each pixel consists of 10 channel features: including the first 7 bands of the original data (Band1 - Band7 in Landsat8), representing the measurement values of each independent multi-spectral band (OLI): ultra-blue, blue, green, red, NIR, SWIR1, and SWIR2; it also includes three complementary radiation indices (normalized difference vegetation index, normalized difference water index, and brightness index). The specific Landsat 8 band information and the three constructed vegetation indices are shown in Tables 1 and 2.
[0032]
[0033]
[0034] The organizers of TiSeLaC sampled 99,687 pixels from this publicly available dataset. At the same time, referring to the Corine land cover map in 2012 and the local farmers' land cover registration results in 2014, the land cover types in the Reunion Island study area are divided into 9 categories, including a training set of 81,714 pixels and a test set of 17,973 pixels. The proportion of each class in the training set and the test set is approximately 4:1, as shown in Table 3 below.
[0035]
[0036] S2: Input the multi-source time-series remote sensing data into the MSSTAN algorithm, and output the classification and probability of ground object targets;
[0037] The MSSTAN algorithm is a multi-scale spectral-temporal attention network based on Transformer. In the multi-temporal classification task in the field of land use, only the Encoder encoding layer of MSSTAN participates in classification, and the Decoder decoding layer is discarded. This is because the Encoder can focus on the relationships between different time steps in the sequence through the self-attention mechanism and capture long-term dependencies within the sequence. Since the goal of the classification task is usually to model or extract features from the entire time series, rather than predicting the next time step or generating a new sequence, the Decoder part is not required to play a role in the classification task.
[0038] As Figure 2 shown, the MSSTAN algorithm includes: AM-GTU module, encoding layer, Residual-Gate KAN module, and classifier.
[0039] In the multi-temporal classification task, the input multi-temporal remote sensing data is subjected to feature extraction by the AM-GTU module (Advanced Multi-Scale Gated Tanh Unit). On the basis of retaining important local information, the module enhances the ability to extract global features. The structural diagram of the improved AM-GTU module is as Figure 3 shown. By stacking multiple GTUs, the receptive field in the time dimension can be extended to improve the model's ability to extract long-term temporal correlations in the data. In this embodiment, the dimension of the input data is (10, 1, 23), which means that each sample has 10 band features and 23 time dimensions. Each sample is input into the next layer of gated unit layer through a five-layer multi-scale 1D convolutional layer of 1×3, 1×5, …, 1×12. The gated unit layer uses the activation functions Sigmoid and Tanh units to control the proportion of information passing through the multi-scale data. According to the cross-entropy loss function, the feature part more beneficial to this classification task and the original input are selected for residual connection. Finally, the output obtained after passing through the AM-GTU module has the same dimension as the input, and the dimension is still (10, 1, 23).
[0040] The features extracted by the AM-GTU module are input into the encoding layer, and successively pass through the input embedding layer (InputEmbedding), positional encoding layer (Positional Encoding), multi-head attention layer (Multi-HeadAttention), add & norm layer (Add&Norm), feed-forward network layer (Feed Forward), and add & norm layer (Add&Norm) in the encoding layer, and the process from the multi-head attention layer to the last add & norm layer is repeated N times. The input embedding layer and the positional encoding layer perform encoding operations on the input features. Then, the input features are trained through the multi-head attention mechanism in Transformer to continuously optimize the matching degree between the classification result and the true value, so as to output a more accurate classification result graph.
[0041] After passing through the encoding layer, the features are used by the Residual-Gate KAN module to extract local and global features, and the result obtained through the residual connection and KAN is finally input into the sigmoid gating unit. It should be noted that the input of the Residual-Gate KAN module is the result obtained from the Encoder encoding layer. The structural diagram of the Residual-Gate KAN module is as Figure 4 shown, which is divided into two parts. The input is not only passed to the KAN module, but also passed to the linear layer (Linear), activation function (ReLU), linear layer (Linear), and batch normalization layer (BN, BatchNorm). The final activation function Sigmoid acts as a gating unit to dynamically adjust the importance of the information on these two lines, so as to adjust the extraction ratio of the key information of local and global features.
[0042] The classifier is a linear layer (Linear), which maps the feature dimension of the features fused by the Residual-Gate KAN module to the number of categories, and combines the Softmax function to output the classification and probability of the ground object target.
[0043] It should be further noted that:
[0044] 1. The MSSTAN algorithm improves the local relevance of the context:
[0045] In the AM-GTU module, 1D causal convolutions with 5 different scales are used to extract features of the multivariate time series input X. Since the importance of the bands or time step feature intervals concerned by each scale is different, the multi-scale features extracted by the 5 1D causal convolutions cannot be directly concatenated. It is necessary to control the proportion of the information flow to the next module through the gating unit to retain important local information.
[0046] 2. The MSSTAN algorithm dynamically screens important local and global information:
[0047] The Residual-Gate KAN can effectively aggregate multi-level context features. By replacing the single linear layer in the traditional Transformer network with the Residual-Gate KAN, on the one hand, the KAN module is different from the classical MLP network in that it has a fixed activation function at the neuron nodes, while the KAN has a learnable weight activation function on the edges. On the other hand, the Residual-Gate KAN adds a new linear network layer on the basis of the KAN module. The linear network layer includes a linear layer, a ReLU activation function, and a BatchNorm normalization layer, and the linear network layer can well retain and transmit global feature information. Finally, the result obtained through the residual connection and the KAN is finally input into the sigmoid gating unit to adjust the extraction ratio of the important information of the local and global features.
[0048] The training steps of the MSSTAN algorithm model are as follows:
[0049] 1. Set the random number seed. Setting the same seed makes the pre-training weights unchanged for each training, and the training results can be reproduced each time. In addition, complete the hyperparameter settings required for the MSSTAN model.
[0050] 2. According to the data form of the public dataset, construct your own training data loader and test data loader. This facilitates obtaining relevant information about the dataset from the DataLoader later.
[0051] 3. Create the MSSTAN model. The shape of the original training data in the DataLoader is (B, T, C), where B is the batch size, C = 10 is the number of bands, and T = 23 is the number of time steps. After adjustment, the shape of the input data X becomes (B, C, 1, T), and then it passes through the improved AM-GTU module. This module uses 1D causal convolutions of 5 different scales to extract features from the multivariate time series input X, and controls the proportion of the information flow to the next module through the gating unit to retain important local information. Then, the information of multiple scales is concatenated in the time dimension, and the same shape as the input is restored through a linear layer. To alleviate the problem of loss of important shallow features caused by the deep structure of the AM-GTU module, finally, it passes through a residual connection and a ReLU activation function as the output of this module. After passing through the multi-head attention mechanism of the Transformer, the obtained feature input is constructed into a Residual-Gate KAN. Here, Residual-Gate KAN adds a new linear network layer to the KAN module. The linear network layer includes a linear layer, a ReLU activation function, and a BatchNorm normalization layer. The linear network layer can well retain and transmit global feature information. Here, Residual-Gate KAN also serves as a means of dimensionality reduction, integrating features that are more important for the classification result to improve the final classification accuracy.
[0052] Here, the dataset is the publicly available TiSeLaC multivariate time series remote sensing dataset. The above shows the core part of the application of the MSSTAN algorithm in the TiSeLaC land cover classification task and can improve the classification accuracy. In a more extensive land cover classification scenario, only the dataset corresponding to the scenario needs to be constructed. Specifically:
[0053] 1. First, batch download the Landsat or Sentinel-2 satellite images of the corresponding year through the Google Earth Engine platform as the time series image dataset for extracting the category data features later.
[0054] 2. Then, select the label values of N types of land by visual interpretation or by selecting label points according to the publicly available ground truth dataset (such as the Dynamic World land use publicly available dataset), and generate shp files for each category separately. Note that here it is fused into a single shp polygon vector, otherwise, when creating random points, the same number of random points will be created for each polygon vector. Then calculate the proportion of each category and create an appropriate number of random points.
[0055] 3. Finally, use a Python script to implement multi-value extraction to points and extract the time series features of different category random points from the time series image dataset. In this way, a complete land cover dataset is constructed.
[0056] Thus, a general method for batch generating land cover datasets according to different research areas is realized. Then, the generated datasets are input into the MSSTAN algorithm for classification to achieve the goal of procedural classification. Experiments have proved that this general method can be applied to land classification tasks in different scenarios, such as land classification, arbor classification, etc., and good results have been achieved.
[0057] The training method of the MSSTAN algorithm model is as follows:
[0058] 1. Positional encoding;
[0059] The fused features after the improved AM-GTU module are input into the Encoder in Transformer for training. The features input into the Encoder are first subjected to positional encoding to enable the model to understand the position (order) of each word in the sequence, so as to help the model learn this information.
[0060] 2. Multi-head attention mechanism;
[0061] The input features with positional order are passed through the multi-head attention mechanism. Here, the sizes of q, k, and v are set to be the same as the number of heads, which is set to 8. And 8 layers of such Encoders are stacked, and each layer has its own parameters and weights, so that the model can capture features at different depths through stacking. The shape of the features after being processed by the multi-head attention mechanism changes from the initial input x(B, C, T) to x(B, d_model*C), where d_model is set to 512. Here, d_model is used as the linear mapping parameter of the embedding, and the time step T in the input is mapped to the high-dimensional d_model in the Encoder stage.
[0062] 3. Residual-Gate KAN;
[0063] The features x(B, d_model*C) obtained after the multi-head attention mechanism in the Encoder then pass through the Residual-Gate KAN module constructed by the present invention, and the shape becomes output(B, d_model / / 2). On the one hand, the KAN module replaces the linear mapping process with a learnable non-linear mapping process to reduce the problem of high-dimensional information loss. On the other hand, the network design of the residual and sigmoid gating unit can adjust the extraction ratio of local and global feature important information.
[0064] 4. Classifier;
[0065] Finally, the model maps the high-dimensional d_model / / 2 in the output to the final 9 categories through a linear layer. The detailed settings of each module of the model are shown in Table 4 below.
[0066]
[0067] 5. Calculate the loss;
[0068] Use the Cross-Entropy Loss function to calculate the loss for each round of training. The formula for the cross-entropy loss function is shown in Equation (1) below. Here, C is the total number of categories, and y i,c is the true label (one-hot encoded) of sample i for category c. p i ,c is the probability that sample i is predicted as category c (the result after the model output passes through softmax).
[0069]
[0070] 6. Backpropagation and optimization;
[0071] Backpropagate according to the loss value of each round to automatically adjust and find more suitable training weights for this classification task.
[0072] 7. Precision metrics:
[0073] This embodiment uses four metrics, namely Overall Accuracy (OA), Mean Intersection over Union (mIoU), Recall, and F1-score, to evaluate the performance of the model. OA refers to the overall accuracy, which is the proportion of samples predicted correctly by the model among all samples and can directly reflect the overall performance of the model. However, for datasets with class imbalance, it cannot fully reflect the prediction accuracy of the model. Therefore, mIoU is introduced to measure the accuracy of each category and take the average. Similarly, the Recall and F1 Score metrics are introduced to handle imbalanced datasets and better evaluate whether the model can identify samples of the minority class. The formulas for each precision metric are shown in Table 5 below, where C is the total number of categories, TP is the number of positive samples correctly identified, FP is the number of false-negative samples, TN is the number of negative samples correctly identified, and FN is the number of missed positive samples. Finally, according to the MSSTAN method proposed in the present invention and several common time series classification algorithms, mapping is carried out, and the local classification result diagrams of different algorithms on the TiSeLaC land cover public dataset are as Figure 5 shown.
[0074]
[0075] The embodiments of the present invention have the following technical effects:
[0076] The local object classification method is a multi-temporal classification method based on Transformer. The MSSTAN algorithm includes an AM-GTU module, which can effectively solve the problem of insufficient capture of local relevance; in addition, the proposed Residual-Gate KAN can effectively aggregate multi-level context features. Combining the improvements of the above two parts can achieve high-precision time series classification. The local object classification method verifies the high-precision classification of multi-temporal series through public datasets, and also provides an important reference for subsequent extended applications in land use, electronic health record analysis, human activity recognition, etc.
[0077] A structural block diagram of a device for classifying ground objects of multi-temporal remote sensing data according to an embodiment of the present invention. The device includes:
[0078] An acquisition module for acquiring multi-temporal remote sensing data, where the multi-temporal remote sensing data includes multi-band spectral features and time series;
[0079] A detection module that inputs the multi-temporal remote sensing data into the MSSTAN algorithm model and outputs the classification and probability of ground object targets;
[0080] Among them, the MSSTAN algorithm includes an AM-GTU module, an encoding layer, a Residual-Gate KAN module, and a classifier; the AM-GTU module includes a multi-scale 1D convolutional layer and a gated unit layer. The multi-scale 1D convolutional layer is used to extract spectral and temporal information at different scales, and the gated unit layer controls the incoming proportion of information; the encoding layer pays attention to the relationship between different time steps in the sequence through a multi-head attention mechanism to capture long-term dependencies within the sequence; the Residual-Gate KAN module is used to aggregate multi-level context features, including a KAN module and a linear network layer. The KAN module is used to extract local features, and the linear network layer is used to extract global features, and the extraction ratio of local and global features is adjusted through a gated unit; the classifier is used to output the classification and probability of ground object targets.
[0081] The functions of the modules in the device embodiment correspond to the content in the corresponding method embodiment, and will not be elaborated here.
[0082] It should be noted that arranging each module in a flow layout is only one embodiment of the present invention, and other arrangements can also be adopted. The present invention does not make any limitations in this regard.
[0083] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A method for classifying objects in multivariate time series remote sensing data, characterized in that: The following steps are involved: S1: Acquire multivariate time series remote sensing data, wherein the multivariate time series remote sensing data includes multi-band spectral features and time series; S2: Input the multivariate time series remote sensing data into the MSSTAN algorithm, and output the classification and probability of ground objects; Among them, the MSSTAN algorithm includes an AM-GTU module, a coding layer, a Residual-Gate KAN module and a classifier; the AM-GTU module includes a multi-scale 1D convolution layer and a gated unit layer, the multi-scale 1D convolution layer is used to extract spectral and temporal information at different scales, and the gated unit layer controls the incoming information ratio; the coding layer uses a multi-head attention mechanism to focus on the relationship between different time steps in the sequence and capture the long-term dependency within the sequence; the Residual-GateKAN module is used to aggregate multi-level contextual features, including a KAN module and a linear network layer, the KAN module is used to extract local features, the linear network layer is used to extract global features, and the extraction ratio of local and global features is adjusted through the gated unit; the classifier is used to output the classification and probability of ground objects.
2. The method for classifying land features according to claim 1, characterized in that: The multi-scale 1D convolution layer includes 1D causal convolution kernels of five scales, namely 1×3, 1×5, 1×7, 1×9, and 1×12.
3. The method for classifying land features according to claim 1, characterized in that: The gated unit layer of the AM-GTU module uses activation functions Sigmoid and Tanh to control the proportion of information input, and then performs a residual connection with the original input.
4. The method for classifying land features according to claim 1, characterized in that: The encoding layer includes an input embedding layer, a position encoding layer, a multi-head attention layer, a cross-layer and normalization layer, a feedforward network layer, and a cross-layer and normalization layer.
5. The method for classifying land features according to claim 1, characterized in that: The linear network layer of the Residual-Gate KAN module includes a linear layer, a ReLU activation function, and a BN normalization layer.
6. The method for classifying land features according to claim 5, characterized in that: The gate control unit in the Residual-Gate KAN module adopts a sigmoid function to adjust the fusion ratio of local and global features.
7. The method for classifying land features according to claim 1, characterized in that: The classifier is a linear layer that maps the feature dimension from d_model / / 2 to the number of categories.
8. A device for classifying objects in multivariate time series remote sensing data, characterized in that: The method for classifying land features as claimed in any one of claims 1 to 7 comprises: An acquisition module is used to acquire multivariate time series remote sensing data, wherein the multivariate time series remote sensing data includes multi-band spectral features and time series; A detection module, inputting the multivariate time series remote sensing data into the MSSTAN algorithm, and outputting the classification and probability of ground objects; Among them, the MSSTAN algorithm includes an AM-GTU module, a coding layer, a Residual-Gate KAN module and a classifier; the AM-GTU module includes a multi-scale 1D convolution layer and a gated unit layer, the multi-scale 1D convolution layer is used to extract spectral and temporal information at different scales, and the gated unit layer controls the incoming information ratio; the coding layer uses a multi-head attention mechanism to focus on the relationship between different time steps in the sequence and capture the long-term dependency within the sequence; the Residual-GateKAN module is used to aggregate multi-level contextual features, including a KAN module and a linear network layer, the KAN module is used to extract local features, the linear network layer is used to extract global features, and the extraction ratio of local and global features is adjusted through the gated unit; the classifier is used to output the classification and probability of ground objects.
Citation Information
Patent Citations
Hyperspectral remote sensing image classification method based on self-attention context network
CN112287978A
High-resolution remote sensing image land coverage classification method and device and storage medium
CN117036936A
Voice spoofing detection method based on feature-enhanced attention mechanism
CN118298832A
Constellation-based distributed collaborative remote sensing determination method and apparatus, storage medium and satellite
EP4390860A1