Time sequence classification method, system and device based on multi-dimensional and multi-scale space-time spectrum characteristics
By using the multi-dimensional and multi-scale spatiotemporal spectral characteristics in remote sensing time series data processing, data information is enhanced and feature extraction module is constructed, which solves the problem of difficulty in extracting spatiotemporal information in the prior art and achieves higher classification accuracy.
Patent Information
- Application Number
- CN202510063693.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-30
AI Technical Summary
When the prior art processes satellite image time series data in complex and large-scale areas, it is difficult to extract rich spatio-temporal information, and it is difficult to synergistically utilize information from different scales and dimensions, resulting in a decrease in classification accuracy.
The time series classification method based on multi-dimensional and multi-scale spatiotemporal spectral features is adopted. By obtaining remote sensing time series data, time, space and spectral information is enhanced, local feature extraction modules and global feature extraction modules are constructed, multi-dimensional and multi-scale spatiotemporal spectral features are extracted, and multi-layer perceptrons are classified.
The spatiotemporal spectral information in remote sensing time series data is effectively extracted, which improves classification accuracy, especially in complex dynamic scenarios, and can better use spatiotemporal spectral composite features in coordinated use.
Smart Images

Figure CN120070957A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing data processing, and particularly to a temporal classification method, system and device based on multi-dimensional and multi-scale spatio-temporal spectral features. Background Art
[0002] Satellite image time series (SITS) classification involves classifying the land cover of a specific area using a series of images acquired at different time points. Given the inherent dynamic changes and temporal correlations in time series images, SITS has been widely used in remote sensing tasks, including crop type identification, growth monitoring, large-scale geographical surveys, and disaster detection and assessment. However, traditional time series classification tasks mainly focus on SITS data arranged in chronological order, mainly relying on recurrent neural networks (RNNs), especially long short-term memory (LSTM) networks, or Transformer models that adopt self-attention mechanisms. These methods aim to model long-term temporal dependencies to obtain high-accuracy classification results. Although in scenarios with small target areas and simple surface coverages, these methods can effectively extract useful classification features, in regions with high spatio-temporal dynamics, large application scales, or significant cross-sensor variability, it is difficult for them to extract key features from complex data. Therefore, in the field of time series classification, it has become increasingly important to develop algorithms that can extract rich information from SITS data and simultaneously consider large-scale and multi-dimensional features. To address this challenge, researchers have emphasized the importance of hybrid neural networks, which combine the advantages of different neural architectures and aim to effectively utilize multi-dimensional SITS data of various scales.
[0003] To systematically and strategically address the challenges faced in achieving accurate classification using SITS data in complex, large-scale regions, the following two core requirements must be re-examined: Difficult-to-extract raw SITS data information: Compared with single remote sensing images, SITS data contains richer temporal, spectral, and positional information. Effectively extracting this information is a major challenge. If large-scale SITS data is directly input into a deep learning network, it is inefficient and non-optimal because this usually leads to the accumulation of data noise, increases the risk of model overfitting, and severely consumes computing resources; Difficulties in synergistically utilizing information at different scales and dimensions: Classification networks focusing on a single scale or dimension usually have difficulty achieving accurate land cover classification. In complex regions, land elements and surface coverages often exhibit high spatio-temporal dynamics, which may reduce the model performance. Collaboratively using spatio-temporal spectral composite features at global and local scales is crucial for improving classification accuracy in complex dynamic scenarios. Summary of the Invention
[0004] The present invention aims at the disadvantages in the prior art and provides a time series classification method, system and device based on multi-dimensional and multi-scale spatio-temporal spectral features.
[0005] To solve the above technical problems, the present invention is solved by the following technical solutions: A time series classification method based on multi-dimensional and multi-scale spatio-temporal spectral features, comprising the following steps: Obtain remote sensing time series data, where the remote sensing time series data includes time information, spatial information and spectral information; Based on the length of the time series and the change of pixels within the length of the time series, obtain a time influence factor, enhance the time information based on the time influence factor, and then perform normalization processing to obtain enhanced time features; Encode the remote sensing time series data through a position encoder to obtain a position encoding, and then enhance the spatial information to obtain enhanced spatial features; Based on a linear transformation, scale the dimension of the spectral information to a preset dimension to obtain enhanced spectral features; Construct a local feature extraction module, and the local feature extraction module respectively extracts local features from the enhanced time features, enhanced spatial features and enhanced spectral features to obtain multi-dimensional and multi-scale spatio-temporal spectral features corresponding to the enhanced time features, enhanced spatial features and enhanced spectral features; Construct a global feature extraction module, and the global feature extraction module associates the multi-dimensional and multi-scale spatio-temporal spectral features with the position encoding to obtain encoded features, processes the encoded features to obtain spatio-temporal spectral comprehensive features, and adjusts the channel dimension of the spatio-temporal spectral comprehensive features based on a multi-layer perceptron as a classification head to make the classification categories and quantities match each other.
[0006] As an implementable manner, the time influence factor is expressed as follows:
[0007] Wherein, represents the length of the time series, represents the change of pixels at different time points, , represents the pixel value at time represents the pixel value at time The normalization processing is expressed as:
[0008] Wherein, represents the normalization result, represents the time influence factor.
[0009] As an implementable manner, encoding the remote sensing time series data through a position encoder, and then enhancing the spatial information to obtain enhanced spatial features, includes the following steps: Encoding the two-dimensional coordinates of the remote sensing time series data based on a sine function position encoder and a cosine function position encoder to obtain position encoding, and determining the relative spatial relationship between the remote sensing time series data through the position encoding; Taking the position encoding as deep spatial features, and then extracting features through residual connection processing to obtain enhanced spatial features; The position encoding is obtained through the following method:
[0010]
[0011] wherein, represents the coordinate component corresponding to the position corresponding coordinate component, represents the size of the hidden layer, and respectively represent the corresponding sine parity check dimension index and cosine parity check dimension index.
[0012] As an implementable manner, the linear transformation is expressed as follows:
[0013] wherein, represents enhanced spectral information, represents spectral information, represents transpose, represents a weight matrix.
[0014] As an implementable manner, the local feature extraction module includes a linear amplification layer, a first 1×1 convolution module, a plurality of 3×3 convolution modules, and a second 1×1 convolution module; The linear amplification layer matches the dimensions of the enhanced time features, enhanced spatial features, and enhanced spectral features to obtain first features with a unified dimension; The first 1×1 convolution module reduces the first features from the first dimension to the second dimension to obtain second features, and performs batch normalization processing and ReLU activation on the second features to obtain an output result; A plurality of 3×3 convolution modules extract features from the output result to obtain multi-scale local features, and connect the multi-scale local features with the original features to form combined features; The second 1×1 convolution module reduces the dimensions of the combined features to obtain output features; Residual connection and ReLU activation are performed on the enhanced temporal features, enhanced spatial features, and enhanced spectral features with the output features to obtain multi-dimensional and multi-scale local features; The output result is expressed as: ; The multi-scale local features are expressed as: ; The output features are expressed as: ; The multi-dimensional and multi-scale local features are expressed as:
[0015] Among them, represents the second feature, represents the output result, represents the multi-scale local features, represents the combined features, represents the output features, represents the multi-dimensional and multi-scale local features.
[0016] As an implementable manner, the global feature extraction module includes multiple layers of encoders; The multi-dimensional and multi-scale local features are associated with the positional encoding to obtain the encoded positioning features, and the encoded positioning features are expressed as: , where B represents the batch size, L represents the sequence length, and d represents the feature dimension; The encoded positioning features and the relative position information are input into the multiple layers of encoders for multi-layer encoding to obtain the encoded features; The temporal features, spatial features, and spectral features of the encoded features are respectively used as query embeddings , key embeddings and value embeddings to obtain the cross-feature global attention; The cross-feature global attention is normalized and averaged on the preset sequence dimension to obtain the spatio-temporal-spectral comprehensive features; The cross-feature global attention is expressed as:
[0017] The spatio-temporal-spectral comprehensive features are expressed as:
[0018] Among them, , represents the temporal feature, represents the spatial feature, represents the spectral feature.
[0019] As an implementable manner, inputting the encoded positioning feature and the relative position information into a multi-layer encoder for multi-layer encoding to obtain an encoded feature includes the following steps: Each layer of the encoder performs a linear transformation on the encoded positioning feature through a weight matrix , and to respectively generate a query matrix Q, a key matrix K, and a value matrix V, that is , , where represents the dimension of each attention head, and the attention weight is calculated through the dot product of the query matrix Q and the key matrix K, expressed as: ; Split the query, key, and value matrices into multiple attention heads, calculate based on each attention head respectively, obtain the attention results, and connect each attention result to obtain the total attention, expressed as: , where , represents the number of attention heads, represents the linear transformation matrix; Combine the total attention with the multi-dimensional multi-scale local features to obtain a combined feature, and perform residual connection and batch normalization processing on the combined feature to obtain an encoded feature; The combined feature is expressed as:
[0020] The encoded feature is expressed as:
[0021] Where is a form of the transformed feed-forward neural network, including two linear transformations and a non-linear activation function, , represents the input feature vector; respectively represent the weight matrix of the first linear transformation and the weight matrix of the second linear transformation, respectively represent the bias vector of the first linear transformation and the bias vector of the second linear transformation.
[0022] A time series classification system based on multi-dimensional multi-scale spatio-temporal spectral features includes a data acquisition module, a feature enhancement module, a local feature extraction module, and a global classification module; The data acquisition module is used to acquire remote sensing time series data, where the remote sensing time series data includes time information, space information, and spectral information; The feature enhancement module obtains a time influence factor based on the length of the time series and the changes in pixels within the length of the time series, enhances the time information based on the time influence factor, and then performs normalization processing to obtain enhanced time features; encodes the remote sensing time series data through a position encoder to obtain position encoding, and then enhances the spatial information to obtain enhanced spatial features; scales the dimension of the spectral information to a preset dimension based on a linear transformation to obtain enhanced spectral features; The local feature extraction module is used to construct a local feature extraction module. The local feature extraction module respectively performs local feature extraction on the enhanced time features, enhanced spatial features, and enhanced spectral features to obtain multi-dimensional and multi-scale spatio-temporal spectral features corresponding to the enhanced time features, enhanced spatial features, and enhanced spectral features; The global classification module is used to construct a global feature extraction module. The global feature extraction module associates the multi-dimensional and multi-scale spatio-temporal spectral features with the position encoding to obtain encoded features, processes the encoded features to obtain spatio-temporal spectral comprehensive features, and adjusts the channel dimension of the spatio-temporal spectral comprehensive features based on a multi-layer perceptron as a classification head to make the classification categories and quantities match each other.
[0023] A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the following method is implemented: Obtain remote sensing time series data, where the remote sensing time series data includes time information, spatial information, and spectral information; Based on the length of the time series and the changes in pixels within the length of the time series, obtain a time influence factor, enhance the time information based on the time influence factor, and then perform normalization processing to obtain enhanced time features; Encode the remote sensing time series data through a position encoder to obtain position encoding, and then enhance the spatial information to obtain enhanced spatial features; Scale the dimension of the spectral information to a preset dimension based on a linear transformation to obtain enhanced spectral features; Construct a local feature extraction module. The local feature extraction module respectively performs local feature extraction on the enhanced time features, enhanced spatial features, and enhanced spectral features to obtain multi-dimensional and multi-scale spatio-temporal spectral features corresponding to the enhanced time features, enhanced spatial features, and enhanced spectral features; Construct a global feature extraction module. The global feature extraction module associates the multi-dimensional and multi-scale spatio-temporal spectral features with the position encoding to obtain encoded features, processes the encoded features to obtain spatio-temporal spectral comprehensive features, and adjusts the channel dimension of the spatio-temporal spectral comprehensive features based on a multi-layer perceptron as a classification head to make the classification categories and quantities match each other.
[0024] A time series classification device based on multi-dimensional and multi-scale spatio-temporal spectral features, comprising a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, the following method is implemented: Obtain remote sensing time series data, where the remote sensing time series data includes time information, spatial information, and spectral information; Based on the length of the time series and the changes of pixels within the length of the time series, obtain a time influence factor, enhance the time information based on the time influence factor, and then perform normalization processing to obtain enhanced time features; Encode the remote sensing time series data through a position encoder to obtain a position encoding, and then enhance the spatial information to obtain enhanced spatial features; Based on a linear transformation, scale the dimension of the spectral information to a preset dimension to obtain enhanced spectral features; Construct a local feature extraction module, and the local feature extraction module respectively extracts local features from the enhanced time features, enhanced spatial features, and enhanced spectral features to obtain multi-dimensional and multi-scale spatio-temporal spectral features corresponding to the enhanced time features, enhanced spatial features, and enhanced spectral features; Construct a global feature extraction module, and the global feature extraction module associates the multi-dimensional and multi-scale spatio-temporal spectral features with the position encoding to obtain encoded features, processes the encoded features to obtain spatio-temporal spectral comprehensive features, and adjusts the channel dimension of the spatio-temporal spectral comprehensive features based on a multi-layer perceptron as a classification head to make the classification categories and quantities match each other.
[0025] Due to the adoption of the above technical solutions, the present invention has significant technical effects: The present invention proposes a spatio-temporal spectral residual converter network, namely the method involved in this application. The spatio-temporal spectral information is enhanced through three independent channels, and then local feature extraction modules and global feature extraction modules composed of a residual network and a self-attention mechanism are used to extract features. Pixel coordinates, spectral band information, and time changes are extracted from the original remote sensing time series data to obtain enhanced time features, enhanced spatial features, and enhanced spectral features. Res2Net is used to divide the features into multiple sub-feature groups to capture local features and form spatio-temporal spectral composite features.
[0026] In summary, from the data perspective, information representation is enhanced in the information enhancement module. The present invention has three different enhancement channels to specifically enhance the information in the time, space, and spectral dimensions. By calculating the OPE from the coordinates of the original SITS pixels, extracting high-dimensional spectral features from the spectral bands, and deriving time factors from the time series, the spatio-temporal spectral information in the SITS data is maximally utilized, thereby improving the information representation; From the perspective of the model, the process of feature extraction is improved by combining the advantages of the residual network and the self-attention mechanism, ensuring focus on local features while incorporating OPE to capture long-range dependencies, enabling the extraction of comprehensive features of spatio-temporal spectral information at both global and local scales from the enhanced multi-dimensional representation. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0028] Figure 1 is the overall flowchart of the method of the present invention; Figure 2 is the overall structural diagram of the system of the present invention; Figure 3 is the overall flowchart of a specific embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0029] The following further elaborates on the present invention in conjunction with embodiments. The following embodiments are explanations of the present invention, and the present invention is not limited to the following embodiments.
[0030] Embodiment 1: A time series classification method based on multi-dimensional multi-scale spatio-temporal spectral features, as Figure 1 shown, includes the following steps: S100. Obtain remote sensing time series data, where the remote sensing time series data includes time information, spatial information, and spectral information; S200. Based on the length of the time series and the changes in pixels within the length of the time series, obtain a time influence factor, enhance the time information based on the time influence factor, and then perform normalization processing to obtain enhanced time features; S300. Encode the remote sensing time series data through a position encoder to obtain a position encoding, and then enhance the spatial information to obtain enhanced spatial features; S400. Scale the dimension of the spectral information to a preset dimension based on a linear transformation to obtain enhanced spectral features; S500. Construct a local feature extraction module, and the local feature extraction module respectively extracts local features from the enhanced time features, enhanced spatial features, and enhanced spectral features to obtain multi-dimensional multi-scale spatio-temporal spectral features corresponding to the enhanced time features, enhanced spatial features, and enhanced spectral features; S600. Construct a global feature extraction module. The global feature extraction module correlates multi-dimensional and multi-scale spatio-temporal spectral features with positional encoding to obtain encoded features, processes the encoded features to obtain spatio-temporal spectral comprehensive features, and adjusts the channel dimension of the spatio-temporal spectral comprehensive features based on a multi-layer perceptron as a classification head to make the classification categories and quantities match each other.
[0031] Remote sensing time series data is the SITS data mentioned in the background technology of the present invention. A time series classification method based on multi-dimensional and multi-scale spatio-temporal spectral features is the core content involved in the spatio-temporal spectral residual transformer network mentioned in this application. The method based on the present invention also adjusts the classification categories and quantities, and finally obtains spatio-temporal spectral comprehensive features with rich details and interdependent relationships between three different feature dimensions.
[0032] Different natural land cover types usually exhibit different seasonal or phenological characteristics, which change over time and thus form the basis for their identification, which can be called time features. Therefore, enhancing the representation of time features is crucial for improving classification accuracy. Here, a time influence factor is introduced to quantify the time variation of pixel values. The time influence factor is expressed as follows:
[0033] where, represents the length of the time series, represents the change of pixels at different time points, , represents the pixel value at time represents the pixel value at time To ensure effective model processing, the time influence factor It is normalized to the range of 0 to 1 through the following regularization:
[0034] where, represents the normalization result, represents the time influence factor.
[0035] This approach can map the changes of different object types over the entire time series, enhance the ability to capture time features, and thus improve the accuracy of time classification.
[0036] Encoding the remote sensing time series data through a position encoder to enhance the spatial information and obtain enhanced spatial features includes the following steps: Encode the two-dimensional coordinates of remote sensing time series data based on a sine function position encoder and a cosine function position encoder to obtain position encoding, and determine the relative spatial relationship between remote sensing time series data through the position encoding; Take the position encoding as a deep spatial feature, and extract features through residual connection processing to obtain enhanced spatial features; The position encoding is obtained through the following method:
[0037]
[0038] wherein, represents the coordinate component corresponding to the position corresponding coordinate component, represents the size of the hidden layer, and respectively represent the corresponding sine parity check dimension index and cosine parity check dimension index.
[0039] Spatial texture information has been crucial in the SITS data classification task for a long time. The spatial relationship between pixels plays an equally important role in learning the correlation between adjacent pixels and understanding the spatial distribution pattern of pixels on different features. The current self-attention mechanism mainly encodes the absolute position relationship after the input enters the model, without considering the original relative position relationship of the sequence images along the time axis. Therefore, enhancing the spatial position relationship between pixels is crucial for improving the classification accuracy. Compared with a single remote sensing image, SITS data not only involves the absolute position relationship between pixels in the same view, but also involves a unique relative position association along the time series dimension. Therefore, the present invention uses a position encoder based on sine and cosine functions to encode the two-dimensional coordinates of the original SITS data, effectively capturing the relative spatial relationship between SITS data. This encoding is represented as the relative position encoding of the original SITS data, acting as a deep spatial feature, and then integrated into the feature extractor through residual connection to obtain enhanced spatial features. By generating unique sine and cosine position encodings for the original coordinates of each pixel, the spatial information is enhanced. It can ensure that the model effectively captures spatial texture features while emphasizing the inherent relative spatial dependence in SITS data.
[0040] The spectral information of different natural land cover types usually exhibits significant variations. Therefore, capturing the spectral differences between bands and their change patterns is crucial for land cover classification. Since remote sensing time series data is pixel-based, represented as a tensor with time and spectral dimensions and containing multiple spectral bands, it is inherently multi-dimensional, different from the traditional word embeddings used in natural language processing. To adapt to this complexity and enable the model to better utilize spectral information, a linear transformation layer is applied to scale the data features to a specified dimension Cin. This can more effectively utilize the rich details contained in multiple spectral bands of SITS data. Here, the linear transformation in the enhanced spectral features is represented as follows:
[0041] where, represents the enhanced spectral information, represents the spectral information, represents the transpose, represents the weight matrix.
[0042] Residual networks are particularly effective in mining local information in deep neural networks. By integrating skip connections that enable the network to directly learn residuals, the training efficiency and overall model performance are improved. Among these architectures, Res2Net is good at simultaneously capturing fine-grained local features in different receptive fields, further enhancing the network's ability to express multi-scale local information. The present invention improves the residual network and constructs a local feature extraction module, including a linear amplification layer, a first 1×1 convolution module, multiple 3×3 convolution modules, and a second 1×1 convolution module; The linear amplification layer matches the dimensions of the enhanced time feature, enhanced spatial feature, and enhanced spectral feature to obtain a first feature with a unified dimension; The first 1×1 convolution module reduces the first feature from the first dimension to the second dimension to obtain a second feature, performs batch normalization processing and ReLU activation on the second feature to obtain an output result; Multiple 3×3 convolution modules perform feature extraction on the output result to obtain multi-scale local features, connect the multi-scale local features with the original features to form a combined feature; The second 1×1 convolution module reduces the dimension of the combined feature to obtain an output feature; Perform residual connection and ReLU activation on the enhanced time feature, enhanced spatial feature, and enhanced spectral feature with the output feature to obtain multi-dimensional multi-scale local features; The output result is represented as: ; The multi-scale local feature is represented as: is represented as: ; The output feature is represented as: ; The multi-dimensional and multi-scale local features are expressed as:
[0043] Wherein, represents the second feature, represents the output result, represents the multi-scale local features, represents the combined features, represents the output features, represents the multi-dimensional and multi-scale local features. At this stage, Res2Net subdivides the spatio-temporal spectral features into multiple sub-feature groups, enabling each sub-feature group to interact through a cascading operation and perform convolutions within itself to gradually capture local details.
[0044] In one embodiment, the global feature extraction module includes multiple layers of encoders; each layer of the encoder consists of a self-attention mechanism and a feed-forward network. The self-attention block originated from the Transformer architecture and mainly utilizes the multi-head attention mechanism. Different from CNNs and RNNs that may rely on various forms of sequence encoding, the self-attention mechanism usually adopts absolute position encoding to capture the sequential position information of elements.
[0045] The multi-dimensional and multi-scale local features are associated with the position encoding to obtain the encoded location features, and the encoded location features are expressed as: , where B represents the batch size, L represents the sequence length, and d represents the feature dimension; The encoded location features and the relative position information are input into the multiple layers of encoders for multi-layer encoding to obtain the encoded features; The temporal feature, spatial feature, and spectral feature of the encoded features are respectively used as query embeddings , key embeddings and value embeddings to obtain the cross-feature global attention; The cross-feature global attention is normalized and averaged over the preset sequence dimension to obtain the spatio-temporal-spectral comprehensive feature; The cross-feature global attention is expressed as:
[0046] The spatio-temporal-spectral comprehensive feature is expressed as:
[0047] Wherein, , represents the temporal feature, represents the spatial feature, represents the spectral feature.
[0048] In one embodiment, the step of inputting the encoded positioning feature and the relative position information into a multi-layer encoder for multi-layer encoding to obtain an encoded feature includes the following steps: Each layer of the encoder performs a linear transformation on the encoded positioning feature through a weight matrix , and to respectively generate a query matrix Q, a key matrix K, and a value matrix V, that is, , , where represents the dimension of each attention head, and the attention weight is calculated by the dot product of the query matrix Q and the key matrix K, expressed as: ; The query, key, and value matrices are split into multiple attention heads, and calculations are respectively performed based on each attention head to obtain attention results, and each attention result is concatenated to obtain the total attention, expressed as: , where , represents the number of attention heads, represents the linear transformation matrix; The total attention is combined with the multi-dimensional multi-scale local features to obtain a combined feature, and the combined feature is subjected to residual connection and batch normalization processing to obtain an encoded feature; The combined feature is expressed as:
[0049] The encoded feature is expressed as:
[0050] where is a form of the transformed feed-forward neural network, which includes two linear transformations and a non-linear activation function, , represents the input feature vector; respectively represent the weight matrix of the first linear transformation and the weight matrix of the second linear transformation, respectively represent the bias vector of the first linear transformation and the bias vector of the second linear transformation, is a non-linear activation function that compares the result of the linear transformation with 0 and takes the larger value; The calculation process of the entire FFN can be divided into the following steps: The first linear transformation: Multiply the input feature vector by the weight matrix of the first linear transformation and add the bias vector of the first linear transformation to obtain the result of the linear transformation.
[0051] Non-linear activation: Apply the ReLU function to the result of the linear transformation to introduce non-linearity; The second linear transformation: Multiply the result of the non-linear activation by the weight matrix of the second linear transformation and add the bias vector of the second linear transformation to obtain the final output.
[0052] This application can be understood as two major modules. As shown in Figure 3, the first module is the spatio-temporal-spectral information enhancement module: Recognizing the complexity and richness of the mixed information in SITS data, three independent enhancement channels are adopted to target the unique features of the time, space, and spectral dimensions. The aim is to facilitate the extraction of deep spatio-temporal-spectral information and lay a solid foundation for capturing comprehensive local and global features in subsequent stages.
[0053] The second module is the global-local feature extraction module: To effectively integrate the enhanced spatio-temporal-spectral information, a global-local feature extraction module is constructed based on the Res2Net residual network and the self-attention mechanism of ViT, enabling the model to extract multi-scale spatio-temporal-spectral composite features. Specifically, the enhanced spatio-temporal-spectral representation, including time factors, pixel coordinates, and spectral information, is first input into the Res2Net module as three different dimensions. Then, the generated local features are combined with the OPE residuals and passed into the self-attention mechanism. Finally, the multi-scale spatio-temporal-spectral composite features are input into the MLP classification head to generate the classification results.
[0054] In a specific embodiment, relevant experiments were conducted using several datasets and several different models. The datasets used were the TiSeLaC dataset, the TiSeLaC dataset, and the TimeSen2Crop dataset. The models used were STM (1997), TempCNN (2023), DuPLO (2019), UNET2D-CLSTM (2019), ImprovedTransformer (2023), UTAE (2021), TSViT (2023), and TSSRes2Former. Among them, TSSRes2Former is the spatio-temporal-spectral residual transformer network, that is, the method of this application. That is to say, the process implemented by the method of this application is the spatio-temporal-spectral residual transformer network.
[0055] The TiSeLaC dataset is a time series land cover classification dataset. The TiSeLaC dataset contains data from Reunion Island. Each case is a pixel, and the measurement data is taken at 23 time points (days) and has 10 dimensions: 7 surface reflectances (ultra-blue, blue, green, red, near-infrared, short-wave infrared 1, and short-wave infrared 2) plus 3 indices (NDVI, NDWI, and BI). The class values are related to 9 land cover types.
[0056] The TimeSen2Crop dataset contains over 1 million Sentinel 2 time series (TS) samples related to 16 crop types. This dataset aims to facilitate research on supervised classification of Sentinel 2 data time series globally for crop type mapping. The dataset includes atmospherically corrected Sentinel 2 images and reports snow, shadow, and cloud information for each labeled unit. To generate the dataset, an openly available Austrian crop type map based on farmer declarations was considered, and an automated procedure based on a multi-temporal deep learning model was defined to extract the training set. The TimeSen2Crop dataset also includes time series of Sentinel 2 images collected during the next agronomic year (i.e., from September 2018 to August 2019). The PASTIS dataset contains 2,433 plots in mainland France, each with panoramic annotations (instance index and semantic label for each pixel). Each plot is a time series of Sentinel-2 multi-spectral images with variable length. Additionally, PASTIS was extended with radar data to form the PASTIS-R dataset for evaluating the application of optical and radar fusion methods in plot classification, semantic segmentation, and panoramic segmentation.
[0057] OA (Overall Accuracy) and Kappa coefficient are two commonly used classification accuracy evaluation metrics, both of which are calculated based on the confusion matrix. The overall classification accuracy (OA) refers to the proportion of correctly classified samples in the total number of samples. The calculation formula is: where, is the number of true positives in the i-th class (i.e., the number of samples correctly classified in this class), is the total number of samples. OA is used to measure the overall performance of the classification model and is a simple and intuitive metric, but it is not sensitive enough to datasets with class imbalance.
[0058] The Kappa coefficient is used to measure the degree of difference between the classification result and the random classification result. Its calculation formula is: , is the observed agreement (i.e., the overall classification accuracy OA), is the expected agreement (i.e., the probability of agreement under random classification). The Kappa coefficient takes into account the agreement due to random factors in the classification result, so it can more accurately reflect the actual performance of the classification model than OA, especially in the case of class imbalance.
[0059] The value of the Kappa coefficient ranges from -1 to 1 and usually falls between 0 and 1. The closer its value is to 1, the higher the agreement; a value of 0 indicates agreement with random classification; a negative value indicates agreement lower than random classification.
[0060] Table 1 shows the comparison of different datasets and different models. It can be seen from Table 1 that the OA / Kappa obtained by using the method of this application is relatively high, which indicates a higher degree of consistency.
[0061] Table 1 shows the results of different datasets and different models
[0062] In addition, in one embodiment, ablation experiments were conducted on the TSSRes2Former, which is a spatio-temporal spectral residual transformer network, and the TiSeLaC dataset, and specific results were given. An ablation experiment is an experimental method used to verify the importance of each component in a model or system. Its main purpose is to observe the impact of these changes on the overall performance by consciously removing or modifying certain parts of the model, so as to understand the specific contributions of each part to the model performance. In this embodiment, ablation experiments were carried out through the TSS Low-dinmesion Features Generation Module, which is a low-dimensional feature generation module, and the Global-local Feature Extraction Module. When the global-local feature extraction module of this application was used, the OA / Kappa results obtained were very ideal, which indicates a higher degree of consistency.
[0063] Table 2 shows the ablation results without some modules
[0064] It can be seen from the above two tables that the results obtained by the method of this application are better than those obtained by using other models in the prior art.
[0065] Embodiment 2: A time series classification system based on multi-dimensional multi-scale spatio-temporal spectral features, as Figure 2 shown, includes a data acquisition module 100, a feature enhancement module 200, a local feature extraction module 300, and a global classification module 400; The data acquisition module 100 is used to acquire remote sensing time series data, where the remote sensing time series data includes time information, spatial information, and spectral information; The feature enhancement module 200 obtains a time influence factor based on the length of the time series and the changes in pixels within the length of the time series, enhances the time information based on the time influence factor, and then performs normalization processing to obtain enhanced time features; encodes the remote sensing time series data through a position encoder to obtain position encoding, and then enhances the spatial information to obtain enhanced spatial features; scales the dimension of the spectral information to a preset dimension based on a linear transformation to obtain enhanced spectral features; The local feature extraction module 300 is used to construct a local feature extraction module. The local feature extraction module respectively performs local feature extraction on the enhanced temporal feature, enhanced spatial feature, and enhanced spectral feature to obtain multi-dimensional and multi-scale spatio-temporal spectral features corresponding to the enhanced temporal feature, enhanced spatial feature, and enhanced spectral feature. The global classification module 400 is used to construct a global feature extraction module. The global feature extraction module associates the multi-dimensional and multi-scale spatio-temporal spectral features with position encoding to obtain encoded features, processes the encoded features to obtain spatio-temporal spectral comprehensive features, and adjusts the channel dimension of the spatio-temporal spectral comprehensive features based on a multi-layer perceptron as a classification head to make the classification categories and quantities match each other.
[0066] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other.
[0067] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, devices, or computer program products. Therefore, the present invention can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0068] The present invention is described with reference to the flowcharts and / or block diagrams of methods, terminal devices (systems), and computer program products according to the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing terminal devices to generate a machine, so that the instructions executed by the processors of the computer or other programmable data processing terminal devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0069] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0070] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, so that a series of operation steps are executed on the computer or other programmable terminal device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable terminal device provide for implementing the steps of the function specified in one process or multiple processes and / or blocks Figure 1 One process or multiple processes and / or blocks Figure 1 The steps of the function specified in one block or multiple blocks.
[0071] It should be noted that: The "one embodiment" or "embodiment" mentioned in the specification means that the specific features, structures or characteristics described in connection with the embodiment are included in at least one embodiment of the present invention. Therefore, the phrases "one embodiment" or "embodiment" that appear throughout the specification do not necessarily all refer to the same embodiment.
[0072] In addition, it should be noted that for the specific embodiments described in this specification, the shapes of their components, the names taken, etc. can be different. Any equivalent or simple changes made according to the structure, features and principles described in the inventive concept of the present invention are included in the protection scope of the present invention. Those skilled in the art to which the present invention pertains can make various modifications or supplements to the described specific embodiments or use similar ways to substitute them, as long as they do not deviate from the structure of the present invention or exceed the scope defined by the claims, they should fall within the protection scope of the present invention.
Claims
1. A time series classification method based on multi-dimensional and multi-scale spatiotemporal spectral features, characterized in that: The following steps are involved: Acquiring remote sensing time series data, wherein the remote sensing time series data includes time information, spatial information and spectral information; Based on the length of the time series and the change of pixels within the length of the time series, a time influence factor is obtained, and the time information is enhanced based on the time influence factor, and then normalized to obtain enhanced time features; The remote sensing time series data is encoded by a position encoder to obtain position coding, and then the spatial information is enhanced to obtain enhanced spatial features; The dimension of the spectral information is scaled to a preset dimension based on a linear transformation to obtain enhanced spectral features; Constructing a local feature extraction module, the local feature extraction module extracts local features of the enhanced time feature, the enhanced space feature and the enhanced spectrum feature respectively, and obtains multi-dimensional and multi-scale spatiotemporal spectral features corresponding to the enhanced time feature, the enhanced space feature and the enhanced spectrum feature; A global feature extraction module is constructed. The global feature extraction module associates the multi-dimensional and multi-scale spatiotemporal spectral features with the position coding to obtain the coding features. The coding features are processed to obtain the spatiotemporal spectral comprehensive features. The channel dimension of the spatiotemporal spectral comprehensive features is adjusted based on the multi-layer perceptron as the classification head so that the classification categories and quantities match each other.
2. The time series classification method based on multi-dimensional and multi-scale spatiotemporal spectral features according to claim 1 is characterized in that: The time impact factor is expressed as follows: in, represents the length of the time series, Represents the change of pixels at different time points, , express The pixel value at the moment, express The pixel value at the moment; The normalization process is expressed as: in, represents the normalized result, Represents the time impact factor.
3. The time series classification method based on multi-dimensional and multi-scale spatiotemporal spectral features according to claim 1 is characterized in that: The method of encoding the remote sensing time series data by a position encoder and then enhancing the spatial information to obtain enhanced spatial features includes the following steps: Encoding the two-dimensional coordinates of the remote sensing time series data based on the sine function position encoder and the cosine function position encoder to obtain the position code, and determining the relative spatial relationship between the remote sensing time series data through the position code; The position encoding is used as the deep spatial feature, and the feature is extracted through residual connection processing to obtain enhanced spatial features; The position code is obtained in the following way: in, Representation and location Corresponding Coordinate components, represents the size of the hidden layer, and They represent the corresponding sine parity dimension index and cosine parity dimension index respectively.
4. The time series classification method based on multi-dimensional and multi-scale spatiotemporal spectral features according to claim 1 is characterized in that: The linear transformation is expressed as follows: in, represents enhanced spectral information, Represents spectral information, represents transpose, represents the weight matrix.
5. The time series classification method based on multi-dimensional and multi-scale spatiotemporal spectral features according to claim 1 is characterized in that: The local feature extraction module includes a linear amplification layer, a first 1×1 convolution module, a plurality of 3×3 convolution modules, and a second 1×1 convolution module; The linear amplification layer matches the dimensions of the enhanced time feature, the enhanced space feature, and the enhanced spectrum feature to obtain a first feature of unified dimension; The first 1×1 convolution module reduces the first feature from the first dimension to the second dimension to obtain the second feature, performs batch normalization and ReLU activation on the second feature to obtain the output result; Multiple 3×3 convolution modules extract features from the output results to obtain multi-scale local features, which are then connected with the original features to form combined features. The second 1×1 convolution module performs dimensionality reduction processing on the combined features to obtain output features; The enhanced temporal features, enhanced spatial features and enhanced spectral features are residually connected and ReLU activated with the output features to obtain multi-dimensional and multi-scale local features; The output result is expressed as: ; The multi-scale local feature is expressed as: ; The output feature is expressed as: ; The multi-dimensional and multi-scale local features are expressed as: in, Represents the second feature, Indicates the output result, Represents multi-scale local features, represents the combined features, represents the output features, Represents multi-dimensional and multi-scale local features.
6. The time series classification method based on multi-dimensional and multi-scale spatiotemporal spectral features according to claim 1 is characterized in that: The global feature extraction module includes a multi-layer encoder; The multi-dimensional and multi-scale local features are associated with the position coding to obtain the coding positioning features, which are expressed as: , where B represents the batch size, L represents the sequence length, and d represents the feature dimension; Inputting the coded positioning features and the relative position information into a multi-layer encoder for multi-layer encoding to obtain coded features; The temporal, spatial and spectral features of the encoded features are respectively embedded as queries , key embedding and value embedding , get global attention across features; The global attention across features is normalized and averaged over a preset sequence dimension to obtain a spatiotemporal comprehensive feature; The global attention across features is expressed as: The comprehensive characteristics of the space-time spectrum are expressed as: in, , Represents time characteristics, Represents spatial features, Represents spectral characteristics.
7. The time series classification method based on multi-dimensional and multi-scale spatiotemporal spectral features according to claim 1 is characterized in that: The step of inputting the coded positioning feature and the relative position information into a multi-layer encoder for multi-layer encoding to obtain the coded feature comprises the following steps: Each layer of encoder encodes the positioning features through the weight matrix , and Perform linear transformation to generate query matrix Q, key matrix K and value matrix V respectively, that is, , ,in, represents the dimension of each attention head. The attention weight is calculated by the dot product of the query matrix Q and the key matrix K, expressed as: ; Split the query, key, and value matrices into multiple attention heads, perform calculations based on each attention head, get the attention results, and concatenate each attention result to get the total attention, expressed as: ,in, , represents the number of attention heads, represents a linear transformation matrix; The total attention is combined with the multi-dimensional and multi-scale local features to obtain a combined feature, and the combined feature is subjected to residual connection and batch normalization processing to obtain a coding feature; The combination feature is expressed as: The coding feature is expressed as: in, It is a form of transformed feedforward neural network, which contains two linear transformations and a nonlinear activation function. , represents the input feature vector; Represent the weight matrix of the first linear transformation and the weight matrix of the second linear transformation respectively, Represent the bias vector of the first linear transformation and the bias vector of the second linear transformation respectively.
8. A time series classification system based on multi-dimensional and multi-scale spatiotemporal spectral features, characterized in that: It includes data acquisition module, feature enhancement module, local feature extraction module and global classification module; The data acquisition module is used to acquire remote sensing time series data, wherein the remote sensing time series data includes time information, spatial information and spectral information; The feature enhancement module obtains a time influence factor based on the length of the time series and the change of pixels within the length of the time series, enhances the time information based on the time influence factor, and then performs normalization processing to obtain enhanced time features; encodes the remote sensing time series data through a position encoder to obtain a position code, and then enhances the spatial information to obtain enhanced spatial features; scales the dimension of the spectral information to a preset dimension based on a linear transformation to obtain an enhanced spectral feature; The local feature extraction module is used to construct a local feature extraction module, which performs local feature extraction on the enhanced time feature, enhanced space feature and enhanced spectrum feature respectively to obtain multi-dimensional and multi-scale spatiotemporal spectral features corresponding to the enhanced time feature, enhanced space feature and enhanced spectrum feature; The global classification module is used to construct a global feature extraction module, which associates multi-dimensional and multi-scale spatiotemporal spectral features with position coding to obtain coding features, processes the coding features to obtain spatiotemporal spectral comprehensive features, and adjusts the channel dimensions of the spatiotemporal spectral comprehensive features based on a multi-layer perceptron as a classification head so that the classification categories and quantities match each other.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
10. A time series classification device based on multi-dimensional and multi-scale spatiotemporal spectral features, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Sequence data time-spectrum radiation unification processing method and system under constraint of reference data
CN121074441A
Method and system for time spectrum radiation homogenization processing of sequence data under reference data constraint
CN121074441B