Multivariate time series prediction method based on GLoSyform
Through the GLoSyformer model, the problem of insufficient dependence and multivariate correlation in medium and long-term dependence and multivariate correlation through the GLoSyformer model is solved by using dual sampling embedding, TimesBlock module and sparse variable attention mechanism, and more efficient prediction performance is achieved, especially significant improvements in high-dimensional and low-dimensional datasets.
Patent Information
- Application Number
- CN202510325446.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-07-22
AI Technical Summary
The existing multivariate time series prediction model has shortcomings in capturing long-term dependencies and multivariate correlations, limiting the improvement of prediction performance, especially in capturing dynamic temporal changes and deep temporal features between different sub-sequences.
The GLoSyformer model is adopted to enhance the time feature representation of variable tokens through dual sampling embedding technology. The TimesBlock module is designed to capture global and local time dependencies, and through adaptive gating fusion units and sparse variable attention mechanisms, flexible fusion of global and local time features and effective capture of multivariate correlations.
It improves the accuracy and reliability of multivariate time series prediction, significantly improves prediction performance in multiple fields, especially prediction results on high-dimensional and low-dimensional datasets, reduces interference from irrelevant variables, and enhances the ability to capture complex time patterns.
Smart Images

Figure CN120354886A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of deep learning and time series prediction, and particularly relates to a multivariate time series prediction method based on GLoSyformer. Background Art
[0002] Multivariate time series (MTS) prediction has been widely applied in various fields such as finance, energy, weather, and transportation. Its core objective is to accurately predict the change trend in a future period by deeply analyzing historical observation data, thereby providing important references for people's decision-making. With the development of deep learning technology, many advanced models have been proposed and achieved remarkable results in MTS prediction. Nowadays, MTS prediction has become a research hotspot jointly concerned by academia and industry.
[0003] MTS prediction models can be roughly divided into statistical models and deep learning models. As typical statistical models, the vector autoregressive (VAR) model and the autoregressive integrated moving average (ARIMA) model have played important roles in past MTS prediction models. However, with the development of deep learning, these models have gradually been replaced by deep network prediction models with better prediction performance. In deep learning methods, prediction models based on graph neural networks (GNNs) make predictions by capturing the spatial dependencies in MTS. Prediction models based on recurrent neural networks (RNNs) have excellent time-dependent modeling capabilities. SegRNN divides the input sequence into several segments to improve training efficiency and adopts a parallel multi-step prediction strategy to enhance prediction performance. Prediction models based on convolutional neural networks (CNNs) use temporal convolutions to extract sequence features. MICN utilizes convolutions to extract the local features of time series and the correlations between local features, and then obtains global features. TimesNet converts one-dimensional time series into two-dimensional space and captures multi-period features through convolutions. However, although the above models have achieved remarkable results, they still face challenges in capturing long-term dependencies, which to a certain extent restricts the further improvement of their prediction performance.
[0004] The Transformer architecture was initially applied in the field of natural language processing (NLP) and later extended to other fields such as computer vision (CV). Due to its great potential in capturing long-term dependencies, the Transformer has become the dominant architecture for time series prediction in recent years. For the MTS prediction task, in addition to considering the dependencies in time, the correlations between variables are also crucial. In traditional Transformer-based prediction models, multiple variables at the same time step are usually embedded into indistinguishable feature channels, and the time dependencies are captured by processing these time tokens. Although this method performs well in capturing time dependencies, embedding multiple variables into a single feature vector at the same time makes it difficult for the model to capture the correlations between variables, thus limiting its prediction performance. In addition, the time tokens formed by a single time step have been proven to be difficult to reveal more useful information. To improve the effectiveness of prediction and solve the problem of capturing the correlations of multiple variables, iTransformer proposed an innovative series-wise tokenization method. It regards the entire sequence of each variable as an independent variable token, and captures the correlations between variables by using the self-attention mechanism on the variable tokens. Although iTransformer has made significant progress in capturing the correlations of multiple variables, its special embedding method limits the model's ability to capture the temporal dynamics between different subsequences. In addition, the variable tokens formed by simple linear projection are limited to shallow temporal feature representations, which makes it challenging for the model to extract deeper temporal features and learn complex temporal patterns.
[0005] Inspired by these works, the present invention proposes a global-local collaborative enhancement strategy to enhance the temporal features of variable tokens from both global and local perspectives. While ensuring the capture of the correlations of multiple variables, it better captures the time dependencies. Summary of the Invention
[0006] The present invention aims to study a method for multivariate time series prediction, which is applied in multiple fields such as finance, energy, weather, and transportation. By deeply analyzing historical observation data, it accurately predicts the change trend in a future period of time, thus providing an important reference for people's decision-making. This method uses the GLoSyformer model to predict multivariate time series and compares the prediction results with those of other models.
[0007] The specific implementation of the invention is as follows:
[0008] (1) The reversible instance normalization RevIN technology is adopted to process the training data, reduce the non-stationarity of the sequence, and improve the accuracy and robustness of the model.
[0009] (2) From a global and local perspective, a dual-sampling embedding technique is adopted to enhance the representation of variable token time features. Among them, down-sampling embedding is used to enhance global time information, and segmented sampling embedding is used to enhance local time information.
[0010] (3) Design the TimesBlock module to capture global and local time dependencies respectively. This module uses Cross-Attention to centrally process the enhanced time features to capture global and local time dependencies, and further extracts and enhances the representations of global and local time features for the next-layer encoder.
[0011] (4) To better fuse the captured global and local time features, an adaptive gating fusion unit is designed, which can assign appropriate weights to the information at each time step to achieve flexible and adaptive fusion of global and local time features.
[0012] (5) Multivariate correlation is crucial in MTS prediction, but not all variables have effective correlation information. To avoid the interference of irrelevant variables, a sparse variable attention mechanism is proposed. It constructs a correlation matrix by calculating the Pearson correlation coefficient between variables and introduces this matrix externally to sparse the attention scores, thereby reducing the impact of irrelevant variables on the prediction results.
[0013] (6) In the training stage, GLoSyformer is used to predict multivariate time series. The dataset is indexed in chronological order and divided into training set, validation set, and test set according to the ratio of 6:2:2. Determine the input length and prediction length of the model. After training, the model can fully learn the complex dependencies in the sequence and improve the prediction accuracy.
[0014] (7) Compare GLoSyformer with other time series prediction methods on seven real-world datasets, test the performance of these methods in prediction, and compare the predicted values of the prediction model with the true values. Description of the Drawings
[0015] The drawings are used to provide further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention.
[0016] Figure 1 . The overall structure of GLoSyformer of the present invention
[0017] Figure 2 . The down-sampling, segmented sampling and sequence autocorrelation methods of the present invention
[0018] Figure 3. The TimesBlock module of the present invention captures global and local time features
[0019] Figure 4 . The adaptive gating fusion unit of the present invention
[0020] Figure 5 . The sparse attention block of the present invention
[0021] Figure 6 . The prediction results of the GLoSyformer of the present invention and all baselines on eight datasets
[0022] Figure 7 . The data fitting graph of the present invention Detailed implementation manners
[0023] The following combines the attached Figure 1-7 , and makes a further detailed description of the present invention. The specific implementation steps are as follows:
[0024] Step 01. Adopt the Reversible Instance Normalization (RevIN) method to normalize the training data and perform the inverse normalization operation in the prediction stage. Specifically, first calculate the statistical features of the input time series, including the mean and standard deviation, etc., and normalize the time series based on this statistical information to weaken the non-stationarity of the series and improve the stability of the data. Subsequently, after the model generates the prediction results, use the original statistical features to inverse normalize the output data to restore the original scale of the data and ensure the physical interpretability of the prediction results.
[0025] Step 02. From the global and local perspectives, adopt double-sampling embedding to enhance the representation of the variable token time features. Specifically, the double-sampling enhanced embedding layer includes three different embedding methods, namely variable embedding, down-sampling embedding, and segment-sampling embedding. Variable embedding only performs a linear dimension expansion on the original sequence, and the variable token only contains the shallow and coarse-grained time information of the sequence; down-sampling embedding samples the original sequence at intervals, combines the information of time steps with a relatively large interval into a subsequence, and each subsequence contains certain global time information. Subsequently, each subsequence is embedded to enhance the representation of the global time features of the sequence; segment-sampling embedding converts the original sequence into continuous subsequences of the same length. Each subsequence contains the local time information within a period of historical time. Subsequently, each subsequence is embedded to enhance the representation of the local time features of the sequence.
[0026] Variable embedding aims to take the entire time series of each variable as a token and independently embed it into each variable token through a temporal linear projection. It can be expressed as:
[0027] Vi = VarEmbedding(x (i) )
[0028] where the variable embedding layer VarEmbedding: R L -> R D , represents a trainable linear projection, where V i represents the variable token obtained after the embedding process of the i-th variable sequence.
[0029] In the present invention, the input sequence of length L is sampled at intervals of step size K, and the information of time steps with a large interval is combined together to form a subsequence. Such a subsequence contains certain global time information. After multiple samplings, the obtained subsequences are put together, and each subsequence is projected onto a time token. The downsampling embedding can be expressed as:
[0030]
[0031] where, is the number of subsequences, is the j-th subsequence in the downsampled sequence. The downsampling embedding layer DsEmbedding maps each subsequence to D dimensions through a trainable linear projection and a positional embedding.
[0032] In the present invention, the input sequence of length L is sampled continuously and non-overlappingly according to the window size W, and the data of each window is used as a subsequence. Each subsequence contains local information within a certain historical time period. Multiple subsequences are combined together, and each subsequence is projected onto a time token. The segmented sampling embedding can be expressed as:
[0033]
[0034] where N w is the number of subsequences, and the length of each subsequence is W, is the j-th subsequence in the segmented sampling sequence. The segmented sampling embedding layer PsEmbedding maps each subsequence of length W to D dimensions through a trainable linear projection and a positional embedding.
[0035] Step 03. As Figure 3As shown, the TimesBlock consists of a Global TimesBlock and a Local TimesBlock with the same network structure. In the Global TimesBlock, the global temporal information of the variable tokens is enhanced by integrating the downsampled embedding vectors into the variable tokens. Cross-Attention uses the original variable tokens as queries, the enhanced tokens as keys and values, and establishes the connection between the two types of tokens. By integrating the global temporal information of the enhanced tokens into the variable tokens, the global temporal dependencies are captured. In addition, to further enhance the global temporal features contained in the downsampled embedding vectors, we capture the global temporal relationships within the subsequences through one-dimensional convolution. The above process can be expressed as follows.
[0036]
[0037] In the formula, Norm is normalization, and l is the l-th layer encoder in the GLoSyformer. [,] represents concatenation along the sequence dimension. After Cross-Attention and normalization, we obtain It is a vector with the ability to represent global temporal features and will subsequently enter the adaptive gating fusion unit for fusion. The downsampled embedding vector Enters the next layer encoder after convolution to enhance the global temporal information of the next layer.
[0038] In the Local TimesBlock, the segmented sampled embedding vectors are integrated into the variable tokens to enhance the local temporal information of the variable tokens. Subsequently, the same processing method as in the Global TimesBlock is adopted to capture the local temporal dependencies and enhance the local temporal features contained in the segmented sampled embedding vectors. The above can be expressed as follows.
[0039]
[0040] In the formula, Is a vector with the ability to represent local temporal features and will subsequently enter the adaptive gating fusion unit for fusion. The segmented sampled embedding vector Enters the next layer encoder after convolution to enhance the local temporal information of the next layer.
[0041] Step 04. Due to the complexity of the time series, the sequences at different times exhibit different distributions, and the contributions of the global and local features to the prediction results will also vary greatly. To learn this contribution and avoid the loss of local details or the over-amplification of global information caused by simple weighted fusion. The present invention designs an adaptive gating fusion unit, as shown in Figure 4As shown. The gating unit can learn weight allocation according to the content of the input features, making the fusion process more sensitive to the differences of different features, so as to flexibly adapt to different dependencies changing with time in the sequence. The fusion process can be expressed as.
[0042]
[0043] In the formula, [;] represents concatenation along the embedding dimension, and σ is the Sigmoid activation function. The concatenated vector generates a weight vector g after passing through a linear transformation and the Sigmoid function. The time vector adapts to fuse the global and local temporal features according to the weight vector g and outputs the fused feature vector so as to effectively cope with the complex temporal patterns in the time series.
[0044] Step 05. As Figure 5 shown, the Sparse Variable Attention (SVA) reduces the interference of irrelevant variables by introducing a Pearson correlation weight matrix between external variables to sparsify the attention scores. The calculation process of the Pearson correlation coefficient between variables is as follows:
[0045]
[0046] In the formula, represents the value of the i-th variable at the t-th time step, represents the mean value of the i-th variable over the entire time series L. The present invention sets a hyperparameter τ to control the construction of the correlation matrix R. When |r i,j | >= τ, r i,j = r i,j ; when |r i,j | < τ, r i,j = -∞. The matrix R forms a Pearson correlation weight matrix W r through the softmax function for sparsifying the attention scores. The formula of SVA can be expressed as.
[0047] W r = Softmax(R)
[0048]
[0049] The present invention concatenates the variable tokens of multiple variables that have fused the global and local temporal features along the variable dimension to form and captures the dependencies between variables through sparse variable attention. After residual connection and normalization, the output V l+1 is obtained, which will enter the next layer of the encoder after passing through the subsequent FFN layer. The specific process is as follows.
[0050]
[0051]
[0052] Step 06. The present invention uses the mean squared error (MSE) as the loss function to measure the difference between the predicted value and the true value. The MSE loss is expressed as:
[0053]
[0054] where represents the true value of variable i in the next T time steps, and represents the predicted value of variable i in the next T time steps.
[0055] Step 07. To evaluate the performance of GLoSyformer, we conducted experiments on eight real-world datasets, including: ETTh1, ETTm1, Exchange, Weather, Solar (solar power generation of a photovoltaic power station), Electricity, Traffic, and Drift (drift trajectory of an autonomous underwater vehicle). The detailed information of the datasets is shown in Table 1.
[0056] Table 1 Details of Datasets
[0057]
[0058] The present invention selected eight advanced prediction models as baselines, including Transformer-based models: iTransformer, PatchTST, Crossformer, FEDformer, Autoformer; CNN-based model TimesNet, MICN; linear method-based model: DLinear. All baselines follow the same input length L = 96 and prediction length T ∈ {96, 192, 336, 720}, and the mean absolute error (MAE) and mean squared error (MSE) are selected to evaluate the performance of these models. The experiments of the present invention were conducted on a single Nvidia GeForce RTX 2080Ti GPU.
[0059] Figure 6Shows the prediction results of GLoSyformer and all baselines on eight datasets. The best results are presented in bold format, and the sub-optimal results are presented in underlined format. The smaller the MSE and MAE, the more accurate the prediction results. GLoSyformer achieved advanced results in experiments on different datasets, obtaining 48 best results and 13 sub-optimal results out of 64 results. In the prediction of low-dimensional datasets, the baseline iTransformer performed worse than PatchTST and MICN because the shallow temporal information contained in the variable tokens was insufficient to fully extract temporal features. Compared with iTransformer, GLoSyformer benefited from global and local temporal feature enhancement, with the average MSE decreasing by 4.07% and the average MAE decreasing by 2.96% in the prediction results of low-dimensional datasets. Compared with other SOTA models, GLoSyformer obtained the most best results in the prediction of low-dimensional datasets and significantly outperformed the second-ranked model. In the prediction of high-dimensional datasets, the prediction results of GLoSyformer were almost all the best and surpassed iTransformer, which is good at predicting high-dimensional time series. Compared with iTransformer, GLoSyformer not only reduced the interference of irrelevant variables while capturing multivariate correlations, but also further deeply extracted complex temporal features, with the average MSE decreasing by 2.91% and the average MAE decreasing by 2.4% in the prediction results of high-dimensional datasets. In addition, the present invention also visualizes the prediction results of several representative models together with the prediction results of GLoSyformer on the ETTh1 dataset. As Figure 7 shown, GLoSyformer is superior to other models in predicting both the overall trend and local details. Therefore, GLoSyformer can effectively handle real-world multivariate time series prediction problems and achieve remarkable results.
[0060] The above content is a further detailed description of the present invention in combination with the preferred technical solutions, and it cannot be determined that the specific implementation of the invention is limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, simple deductions and substitutions should be regarded as within the protection scope of the present invention.
Claims
1. A multivariate time series prediction method based on the GLoSyformer model, characterized in that from the global and local perspectives, double-sampling embedding is adopted to enhance the representation of sequence time features; the enhanced variable tokens are centrally processed by the TimesBlock module to capture global and local time dependencies in the sequence; a sparse attention block is used to sparse the attention scores and reduce the interference of irrelevant variables.
2. The multivariate time series prediction method based on the GLoSyformer model according to claim 1, wherein The double-sampling embedding module includes: The double-sampling embedding module includes three different embedding methods, namely variable embedding, down-sampling embedding, and segmented sampling embedding.
3. The multivariate time series prediction method based on the GLoSyformer model according to claim 1, wherein The TimesBlock module includes: TimesBlock is composed of a Global TimesBlock and a Local TimesBlock with the same network structure, which capture global and local time dependencies respectively.
4. The multivariate time series prediction method based on the GLoSyformer model according to claim 1, wherein, The sparse attention module includes: Calculate the Pearson correlation coefficient between variables, and introduce the matrix composed of this coefficient externally to sparse the attention scores between variables, thereby reducing the interference of irrelevant variables.
5. The method according to claim 2, wherein, The variable embedding method, characterized in that variable embedding only performs a linear dimension expansion on the original sequence, and the variable tokens only contain the shallow and coarse-grained time information in the sequence.
6. The method according to claim 2, wherein The down-sampling embedding method, characterized in that down-sampling embedding samples the original sequence at intervals, combines the information of time steps with a relatively large interval into a subsequence, and each subsequence contains certain global time information. Subsequently, each subsequence is embedded to enhance the representation of the global time features of the sequence.
7. The method according to claim 2, wherein The segmented sampling embedding method, characterized in that segmented sampling embedding converts the original sequence into continuous subsequences of the same length. Each subsequence contains local time information within a period of historical time. Subsequently, each subsequence is embedded to enhance the representation of the local time features of the sequence.
8. The method according to claim 3, wherein, The Global TimesBlock module, characterized in that establish a connection between the enhanced tokens and the original variable tokens, and capture the global time dependency by integrating the global time information of the enhanced tokens into the variable tokens. In addition, convolution is also used to further enhance the representation of the global time features contained in the down-sampling embedding vector.
9. The method according to claim 3, wherein The Local TimesBlock module, characterized in that establish a connection between the enhanced tokens and the original variable tokens. Capture the local time dependency by integrating the local time information of the enhanced tokens into the variable tokens. In addition, convolution is also used to further enhance the representation of the local time features contained in the segmented sampling embedding vector.
Citation Information
Cited By
Aerial material demand prediction method and system based on cooperation of sequence decomposition and double-feature channel
CN121599429A