Segmented improved attention-based multivariable long-term time sequence prediction method
Through segmented processing and improved attention calculation methods, the problem of high computational complexity in large-scale, long-sequence time series prediction is solved, the prediction accuracy and calculation efficiency are improved, and a larger timing look-up window is supported.
Patent Information
- Application Number
- CN202510309922.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-06-13
AI Technical Summary
When the prior art deals with large-scale, long-sequence time series prediction, the calculation complexity is high, making it difficult to capture long-distance dependencies, resulting in low prediction accuracy and low computing efficiency.
The idea of segmentation processing is adopted, multiple timestamps are composed of small segments as input to the Transformer encoder, and through improved attention calculation methods, only the parts with high probability are calculated, and the sparsity is measured using Kullback-Leibler(KL) divergence to reduce the calculation amount.
It improves prediction accuracy, reduces calculation amount and improves calculation efficiency, supports a larger timing look-up window, and is suitable for multivariate long-term timing prediction tasks.
Smart Images

Figure CN120144963A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular, to a segmented multi-variable long-term time series prediction method based on improved attention. Background Art
[0002] Time series prediction is a technology for predicting future development trends based on historical data. In this field, accurate and efficient prediction models are crucial for multiple industries such as financial analysis, energy consumption prediction, traffic flow prediction, weather forecasting, etc.
[0003] There are three main methods for time series prediction. The first is the time series prediction method based on traditional mathematical models, such as the autoregressive moving average model (ARMA) and the autoregressive integrated moving average model (ARIMA); the second is the deep autoregressive model, such as the recurrent neural network (RNN) and the long short-term memory network LSTM, etc. Although the above two methods perform well in some scenarios, they perform poorly in dealing with large-scale data and long-term prediction: on the one hand, public datasets in the field of time series prediction, such as Traffice, have 862 variables and 17,544 time steps, and it is difficult for traditional models to handle such complex variables and timestamps. On the other hand, when the prediction time step is very long, the error of the autoregressive model will accumulate continuously, resulting in a deterioration of the prediction result; the third is the time series prediction model based on Transformer, which is one of the main research directions in current time series prediction.
[0004] Long-term time series prediction is crucial in time series prediction because it can provide early warnings for decision-makers, optimize resource allocation, support scientific research, enhance the interpretability and robustness of the model, improve prediction performance, and achieve precise and efficient management and services in multiple fields by analyzing long-term trends and periodic patterns in historical data. Due to the existence of the full attention calculation method, the spatio-temporal complexity of Transformer-based models is proportional to the square of the sequence length L. Therefore, when dealing with long sequences, the performance may decline and the computational cost will increase significantly. This is because Transformer needs to calculate the correlation between each position in the sequence and other positions, resulting in an increase in the number of parameters as the input sequence length increases. Therefore, in order to enable the model to efficiently perform long-term time series prediction, many works have made improvements in this regard.
[0005] In recent years, many works have tried to apply the Transformer model to long-term time series prediction. For example, LogTrans uses a convolutional self-attention layer with a LogSparse design to capture local information and reduce the spatial complexity. In addition, LogTrans introduces the LogSparse Self-Attention mechanism, allowing each time step to only focus on its previous exponential number of time steps.
[0006] LogTrans first uses causal convolution to convert the input data into queries and keys, thus generating the input for the self-attention layer. This step helps to extract feature representations that are more conducive to the self-attention mechanism while maintaining the order of the input sequence unchanged. Next, in the self-attention calculation process, LogTrans adopts the LogSparse self-attention mechanism. This design not only enhances the local perception ability of the self-attention layer but also significantly reduces the computational complexity and memory usage. This makes LogTrans perform better than ordinary Transformer models in processing large-scale, long-sequence time series prediction tasks. However, this attention mechanism allows each time step to only see time steps exponentially far back for attention calculation, which makes the attention calculation have a certain degree of uncertainty and is likely to miss important information. Therefore, this attention mechanism will largely limit the model's ability to capture long-range dependencies. If two points that influence each other greatly in the time series data are too far or too close, then LogTrans's self-attention mechanism cannot capture this relationship, resulting in the loss of important information and affecting the prediction accuracy. In addition, LogTrans performs poorly in processing some time series with complex dependencies. For this kind of data, it is also difficult for models using the full attention mechanism to handle, and the sparse attention mechanism is even less able to extract the implicit information in the time series data.
[0007] In addition to LogTrans, there are also some works that modify the attention module. Autoformer takes a different approach and modifies the way of calculating attention by introducing relevant theories in the field of communication. Autoformer borrows the ideas of decomposition and autocorrelation from the traditional signal processing field and improves the computational efficiency and prediction accuracy by introducing the autocorrelation mechanism (Auto-Correlation) and the decomposition architecture. The core idea of Autoformer is the autocorrelation mechanism (Auto-Correlation), which is a method for discovering subsequence similarity based on the periodicity of the sequence, aiming to discover and represent aggregations of dependencies at the subsequence level, thereby improving the computational efficiency and information utilization rate. Through the autocorrelation mechanism, Autoformer can process sequences of length L within a time complexity of O(LlogL), which significantly reduces the computational cost. However, this autocorrelation mechanism may not be able to fully capture all the complex dependencies in the sequence, especially when dealing with highly non-linear or non-periodic time series. For example, financial data such as stock markets and exchange rates usually have highly non-linear and non-periodic characteristics, and these markets are affected by many unpredictable factors such as political events, economic data releases, and company performances, which makes it difficult to describe price fluctuations with simple linear models. Autoformer is not suitable for predicting such data.
[0008] In addition, both LogTrans and Autoformer essentially encode each time step of the time series as a vector as the input to the Transformer. That is, under the same input sequence, the same number of tokens are encoded as the input, and then various methods are used to optimize the full attention calculation to reduce the computational amount. However, even when using optimized attention calculation methods, when the lookback window required for actual prediction is very long, the overhead of the model is still very large, which cannot meet the production needs of long-term time series prediction.
[0009] In addition, there are other works such as FEDformer that use a Fourier enhancement structure to obtain linear complexity. Some of the above-mentioned derivative models modify the Transformer model architecture, and some modify the self-attention calculation method. However, no matter how they are modified, their purpose is to address the problem of the excessively high spatio-temporal complexity of the Transformer when performing long-term time series prediction. Summary of the Invention
[0010] The technical problem to be solved by the present invention is to provide a segmented multi-variable long-term time series prediction method based on improved attention aiming at the deficiencies of the above-mentioned prior art, which supports a larger time series lookback window and uses an efficient attention calculation method, thereby improving the prediction accuracy, reducing the computational amount, and improving the computational efficiency.
[0011] To solve the above technical problems, the technical solution adopted by the present invention is as follows:
[0012] A segmented multi-variable long-term time series prediction method based on improved attention. After preprocessing the time series, using the idea of segmented processing, multiple timestamps are grouped into a small segment and embedded as tokens as the input of the Transformer encoder; then the segmented segments are tokenized, and for different attention heads, different weight matrices W Q 、W K 、W V are used to transform the input into different query matrices Q, key matrices K, and value matrices V for improved attention calculation; in the improved attention calculation, through the probability sparsity hypothesis, only the parts with high probability in the self-attention matrix are calculated. The interpretation of probability is that the self-attention mechanism calculates the similarity scores between queries (Query) and keys (Key), and converts them into a probability distribution through the Softmax function; the probability distribution reflects the correlation between each position and other positions, and the parts with high probability mean that in the context of the current time series, the potential association between this time series and other time series is stronger; the "active" query matrix Q with probability close to 1 is selected to calculate the attention mechanism, and the "lazy" query matrix Q with probability close to 0 is discarded and not involved in the attention calculation.
[0013] Furthermore, the specific method for preprocessing the time series is as follows:
[0014] First, for duplicate values, outliers, and missing values, duplicate time series records are detected and deleted respectively, outliers are detected and corrected through box plots, and the linear interpolation method is used to fill in the missing values; then, to ensure the consistency of the timestamp format, the timestamp is converted to the standard time format ISO 8601, and at the same time, the timestamp is set as the index of the data to facilitate data operations and analysis in chronological order; then, normalization is performed to scale the data of some data sets to a smaller range; finally, the time series data is divided into a training set and a test set, and the first 80% of the time series is used as the training set, and the last 20% is used as the test set.
[0015] Furthermore, the specific method for segmented processing is as follows:
[0016] Given a set of multi-variable time series samples (x 1 ,…,x L ) with a backtracking window L, where the time series input x t at each time step t is a vector with dimension M, and M is the number of variables; to predict the future T values (x L+1 ,…,x L+T ), the standard Transformer encoder, i.e., the Encoder-Only architecture, is used.
[0017] Design the forward propagation mechanism: Denote the i-th univariate time series starting from time index 1 and with length L, where i = 1, 2, …, M; the input (x 1 , …, x L ) is divided into M univariate time series R 1×L , where R is a univariate time series with 1 variable and length L, and then M x (i) are obtained. After that, each sequence is independently input into the Transformer encoder backbone network according to channel independence, and then the standard Transformer encoder backbone network provides corresponding prediction results
[0018] Each input univariate time series x (i) is first divided into multiple segments, which can be overlapping or non-overlapping; let the segment length be P, and the non-overlapping region (i.e., the step size) between two consecutive segments be S, then the segmentation process will generate a sequence of segments where N is the number of segments, Before segmentation, the last value of the i-th univariate time series is filled at the end of the original sequence repeated S times.
[0019] Furthermore, the step of tokenizing the segmented segments is as follows: Using the Transformer encoder, map the divided small segments to the Transformer input size, perform positional embedding, and use multiple attention heads to capture information in different time dimensions; apply different weight matrices W Q , W K , W V to transform the input into different query matrices Q, key matrices K, and value matrices V, which are used as the input of the Transformer encoder to prepare for the improved attention calculation.
[0020] Furthermore, the improved attention calculation is as follows:
[0021] Sample the key matrix K, sampling out 1 / 8 of K. According to statistical principles, the sample is distributed identically to the population, so it is considered that sampling out 1 / 8 of K represents the overall K; for each query matrix Q, calculate the correlation with the sampled 1 / 8 of K to obtain the activity ranking list of the query matrix Q, and then find the top 1 / 8 of Q with the largest difference, and calculate the distribution of these Q with all K;
[0022] Through the above calculations, obtain the query q i and all keys k jThe attention distribution p(k j |q i ), where q i is an element of the query matrix Q, and k j is an element of the key matrix K.
[0023] Furthermore, in the improved attention calculation, the correlation between two distributions is used to replace the correlation after matrix multiplication. The activity level of a query matrix Q is measured by comparing the deviation between the QK similarity distribution and the uniform distribution; the greater the difference between the two, the more active the query matrix Q is, and vice versa, indicating that the query matrix Q is not active;
[0024] To measure the degree of difference between two distributions, calculate the attention distribution p(k i |q j ) between the query q j and all keys k i and the Kullback-Leibler divergence, i.e., KL divergence, between the uniform distribution q(k j |q i ). The calculation formula is:
[0025]
[0026] where k(q i , k j ) represents the attention score between the query q i and each key k j , ∑ l k(q i , k l ) represents the sum of the attention scores between q i and all keys k l , l = 1, 2,..., L k ; L k represents the number of all keys;
[0027] By calculating and simplifying the KL divergence between distributions, the sparsity metric formula for the i-th query is obtained:
[0028]
[0029] where M(q i , K) represents the KL divergence value of the two distributions, and d is the feature dimension of each element; the first term on the right side of the equation is the logarithmic sum exponential function of the inner product of q i and all k j , and the second term is the arithmetic mean of the inner product of q i and all k j . The greater the divergence, the more diverse the probability distribution and the more likely there is a focus of attention.
[0030] Further, after the KL divergence is calculated, in order to reduce the computational complexity of calculating the sparsity of the i-th query, the difference between the maximum value of the distribution and the uniform distribution is used to calculate the difference, so as to further speed up the calculation process, that is, the sparsity metric formula of the i-th query is simplified to:
[0031]
[0032] The beneficial effects of adopting the above technical solutions are as follows: The segmented multi-variable long-term time series prediction method based on improved attention provided by the present invention adopts a segmentation strategy, combines multiple consecutive time points into a small segment, and then embeds this small segment as a whole into the input of the Transformer. This method aims to control the complexity of the model by reducing the time span represented by each token, while still being able to utilize long historical information to improve the prediction accuracy. The present invention proposes an improved attention calculation method. Before performing the self-attention operation, the sparsity degree of the input data is first evaluated, and the difference between two probability distributions is used to replace the correlation obtained directly through matrix multiplication. The activity level of a Q is measured by comparing the deviation degree between the QK similarity distribution and the uniform distribution; if the difference between the two is larger, it indicates that the Q is more active. Then, for each query q i , the Kullback-Leibler (KL) divergence is used as an index to measure its sparsity. This method can more accurately identify which Qs are truly important and only let these active Qs participate in the final attention mechanism calculation, thereby effectively improving the model efficiency and performance. Through these two improvements, namely, segmenting the time series data and improving the attention calculation method, the present invention greatly increases the size of the model's lookback window. Compared with LogTrans and Autoformer, it improves the prediction efficiency while improving the prediction accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 It is a flowchart of a segmented multi-variable long-term time series prediction method based on improved attention provided by an embodiment of the present invention;
[0034] Figure 2 It is a heat map provided by an embodiment of the present invention;
[0035] Figure 3 It is an attention map provided by an embodiment of the present invention;
[0036] Figure 4 It is a segmented size result map provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0037] The following will further describe in detail the specific implementation manners of the present invention in conjunction with the accompanying drawings and embodiments. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention.
[0038] This embodiment proposes a segmented multi-variable long-term time series prediction method based on improved attention. This method supports a larger time series lookback window and uses an efficient attention calculation method, thereby improving the prediction accuracy, reducing the computational amount, and improving the computational efficiency.
[0039] In time series prediction, the size of the lookback window will significantly affect the prediction effect. A larger lookback window is beneficial for the model to better capture the past time features and patterns, thereby improving the prediction accuracy. This is because the value at the current moment may depend on the data at a relatively distant past moment, and an appropriate lookback window size can help the model capture these long-term dependencies. Therefore, when performing time series prediction, people will try to increase the size of the lookback window to improve the prediction accuracy.
[0040] Due to the existence of the full attention calculation method in Transformer-based models, the spatio-temporal complexity is proportional to the square of the sequence length L. Therefore, it is usually difficult for Transformer-based time series prediction models to set a too large lookback window because a too large lookback window will significantly increase the computational amount and slow down the prediction efficiency. Therefore, there has always been a contradiction between the prediction efficiency and prediction accuracy of Transformer-based time series models.
[0041] To solve the problem of the lookback window, this embodiment proposes a segmented idea: instead of embedding a single variable as a token as the input of the Transformer, the segmented idea is used to embed multiple timestamps into a small segment as a token as the input of the Transformer.
[0042] This embodiment proposes a time series prediction model based on the Transformer architecture, aiming to improve the prediction accuracy and efficiency by splitting the time series data into multiple small segments. Splitting the long sequence into multiple shorter small segments can reduce the computational complexity of the model and enable the model to better capture local features.
[0043] The structure of the time series prediction model based on the Transformer architecture is as follows:
[0044] Consider the following problem: Given a set of multi-variable time series samples (x 1 , …, x L ) with a lookback window L, where x t at each time step t is a vector of dimension M, and we hope to predict the next T values (x L+1 , …, x L+T ). The model is asFigure 1 As shown, a standard Transformer encoder is used as its core architecture.
[0045] Forward propagation: represents the i-th univariate sequence of length L starting from time index 1, where i = 1, 2, …, M. The input (x 1 , …, x L ) is split into M univariate sequences R 1×L , and each sequence is independently input into the Transformer backbone network according to the channel independence setting. Then, the Transformer backbone network will accordingly provide prediction results
[0046] Segment processing: Each input univariate time series x (i) is first split into multiple segments, which can be overlapping or non-overlapping. Let the segment length be P and the stride (i.e., the non-overlapping region between two consecutive segments) be S, then the segmentation process will generate a segment sequence where N is the number of segments, here, the last value is repeated S times and filled to the end of the original sequence. By using segments, the number of input tokens can be reduced from L to approximately L / S. This means that the memory usage and computational complexity of the attention map are reduced quadratically by the S factor. Therefore, in the case of limited training time and GPU memory, the segment design can enable the model to see longer historical sequences, which can significantly improve the prediction performance.
[0047] Transformer encoder: Map the divided segments to the Transformer input size, perform positional embedding and then use them as the input x of the Transformer encoder. After that, for different attention heads, different weight matrices W Q , W K , W V are used to transform the input into different query matrices Q, key matrices K, and value matrices V for improved attention calculation and then used as the input x of the Transformer encoder.
[0048] For the attention calculation, this embodiment proposes a calculation method different from LogSparse Self-Attention.
[0049] In a long sequence, not every position of the attention is important. Before calculating between Q, K, and V, first obtain the heat map of the dot product of Q and K, as Figure 2As shown, the brighter part indicates a higher Q-K correlation. Most of the heatmap is black. Experiments have found that for each Q, only a small part of K has a strong relationship with it.
[0050] As Figure 2 shown, when calculating Q and K for 8,000 samples of the Traffic dataset, there are less than 2,000 with relatively high correlations. Most of the time, the relationship between Q and K is close to 0.
[0051] As Figure 3 shown in the attention map, the vertical axis represents the query vector (Q), and the horizontal axis represents the key vector (K). Each row in the figure corresponds to the correlation distribution of a specific query vector with all key vectors. Among them, the upper line chart identifies an "active" query vector, and through the significant highlighting of this area, it can be clearly identified that there is a relatively high correlation between it and a specific key vector. Relatively speaking, the lower line chart represents a "lazy" query vector, and the correlations of this vector with all key vectors are in a relatively "flat" state, lacking significant correlation peaks. In actual calculations, relatively "lazy" Q not only fails to provide effective value, but most of the Qs are relatively "lazy". The same is true in time series prediction. For a time step, the main influence on it comes from several previous time steps, and most time steps have no strong connection with the current time step. Only selecting "active" Qs to calculate the attention mechanism and discarding "lazy" Qs from participating in the attention calculation is the attention calculation method proposed in this embodiment, which is different from LogTrans.
[0052] Through the probabilistic sparsity assumption, only calculate the parts with relatively high probabilities in the self-attention matrix, thereby reducing the computational complexity from L 2 to L * logL, which significantly reduces the inference time and enables the model to better handle uncertainty and noise, improving robustness.
[0053] Specifically, when calculating self-attention, first measure the sparsity of the input. For each query q i , use the Kullback-Leibler (KL) divergence to measure its sparsity, and calculate the KL divergence between the attention distribution p(k i | q j ) between query q j and all keys k i and the uniform distribution q(k j | q i ).
[0054] For the case where the lookback window is 966, traditional attention would calculate 966 * 966 attentions, but the improved method proposed in this embodiment is as follows:
[0055] Step 1: Sample the key vectors K, sampling out 1 / 8 of K, that is, 120 Ks. These 120 Ks are a sample of 966 Ks. According to the principle of statistics, the sample is distributed the same as the population, so it can be considered that the 120 Ks can represent the overall K;
[0056] Step 2: For each query vector Q, calculate the correlation with the sampled 1 / 8 of K, and then change the full attention of 966×966 to 966×120, that is, 966 Qs × 120 Ks;
[0057] Step 3: Through the calculation in Step 2, the activity ranking list of Q can be obtained. Since all 966 Qs have been calculated, a list with 966 elements is obtained. By sorting, the top 120 Qs with the largest differences are found, and the distribution of these 120 Qs and all Ks is calculated, that is, the most active 120 Qs × 966 Ks.
[0058] Step 4: q can be obtained through the above calculations i and all keys k j the attention distribution p(k j |q i ). After that, calculate the KL divergence between the attention distribution p(k i |q j ) of query q j and all keys k i and the uniform distribution q(k j |q i ).
[0059] Step 5: Based on the sparsity measure, calculate the weighted sum of the values v j corresponding to each query, and the weights are given by the attention distribution p(k j |q i ). In this way, only those key-value pairs that are highly correlated with the query q i will be taken into account, thus reducing the computational amount.
[0060] Among them, replacing the correlation after matrix multiplication with the correlation between the two distributions, after obtaining the distribution of qk, calculate the KL divergence between the distributions, and the calculation formula is:
[0061]
[0062] Among them, k(q i ,k j ) represents the attention score between the query q i and each key k j , and ∑ l k(q i ,k l ) represents qi and the sum of the attention scores for all keys k l , where l = 1, 2, …, L k ; L k represents the number of all keys.
[0063] Finally, through the calculation and simplification of the KL divergence between distributions, the sparsity metric formula for the i-th query is obtained:
[0064]
[0065] where M(q i , K) represents the KL divergence value of the two distributions, d is the feature dimension of each element, and the first term on the right side of the equation is the log-sum-exp (LSE) of the inner products of q i and all k j , and the second term is the arithmetic mean of the inner products of q i and all k j . The greater the divergence, the more diverse the probability distribution and the more likely there are areas of focus.
[0066] To reduce the computational complexity of calculating the sparsity of the i-th query, is used to replace the first term in the original formula , that is, the difference between the maximum value of the distribution and the uniform distribution is calculated to further speed up the calculation process. The sparsity metric formula for the i-th query is thus simplified to:
[0067]
[0068] In the improved attention calculation method proposed in this embodiment, before performing self-attention operations, the sparsity degree of the input data is first evaluated, and the difference between two probability distributions is used to replace the correlation obtained directly through matrix multiplication. The activity level of a Q is measured by comparing the deviation between the QK similarity distribution and the uniform distribution; if the difference between the two is greater, it indicates that the Q is more active. Then, for each query q i , the Kullback-Leibler (KL) divergence is used as an indicator to measure its sparsity. This method can more accurately identify which Qs are truly important and only let these active Qs participate in the final attention mechanism calculation, thus effectively improving the model efficiency and performance.
[0069] Next, experiments were conducted on the time series dataset Traffic, using the methods of segmented embedding and direct embedding respectively. Under the condition that the lookback window is 336, the MSE of directly embedding 336 tokens is 0.397, while the MSE of first taking 42 segments for 336 time steps and then embedding them into 42 tokens is 0.367. After using segmentation, the computational cost is reduced to 1 / 64 of that without using segmentation, and the accuracy is improved. This fully proves that using segmentation is very effective in time series prediction and also verifies the rationality of the design of this embodiment.
[0070] Table 1 Comparative experiments on segmentation of Traffic dataset
[0071]
[0072] The segmentation experiment on Traffic proves that segmenting the input is reasonable:
[0073] From the experimental results in Table 1, it can be observed that when the lookback window is 336, after segmentation, the number of tokens is reduced to 42, and the running time is 1 unit. Without segmentation, the running time is 64 units. In terms of running efficiency, segmentation can greatly improve the running efficiency. Therefore, it can be proved that segmentation effectively improves the prediction efficiency and results.
[0074] The datasets used in this embodiment are the time series prediction standard datasets Weather (21 variables and 52696 time steps), Traffic (862 variables and 17544 time steps), Electricity (321 variables and 26304 time steps), ETTh1 (7 variables and 17420 time steps), ETTh2 (7 variables and 17420 time steps), and the wind power dataset Wind Energy Generation (WEG) (12 variables and 49366 time steps).
[0075] The comparison models selected are Dlinear, Autoformer in 2023, FEDformer and LogTrans in 2022. The main evaluation indicators are MSE and MAE. The lookback window length is 512, the segment length is 12, and it is finally divided into 42 segments. The prediction time step lengths of 96 and 720 are selected. The final results are shown in Table 2. Under the condition of the same running time, this embodiment obtains the best MSE and MAE.
[0076] Table 2 Model evaluation comparison
[0077]
[0078]
[0079] Experiments show that in this embodiment, by segmenting the input and using an improved attention mechanism, good results have been achieved in the mainstream datasets for time series prediction.
[0080] In addition, relevant experiments on the segment size have been conducted in this embodiment. An experiment with a prediction step of 96 is selected on the Weather dataset, and the experimental results are as Figure 4 shown. It is found that the choice of the segment size S has no significant impact on the score of MSE, which indicates that this embodiment is robust in the face of the hyperparameter of the segment length S. Generally speaking, the prediction accuracy of the method in this embodiment benefits from the increase in the segment length, not only improving the prediction performance but also reducing the computational cost. The ideal segment size may depend on the specific dataset, but a good general choice is between 8 and 32.
[0081] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope defined by the claims of the present invention.
Claims
1. A segmented multivariate long-term time series prediction method based on improved attention, characterized by: After preprocessing the time series, we use the idea of segmentation to group multiple timestamps into a small segment and embed them into a word as the input of the Transformer encoder. Then we tokenize the segmented segments and use different weight matrices W for different attention heads. Q , W K , W V The input is changed into different query matrices Q, key matrices K, and value matrices V for improved attention calculation; the improved attention calculation, through the probability sparsity assumption, only calculates the high-probability part in the self-attention matrix, and the explanation of probability is as follows: the self-attention mechanism calculates the similarity score between the query (Query) and the key (Key), and converts it into a probability distribution through the Softmax function; the probability distribution reflects the correlation between each position and other positions, and the high-probability part means that in the current time series context, the potential correlation between this time series and other time series is stronger; The "active" query matrix Q with a probability close to 1 is selected to calculate the attention mechanism, and the "lazy" query matrix Q with a probability close to 0 is discarded and does not participate in the attention calculation.
2. A segmented multivariate long-term time series prediction method based on improved attention according to claim 1, characterized in that: The specific method of the time series preprocessing is: First, for duplicate values, outliers, and missing values, duplicate time series records are detected and deleted, outliers are detected and corrected through box plots, and missing values are filled using linear interpolation methods. Then, to ensure the consistency of the timestamp format, the timestamp is converted to the standard time format ISO 8601. At the same time, the timestamp is set as the index of the data to facilitate data operation and analysis in chronological order. After that, normalization is performed to scale the data of some data sets to a smaller range; finally, the time series data is divided into training and test sets, using the first 80% of the time series as the training set and the last 20% as the test set.
3. A segmented multivariate long-term time series prediction method based on improved attention according to claim 2, characterized in that: The specific method of the segmentation process is as follows: Given a set of multivariate time series samples (x1,…,x L ), where the time series input x at each time step t is t is a vector of dimension M, where M is the number of variables; in order to predict the future T values (x L+1 ,…,x L+T ), using the standard Transformer encoder, i.e., Encoder-Only architecture; Design the forward propagation mechanism: represents the i-th univariate time series starting from time index 1 and of length L, where i = 1, 2, …, M; input (x1, …, x L ) is split into M univariate time series R 1×L , R is a univariate time series with 1 variable and L length, and then M x (i) After that, each sequence is independently input into the Transformer encoder backbone network according to the channel independence setting, and then the standard Transformer encoder backbone network provides the corresponding prediction results Each input univariate time series x (i) First, it is segmented into multiple segments, which are overlapping or non-overlapping. Let the segment length be P, and the non-overlapping area between two consecutive segments, i.e., the step size, be S. Then, the segmentation process will generate a segment sequence where N is the number of segments, Before segmentation, the last value of the i-th univariate time series is Repeat S times to fill the original sequence The end of .
4. A segmented multivariate long-term time series prediction method based on improved attention according to claim 3, characterized in that: The steps of tokenizing the divided segments are as follows: using the Transformer encoder, mapping the divided segments to the Transformer input size, performing position embedding, and using multiple attention heads to capture information in different time dimensions; using different weight matrices W Q , W K , W V The input is transformed into a different query matrix Q, key matrix K, and value matrix V as the input of the Transformer encoder, ready for improved attention calculation.
5. A segmented multivariate long-term time series prediction method based on improved attention according to claim 4, characterized in that: The improved attention calculation is specifically as follows: Sampling the key matrix K, sampling 1 / 8 of K. According to the statistical principle, the sampling is the same as the overall distribution, so it is considered that the sampled 1 / 8 of K represents the overall K; For each query matrix Q, calculate the correlation with the sampled 1 / 8 of K to obtain the activity ranking of the query matrix Q, and then find the top 1 / 8 of Q with the largest difference, and calculate the distribution of these Q and all K; Through the above calculation, we get the query q i With all keys k j The attention distribution between j |q i ), where q i is the element of the query matrix Q, k j are the elements of the key matrix K.
6. A segmented multivariate long-term time series prediction method based on improved attention according to claim 5, characterized in that: In the improved attention calculation, the correlation after matrix multiplication is replaced by the correlation of the two distributions, and the activity level of a query matrix Q is measured by comparing the deviation between the QK similarity distribution and the uniform distribution; if the difference between the two is greater, it indicates that the query matrix Q is more active, otherwise it indicates that the query matrix Q is not active; To measure the difference between the two distributions, we calculate the query q i With all keys k j The attention distribution between j |q i ) and uniform distribution q(k j |q i ) is the Kullback-Leibler divergence between , which is the KL divergence, and the calculation formula is: Among them, k(q i ,k j ) represents the query q i and each key k j The attention score, ∑ l k(q i ,k l ) indicates q i and all keys k l The sum of attention scores, l = 1, 2, ..., L k ; L k Represents the number of all keys; By calculating and simplifying the KL divergence between distributions, we get the sparsity measurement formula for the i-th query: Among them, M(q i ,K) represents the KL divergence value of the two distributions, d is the characteristic dimension of each element; the first term on the right side of the equation is q i and all k j The logarithmic sum exponential function of the inner product, the second term is q i and all k j The arithmetic mean of the inner product, the larger the divergence, the more diverse the probability distribution, and the more likely it is that there is a focus.
7. A segmented multivariate long-term time series prediction method based on improved attention according to claim 6, characterized in that: After the KL divergence is calculated, in order to reduce the amount of calculation for calculating the sparsity of the i-th query, the maximum value of the distribution and the uniform distribution are used to calculate the difference to further speed up the calculation process, that is, the sparsity measurement formula for the i-th query is simplified to:
Citation Information
Cited By
Multivariable time sequence prediction interpretation method and system based on information theory and causal reasoning
CN120373478A