Transform-based time sequence prediction method
By decomposing multivariate time series into seasonal terms and trend terms, and utilizing MLP, GCN, and convolution kernels of different scales combined with sparse-local attention mechanism, the problem of time series feature separation and collaborative modeling in multivariate scenarios based on the Transformer model is solved, achieving higher accuracy prediction.
Patent Information
- Application Number
- CN202510690541.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-09-05
AI Technical Summary
Existing Transformer-based time series forecasting models have difficulty explicitly separating the seasonal and trend characteristics of time series data in multivariate scenarios, and lack the ability to collaboratively model multi-scale time series features, resulting in insufficient prediction accuracy.
The multivariate time series is decomposed into seasonal terms and trend terms. A linear model is used to predict the trend terms. The dynamic correlation between seasonal term variables is modeled using multi-layer perceptrons and graph convolutional networks. Convolution kernels of different scales are combined to extract the multi-granularity time series patterns of seasonal terms. The global and local dependencies are captured through a sparse-local hybrid attention mechanism to finally generate prediction results.
It improves the modeling ability of complex time series patterns, improves prediction accuracy, reduces the computational complexity of long sequence predictions, and improves inference speed.
Smart Images

Figure CN120596834A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence time series prediction and relates to a time series prediction method based on Transformer. Background Art
[0002] Time series forecasting is a key technology for inferring future trends based on historical data, and has a wide range of applications in fields such as energy management, traffic scheduling, and financial markets. Traditional time series forecasting methods such as moving average and exponential smoothing rely primarily on modeling the statistical characteristics of time series, making it difficult to capture complex nonlinear relationships and the influence of external factors. With the development of deep learning technology, time series forecasting methods based on recurrent neural networks and time series convolutional networks have improved nonlinear modeling capabilities, but still face problems such as difficulty in modeling long-term dependencies and insufficient multivariate collaborative analysis capabilities. In recent years, Transformer-based forecasting models have attracted widespread attention for their use of self-attention mechanisms to dynamically capture global dependencies, significantly improving forecast accuracy. However, Transformer-based forecasting models still lack the ability to explicitly separate the seasonal and trend characteristics of time series data, are insufficiently capable of modeling complex cross-dimensional correlations in multivariate scenarios, and lack the ability to collaboratively model multi-scale time series features. Summary of the Invention
[0003] In view of this, the purpose of the present invention is to provide a multivariate time series prediction method based on Transformer, which decomposes the original time series into seasonal terms and trend terms, adopts a linear model to predict the trend terms, utilizes a multi-layer perceptron (MLP) and a graph convolutional network (GCN) to model the dynamic correlation between seasonal term variables, and uses convolution kernels of different scales to extract multi-granularity time series patterns in seasonal terms. The position encoding and timestamp encoding are combined to enhance the time series context expression, and a sparse-local hybrid attention mechanism is used in the encoder to capture global dependencies and local dependencies respectively. Finally, the seasonal term is predicted by the decoder and combined with the trend term to generate the final prediction, which effectively improves the modeling ability of complex time series patterns.
[0004] In order to achieve the above object, the present invention provides the following technical solutions:
[0005] A Transformer-based time series prediction method specifically includes the following steps:
[0006] S1. Preprocess the multivariate time series data and perform seasonal trend decomposition to obtain seasonal terms and trend terms;
[0007] S2, using linear models to predict trend items;
[0008] S3, using multi-layer perceptron (MLP) and graph convolutional network (GCN) to model the correlation between seasonal variables;
[0009] S4, using convolution kernels of different scales to extract multi-scale features of seasonal terms;
[0010] S5. Constructing position code and timestamp code;
[0011] S6. Build a Transformer-based encoder that uses a sparse attention mechanism to capture the global dependencies of time series.
[0012] S7, using local attention mechanism to capture short-term dependencies of time series;
[0013] S8, initialize decoder input;
[0014] S9. Build a Transformer-based decoder to implement prediction.
[0015] Furthermore, in step S1, the multivariate time series data is preprocessed and seasonal trend decomposition is performed to obtain seasonal terms and trend terms, specifically including: setting X = [x1, x2, ..., x T ] T ∈R T×M represents the original multivariate time series, x t Represents the characteristics of X at time step t, 1≤t≤T, T is the number of time steps, M is the number of variables; let τ t Represents x t timestamp; let X m =[x 1,m ,x 2,m ,…,x T,m ] T ∈R T×1 represents the time series of the mth variable, 1≤m≤M, x t,m represents X at time step t m The characteristics of μ m Represents X m The mean of , modeled as: make Represents X m The variance of , is modeled as: Let x′ t,m Represents x t,m Normalized features of: Modeled as: Let X′ m =[x′ 1,m ,x′ 2,m ,…,x′ T,m ] T ∈R T×1; Normalize all variables to obtain the normalized multivariate time series X′=[X′1,X′2,…,X′ M ]∈R T×M ;
[0016] Perform seasonal trend decomposition on X′, let Represents the trend item of X′, and uses the sliding average method to extract the trend item. Specifically, let k represent the sliding window size, where k is an odd number, and add the number of Zero; set the sliding step to 1, take the first time step as the starting point, intercept the features of length k in X′ in sequence and perform average pooling to obtain the trend value of X′ at each time step. The calculation formula is: X t =AvgPool(Padding(X′)); represents the seasonal term of X′, we can get X s =X′-X t ,in, represents X at time t s characteristics.
[0017] Furthermore, in step S2, the trend item is predicted based on the linear model, specifically including: let Y t ∈R F×M Represents the trend item prediction result, F is the number of time steps, model Y t For: Y t =W t X t +b t , where W t represents the weight matrix, b t represents the bias term.
[0018] Furthermore, in step S3, MLP and GCN are used to model the correlation between seasonal variables, specifically including: using MLP to calculate the correlation score between seasonal variables at each time step, let represents X at time step t s The characteristics of the mth variable in , where 1≤m≤M; let express and The correlation score of , where 1≤m′≤M, m≠m′, modeling for: Among them, σ represents the activation function;
[0019] GCN is used to aggregate seasonal multivariate features; specifically, let J represent the number of layers of GCN and the nodes in GCN represent X s The variables of The adjacency matrix of the graph is defined as Add a self-loop to the graph, that is, build a connection between each node and itself, let Represents the adjacency matrix after adding the self-loop; let Represents the degree matrix, which is modeled as a diagonal matrix. The calculation formula of its diagonal elements is: make Represents the node feature matrix of GCN at layer j at time step t, where As the input of the first layer, modeling for:
[0020]
[0021] in, represents the learnable GCN j-th layer weight matrix, C j represents the output dimension of the jth layer; let C1 = 1, C J =C, output of GCN Flattening is performed to obtain row vectors And splice the outputs of all time steps to obtain the fusion features of each time step
[0022] Furthermore, in step S4, convolution kernels of different scales are used to extract multi-scale features of seasonal terms, specifically including: using Z groups of parallel convolution layers to extract the time series Perform feature extraction; represents the d-th convolution kernel of the z-th scale, 1≤z≤Z, 1≤d≤d m , d m Indicates the number of convolution kernels shared by each scale, K z Indicates the length of the convolution kernel at the z-th scale; in the time series The number of additions at both ends is Zero, we get make Represents the features extracted by the d-th convolution kernel in the z-th scale, modeling for:
[0023]
[0024] in, represents a one-dimensional convolution operation, b z,d Represents the bias term of the d-th convolution kernel in the z-th scale; the features extracted by the d convolution kernels in the z-th scale are concatenated to obtain The features extracted by the convolution kernels of Z scales are averaged to obtain the fused feature representation make Represents the features at time step t.
[0025] Furthermore, in step S5, the position code and the timestamp code are constructed, specifically including: t represents the position encoding of the t-th time step, PE t,d′ It represents the position code of the d′th element in the position vector at the tth time step. The position code calculation formula is as follows:
[0026]
[0027] in,
[0028] Let A denote τ t The number of time features, A time features are encoded into vectors to obtain s t ∈R 1×A ; Use linear projection to transform s t Mapped to a high-dimensional semantic space, the timestamp encoding at time step t is obtained Modeled as: SE t =s t W e +b e , where W e represents the learnable weight matrix, b e represents the bias term; let The features obtained by adding position encoding and timestamp encoding at time step t are modeled as:
[0029] Furthermore, in step S6, a Transformer-based encoder is constructed, and a sparse attention mechanism is used to capture the global dependencies of the time series. Specifically, let L represent the number of encoder layers, represents the encoder input features, Represents X e The query matrix, key matrix, and value matrix are obtained by linear transformation of the g0th sparse attention head in the lth layer encoder, where 1≤l≤L, 1≤g0≤G0, G0 represents the number of sparse attention heads, Model separately for: in, They represent the query weight matrix, key weight matrix, and value weight matrix of the g0th sparse attention head in the lth layer encoder respectively;
[0030] make Indicates time step t The query vector, represents the key matrix composed of the key vectors of S randomly selected time steps, express and The sparsity evaluation value of is obtained using the approximate KL divergence formula Modeled as:
[0031]
[0032] in, express The Sth value vector in Represents A sparse matrix consisting of the first S key vectors with the largest sparsity evaluation value, represents the sparse attention output of the selected S time steps, modeled as: The output vectors for the remaining time steps are defined as The average value of all value vectors in the l-th layer encoder is used to obtain the output of the g0-th sparse attention head Concatenate the outputs of G0 sparse attention heads in the l-th layer encoder, and let Represents the sparse multi-head attention output after splicing, modeled as Among them, Concat(·) represents the concatenation operation, Represents the weight matrix of the l-th layer encoder.
[0033] Furthermore, in step S7, a local attention mechanism is used to capture the short-term dependencies of the time series, specifically including: e Divide into multiple segments, use the sliding window method, let P represent the window length and sliding step length, each window as a segment, take the first time step as the starting point, intercept X in turn e The feature with length P is obtained fragments, recorded as: in, represents the rth fragment, 1≤r≤R; let Respectively The key matrix and value matrix obtained by linear transformation of the g1th local attention head in the lth layer encoder, where 1≤g1≤G1, G1 represents the number of local attention heads, They are modeled as: in, denote the key weight matrix and value weight matrix of the g1th local attention head in the lth layer encoder respectively; let represents the g1th local attention head in the lth layer encoder The trainable query vector Input into the local multi-head attention layer, let represents the g1th local attention head in the lth layer encoder The attention output of is modeled as:
[0034]
[0035] Concatenate the outputs of G1 local attention heads in the l-th layer encoder, and let Represents the local multi-head attention output after splicing, modeled as make Represents the local multi-head attention output obtained by splicing the output of all segments, modeled as: Using linear transformation Mapping to T×d m Dimension, get
[0036] Further, in step S8, the decoder input is initialized, specifically including: setting X s,0 ∈R (T / 2)×M Represents X s The second half of the sequence, X 0 ∈R F×M Represents an all-zero matrix, and X s,0 and X 0 Splice along the time dimension and project the features of the spliced matrix into the latent space d through a learnable embedding layer m ,make Denotes the decoder input, modeled as: X d,0 =Embed(Concat(X s,0 ,X 0 )); for X d,0 Perform position encoding and timestamp encoding to obtain X d ;
[0037] Furthermore, in step S9, a Transformer-based decoder is constructed to achieve prediction, specifically including: let I represent the number of decoders, and Represents X d The query matrix, key matrix and value matrix are obtained by linear transformation of the g2th masked sparse attention head in the i-th layer decoder, where 1≤i≤I,1≤g2≤G2, G2 represents the number of cross attention heads; All key vectors in are evaluated for sparsity, and express A sparse matrix composed of the first S key vectors with the largest sparsity evaluation value is constructed, and a mask matrix is constructed. The mask operation is performed on the attention score matrix. Let H0 represent the mask matrix and H0(i,j) represent the element in the i-th row and j-th column. The model is:
[0038]
[0039] Among them, t irepresents the time step position of the i-th active query vector; let represents the masked sparse attention output of the first S time steps of the g2-th masked sparse attention head in the i-th layer decoder, which is modeled as: in, The output vectors for the remaining time steps are defined as The average value of all value vectors before the current time step in , concatenate the outputs of G2 masked sparse attention heads in the i-th layer decoder, and let Represents the masked multi-head attention output after splicing, modeled as
[0040] The mask information matrix and the coded information matrix Input the multi-head cross attention layer of the decoder, let express The query matrix obtained by linear transformation of the g3th cross attention head is and Respectively The key matrix and value matrix obtained by linear transformation of the g3th cross attention head in the i-th layer decoder, where 1≤g3≤G3, G3 represents the number of cross attention heads, let Represents the output of the cross-attention calculation, modeled as:
[0041]
[0042] in, The outputs of G3 cross-attention heads in the i-th layer decoder are concatenated to obtain the multi-head cross-attention output in the i-th layer decoder After the I-layer decoder, interception The feature matrix of the target sequence is obtained right Use linear transformation to get the seasonal item prediction result Y s ∈R F×M , add the predicted trend term and seasonal term to get the final prediction result Y=Y t +Y s ; Denormalize Y to obtain the prediction result of the original sequence.
[0043] The beneficial effects of the present invention are: through MLP and GCN, the dynamic relationship between variables is adaptively captured, further improving the prediction performance of the model; through multi-branch convolutional layers, high-frequency details and low-frequency trends are simultaneously captured, avoiding the information loss problem of traditional single-scale modeling, thereby improving the prediction accuracy of the model; by combining local attention and global sparse attention, local and global correlations in time series are simultaneously captured, effectively solving complex dependency modeling problems, reducing computational complexity and improving inference speed in long sequence prediction.
[0044] Other advantages, objects, and features of the present invention will be described in part in the following description and, in part, will be apparent to those skilled in the art upon examination of the following description or may be learned from practice of the present invention. The objects and other advantages of the present invention may be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below with reference to the accompanying drawings, in which:
[0046] Figure 1 This is a schematic diagram of the timing prediction proposed by the present invention;
[0047] Figure 2 Flowchart of the time series prediction method. DETAILED DESCRIPTION
[0048] The following describes the embodiments of the present invention by means of specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present invention, and the following embodiments and features in the embodiments can be combined with each other without conflict.
[0049] See also Figure 1 and Figure 2The present invention discloses a Transformer-based time series prediction method, which mainly includes: first, normalizing and preprocessing the original time series data, and decomposing the multivariate sequence into seasonal terms and trend terms through the sliding window averaging method. The trend term component is predicted using a linear regression model. For the seasonal term component, a hybrid model including MLP and GCN is constructed, and the time-varying correlation between multiple variables is modeled through node feature aggregation and adjacency matrix dynamic update mechanism. A multi-scale feature extraction module is designed, and convolution kernels with different window sizes are used to capture time series features of different granularities, and the time series context expression is enhanced by fusing with position coding and timestamp coding. In the encoding stage, the sparse attention mechanism is used to reduce the computational complexity of long sequences, and the local sliding window attention is combined to capture short-term dependency patterns. Finally, the predicted value is generated by a multi-layer Transformer decoder.
[0050] The method of the present invention specifically comprises the following steps:
[0051] Step 1: Preprocess the multivariate time series data and perform seasonal trend decomposition to obtain seasonal terms and trend terms;
[0052] Preprocess the multivariate time series data and perform seasonal trend decomposition to obtain seasonal terms and trend terms. Specifically, let X = [x1, x2,…, x T ] T ∈R T×M represents the original multivariate time series, x t Represents the characteristics of X at time step t, 1≤t≤T, T is the number of time steps, M is the number of variables; let τ t Represents x t timestamp; let X m =[x 1,m ,x 2,m ,…,x T,m ] T ∈R T×1 represents the time series of the mth variable, 1≤m≤M, x t,m represents X at time step t m The characteristics of μ m Represents X m The mean of , modeled as: make Represents X m The variance of , is modeled as: Let x′ t,m Represents x t,m Normalized features of: Modeled as: Let X′ m =[x′ 1,m ,x′ 2,m ,…,x′ T,m ]T ∈R T×1 ; Normalize all variables to obtain the normalized multivariate time series X′=[X′1,X′2,…,X′ M ]∈R T×M ;
[0053] Perform seasonal trend decomposition on X′, let Represents the trend item of X′, and uses the sliding average method to extract the trend item. Specifically, let k represent the sliding window size, where k is an odd number, and add the number of Zero; set the sliding step to 1, take the first time step as the starting point, intercept the features of length k in X′ in sequence and perform average pooling to obtain the trend value of X′ at each time step. The calculation formula is: X t =AvgPool(Padding(X′)); represents the seasonal term of X′, we can get X s =X′-X t ,in, represents X at time t s characteristics.
[0054] Step 2: Predict the trend item based on the linear model;
[0055] The trend item is predicted based on the linear model, specifically: let Y t ∈R F×M Represents the trend item prediction result, F is the number of time steps, model Y t For: Y t =W t X t +b t , where W t represents the weight matrix, b t represents the bias term.
[0056] Step 3: Use MLP and GCN to model the correlation between seasonal variables;
[0057] MLP and GCN are used to model the correlation between seasonal variables. Specifically, MLP is used to calculate the correlation score between seasonal variables at each time step. represents X at time step t s The characteristics of the mth variable in , where 1≤m≤M; let express and The correlation score of , where 1≤m′≤M, m≠m′, modeling for: Among them, σ represents the activation function;
[0058] GCN is used to aggregate seasonal multivariate features; specifically, let J represent the number of layers of GCN and the nodes in GCN represent X s The variables of The adjacency matrix of the graph is defined as Add a self-loop to the graph, that is, build a connection between each node and itself, let Represents the adjacency matrix after adding the self-loop; let Represents the degree matrix, which is modeled as a diagonal matrix. The calculation formula of its diagonal elements is: make Represents the node feature matrix of GCN at layer j at time step t, where As the input of the first layer, modeling for:
[0059]
[0060] in, represents the learnable GCN j-th layer weight matrix, C j represents the output dimension of the jth layer; let C1 = 1, C J =C, output of GCN Flattening is performed to obtain row vectors And splice the outputs of all time steps to obtain the fusion features of each time step
[0061] Step 4: Use convolution kernels of different scales to extract multi-scale features of seasonal terms;
[0062] Convolution kernels of different scales are used to extract multi-scale features of seasonal terms. Specifically, Z groups of parallel convolution layers are used to extract the seasonal features of the time series. Perform feature extraction; represents the d-th convolution kernel of the z-th scale, 1≤z≤Z, 1≤d≤d m , d m Indicates the number of convolution kernels shared by each scale, K z Indicates the length of the convolution kernel at the z-th scale; in the time series The number of additions at both ends is Zero, we get make Represents the features extracted by the d-th convolution kernel in the z-th scale, modeling for:
[0063]
[0064] in, represents a one-dimensional convolution operation, b z,dRepresents the bias term of the d-th convolution kernel in the z-th scale; the features extracted by the d convolution kernels in the z-th scale are concatenated to obtain The features extracted by the convolution kernels of Z scales are averaged to obtain the fused feature representation make Represents the features at time step t.
[0065] Step 5: Construct position code and timestamp code;
[0066] Construct position code and timestamp code, specifically: let PE t represents the position encoding of the t-th time step, PE t,d′ It represents the position code of the d′th element in the position vector at the tth time step. The position code calculation formula is as follows:
[0067]
[0068] in,
[0069] Let A denote τ t The number of time features, A time features are encoded into vectors to obtain s t ∈R 1×A ; Use linear projection to transform s t Mapped to a high-dimensional semantic space, the timestamp encoding at time step t is obtained Modeled as: SE t =s t W e +b e , where W e represents the learnable weight matrix, b e represents the bias term; let The features obtained by adding position encoding and timestamp encoding at time step t are modeled as:
[0070] Step 6: Build a Transformer-based encoder and use a sparse attention mechanism to capture the global dependencies of the time series;
[0071] Construct a Transformer-based encoder and use a sparse attention mechanism to capture the global dependencies of time series. Specifically, let L represent the number of encoder layers, represents the encoder input features, Represents X e The query matrix, key matrix, and value matrix are obtained by linear transformation of the g0th sparse attention head in the lth layer encoder, where 1≤l≤L, 1≤g0≤G0, G0 represents the number of sparse attention heads, Model separately for: in, They represent the query weight matrix, key weight matrix, and value weight matrix of the g0th sparse attention head in the lth layer encoder respectively;
[0072] make Indicates time step t The query vector, represents the key matrix composed of the key vectors of S randomly selected time steps, express and The sparsity evaluation value of is obtained using the approximate KL divergence formula Modeled as:
[0073]
[0074] in, express The Sth value vector in Represents A sparse matrix consisting of the first S key vectors with the largest sparsity evaluation value, represents the sparse attention output of the selected S time steps, modeled as: The output vectors for the remaining time steps are defined as The average value of all value vectors in the l-th layer encoder is used to obtain the output of the g0-th sparse attention head Concatenate the outputs of G0 sparse attention heads in the l-th layer encoder, and let Represents the sparse multi-head attention output after splicing, modeled as Among them, Concat(·) represents the concatenation operation, Represents the weight matrix of the l-th layer encoder.
[0075] Step 7: Use the local attention mechanism to capture the short-term dependencies of the time series;
[0076] The local attention mechanism is used to capture the short-term dependencies of the time series. Specifically, X e Divide into multiple segments, use the sliding window method, let P represent the window length and sliding step length, each window as a segment, take the first time step as the starting point, intercept X in turn e The feature with length P is obtained fragments, recorded as: in, represents the rth fragment, 1≤r≤R; let Respectively The key matrix and value matrix obtained by linear transformation of the g1th local attention head in the lth layer encoder, where 1≤g1≤G1, G1 represents the number of local attention heads, They are modeled as: in, denote the key weight matrix and value weight matrix of the g1th local attention head in the lth layer encoder respectively; let represents the g1th local attention head in the lth layer encoder The trainable query vector Input into the local multi-head attention layer, let represents the g1th local attention head in the lth layer encoder The attention output of is modeled as:
[0077]
[0078] Concatenate the outputs of G1 local attention heads in the l-th layer encoder, and let Represents the local multi-head attention output after splicing, modeled as make Represents the local multi-head attention output obtained by splicing the output of all segments, modeled as: Using linear transformation Mapping to T×d m Dimension, get
[0079] Step 8: Initialize decoder input;
[0080] Initialize the decoder input, specifically: let X s,0 ∈R (T / 2)×M Represents X s The second half of the sequence, X 0 ∈R F×M Represents an all-zero matrix, and X s,0 and X 0 Splice along the time dimension and project the features of the spliced matrix into the latent space d through a learnable embedding layer m ,make Denotes the decoder input, modeled as: X d,0 =Embed(Concat(X s,0 ,X 0 )); for X d,0 Perform position encoding and timestamp encoding to obtain X d ;
[0081] Step 9: Build a Transformer-based decoder to implement prediction;
[0082] Construct a Transformer-based decoder to achieve prediction, specifically: let I represent the number of decoders, and Represents X d The query matrix, key matrix and value matrix are obtained by linear transformation of the g2th masked sparse attention head in the i-th layer decoder, where 1≤i≤I,1≤g2≤G2, G2 represents the number of cross attention heads; All key vectors in are evaluated for sparsity, and express A sparse matrix composed of the first S key vectors with the largest sparsity evaluation value is constructed, and a mask matrix is constructed. The mask operation is performed on the attention score matrix. Let H0 represent the mask matrix and H0(i,j) represent the element in the i-th row and j-th column. The model is:
[0083]
[0084] Among them, t i represents the time step position of the i-th active query vector; let represents the masked sparse attention output of the first S time steps of the g2-th masked sparse attention head in the i-th layer decoder, which is modeled as: in, The output vectors for the remaining time steps are defined as The average value of all value vectors before the current time step in , concatenate the outputs of G2 masked sparse attention heads in the i-th layer decoder, and let Represents the masked multi-head attention output after splicing, modeled as
[0085] The mask information matrix and the coded information matrix Input the multi-head cross attention layer of the decoder, let express The query matrix obtained by linear transformation of the g3th cross attention head is and Respectively The key matrix and value matrix obtained by linear transformation of the g3th cross attention head in the i-th layer decoder, where 1≤g3≤G3, G3 represents the number of cross attention heads, let Represents the output of the cross-attention calculation, modeled as:
[0086]
[0087] in, The outputs of G3 cross-attention heads in the i-th layer decoder are concatenated to obtain the multi-head cross-attention output in the i-th layer decoder After the I-layer decoder, interception The feature matrix of the target sequence is obtained right Use linear transformation to get the seasonal item prediction result Y s ∈R F×M , add the predicted trend term and seasonal term to get the final prediction result Y=Y t +Y s ; Denormalize Y to obtain the prediction result of the original sequence.
[0088] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solutions, which should all be included in the scope of the claims of the present invention.
Claims
1. A Transformer-based time series prediction method, characterized by: The method comprises the following steps: S1. Preprocess the multivariate time series data and perform seasonal trend decomposition to obtain seasonal terms and trend terms; S2, using linear models to predict trend items; S3, using multi-layer perceptron (MLP) and graph convolutional network (GCN) to model the correlation between seasonal variables; S4, using convolution kernels of different scales to extract multi-scale features of seasonal terms; S5. Constructing position code and timestamp code; S6. Build a Transformer-based encoder that uses a sparse attention mechanism to capture the global dependencies of time series. S7, using local attention mechanism to capture short-term dependencies of time series; S8, initialize decoder input; S9. Build a Transformer-based decoder to implement prediction.
2. The Transformer-based time series prediction method according to claim 1, characterized in that: In step S1, the multivariate time series data is preprocessed and seasonal trend decomposition is performed to obtain seasonal terms and trend terms, specifically including: setting X = [x1, x2, ..., x T ] T ∈R T×M represents the original multivariate time series, x t Represents the characteristics of X at time step t, 1≤t≤T, T is the number of time steps, M is the number of variables; let τ t Represents x t timestamp; let X m =[x 1,m ,x 2,m ,...,x T,m ] T ∈R T×1 represents the time series of the mth variable, 1≤m≤M, x t,m represents X at time step t m The characteristics of μ m Represents X m The mean of , modeled as: make Represents X m The variance of , is modeled as: Let x′ t,m Represents x t,m Normalized features of: Modeled as: Let X′ m =[x′ 1,m ,x′ 2,m ,...,x′ T,m ] T ∈R T×1 ; Normalize all variables to obtain the normalized multivariate time series X′=[X′1,X′2,...,X′ M ]∈R T×M ; Perform seasonal trend decomposition on X′, let Represents the trend item of X′, and uses the sliding average method to extract the trend item. Specifically, let k represent the sliding window size, where k is an odd number, and add the number of Zero; set the sliding step to 1, take the first time step as the starting point, intercept the features of length k in X′ in sequence and perform average pooling to obtain the trend value of X′ at each time step. The calculation formula is: X t =AvgPool(Padding(X′)); represents the seasonal term of X′, we can get X s =X′-X t ,in, represents X at time t s characteristics.
3. The Transformer-based time series prediction method according to claim 1, characterized in that: In step S2, the trend item is predicted based on the linear model, specifically including: let Y t ∈R F×M Represents the trend item prediction result, F is the number of time steps, model Y t For: Y t =W t X t +b t , where W t represents the weight matrix, b t represents the bias term.
4. The Transformer-based time series prediction method according to claim 1, characterized in that: In step S3, MLP and GCN are used to model the correlation between seasonal variables, specifically including: using MLP to calculate the correlation score between seasonal variables at each time step, let represents X at time step t s The characteristics of the mth variable in , where 1≤m≤M; let express and The correlation score of , where 1≤m′≤M, m≠m′, modeling for: Where σ represents the activation function; GCN is used to aggregate seasonal multivariate features; specifically, let J represent the number of layers of GCN and the nodes in GCN represent X s The variables of The adjacency matrix of the graph is defined as Add a self-loop to the graph, that is, build a connection between each node and itself, let Represents the adjacency matrix after adding the self-loop; let Represents the degree matrix, which is modeled as a diagonal matrix. The calculation formula of its diagonal elements is: make Represents the node feature matrix of GCN at layer j at time step t, where As the input of the first layer, modeling for: in, represents the learnable GCN j-th layer weight matrix, C j represents the output dimension of the jth layer; let C1 = 1, C J =C, output for GCN Flattening is performed to obtain row vectors And splice the outputs of all time steps to obtain the fusion features of each time step 5. The Transformer-based time series prediction method according to claim 1, characterized in that: In step S4, convolution kernels of different scales are used to extract multi-scale features of seasonal terms, specifically including: using Z groups of parallel convolution layers to extract the time series Perform feature extraction; represents the d-th convolution kernel of the z-th scale, 1≤z≤Z, 1≤d≤d m , d m Indicates the number of convolution kernels shared by each scale, K z Indicates the length of the convolution kernel at the z-th scale; in the time series The number of additions at both ends is Zero, we get make Represents the features extracted by the d-th convolution kernel in the z-th scale, modeling for: in, represents a one-dimensional convolution operation, b z,d Represents the bias term of the d-th convolution kernel in the z-th scale; the features extracted by the d convolution kernels in the z-th scale are concatenated to obtain The features extracted by the convolution kernels of Z scales are averaged to obtain the fused feature representation make Represents the features at time step t.
6. The Transformer-based time series prediction method according to claim 1, characterized in that: In step S5, the position code and timestamp code are constructed, specifically including: t represents the position encoding of the t-th time step, PE t,d′ It represents the position code of the d′th element in the position vector at the tth time step. The position code calculation formula is as follows: in, Let A denote τ t The number of time features, A time features are encoded into vectors to obtain s t ∈R 1×A ; Use linear projection to transform s t Mapped to a high-dimensional semantic space, the timestamp encoding at time step t is obtained Modeled as: SE t =s t W e +b e , where W e represents the learnable weight matrix, b e represents the bias term; let The features obtained by adding position encoding and timestamp encoding at time step t are modeled as:
7. The Transformer-based time series prediction method according to claim 1, characterized in that: In step S6, a Transformer-based encoder is constructed, and a sparse attention mechanism is used to capture the global dependencies of the time series. Specifically, let L represent the number of encoder layers, represents the encoder input features, Represents X e The query matrix, key matrix, and value matrix are obtained by linear transformation of the g0th sparse attention head in the lth layer encoder, where 1≤l≤L, 1≤g0≤G0, G0 represents the number of sparse attention heads, Model separately for: in, They represent the query weight matrix, key weight matrix, and value weight matrix of the g0th sparse attention head in the lth layer encoder respectively; make Indicates time step t The query vector, represents the key matrix composed of the key vectors of S randomly selected time steps, express and The sparsity evaluation value of is obtained using the approximate KL divergence formula Modeled as: in, express The Sth value vector in Represents A sparse matrix consisting of the first S key vectors with the largest sparsity evaluation value, represents the sparse attention output of the selected S time steps, modeled as: The output vectors for the remaining time steps are defined as The average value of all value vectors in the l-th layer encoder is used to obtain the output of the g0-th sparse attention head Concatenate the outputs of G0 sparse attention heads in the l-th layer encoder, and let Represents the sparse multi-head attention output after splicing, modeled as Among them, Concat(·) represents the concatenation operation, Represents the weight matrix of the l-th layer encoder.
8. The Transformer-based time series prediction method according to claim 1, characterized in that: In step S7, a local attention mechanism is used to capture the short-term dependencies of the time series, specifically including: e Divide into multiple segments, use the sliding window method, let P represent the window length and sliding step length, each window as a segment, take the first time step as the starting point, intercept X in turn e The feature with length P is obtained fragments, recorded as: in, represents the rth fragment, 1≤r≤R; let Respectively The key matrix and value matrix obtained by linear transformation of the g1th local attention head in the lth layer encoder, where 1≤g1≤G1, G1 represents the number of local attention heads, Modeled as: in, denote the key weight matrix and value weight matrix of the g1th local attention head in the lth layer encoder respectively; let represents the g1th local attention head in the lth layer encoder The trainable query vector Input into the local multi-head attention layer, let represents the g1th local attention head in the lth layer encoder The attention output of is modeled as: Concatenate the outputs of G1 local attention heads in the l-th layer encoder, and let Represents the local multi-head attention output after splicing, modeled as make Represents the local multi-head attention output obtained by splicing the output of all segments, modeled as: Using linear transformation Mapping to T×d m Dimension, get Output the sparse multi-head attention and local multi-head attention output Add together to get the output of the encoder at layer l, which is modeled as: After L layers of encoders, the encoding information matrix is obtained 9. The Transformer-based time series prediction method according to claim 1, characterized in that: In step S8, the decoder input is initialized, specifically including: setting X s,0 ∈R (T / 2)×M Represents X s The second half of the sequence, X 0 ∈R F×M Represents an all-zero matrix, and X s,0 and X 0 Splice along the time dimension and project the features of the spliced matrix into the latent space d through a learnable embedding layer m ,make Denotes the decoder input, modeled as: X d,0 =Embed(Concat(X s,0 ,X 0 )); for X d,0 Perform position encoding and timestamp encoding to obtain X d .
10. The Transformer-based time series prediction method according to claim 1, characterized in that: In step S9, a Transformer-based decoder is constructed to achieve prediction, specifically including: let I represent the number of decoders, and Represents X d The query matrix, key matrix and value matrix are obtained by linear transformation of the g2th masked sparse attention head in the i-th layer decoder, where 1≤i≤I,1≤g2≤G2, G2 represents the number of cross attention heads; All key vectors in are evaluated for sparsity, and express A sparse matrix composed of the first S key vectors with the largest sparsity evaluation value is constructed, and a mask matrix is constructed. The mask operation is performed on the attention score matrix. Let H0 represent the mask matrix and H0(i,j) represent the element in the i-th row and j-th column. The model is: Among them, t i represents the time step position of the i-th active query vector; let represents the masked sparse attention output of the first S time steps of the g2-th masked sparse attention head in the i-th layer decoder, which is modeled as: in, The output vectors for the remaining time steps are defined as The average value of all value vectors before the current time step in , concatenate the outputs of G2 masked sparse attention heads in the i-th layer decoder, and let Represents the masked multi-head attention output after splicing, modeled as The mask information matrix and the coded information matrix Input the multi-head cross attention layer of the decoder, let express The query matrix obtained by linear transformation of the g3th cross attention head is and Respectively The key matrix and value matrix obtained by linear transformation of the g3th cross attention head in the i-th layer decoder, where 1≤g3≤G3, G3 represents the number of cross attention heads, let Represents the output of the cross-attention calculation, modeled as: in, The outputs of G3 cross attention heads in the i-th layer decoder are concatenated to obtain the cross multi-head cross multi-attention output in the i-th layer decoder After the I-layer decoder, interception The feature matrix of the target sequence is obtained right Use linear transformation to get the seasonal item prediction result Y s ∈R F×M , add the predicted trend term and seasonal term to get the final prediction result Y=Y t +Y s ; Denormalize Y to obtain the prediction result of the original sequence.
Citation Information
Cited By
Traffic flow prediction method and system based on multi-scale dynamic decomposition and space-time Transform
CN121034085A
Dynamic multi-scale coding source load prediction method and system based on prompt
CN121124038A
Intelligent gait plantar pressure prediction system based on Transform structure
CN121483654A
Distributed photovoltaic access area-oriented net load prediction method
CN121546548A