Prediction method based on time sequence data containing multiple variables
By adopting multivariate attention network and batch embedding technology in time series prediction, the existing deep learning methods are solved, and the problems of high computational cost, high memory consumption and lack of generalization and scalability are achieved, and more efficient and accurate time series prediction is achieved.
Patent Information
- Application Number
- CN202510259612.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-06-24
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing deep learning methods have problems such as high computing costs, high memory consumption, and lack of generalization and scalability in time series prediction.
Using a prediction method based on a multivariate attention network, a multivariate variable affecting the traffic flow in the area is separated and discrete, batch embedding and reprogramming is performed, and label embedding and fusion is performed in combination with label information, and finally prediction is performed through a multi-cycle training module.
It improves the efficiency and accuracy of time series prediction, reduces computational cost and memory consumption, and enhances the generalization ability and scalability of the model.
Smart Images

Figure CN120196946A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of time series prediction, and specifically refers to a prediction method based on time series data containing multiple variables. Background Art
[0002] Time series prediction is an important research direction in the field of data analysis. It aims to predict future trends, patterns or values based on past historical data. With the rapid development of social economy, time series data has been widely used in various fields, such as financial market analysis, energy demand prediction, weather forecasting, inventory management, traffic flow prediction, etc. Accurate time series prediction can help enterprises and governments make more scientific decisions, improve resource utilization efficiency and reduce risks.
[0003] Currently, time series prediction methods are mainly divided into classical statistical methods, non-linear time series prediction and deep learning methods, etc. Classical statistical methods, such as ARIMA (Autoregressive Integrated Moving Average Model), exponential smoothing method, etc., mainly predict based on the linear relationship of time series, can handle limited problem complexity, and the effect is also average. Non-linear time series prediction methods, such as neural networks, support vector machines, etc., predict by learning the complex non-linear relationship of time series, and further improve the prediction performance for complex prediction problems. The currently widely used deep learning method, which is also the best prediction method, has the characteristics of high time and computational cost but good prediction effect. Classical methods include recurrent neural network (RNN), long short-term memory network (LSTM), etc., which can handle more complex time series data.
[0004] One of the main challenges in current research problems is that as a method with extremely high time and computational cost, deep learning methods consume too much memory. In addition, the current deep learning methods are too targeted at a specific field, and the models trained by them often lack generalization and scalability. How to make the proposed method more feasible and scalable are issues that urgently need further research. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a prediction method based on time series data containing multiple variables in view of the deficiencies mentioned in the above background art.
[0006] To solve the above technical problem, the technical solution provided by the present invention is: a prediction method based on time series data containing multiple variables, including the following steps:
[0007] S1: Separate multiple variables (including time factor, cycle factor, route factor) affecting the traffic flow in the area, and discretize them according to different needs;
[0008] S2: Batch-embed the discretized regional traffic flow data, and then perform reprogramming (input it into a multi-variable attention network and then linearize it) to obtain a reprogrammed embedding batch;
[0009] S3: Tokenize the label information in the regional traffic flow data and obtain a token embedding batch through token embedding;
[0010] S4: Fuse the reprogrammed embedding batch obtained in S2 with the token embedding batch obtained in S3 to obtain a training input embedding batch;
[0011] S5: Input the training input embedding batch into a training module, and this training module needs to loop for multiple epochs;
[0012] S6: After training, flatten and linearize the output batch, and finally fuse the prediction results of multiple variables respectively to obtain the final prediction result.
[0013] Further, in step S1, separating the various variables that affect regional traffic flow - time factor, cycle factor, route factor, is expressed as
[0014]
[0015] In the formula, n represents the dimension that the time series can be separated into, where n = 3; L represents the time sampling number of the time series; t is the time, 1 ≤ t ≤ L.
[0016] Further, in step S1, discretizing it according to different needs is expressed as
[0017]
[0018] In the formula, X A is the time series of a certain variable separated in the previous step, and A is one of T, P, S; X' A is the new sequence composed of the sampling points of the selected part in X A . When selecting sampling points, different discrete intervals are adopted according to different attributes. The time factor is divided by hours to determine the traffic peak period at that time; the cycle factor is divided by days to determine the regular time distributions such as weekends, holidays, winter and summer; the route factor is divided by minutes to determine whether the route traveled by this journey is detoured or passes through congested sections, etc. For the sake of convenience of expression, all Xs appearing in the following text are the new sequences generated in this step.
[0019] Further, in step S2, batch-embedding the discretized time series data. It is specifically expressed as:
[0020] PEX = PatchEmbed(X);
[0021] In the formula, the PatchEmbed() method represents batch embedding of the time series X. Here, the time series corresponding to each variable is divided into 8 batches, and one batch from each variable is selected and embedded together. PEX represents the batch embedding sequence.
[0022] Furthermore, reprogramming is performed as described in step S2. The execution process is as follows:
[0023] REP1, REP2, …, REP8 = Linear(MultivariateAttention(PEX1, PEX2, …, PEX8));
[0024] Among them, MultivariateAttention() is a multi-variable attention network, Linear() is a linearization function, PEX represents the batch embedding sequence, and REP represents the reprogramming embedding batch.
[0025] Furthermore, the label information in the regional traffic flow data is tokenized as described in step S3, and through token embedding, a token embedding batch is obtained:
[0026] TEP1, TEP2, …, TEP4 = TokenEmbed(Split(label));
[0027] In the formula, label is the name of the data label, Split() is a separation function used to extract the valid information in label as a token, TokenEmbed() embeds the token to obtain the token embedding batch TEP. The number of the token embedding batch is set to half of the reprogramming embedding batch, which is 4. In this way, different label names used in different datasets can be compatible, improving the generalization ability of the model.
[0028] Furthermore, the reprogramming embedding batch obtained in S2 and the token embedding batch obtained in S3 are fused as described in step S4 to obtain a training input embedding batch:
[0029] InEP1, InEP2, …, InEP 12 = Fusion(REP1, REP2, …, REP8, TEP1, TEP2, …, TEP4);
[0030] In the formula, Fusion() represents the fusion function, and InEP represents the training input embedding batch.
[0031] Furthermore, the training module in step S5 consists of S51 - S54:
[0032] S51: A multivariate attention network extracts key - values from the separated multivariate variables respectively, multiplies the key - values and then scales them to obtain a relationship graph;
[0033] S52: Perform layer normalization;
[0034] S53: Perform feed - forward propagation learning;
[0035] S54: Perform layer normalization again;
[0036] In the multivariate attention network described in step S51, key - values are extracted from the separated multivariate variables respectively, the key - values are multiplied and then scaled to obtain a relationship graph. Among them, the execution process is as follows:
[0037] Key, Value = MultivariateAttention(InEP1, InEP2, …, InEP 12 )
[0038] MCMap = Scale(Key × Value)
[0039] In the formula, MultivariateAttention() is the multivariate attention network; Key, Value are the output tensors of the multivariate attention network, Scale() is the scaling function, and MCMap is the multivariate relationship graph.
[0040] Furthermore, the layer normalization and feed - forward propagation network described in steps S52, S53, and S54 can be expressed as:
[0041] S52: LayerNorm(MCMap)
[0042] S53: FeedForward(·)
[0043] S54: OutEP = LayerNorm(·)
[0044] In the formula, LayerNorm() is the layer normalization; FeedForward() is the feed - forward propagation network, and OutEP is the output embedding batch, and this output embedding batch contains the output embedding batches of each separated variable.
[0045] Furthermore, the flattening and linearization of the output batch described in step S6, and finally the fusion of the prediction results of each multivariate variable to obtain the final prediction result can be expressed as:
[0046] Pred = Fusion(Linear(Flatten(OutEP)))
[0047] Where Flatten() is the flattening function; Linear() is the linearization function, and Fusion() is the fusion function that fuses the prediction results obtained by embedding the outputs of each variable in batches to obtain the final prediction result.
[0048] After adopting the above method, the present invention has the following advantages: The multi-variable time series prediction method based on TEMPO-FORMER provided by the present invention fully utilizes multiple factors and label information that affect the prediction result, separately predicts each factor, and finally performs fusion prediction, solving the problem of weak correlation between multi-variable time series prediction and multiple factors; By batch processing the label information, it is compatible with different data sets, improves the training efficiency and model scalability, and enhances the generalization ability; Through the batch processing technology, large-scale data is loaded and trained in batches, saving resource overhead. Description of the Drawings
[0049] Figure 1 It is a schematic diagram of the overall framework of a prediction method based on time series data containing multiple variables. Detailed Embodiments
[0050] The present invention will be further described in detail below with reference to the accompanying drawings.
[0051] Combined with the attached Figure 1 , a prediction method based on time series data containing multiple variables includes the following steps:
[0052] S1: Separate multiple variables (including time factor, cycle factor, route factor) that affect the traffic flow in the area, and discretize them according to different needs;
[0053] The separation of multiple variables - time factor, cycle factor, route factor that affect the traffic flow in the area in step S1 is expressed as
[0054]
[0055] Where n represents the dimension that the time series can be separated into, where n = 3 here; L represents the time sampling number of the time series; t is the time, 1 ≤ t ≤ L.
[0056] Specifically, the discretization according to different needs in step S1 is expressed as
[0057]
[0058] Where X A is the time series of a certain variable separated in the previous step, A is one of T, P, S; X' A is X AA new sequence composed of sampling points of the selected part in []. When selecting sampling points, different discrete intervals are adopted according to different attributes. The time factor is divided by hours to determine the traffic peak period at that time; the cycle factor is divided by days to determine regular time distributions such as weekends, holidays, winter and summer; the route factor is divided by minutes to determine whether the route traveled by this journey is a detour or passes through congested sections, etc. For the convenience of representation, X that appears in the following text is the new sequence generated in this step.
[0059] S2: Batch-embed the discretized regional traffic flow data, and then perform reprogramming (input into a multi-variable attention network and then linearize) to obtain a reprogrammed embedding batch;
[0060] The discretized time series data is batch-embedded as described in step S2. Specifically, it is expressed as:
[0061] PEX = PatchEmbed(X);
[0062] In the formula, the PatchEmbed() method represents batch-embedding the time series X. Here, the time series corresponding to each variable is divided into 8 batches, and one batch from each variable is selected and embedded together. PEX represents the batch-embedded sequence.
[0063] Specifically, for the reprogramming described in step S2, the execution process is as follows:
[0064] REP1, REP2, …, REP8 = Linear(MultivariateAttention(PEX1, PEX2, …, PEX8));
[0065] Among them, MultivariateAttention() is a multi-variable attention network, Linear() is a linearization function, PEX represents the batch-embedded sequence, and REP represents the reprogrammed embedding batch.
[0066] S3: Tokenize the label information in the regional traffic flow data and obtain a token-embedded batch through token embedding;
[0067] The label information in the regional traffic flow data is tokenized and a token-embedded batch is obtained as described in step S3:
[0068] TEP1, TEP2, …, TEP4 = TokenEmbed(Split(label));
[0069] In the formula, label is the name of the data label, Split() is the separation function used to extract the valid information in label as the identifier, and TokenEmbed() embeds the identifier to obtain the identifier embedding batch TEP. The number of identifier embedding batches is set to half of the reprogramming embedding batch, which is 4. In this way, different label names used in different datasets can be compatible, improving the generalization ability of the model.
[0070] S4: Fuse the reprogramming embedding batch obtained in S2 and the identifier embedding batch obtained in S3 to obtain the training input embedding batch;
[0071] The step S4 of fusing the reprogramming embedding batch obtained in S2 and the identifier embedding batch obtained in S3 to obtain the training input embedding batch:
[0072] InEP1, InEP2, …, InEP 12 = Fusion(REP1, REP2, …, REP8, TEP1, TEP2, …, TEP4);
[0073] In the formula, Fusion() represents the fusion function, and InEP represents the training input embedding batch.
[0074] S5: Input the training input embedding batch into the training module, and this training module needs to loop for multiple cycles repeatedly;
[0075] The training module in step S5 consists of S51 - S54:
[0076] S51: A multivariate attention network extracts the key and value of the separated multivariate respectively, multiplies the key and value and then scales them to obtain the relationship graph;
[0077] S52: Perform layer normalization;
[0078] S53: Perform feed-forward propagation learning;
[0079] S54: Perform layer normalization again;
[0080] For the multivariate attention network described in step S51, the separated multivariate are respectively used to extract the key and value, multiply the key and value and then scale them to obtain the relationship graph. Among them, the execution process is as follows:
[0081] Key, Value = MultivariateAttention(InEP1, InEP2, …, InEP 12 )
[0082] MCMap = Scale(Key × Value)
[0083] Wherein, MultivariateAttention() is a multivariate attention network; Key and Value are output tensors of the multivariate attention network, Scale() is a scaling function, and MCMap is a multivariate relationship graph.
[0084] The layer normalization and feed-forward propagation network described in steps S52, S53, and S54 can be expressed as:
[0085] S52: LayerNorm(MCMap)
[0086] S53: FeedForward(·)
[0087] S54: OutEP = LayerNorm(·)
[0088] Wherein, LayerNorm() is layer normalization; FeedForward() is a feed-forward propagation network, and OutEP is an output embedding batch, which contains the output embedding batches of each separated variable.
[0089] S6: After the training is completed, flatten and linearize the output batch, and finally fuse the prediction results of each multivariate variable to obtain the final prediction result;
[0090] The flattening and linearizing of the output batch described in step S6, and finally fusing the prediction results of each multivariate variable to obtain the final prediction result can be expressed as:
[0091] Pred = Fusion(Linear(Flatten(OutEP)))
[0092] Wherein, Flatten() is a flattening function; Linear() is a linearizing function, and Fusion() is a fusion function, which fuses the prediction results obtained from the output embedding batches of each variable to obtain the final prediction result.
[0093] The above describes the present invention and its implementation manners, and this description is not restrictive. The actual structure is not limited thereto. Generally speaking, if those of ordinary skill in the art are inspired by it and design similar structural manners and embodiments without creative efforts without departing from the purpose of the present invention, they shall fall within the protection scope of the present invention.
Claims
1. A prediction method based on time series data containing multiple variables, characterized by: The following steps are involved: S1: Separate the various variables that affect regional traffic flow (including time factors, cycle factors, and route factors) and discretize them according to different needs; S2: The discretized regional traffic flow data is batch-embedded and then reprogrammed (input into the multivariate attention network and then linearized) to obtain a reprogrammed embedding batch; S3: Label information in the regional vehicle flow data is identified and embedded to obtain an identification embedding batch; S4: Fuse the reprogramming embedding batch obtained in S2 with the identification embedding batch obtained in S3 to obtain the training input embedding batch; S5: embed the training input into the batch input into the training module, which needs to be repeatedly cycled for multiple cycles; S6: After training, the output batch is flattened and linearized, and finally the prediction results of multiple variables are fused to obtain the final prediction result.
2. A prediction method based on time series data containing multiple variables according to claim 1, characterized in that: Step S1 separates the multiple variables that affect regional traffic flow, namely, time factor, cycle factor, and route factor, and expresses them as In the formula, n represents the dimension in which the time series can be separated, where n=3; L represents the number of time samples of the time series; t is time, 1≤t≤L.
3. The prediction method based on time series data containing multiple variables according to claim 1, characterized in that: Step S1 discretizes it according to different needs and is expressed as Where, X A is the time series of a variable separated in the previous step, A is one of T, P, S; X′ A For X A The new sequence is composed of the sampling points selected from the above. When selecting sampling points, different discrete intervals are adopted according to different attributes. The time factor is divided by hours to determine the traffic peak at that time. The cycle factor is divided by days to determine the regular time distribution such as weekends, holidays, winter and summer. The route factor is divided by minutes to determine whether the route of the trip takes a detour or passes through a congested section. For convenience, the Xs that appear in the following text are all new sequences generated in this step.
4. The prediction method based on time series data containing multiple variables according to claim 1, characterized in that: Step S2 is to batch embed the discretized time series data, which is specifically expressed as follows: PEX = PatchEmbed(X); In the formula, the PatchEmbed() method indicates that the time series X is embedded in batches. Here, the time series corresponding to each variable is divided into 8 batches, and one batch of each variable is selected and embedded together. PEX represents the batch embedding sequence.
5. The prediction method based on time series data containing multiple variables according to claim 1, characterized in that: Step S2 is described as reprogramming, wherein the execution process is as follows: REP1, REP2,…,REP8=Linear(MultivariateAttention(PEX1,PEX2,…,PEX8)); Among them, MultivariateAttention() is a multivariate attention network, Linear() is a linearization function, PEX represents a batch embedding sequence, and REP is a reprogrammed embedding batch.
6. The prediction method based on time series data containing multiple variables according to claim 1, characterized in that: In step S3, the label information in the regional vehicle flow data is identified, and the identification is embedded to obtain an identification embedding batch: TEP1,TEP2,…,TEP4=TokenEmbed(Split(label)); In the formula, label is the name of the data label, Split() is a separation function used to extract valid information from label as an identifier, TokenEmbed() embeds the identifier to obtain an identifier embedding batch TEP, and the number of identifier embedding batches is set to half of the reprogramming embedding batch, that is, 4. In this way, different label names used in different data sets can be compatible, thereby improving the generalization ability of the model.
7. The prediction method based on time series data containing multiple variables according to claim 1, characterized in that: Step S4 combines the reprogramming embedding batch obtained in S2 with the identification embedding batch obtained in S3 to obtain a training input embedding batch: InEP1,InEP2,…,InEP 12 =Fusion(REP1,REP2,…,REP8,TEP1,TEP2,…,TEP4); Where Fusion() represents the fusion function and InEP represents the training input embedding batch.
8. The prediction method based on time series data containing multiple variables according to claim 1, characterized in that: Step S5: training module consists of S51-S54 composition: S51: A multivariate attention network that extracts key values from the separated multivariate, multiplies the key values and then scales them to get a relationship graph; S52: perform layer normalization; S53: Perform feedforward propagation learning; S54: perform layer normalization again; In step S51, a multivariate attention network extracts key values of the separated multivariate, multiplies the key values and then scales them to obtain a relationship graph, wherein the execution process is as follows: Key,Value=MultivariateAttention(InEP1,InEP2,…,InEP 12 ) MCMap=Scale(Key×Value) In the formula, MultivariateAttention() is the multivariate attention network; Key, Value are the output tensors of the multivariate attention network, Scale() is the scaling function, and MCMap is the multivariate relationship diagram.
9. A prediction method based on time series data containing multiple variables according to claim 8, characterized in that: The layer normalization and feedforward propagation network described in steps S52, S53, and S54 can be expressed as: S52:LayerNorm(MCMap) S53:FeedForward(·) S54: OutEP = LayerNorm (·) In the formula, LayerNorm() is layer normalization; FeedForward() is the feedforward propagation network, and OutEP is the output embedding batch, which contains the output embedding batch of each variable separated.
10. The prediction method based on time series data containing multiple variables according to claim 1, characterized in that: In step S6, the output batch is flattened and linearized, and finally the prediction results of multiple variables are merged to obtain the final prediction result which can be expressed as: Pred=Fusion(Linear(Flatten(OutEP))) In the formula, Flatten() is the flattening function; Linear() is the linearization function, and Fusion() is the fusion function. The output of each variable is embedded in the prediction results obtained in the batch to obtain the final prediction result.