Time sequence prediction method and device based on MPDLINE model and medium
Through the combination of multi-component decomposition and self-learning weight vector module, multi-channel independent modeling weighted prediction module and feature-time dimension mixed self-attention mechanism module based on the MPDLinear model, the accuracy and stability problems of the existing technology when processing complex multi-dimensional time series data are solved, and time series prediction with high accuracy and high stability are achieved.
Patent Information
- Application Number
- CN202510234715.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-02-28
AI Technical Summary
The prior art has problems with low prediction accuracy and poor stability when processing complex multidimensional time series data, especially in capturing nonlinear trends and medium-long periodic trends and processing the interrelationships between multi-channel features.
A time series prediction method based on the MPDLinear model is proposed. The time series data is decomposed into multiple components through the multi-component decomposition module, and combined with the self-learning weight vector module, the multi-channel independent modeling weighted prediction module and the feature-time dimension mixed self-attention mechanism module to achieve high-precision prediction of the time series.
It improves the accuracy and stability of the model in time series prediction, enhances the modeling and understanding of complex time series patterns, and has high generalization and prediction capabilities.
Smart Images

Figure CN120144960A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of meteorological prediction, and particularly relates to a time series prediction method, device, and medium based on the MPDLinear model. Background Art
[0002] In today's data-driven era, time series data widely exists in various fields, and its prediction task (TSF) is also widely applied to various scenarios, including but not limited to traffic flow prediction, energy scheduling, financial investment, smart wearables, intelligent weather, etc.; for a time series containing N variables (features), given historical data where L is the size of the lookback window (input time step), is the value of variable (feature) i at time step t, X represents the values of features 1 to N from the 1st time step to the Lth time step, and the time series prediction task is to predict the values of the sequence, which represents the values of features 1 to N in the next T time steps (prediction time steps). When T>1, it is multi-step prediction. The iterative multi-step prediction (IMS) method obtains multi-step prediction by learning a single-step prediction and gradually iteratively applying the single-step prediction. The direct multi-step prediction (DMS) directly predicts the values of the prediction target at multiple time steps at once. Compared with the prediction results of the direct multi-step prediction DMS, the mean square error of the IMS prediction is smaller. However, in the iterative process, each step of the IMS prediction depends on the prediction result of the previous step and is inevitably affected by error accumulation. Therefore, when there is a very accurate single-step predictor and T is relatively small, the accuracy of the IMS prediction is the best. On the contrary, when it is difficult to obtain an unbiased single-step predictor or T is very large, the accuracy of the DMS prediction is better.
[0003] Informer, Autoformer, and FEDformer represent the latest advancements in Transformer-based time series prediction models. Informer addresses the computational efficiency issue of ultra-long time series through a sparse attention mechanism. Autoformer captures trend and periodic features in time series with the help of an automatic decomposition mechanism. FEDformer combines frequency domain and time domain modeling to further improve performance and efficiency in complex long sequence prediction. The LTSF-Linear model is a new set of linear models proposed in recent years, and their prediction method is direct multi-step prediction (DMS). The core working method of Dlinear is to directly perform a linear mapping regression task on historical time series data through weighted operation in LTSF-Linear to predict the values of future time steps. However, the design flaws and limitations of the DLinear model, a variant of the LTSF-Linear model, are mainly manifested in its inability to effectively capture complex non-linear trends and medium- and long-term periodic trends and its neglect of the mutual relationships between different time steps and features. First, its ability to capture complex multivariate time series change patterns is limited, ignoring more complex components that may exist in the time series. Second, the way the prediction results are generated by the joint effect of multiple components is not reasonable and robust enough, with low flexibility. Third, in the DLinear model, multi-channel sharing modeling cannot handle the differences between multiple channels. Fourth, DLinear cannot effectively extract the complex mutual correlations between the time and feature dimensions of the dataset, resulting in the prediction accuracy of DLinear being overly dependent on long input sequences and having low prediction accuracy when the input sequence is short because it cannot capture complex change patterns.
[0004] In the prior art, the paper titled "Are Transformers Effective for Time Series Forecasting" (Zeng, A., Proceedings of the 39th International Conference on Machine Learning (ICML) (2022)) discloses a time series prediction model based on a simple linear layer. This model achieves prediction performance superior to that of Transformer-based time series prediction models through simple component decomposition of the input sequence and mapping with a single linear layer, and has the advantages of simplicity and efficiency. However, it has problems such as limited ability to capture time series change patterns, unreasonable and non-robust ways of generating the joint effect of multiple components.
[0005] In the prior art, an invention with the patent publication number CN202411439254 and the name "Ultra-short-term Wind Power Prediction Method Based on Improved Loss Function and Patch Time Series Transformer Network" discloses a method for ultra-short-term wind power prediction, mainly through an improved Transformer network structure and a multi-variable non-linear loss function to perform high-precision prediction of wind power; however, when dealing with multi-dimensional input data, this method has insufficient modeling ability for complex interactions between input features, especially when dealing with high-dimensional and noisy time series data, with low prediction accuracy and poor stability. Summary of the Invention
[0006] In order to overcome the above-mentioned shortcomings of the prior art, the purpose of the present invention is to propose a time series prediction method, device and medium based on the MPDLinear model. The time series prediction method decomposes time series data into different time series components through a multi-component decomposition module. The feature-time dimension hybrid self-attention mechanism module and the multi-channel independent modeling weighted prediction module cooperate to complete and obtain the preliminary prediction results of different time series components after applying attention. The self-learning weight vector module uses the above preliminary prediction results, and then through backpropagation calculation and loop traversal, obtains the component dimension weights and feature dimension weights of different time series components. Finally, the multi-component decomposition module performs weighted calculation on the preliminary prediction results according to the component dimension weights and feature dimension weights of different time series components and decomposes the method to obtain the final prediction result; this method overcomes the limitations of the LTSF-Linear-based model in dealing with complex multi-dimensional data, and at the same time maintains a complexity lower than that of the Transformer-baesd model, overall improving the prediction accuracy and stability of the model in time series.
[0007] To achieve the above object, the technical solutions adopted by the present invention are as follows:
[0008] In the first aspect, an MPDLinear (Multi-Part / Plus-Decomposition / DimensionLinear) model, the MPDLinear model includes a multi-component decomposition module, a self-learning weight vector module, a multi-channel independent modeling weighted prediction module, and a feature-time dimension hybrid self-attention mechanism module;
[0009] The multi-component decomposition module deserializes and normalizes the binary data stream into a structure acceptable to the MPDLinear model. The multi-component decomposition module further decomposes the acceptable structure into different time series components. At the same time, this module can perform weighted calculations on the component dimension weights and feature dimension weights of different time series components respectively, and reverse-sum the weighted calculation results according to the decomposition method of the multi-component decomposition module to obtain the final prediction result;
[0010] The self - learning weight vector module calculates and updates the parameters of the weights of different time - series components and the weights of different time - series features in the self - learning weight vector module during backpropagation using the preliminary prediction results of different time - series components. After looping through the parameters of the weights of different time - series components and the weights of different time - series features, the component - dimension weights and feature - dimension weights of different time - series components are obtained.
[0011] The feature - time dimension hybrid self - attention mechanism module can process different time - series components through the self - attention mechanism layer in the look - back window time dimension in this module to obtain the attention scores and attention weights of different time - series components. At the same time, different time - series components can be processed sequentially through the feature - dimension self - attention mechanism layer and the prediction window time dimension self - attention mechanism layer in this module to obtain the preliminary prediction results of different time - series components after applying attention.
[0012] After using different time - series components and the attention weights of different time - series components, the multi - channel independent modeling weighted prediction module obtains the time - series components with attention through matrix multiplication. This module can perform independent - channel modeling for each time - series component with attention.
[0013] In a second aspect, a time - series prediction method based on the MPDLinear model includes the following steps:
[0014] S1: After receiving the binary data stream, the multi - component decomposition module deserializes and normalizes the binary data stream into a structure that can be received by the MPDLinear model. The multi - component decomposition module decomposes the receivable structure into different time - series components and outputs them to the multi - channel independent modeling weighted prediction module, the self - learning weight vector module, and the feature - time dimension hybrid self - attention mechanism module respectively.
[0015] S2: After receiving the different time - series components in step S1, the feature - time dimension hybrid self - attention mechanism module performs dimension rearrangement processing and generates the differently dimension - rearranged time - series components. Then, the differently dimension - rearranged time - series components are processed through the self - attention mechanism layer in the look - back window time dimension in the feature - time dimension hybrid self - attention mechanism module to obtain the attention scores of different time - series components. The attention scores of different time - series components are further processed through this module to obtain the attention weights of different time - series components and output them to the multi - channel independent modeling weighted prediction module.
[0016] S3: After the multi-channel independent modeling weighted prediction module receives the different time series components described in step S1 and applies the attention weights of the different time series components described in step S2, it obtains the different time series components with attention through matrix multiplication. The multi-channel independent modeling weighted prediction module performs independent channel modeling on each different time series component with attention and outputs the preliminary result sequence to the feature-time dimension hybrid self-attention mechanism module;
[0017] S4: The feature-time dimension hybrid self-attention mechanism module obtains the preliminary result sequence described in step S3, and successively processes it through the feature dimension self-attention mechanism layer and the prediction window time dimension self-attention mechanism layer of this module, and obtains the preliminary prediction results of the different time series components after applying attention and outputs them to the self-learning weight vector module and the multi-component decomposition module;
[0018] S5: After the self-learning weight vector module receives the different time series components described in step S1 and the preliminary prediction results described in step S4, it uses the preliminary prediction results to calculate and update the parameters of the weights of the different time series components and the weights of the different time series features in the self-learning weight vector module during backpropagation. After looping through the parameters of the weights of the different time series components and the weights of the different time series features, it finally obtains the component dimension weights and feature dimension weights of the different time series components and outputs them to the multi-component decomposition module;
[0019] S6: The multi-component decomposition module receives the preliminary prediction results described in step S4, and respectively performs weighted calculations on the preliminary prediction results according to the component dimension weights and feature dimension weights of the different time series components described in step S5, and reversely sums the weighted calculation results according to the decomposition method of the multi-component decomposition module to obtain the final prediction result.
[0020] Further, the different time series components described in step S1 include: linear trend component (Trend), non-linear trend component (Nonlinear Trend), short-cycle component (Short-Cycle-Seasonality), medium and long-cycle component (Long-Cycle-Periodic), and residual noise component (Residual). The steps for extracting different time series components are as follows:
[0021] S11: Smooth the receivable structure described in step S1 of the input through simple moving average (SMA) based on a sliding window, and define a moving average (moving_avg) to implement simple moving average and obtain the linear trend component;
[0022] S12: The extraction of the non - linear trend component in the acceptable structure described in step S1 is completed by using double non - linear mapping. Specifically, first, a linear mapping is performed on the time dimension or the feature dimension respectively through a fully - connected layer (Linear layer), then the ReLU non - linear activation function is applied, and finally, another fully - connected layer (Linear) is used for linear transformation of the remaining dimensions, and the result of the double non - linear mapping, that is, the non - linear trend component, is finally obtained;
[0023] S13: First, the linear trend component and the non - linear trend component extracted in steps S11 and S12 are removed from the acceptable structure (i.e., the original sequence), and the remaining residual sequence is obtained. Then, a simple moving average (SMA) operation is applied to the remaining residual sequence to extract the short - period component. In this step, the value of the sliding window is smaller than the size of the sliding window for extracting the linear trend component in step S11;
[0024] S14: The remaining residual sequence in step S13 is further removed from the short - period component in step S13, so as to obtain the remaining sequence. The discrete Fourier transform (DFT) method is used to extract the medium - and long - period component from the remaining sequence;
[0025] S15: The remaining sequence in step S14 is further removed from the medium - and long - period component in step S14 to obtain the residual noise component.
[0026] Furthermore, in step S3, the multi - channel independent modeling weighted prediction module completes the modeling of independent channels through the feature - independent linear layer. The formula of the feature - independent linear layer is as follows:
[0027] (Y ( bs,pl ) =X ( bs,sl ) ·W ( sl,pl ) +bias ( bs, 1) )*channels
[0028] where, Y ( bs,pl ) is the output of each channel independent model, X ( bs,c,sl ) is the tensor input to the linear layer, W ( sl,pl ) is the weight matrix mapping the input time step, and bias ( bs, 1) is the bias vector of the linear model.
[0029] Furthermore, the processing of the feature dimension self-attention mechanism layer and the prediction window time dimension self-attention mechanism layer described in step S4 is as follows. The input is divided into the time dimension and the feature dimension. The time dimension is further divided into the input time step (look-back window) time dimension before entering the multi-channel independent modeling weighted prediction module and the prediction time step time dimension output from the multi-channel independent modeling weighted prediction module:
[0030] S41: The feature dimension, the input time step time dimension, and the prediction time step time dimension are respectively linearly transformed to generate their own query (Q) vectors, key (K) vectors, and value (V) vectors;
[0031] S42: Transpose each key vector described in step S42 to obtain K T , and then calculate the dot product of the query vector and K T (Q·K T ) to obtain the preliminary attention scores;
[0032] S43: Scale the output dimension of the dot product described in step S42 with the scaling factor to obtain the scaled dot product result, i.e., the attention scores;
[0033] S44: Perform the SoftMax operation on the attention scores described in step S43 to convert them into attention weights. The attention weights represent the attention of each element's query vector on all element key vectors;
[0034] Among them, for each query position in the sequence, SoftMax converts the attention scores into a probability distribution over all key vector positions. For the input vector as ( i,j ) = [as ( i, 1) , as ( i, 2) ,..., as ( i,n ) , the formula of SoftMax is:
[0035]
[0036] Among them, as = [as ( i, 1) , as ( i, 2) ,..., as ( i,n ) is the attention scores containing n elements, representing the correlation degree between qi and [k 1 , k 2 ,..., kn], exp(as (i, j ) ) is the exponent of the j-th element in the input attention score vector, and is the sum of all input attention score exponents;
[0037] S45: Use matrix multiplication (MatMul) to apply the attention weights described in step S44 to the value vector to obtain a preliminary prediction result of different temporal components after applying attention. Through an aggregation operation, output the preliminary prediction results of different temporal components after applying attention to the learning weight vector module and the multi-component decomposition module.
[0038] Furthermore, the specific process of step S5 is as follows:
[0039] S51: Define the preliminary prediction results of different temporal components described in step S4 as Si, where i = 1, 2,..., n. Assign a component dimension weight wi to different temporal components in the component dimension, and assign a feature dimension weight V[i, j] to each feature in Si in the feature dimension, where j = 1, 2,..., channels;
[0040] S52: During the backpropagation training process, update wi and V[i, j] described in step S51 through the backpropagation mechanism, and their gradients will be calculated according to the loss function Loss and thus continuously adjusting the component dimension weight and the feature dimension weight during the training process. The component dimension weight and the feature dimension weight are as shown in the following formula:
[0041]
[0042] where η is the learning rate, and wi and V[i, j] are the component dimension weight and the feature dimension weight of different temporal components respectively.
[0043] Furthermore, the specific process of step S6 is as follows:
[0044] S61: Use the preliminary prediction results of different temporal components described in step S4 and the component dimension weight wi described in step S52, which are different temporal components weighted by the component dimension weight, where By summing all different temporal components weighted by the component dimension weight, obtain the comprehensive output Fseries of the component dimension,
[0045] S62: Use the preliminary prediction results of different temporal components described in step S4, and weight each of its features according to the feature dimension weight described in step S52. The specific formula is as follows:
[0046]
[0047] Among them, is the weighted output of the j-th feature channel of the i-th different time series component, and V[i, j] is the corresponding feature dimension weight;
[0048] Summarize the channels after weighting all feature dimensions to obtain the comprehensive output Fchannels of the feature dimension,
[0049] S63: Obtain the final prediction result Ffinal by weighting the comprehensive output Fseries of the component dimension and the comprehensive output Fchannels of the feature dimension described in steps S61 and S62 respectively on the component dimension and the feature dimension. The specific formula is as follows:
[0050]
[0051] Among them, Si[:, j, :] represents the value of the i-th different time series component on the j-th channel.
[0052] Furthermore, the specific process of step S14 is as follows:
[0053] S141: Use the remaining sequence described in step S14 as the input for extracting medium- and long-term cycle components. The remaining sequence is a time series of length N {x 1 , x 2 ,..., xn}, and we split it into an even part and an odd part:
[0054] Even part: x 2 m, where
[0055] Odd part: x 2 m +1 , where
[0056] Expand using the formula of discrete Fourier transform (DFT) according to the even part and the odd part, and the following formula can be obtained:
[0057]
[0058] Further simplify the above formula to obtain:
[0059]
[0060] Among them
[0061] That is, the DFT of the even part
[0062] That is, the DFT of the odd part
[0063] Therefore, the following formula can be obtained:
[0064]
[0065] Through the above splitting, the discrete Fourier transform of the remaining sequence is decomposed into the discrete Fourier transforms of two subsequences with a length of , and this decomposition process can continue recursively until the length of the sequence is reduced to 1 (the smallest unit). Therefore, the discrete Fourier transform can be expressed by the following formula:
[0066]
[0067] where is a rotation factor (or called "rotation kernel"), and the sequence is recursively split into two subsequences with a length of : the even part and the odd part. At each recursive level, the discrete Fourier transform combines the results of the even and odd parts through addition and the product of the rotation factor to form the frequency-domain data;
[0068] S142: Perform high-pass filtering on the frequency-domain data described in step S141 and obtain the low-frequency component;
[0069] S143: Use the inverse Fourier transform (IFT) to convert the low-frequency component described in step S142 back to the time domain to generate the final medium- and long-period component, where the inverse Fourier transform (IFT) formula is as follows:
[0070]
[0071] where xn is the value of the time-domain signal at the nth time step, Xk is the complex value of the kth frequency component of the frequency-domain data, N is the total length of the signal (i.e., the number of frequency-domain points), is the rotation factor of the inverse Fourier transform, and j is the imaginary unit (j 2 = -1).
[0072] In a third aspect, an electronic device includes a memory and a processor:
[0073] Memory: Used to store a computer program for implementing the time series prediction method based on the MPDLinear model;
[0074] Processor: Used to implement the time series prediction method based on the MPDLinear model when executing the computer program.
[0075] Fourthly, a computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the time series prediction method based on the MPDLinear model is implemented.
[0076] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0077] 1. The multi-component decomposition module in the present invention decomposes the time series into multiple components such as linear trend, non-linear trend, short cycle, medium and long cycle, and residual noise, enhancing the adaptability of the model to diverse patterns in the time series;
[0078] 2. The self-learning weight vector module in the present invention introduces self-learning weight vectors in two dimensions of sequence and feature, realizes the weighted summation of components and features, and can dynamically and accurately balance the contributions of each sequence and feature;
[0079] 3. The multi-channel independent modeling weighted prediction module in the present invention solves the problem of homogenization in feature modeling by independently constructing linear layers for each feature channel, effectively improving the ability of the model to differentially model the characteristics of different channels;
[0080] 4. The feature-time dimension hybrid self-attention mechanism module in the present invention combines the self-attention mechanisms of feature and time dimensions, dynamically adjusts the attention of the prediction output to different time steps and features, and further improves the ability of the model to capture complex patterns.
[0081] In summary, the present invention not only overcomes the limitations of the LTSF-Linear-based model in processing complex multi-dimensional data, but also maintains a complexity lower than that of the Transformer-baesd model. Through the respective functions of the four modules and their coordinated cooperation, the accuracy and stability of the entire model in time series prediction are improved, the ability of the model to model and understand complex time series patterns is enhanced, and the present invention has high generalization and prediction capabilities for complex time series. BRIEF DESCRIPTION OF THE DRAWINGS
[0082] Figure 1 is a schematic flowchart of the time series prediction method based on the MPDLinear model of the present invention.
[0083] Figure 2 is a schematic structural diagram of the MPDLinear full-scale model of the present invention.
[0084] Figure 3 is a schematic structural diagram of a single layer of the double non-linear mapping of the present invention.
[0085] Figure 4 is a schematic structural diagram of the self-learning weight vector module of the present invention.
[0086] Figure 5 It is a schematic diagram of the structure of the multi-channel independent modeling weighted prediction module of the present invention.
[0087] Figure 6 It is a schematic diagram of the internal mapping structure of the feature independent linear layer in the multi-channel independent modeling weighted prediction module of the present invention.
[0088] Figure 7 It is a schematic diagram of the internal structure of the hybrid self-attention mechanism of the feature-time dimension hybrid self-attention mechanism module of the present invention.
[0089] Figure 8 It is a decomposition diagram of the original sequence for the humidity feature of the Beijing weather dataset in the first embodiment of the present invention.
[0090] Figure 9 It is a distribution diagram of the linear trend component for the humidity feature of the Beijing weather dataset in the first embodiment of the present invention.
[0091] Figure 10 It is a distribution diagram of the non-linear trend component for the humidity feature of the Beijing weather dataset in the first embodiment of the present invention.
[0092] Figure 11 It is a distribution diagram of the short-period component for the humidity feature of the Beijing weather dataset in the first embodiment of the present invention.
[0093] Figure 12 It is a distribution diagram of the medium- and long-period component for the humidity feature of the Beijing weather dataset in the first embodiment of the present invention.
[0094] Figure 13 It is a distribution diagram of the residual noise component for the humidity feature of the Beijing weather dataset in the first embodiment of the present invention. Detailed implementation manners
[0095] The following combines Figures 1 to 13 and the first embodiment to further elaborate on the present invention in detail:
[0096] In a first aspect, an MPDLinear (Multi-Part / Plus-Decomposition / DimensionLinear) model, the MPDLinear model includes a multi-component decomposition module, a self-learning weight vector module, a multi-channel independent modeling weighted prediction module, and a feature-time dimension hybrid self-attention mechanism module:
[0097] The multi-component decomposition module deserializes the binary data stream and normalizes it into a structure that can be received by the MPDLinear model. The multi-component decomposition module then decomposes the receivable structure into different temporal components. At the same time, this module can perform weighted calculations on the component dimension weights and feature dimension weights of different temporal components respectively, and reverse-sum the weighted calculation results according to the decomposition method of the multi-component decomposition module to obtain the final prediction result;
[0098] The self-learning weight vector module calculates and updates the parameters of the weights of different temporal components and different temporal feature weights in the self-learning weight vector module during backpropagation using the preliminary prediction results of different temporal components. After looping through the parameters of the weights of different temporal components and different temporal feature weights, the component dimension weights and feature dimension weights of different temporal components are obtained;
[0099] The feature-time dimension hybrid self-attention mechanism module can process different temporal components through the self-attention mechanism layer in the look-back window time dimension in this module to obtain the attention scores and attention weights of different temporal components. At the same time, it can process different temporal components through the feature dimension self-attention mechanism layer and the prediction window time dimension self-attention mechanism layer in this module in sequence to obtain the preliminary prediction results of different temporal components after applying attention;
[0100] The multi-channel independent modeling weighted prediction module uses different temporal components and the attention weights of different temporal components, and then obtains the temporal components with attention through matrix multiplication. This module can perform independent channel modeling for each temporal component with attention;
[0101] In a second aspect, a temporal prediction method based on the MPDLinear model, the MPDLinear model includes a multi-component decomposition module, a self-learning weight vector module, a multi-channel independent modeling weighted prediction module, and a feature-time dimension hybrid self-attention mechanism module, Figure 2 shows the overall structure of the MPDLinear full model, where the multi-component decomposition module decomposes time series data into different temporal components, the self-learning weight vector module introduces self-learning weight vectors in the component and feature dimensions of different temporal components to achieve weighted summation of components and features; the multi-channel independent modeling weighted prediction module independently constructs a linear layer for each feature channel; the feature-time dimension hybrid self-attention mechanism module combines the self-attention mechanisms of the feature and time dimensions to dynamically adjust the attention of the prediction output to different time steps and features.
[0102] As Figure 1 shown is the flow of the temporal prediction method in the present invention. The temporal prediction method includes the following steps:
[0103] Preprocessing: The multivariate time series data (i.e., multivariate time - series data) is processed into a binary data stream through the Jackson framework and output to the multi - component decomposition module. The multivariate time series data contains multiple features (i.e., channels) within multiple time steps, and each feature represents a time - related variable.
[0104] S1: After receiving the binary data stream, the multi - component decomposition module deserializes and normalizes the binary data stream into a structure that can be received by the MPDLinear model. The multi - component decomposition module decomposes the receivable structure into different time - series components and outputs them to the multi - channel independent modeling weighted prediction module, the self - learning weight vector module, and the feature - time dimension hybrid self - attention mechanism module respectively. The different time - series components obtained by decomposition will be used for the calculations of the self - learning weight vector module, the multi - channel independent modeling weighted prediction module, and the feature - time dimension hybrid self - attention mechanism module.
[0105] S2: After receiving the different time - series components described in step S1, the feature - time dimension hybrid self - attention mechanism module performs dimension rearrangement processing. Specifically, a simple dimension rearrangement (permute) is performed using the dimension rearrangement function in pytorch to generate the differently dimension - rearranged time - series components. Then, the differently dimension - rearranged time - series components are processed through the self - attention mechanism layer in the time dimension of the look - back window in the feature - time dimension hybrid self - attention mechanism module to obtain the attention scores of the different time - series components. The attention scores of the different time - series components are the attention scores of the different time - series components in the time dimension of the look - back window. The attention scores of the different time - series components are then processed through the Softmax function of this module to obtain the attention weights of the different time - series components and output them to the multi - channel independent modeling weighted prediction module. The attention weights of the different time - series components are the attention weights at the same batch and the same channel for different time steps.
[0106] S3: After receiving the different time - series components described in step S1 and applying the attention weights of the different time - series components described in step S2, the multi - channel independent modeling weighted prediction module obtains the attention - weighted different time - series components through matrix multiplication. The multi - channel independent modeling weighted prediction module performs independent - channel modeling for each of the attention - weighted different time - series components and outputs a preliminary result sequence to the feature - time dimension hybrid self - attention mechanism module.
[0107] Figure 2 The purple - overlaid square in the middle represents that each channel (feature) in each time - series component is modeled through an independent linear layer (such as Trend_Channel[i]_Linear, Nonlinear_Trend_Channel[i]_Linear, etc., where i is the channel number).
[0108] S4: The hybrid self-attention mechanism module in the feature-time dimension obtains the preliminary result sequence described in step S3, and successively processes it through the self-attention mechanism layer in the feature dimension and the self-attention mechanism layer in the prediction window time dimension of this module to obtain the preliminary prediction results of different time series components after applying attention and output them to the self-learning weight vector module and the multi-component decomposition module;
[0109] S5: After receiving the different time series components described in step S1 and the preliminary prediction results described in step S4, the self-learning weight vector module uses the preliminary prediction results to calculate and update the parameters of the weights of different time series components and the weights of different time series features in the self-learning weight vector module during backpropagation. After looping through the parameters of the weights of different time series components and the weights of different time series features, finally, the component dimension weights and feature dimension weights of different time series components are obtained and output to the multi-component decomposition module; Among them, during the model training stage, this module will update the initial parameters during backpropagation and dynamically adjust the effective weights of different components and channels according to the output during forward propagation;
[0110] S6: The multi-component decomposition module receives the preliminary prediction results described in step S4, and respectively performs weighted calculations on the preliminary prediction results according to the component dimension weights and feature dimension weights of different time series components described in step S5, and reversely sums the weighted calculation results according to the decomposition method of the multi-component decomposition module to obtain the final prediction result.
[0111] The present invention designs a multi-component decomposition module, which uses different decomposition methods to extract and model each component to better capture various change patterns in complex time series.
[0112] Further, the different time series components described in step S1 include: linear trend component (Trend), non-linear trend component (Nonlinear Trend), short-cycle component (Short-Cycle-Seasonality), medium- and long-cycle component (Long-Cycle-Periodic), and residual noise component (Residual). The extraction steps of different time series components are as follows:
[0113] S11: Linear trend component extraction, the purpose of which is to extract the long-term trend in the time series and eliminate short-term fluctuations; Smooth the input receptive structure based on the simple moving average (SMA) of the sliding window. Specifically, define the moving average (moving_avg) to implement the simple moving average, and obtain the linear trend component after the moving average;
[0114] The one-dimensional average pooling layer nn.AvgPool1d is used to apply a sliding window on the time series to calculate the mean value. By adjusting the stride, the calculation frequency is controlled. The calculation frequency determines the time points to be included in the moving average. The formula is as follows:
[0115]
[0116] Among them, SMAt represents the moving average value at the t-th time step; Xt represents the value of the time series at the t-th time step; k is the size of the sliding window, indicating the number of time steps to be considered when calculating the moving average;
[0117] S12: Nonlinear trend component extraction. The purpose of this component is to further capture complex long-term nonlinear changes on the basis of the trend component, so as to fit the mutation points or irregular points in the data as much as possible. The core of this part is double nonlinear mapping;
[0118] As Figure 3 shown, double nonlinear mapping is used to complete the extraction of the nonlinear trend component in the acceptable structure described in step S1. Specifically, first, a linear mapping is performed on the time dimension or feature dimension respectively through a fully connected layer Linear layer to learn the linear relationship between each time step or feature, and then the ReLU nonlinear activation function is applied to enable the model to fit more complex relationships. Finally, another fully connected layer Linear is used for linear transformation of the remaining dimensions, and the result of the double nonlinear mapping, that is, the nonlinear trend component, is finally obtained;
[0119] In actual calculation, first, a nonlinear mapping is performed on the trend sequence in the time step dimension to obtain the nonlinear trend component in the time step dimension, and then the nonlinear trend component in the time step dimension is nonlinearly mapped in the feature dimension, so as to finally obtain the result of the double nonlinear mapping, that is, the nonlinear trend component. This component can model from both the time and feature dimensions simultaneously and is applicable to time series data with complex multi-dimensional relationships and nonlinear trends;
[0120] S13: To identify and extract short-periodic fluctuations in the time series, the seasonal component (SeasonSeries), i.e., the short-cycle component (Short-Cycle-Seasonality), will be extracted. First, the linear trend component and the non-linear trend component already extracted in steps S11 and S12 are removed from the receivable structure (i.e., the original sequence) to obtain the remaining residual sequence. Since the trend sequence we extracted is the long-term trend in the time series and eliminates short-term fluctuations, then subtracting in the reverse direction, what is naturally obtained is the short-term fluctuations and other components. Then, the simple moving average (SMA) operation is applied to the remaining residual sequence to extract the short-cycle component. In this step, the value of the sliding window is smaller than the size of the sliding window for extracting the linear trend component in step S11.
[0121] For extracting short-term and stable seasonal components, the stride is still 1 to extract a complete component, and padding operations will also be performed to fill the data before the moving average. In this process, we extracted the short-cycle component from the residual sequence remaining after removing the long-term trend component.
[0122] S14: Next is the extraction of the medium- and long-cycle component (Periodic Series), with the aim of capturing medium- and long-periodic fluctuations in the time series and, together with the long-term trend, modeling and extracting the long-term pattern in the time series.
[0123] The remaining residual sequence in step S13 is further removed from the short-cycle component in step S13 to obtain the remaining sequence, and the discrete Fourier transform (DFT) method is used to extract the medium- and long-cycle component from the remaining sequence.
[0124] The core of the medium- and long-cycle component extraction is the method of discrete Fourier transform (Discrete Fourier Transform, DFT). The Fourier transform can convert the time-domain signal into frequency-domain data, and use the frequency-domain data to reveal the amplitude and phase characteristics in the signal to extract the periodic characteristics of the time series.
[0125] Furthermore, the specific process of step S14 is as follows:
[0126] S141: The remaining sequence in step S14 is used as the input for extracting the medium- and long-cycle component. The remaining sequence is a time series {x 1 , x 2 ,..., xn} of length N, and we split it into an even part and an odd part:
[0127] Even part: x 2 m, where
[0128] Odd part: x 2 m +1 , where
[0129] Expanding using the formula of the discrete Fourier transform (DFT) according to the even part and the odd part, the following formula can be obtained:
[0130]
[0131] Further simplifying the formula in the previous step, we get:
[0132]
[0133] Here:
[0134] That is, the DFT of the even part
[0135] That is, the DFT of the odd part
[0136] Therefore, the following formula can be obtained:
[0137]
[0138] Through the above splitting, the discrete Fourier transform of the remaining sequence is decomposed into the discrete Fourier transforms of two subsequences with lengths of . This decomposition process can continue recursively until the length of the sequence is reduced to 1 (the smallest unit). Therefore, the discrete Fourier transform can be expressed by the following formula:
[0139]
[0140] Where, is the rotation factor (or called "rotation kernel"), and the sequence is recursively split into two subsequences with lengths of : the even part and the odd part. At each recursive level, the discrete Fourier transform combines the results of the even and odd parts through addition and the product of the rotation factor to form the frequency-domain data;
[0141] S142: Perform high-frequency filtering on the frequency-domain data described in step S141 and obtain the low-frequency components. This module filters out high-frequency noise by setting a frequency filtering threshold and only retains the low-frequency components, and these low-frequency components correspond to the medium- and long-term periodic fluctuations;
[0142] S143: Use the inverse Fourier transform (IFT) to convert the low-frequency components described in step S142 back to the time domain to generate the final medium- and long-term periodic components. The formula of the inverse Fourier transform (IFT) (also called the inverse discrete Fourier transform (IDFT)) is as follows:
[0143]
[0144] Among them, xn is the value of the time-domain signal at the nth time step, Xk is the complex value of the kth frequency component of the frequency-domain data, and N is the total length of the signal (i.e., the number of frequency-domain points). is the rotation factor of the inverse Fourier transform, and j is the imaginary unit (j 2 = -1);
[0145] S15: Remove the medium- and long-period components in the remaining sequence described in step S14 from the remaining sequence described in step S14 to obtain the residual noise component.
[0146] The residual noise component represents the random fluctuations and high-frequency noise in the sequence after removing the non-linear trend, and can identify the high-frequency noise and irregular fluctuations in the time series that cannot be explained by other components. It provides information about the data uncertainty in the prediction process and helps to improve the robustness of the model.
[0147] The MPDLinear model introduces a self-learning weight vector module, which can capture more information with this cross-dimensional adaptive mechanism. By introducing self-learning weight vectors in the sequence dimension and the feature dimension, the model can dynamically adjust these weights according to the data characteristics during the training process, continuously learn the importance of different time series components (such as trends, seasonality, periodicity, and residuals) and different features, and allow the model to dynamically adjust the output contribution of each channel for more fine-grained weight allocation, so as to achieve a more detailed weighted combination, and finally integrate and take effect on the output, reducing the error caused by a single dimension, enabling the model to better process multi-dimensional and multi-channel time series features, and having a stronger modeling ability for the prediction output of complex time series data, especially time series data with dense channels.
[0148] Such as Figure 4 shown, the multiple input sequences in the leftmost part are the multiple sequences decomposed by the multi-component decomposition module, that is, different time series components.
[0149] Further, the specific process of step S5 is as follows:
[0150] S51: Define the preliminary prediction results of the different time series components described in step S4 as Si, where i = 1, 2,..., n. Assign a component dimension weight wi to the different time series components in the component dimension, and assign a feature dimension weight V[i,j] to each feature in Si in the feature dimension, where j = 1, 2,..., channels;
[0151] S52: During the backpropagation training process, update the wi and V[i,j] described in step S51 through the backpropagation mechanism, and their gradients will be calculated according to the loss function Loss and Therefore, during the training process, the weights of the component dimensions and the weights of the feature dimensions are continuously adjusted. The weights of the component dimensions and the weights of the feature dimensions are as shown in the following formula:
[0152]
[0153] Among them, η is the learning rate, and wi and V[i,j] are the weights of the component dimensions and the weights of the feature dimensions of different time series components, respectively.
[0154] Furthermore, the specific process of step S6 is as follows:
[0155] S61: Using the preliminary prediction results of different time series components described in step S4 and the component dimension weight wi described in step S52, It is the different time series components after being weighted by the component dimension weights, where By summing up the different time series components weighted by all component dimension weights, the comprehensive output Fseries of the component dimension is obtained.
[0156] After summarizing the contributions of each component sequence, it is necessary to perform weighting on the feature dimension next. We introduce a feature dimension weight V[i,j] on the feature (channels) dimension, whose shape is (n, channels). The first dimension represents the number of different time series components, and the second dimension represents the number of features of each different time series component, that is, the number of feature channels;
[0157] S62: Using the preliminary prediction results of different time series components described in step S4, weight each of its features according to the feature dimension weight described in step S52. The specific formula is as follows:
[0158]
[0159] Among them, is the weighted output of the j-th feature channel of the i-th different time series component, and V[i,j] is the corresponding feature dimension weight;
[0160] Through backpropagation in step S52, it is learned during the training process. After calculating the weighted feature channels of all different time series components, we sum up all the weighted channels of the feature dimension to obtain the comprehensive output Fchannels of the feature dimension.
[0161] S63: The final prediction result Ffinal is obtained by weighting the comprehensive output Fseries of the component dimension and the comprehensive output Fchannels of the feature dimension described in steps S61 and S62 on the component dimension and the feature dimension, respectively. The specific formula is as follows:
[0162]
[0163] Among them, Si[:, j, :] represents the value of the i-th different time series component on the j-th channel.
[0164] Specifically, first, each different time series component is weighted by wi in the component dimension, and then each feature is weighted by V[i, j] in the feature dimension. Finally, through double weighted summation, multi-dimensional dynamic adjustment and combination of time series data are realized.
[0165] In the DLinear model, all input channels (features) share the same linear layer. This design fails to distinguish the differences between different features, restricting the model's ability to utilize the unique information of each channel. In actual time series data, different channels (features) may exhibit completely different patterns, and the way that all channels share the same linear layer cannot fully adapt to the differences between different channels, resulting in the model being unable to utilize the information of all channels. Instead, a unified fitting is performed on all channels, and the fitting effect is also general and average. Because the model does not have enough parameters to learn the change patterns of multiple features over time, only a general "rough" mapping process is done, and the model prediction performance is poor.
[0166] The multi-channel independent modeling weighted prediction module in this method can better capture the features of each independent channel and improve the model's performance by means of multi-channel independent modeling and weighted prediction in cooperation with the self-learning weight vector module.
[0167] Such as Figure 5 Shown is the internal structure of the multi-channel independent modeling weighted prediction module. The core of this module's design lies in its flexibility and the independent processing ability of channels; through experimental parameters, the model can be controlled to choose between independent modeling and shared modeling. It allows the model to either let all input channels (features) share a linear layer or build a dedicated independent linear layer for each feature channel, thereby providing a customized prediction model for each feature. Among them, the feature sharing linear layer modeling is suitable for the situation where the dataset is small, the number of features is small, and the relationships between features are relatively homogeneous. All channels share the same linear layer to reduce the model complexity and improve the calculation efficiency, which is marked with a red cross in the figure; while the feature independent linear layer is suitable for the situation where the dataset is large, the number of features is large, and there are complex relationships between features, which is marked with a green tick in the figure. When conducting experiments on complex datasets in this paper, the feature independent modeling method of the multi-channel independent modeling weighted prediction module will be adopted.
[0168] Similar to the non-linear trend component extraction part in the multi-component decomposition module, the multi-channel independent linear layer modeling in this module uses the nn.Linear module in pytorch to define a fully connected layer to perform a linear transformation. However, there are slight differences in the shape changes of the feature-sharing linear layer and the feature-independent linear layer for All_Channels. At the same time, different linear layers for each sequence series are separately loaded by the nn.ModuleList container in pytorch.
[0169] Furthermore, in step S3, the multi-channel independent modeling weighted prediction module completes the modeling of independent channels through the feature-independent linear layer. The formula for the feature-independent linear layer is as follows:
[0170] (Y ( bs,pl ) =X ( bs,sl ) ·W ( sl,pl ) +bias ( bs, 1) )*channels
[0171] Where, Y ( bs,pl ) is the output of each channel independent model, X ( bs,c,sl ) is the tensor input to the linear layer. The tensor X ( bs,sl ) has a shape of (batch_size, seq_len) and directly inputs a certain decomposition sequence (such as trend_series) of a certain inputs (including batch_size batches) in train_loader. W ( sl,pl ) is the weight matrix mapping the input time step, and bias ( bs, 1) is the bias vector of the linear model.
[0172] A list of module containers (Module0List) stores independent linear layers for multiple features of a certain sequence (e.g., the independent linear layer for the seasonal sequence is ModuleList as Linear_Seasonal). When we input the data of a certain channel of a certain sequence (seasonal_init) to the linear layer, it is Linear_Seasonal[i](seasonal_init[:,i,:]), representing the Linear layer of the i-th channel, with a shape of (batch_size,seq_len), and will be mapped to the output Y ( bs,pl ) , with a shape of (batch_size,pred_len).
[0173] Figure 6 It shows the mapping structure of each scalar entering the independent linear layer of the multi-channel independent modeling weighted prediction module during this process. The linear layer performs a linear transformation on the input, uses the weight matrix W(12,5), and maps the input of the size of the lookback window (seq_len) to the output of the size of the prediction length (pred_len). After the linear transformation, a bias term is added to adjust the value of each output node. The bias term helps the model to have an output benchmark even without input data, enhancing the expressive ability of MPDLinear. Finally, the output layer contains the prediction results of the model for the next 5 time steps.
[0174] DLinear performs a unified weighted processing on all input features and time steps, without distinguishing the importance of different features at different time steps, nor the importance of different time steps for different features. The lack of a mechanism that can handle the interaction between features and time steps simultaneously makes DLinear only able to learn more information about time series features and make better predictions by inputting more temporal data (a larger lookback window) for each inference. In addition, this fixed and uniform way causes the model to fail to capture the complex feature dependencies in the time series and the importance differences of some features changing over time, so it may miss some key patterns and information.
[0175] To address the above deficiencies, the present invention designs a feature-time dimension hybrid self-attention mechanism module. The self-attention mechanism of this module uses the Scaled Dot-Product Attention. At the same time, we capture the corresponding attention weights in two time step dimensions (input time step, prediction time step) and the feature (channel) dimension, and use them in combination, applying them to the input and output sequences respectively, so that the model can learn the features and time relationships in the data in a biased manner, and more flexibly capture the complex patterns in the multi-dimensional time series.
[0176] Figure 7 Figure 4 shows the hybrid calculation workflow of the hybrid self-attention mechanism in the feature-time dimension hybrid self-attention mechanism module in the feature and time dimensions.
[0177] Furthermore, the processing procedures of the feature dimension self-attention mechanism layer and the prediction window time dimension self-attention mechanism layer in step S4 are as follows. The input is divided into a time dimension and a feature dimension. The time dimension is further divided into the input time step (look-back window) time dimension before entering the multi-channel independent modeling weighted prediction module and the prediction time step time dimension output from the multi-channel independent modeling weighted prediction module:
[0178] S41: The feature dimension, the input time step time dimension, and the prediction time step time dimension respectively generate their own query (Q) vectors, key (K) vectors, and value (V) vectors through linear transformation. Among them, the query vector represents the query of the current element, used to obtain relevant attention information. The key vector represents the features of all elements in the sequence, used to match with the query vector to calculate the similarity between features, obtain the attention score and attention weight. The value vector contains the actual information of the elements, used to perform weighted average matrix multiplication according to the attention weight to generate the final output.
[0179] S42: The Q, K, and V of each dimension will enter the Scaled Dot-Product Attention part. The structure of this part is as shown in Figure 5. In this part, the transpose of each key vector described in step S42 is obtained to get K, and then the dot product (Q·K) of the query vector and K is calculated to capture the relationship between different time steps and features, obtaining the preliminary attention score. Figure 7 as shown, and in this part, the transpose of each key vector described in step S42 is obtained to get K T , and then the dot product of the query vector and K T (Q·K T ) is calculated to capture the relationship between different time steps and features, obtaining the preliminary attention score.
[0180] S43: To make the attention scores distributed in a smaller range, the output dimension of the dot product described in step S42 is scaled by the scaling factor to obtain the scaled dot product result, that is, the attention score.
[0181] S44: Perform a SoftMax operation on the attention scores described in step S43 to convert them into attention weights, where the attention weights represent the attention of the query vector of each element on all element key vectors;
[0182] Among them, for each query position in the sequence, SoftMax converts the attention scores into a probability distribution over all key vector positions. For the input vector as ( i,j ) = [as ( i, 1) , as ( i, 2) ,..., as ( i,n ) , and the formula of SoftMax is:
[0183]
[0184] Among them, as = [as ( i, 1) , as ( i, 2) ,..., as ( i,n ) is the attention score containing n elements, representing the correlation degree between qi and [k 1 , k 2 ,..., kn]. exp(as ( i,j ) ) is the exponent of the j-th element in the input attention score vector, and is the sum of the exponents of all input attention scores;
[0185] This means that for each query Q, the degree of attention of this position to other positions is converted into a probability value, and the sum of all probability values (attention degrees) is 1. In addition, in order to eliminate the influence of previously filled data, we set the attention weights of the padding part of the input to 0 in advance (the attention scores are set to -inf in advance);
[0186] S45: Apply the attention weights (of all query vectors to all key vectors) described in step S44 to the value vectors using matrix multiplication (MatMul), and perform a weighted sum of the information at different positions through the attention weights to extract important information from the Value vectors, thereby obtaining a more focused feature representation, dynamically selecting and emphasizing the important features and time steps relevant to the current prediction task, while weakening the attention to unimportant features and time steps, to obtain the preliminary prediction results of different temporal components after applying attention. Through an aggregation operation, output the preliminary prediction results of different temporal components after applying attention to the learning weight vector module and the multi-component decomposition module.
[0187] Embodiment 1
[0188] In this Embodiment 1, the processed Beijing weather dataset (avg_data_beijing) in the Chinese weather dataset (weather_2013_2017_china) is used, and it is decomposed and visualized using the multi-component decomposition module. In this paper, the training set part in the Beijing weather dataset is used to construct a data loader (DataLoader) required for model training using the Pytorch framework, and the dataset is shuffled. All time steps of a certain batch are randomly selected. Here, when conducting experiments on the Chinese weather dataset, since the known best configuration (BestKnown Config, BKC) of the batch size (batch_sizebatch_size), lookback window size (seq_len), and prediction window size (pred_len) of MPDLinear on this dataset has been obtained through comparative experiments, and the lookback window size (seq_len) = 365 is the best, the lookback window size (seq_len) is 365 during decomposition here, and the sampling interval of this dataset is days.
[0189] In Embodiment 1, the convolution kernel size (kernel_size) is uniformly set to 25 so that the smoothing window contains enough data points and is applicable to various sampling frequencies (such as: minutes, hours, days, months). This window size is most suitable for removing short-term fluctuations and retaining long-term trends; the stride is uniformly set to 1. When stride = 1, there is an average value corresponding to each time step t to obtain a smoothed complete time series as the trend component without skipping any intermediate time steps. At the same time, to ensure that the calculation of the sliding window at the sequence boundary does not lose data, we add padding when calculating the moving average, padding some data at both ends of the sequence so that the sliding window can be smoothed from the first time step to the last time step. The number of elements of the inner padding is:
[0190] In the process of extracting the medium- and long-term periodic components (Periodic Series), high-frequency noise is filtered out by setting a frequency filtering threshold, and only low-frequency components are retained. These low-frequency components correspond to the medium- and long-term periodic fluctuations. In the first embodiment: In the code implementation of this module, the first 10% of the frequency-domain components are retained as low-frequency components, and the remaining 90% of the high-frequency components are set to zero.
[0191] Figure 8 Shows the decomposition of the original sequence by the multi-component decomposition module on the humidity feature of the Beijing weather dataset. Figure 8 The blue part in it is the original sequence, representing the change of the humidity feature in Beijing over the entire time range (365 days). Figure 9 Shows the distribution of the linear trend component on the humidity feature of the Beijing weather dataset. Figure 9 The orange part in it is the linear trend component, which is extracted by the moving average method. It shows the overall change trend of humidity over time, revealing the rising and falling laws of humidity within a year (365 days). Figure 10 Shows the distribution of the non-linear trend component on the humidity feature of the Beijing weather dataset. Figure 10 The green part in it is the non-linear trend component, which is the possible complex pattern extracted by non-linear mapping based on the linear trend component. Figure 10 As can be seen from it, the non-linear trend component has both small and large amplitude changes. These changes are often some subtle trend fluctuations and non-linear fluctuations, reflecting the details that are difficult to capture by simple linear methods in the change of humidity over time. Figure 11 Shows the distribution of the short-period component on the humidity feature of the Beijing weather dataset. Figure 11 The red part in it is the short-period component, which mainly extracts the short-term and repetitive fluctuations in the original sequence. Figure 11 As can be seen from it, the short-period component captures the high-frequency periodic changes of humidity in a relatively short time range. The fluctuation curve has many spikes, representing different changes of humidity in each short time period. Each spike represents a possible change pattern within three days or one week or one month. Figure 12 Shows the distribution of the medium- and long-term periodic component on the humidity feature of the Beijing weather dataset. Figure 12 The purple part in it is the medium- and long-term periodic component, which is mainly obtained through discrete Fourier transform (DFT) and high-frequency filtering and low-frequency extraction. It shows an obvious periodic fluctuation, which repeats multiple times within a year. Compared with the short-period fluctuations of the seasonal component, the periodic component has a longer period, a larger amplitude, and is smoother. This is because the medium- and long-term periodic component pays more attention to the medium- and long-term changes of the sequence and ignores the high-frequency short-term changes, and the high-frequency short-term changes are obtained by the seasonal component.Figure 13 shows the distribution of the residual noise component on the humidity feature of the Beijing weather dataset, Figure 13 where the brown part is the residual noise component, which is the final remaining component after removing all the above components. It contains relatively complex high-frequency fluctuations, which may be caused by noise, data measurement errors, or some short-term and unpredictable emergencies. This component is also used to obtain the subsequent learning of the model.
[0192] Through the multi-component decomposition module, we split the complex original sequence into multiple fine-grained components, including the linear trend component (Trend), the non-linear trend component (Nonlinear Trend), the short-cycle component (Short-Cycle-Seasonality), the medium- and long-cycle component (Long-Cycle-Periodic), and the residual noise component (Residual). Each component can individually represent a specific time series feature, enabling the model to understand the change patterns of the time series at different time scales. The model can not only accurately capture the long-term and short-term changes in the data, but also effectively model the non-linear and complex structures in the data, thereby improving the overall prediction performance of the multivariate time series.
[0193] The self-learning weight vector module realizes dynamic weighting in two dimensions of components and features (channels) by introducing the self-learning weight vector, fully utilizes the information of multiple different time series components, and performs weight adjustment across dimensions within the same component. The dynamic weight adjustment and self-learning and adaptive capabilities of the self-learning weight vector module can automatically learn and adjust the prediction weights according to the characteristics of the time series data. By introducing the multi-dimensional self-learning weight vector, the model can targetedly weight the features of different components and channels, thereby improving the prediction accuracy and flexibility. The dynamic weight adjustment mechanism adaptively allocates weights to ensure that the contributions of each component in the multi-variable time series are reasonably measured. For example, some components have a greater impact on the prediction result, while the role of others is smaller. The dynamic adjustment can amplify the weights of important features while weakening the influence of noise or irrelevant features, thereby enhancing the ability to capture key trends. The self-learning and adaptive capabilities further improve the adaptability of the model. The weight vector does not rely on empirical values, but is automatically adjusted through backpropagation and optimization algorithms. This not only reduces the workload of hyperparameter tuning, but also ensures that the model can maintain high robustness in different datasets and tasks. As the time series changes, the self-learning mechanism can continuously optimize the weight allocation, enabling the model to always maintain sensitivity to changes in the data distribution, thereby improving the accuracy and stability of the prediction. Through the dynamic allocation and self-learning adjustment of weights, the model can better mine the important information in the time series, give full play to the value of each feature, and significantly improve the prediction performance and generalization ability.
[0194] The multi-channel independent modeling weighted prediction module constructs independent linear layers for each feature channel, differentiates the modeling of the characteristics of different channels, and more precisely captures the unique change patterns of each channel. It solves the "averaging" defect that DLinear cannot distinguish different features, making MPDLinear have higher adaptability and performance in scenarios where each feature in multivariate time series data has significant heterogeneity or complex interaction relationships. The multi-channel independent modeling weighted prediction module adopts the method of independent modeling, modeling the characteristics of each channel separately, thus significantly improving the model's ability to capture complex multivariate time series. In actual time series data, the feature distributions and dynamic changes of different channels often vary significantly. Compared with the traditional method that shares a unified model weight, independent modeling can perform more fine-grained modeling of the characteristics of each channel, ensuring that the model fully utilizes the unique information of each channel. In this way, the model can not only more accurately fit the patterns of each channel but also avoid interference between different channels. The advantage of independent modeling lies in higher flexibility and adaptability. When the feature pattern of one channel is relatively simple while that of another channel is complex, the model can adjust the structure and parameters for these two situations separately, so that the complex pattern receives more attention while the simple pattern maintains a lower modeling complexity. This differential modeling strategy can reduce unnecessary computational costs and improve the overall generalization ability and prediction performance of the model. In addition, through independent modeling, the multi-channel independent modeling weighted prediction module can better capture the complex interaction patterns between channels. Although each channel is modeled independently, the model can comprehensively consider the features of these independent channels in subsequent stages through a flexible weighting mechanism or interaction module, ensuring that the dependence relationships and complex patterns between channels are fully explored and utilized. This design effectively enhances the model's performance when facing multi-dimensional and high-complexity time series data, enabling it to retain the independence of channels while fully utilizing the interaction information between channels to improve prediction accuracy and modeling stability.
[0195] The Feature-Time Dimension Hybrid Self-Attention Mechanism Module realizes a more refined modeling of multivariate time series data by introducing the scaled dot-product self-attention mechanism in both the feature and time dimensions, enabling the MPDLinear model to handle complex interactions between the feature and time dimensions. Compared with DLinear, MPDLinear will no longer rely too much on long input sequences (lookback windows) to obtain more accurate prediction results, and the change in prediction accuracy of MPDLinear will be reduced under different input sequence lengths. The Feature-Time Dimension Hybrid Self-Attention Mechanism Module focuses on effectively matching the heterogeneity of the time-feature dimension and enhances the model's adaptability to complex time series data through a dynamic attention mechanism. In multivariate time series, the correlations between features and between time steps are often uneven and constantly changing, while traditional methods usually adopt a unified weight or fixed processing method for all dimensions, making it difficult to fully capture the subtle dynamics in the data. To solve this problem, the Feature-Time Dimension Hybrid Self-Attention Mechanism Module introduces a dynamic hybrid attention mechanism that can dynamically adjust the attention weights in the time dimension and the feature dimension, thereby allocating different degrees of attention between different dimensions according to the changes in the input data, mining short-term shocks and long-term trends in the time series, and capturing complex interaction patterns between features. This dynamic attention method enables the model to flexibly adapt to diverse changes in data characteristics and improves the ability to capture potential structures in the time series. In addition, the Feature-Time Dimension Hybrid Self-Attention Mechanism Module identifies the cases where different features are more important in certain time periods through a self-learning heterogeneity matching scheme and dynamically adjusts the prediction weights. This method reduces the dependence on hyperparameter tuning and improves the self-adaptability and stability of the model. When processing multivariate time series data, the Feature-Time Dimension Hybrid Self-Attention Mechanism Module demonstrates stronger flexibility and accuracy, can effectively capture complex patterns and adapt to different task requirements, and improves the prediction accuracy and generalization ability of the model.
[0196] In a third aspect, an electronic device includes a memory and a processor:
[0197] Memory: for storing a computer program for implementing the time series prediction method based on the MPDLinear model;
[0198] Processor: for implementing the time series prediction method based on the MPDLinear model when executing the computer program.
[0199] Fourthly, a computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the time series prediction method based on the MPDLinear model is implemented. The computer-readable storage medium includes various media that can store program codes, such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs.
[0200] The working principle of the present invention is as follows:
[0201] In this method, a high-precision time series prediction model (MPDLinear) is proposed. The model includes a multi-component decomposition module, a self-learning weight vector module, a multi-channel independent modeling weighted prediction module, and a feature-time dimension hybrid self-attention mechanism module. Among them, the multi-component decomposition module decomposes the time series into four components: linear trend, short cycle, medium-long cycle, residual, and non-linear trend, enhancing the model's adaptability to diverse patterns in the time series; secondly, the self-learning weight vector module introduces self-learning weight vectors in both the time series and feature dimensions to achieve weighted summation of components and features, enabling dynamic and precise balancing of the contributions of each sequence and feature; the multi-channel independent modeling weighted prediction module solves the problem of homogenization in feature modeling of traditional models by independently constructing linear layers for each feature channel, effectively enhancing the model's ability to model the characteristics of different channels differently; the feature-time dimension hybrid self-attention mechanism module combines the self-attention mechanisms of the feature and time dimensions to dynamically adjust the attention of the prediction output to different time steps and features, further improving the model's ability to capture complex patterns. The above-mentioned modules work together, enabling MPDLinear to have high generalization and prediction capabilities.
Claims
1. A MPDLinear (Multi-Part / Plus-Decomposition / Dimension Linear) model, characterized by: It includes a multi-component decomposition module, a self-learning weight vector module, a multi-channel independent modeling weighted prediction module, and a feature-time dimension hybrid self-attention mechanism module; The multi-component decomposition module deserializes the binary data stream and normalizes it into a structure acceptable to the MPDLinear model. The multi-component decomposition module further decomposes the acceptable structure into different time series components. At the same time, this module can perform weighted calculations on the component dimension weights and feature dimension weights of different time series components respectively, and reversely sum the weighted calculation results according to the decomposition method of the multi-component decomposition module to obtain the final prediction result; The self-learning weight vector module uses the preliminary prediction results of different time series components to calculate and update the parameters of different time series component weights and different time series feature weights in the self-learning weight vector module during back propagation, and after looping through the parameters of the different time series component weights and different time series feature weights, the component dimension weights and feature dimension weights of different time series components are obtained; The hybrid self-attention mechanism module of the feature-time dimension can process different time series components through the self-attention mechanism layer of the lookback window time dimension in this module to obtain the attention scores and attention weights of different time series components, and can also process different time series components through the feature dimension self-attention mechanism layer and the prediction window time dimension self-attention mechanism layer of this module in turn to obtain the preliminary prediction results of different time series components after applying attention; The multi-channel independent modeling weighted prediction module uses different temporal components and the attention weights of different temporal components, and obtains different temporal components with attention through matrix multiplication. This module can model independent channels for each different temporal component with attention.
2. A time series prediction method based on the MPDLinear model of claim 1, comprising the following steps: S1: After receiving the binary data stream, the multi-component decomposition module deserializes the binary data stream and normalizes it into a structure acceptable to the MPDLinear model. The multi-component decomposition module decomposes the acceptable structure into different time series components and outputs them to the multi-channel independent modeling weighted prediction module, the self-learning weight vector module and the feature-time dimension hybrid self-attention mechanism module respectively; S2: The feature-time dimension hybrid self-attention mechanism module receives the different time series components described in step S1 and performs dimension rearrangement processing to generate different time series components that have been dimensionally rearranged, and then processes the different time series components that have been dimensionally rearranged through the self-attention mechanism layer of the lookback window time dimension in the feature-time dimension hybrid self-attention mechanism module to obtain attention scores of different time series components, and the attention scores of different time series components are further processed by this module to obtain attention weights of different time series components and output to the multi-channel independent modeling weighted prediction module; S3: After receiving the different time series components described in step S1 and applying the attention weights of the different time series components described in step S2, the multi-channel independent modeling weighted prediction module obtains the different time series components with attention through matrix multiplication. The multi-channel independent modeling weighted prediction module models the independent channels for each different time series component with attention and outputs the preliminary result sequence to the feature-time dimension hybrid self-attention mechanism module; S4: The hybrid self-attention mechanism module of feature-time dimension obtains the preliminary result sequence described in step S3, and processes it in turn through the feature dimension self-attention mechanism layer and the prediction window time dimension self-attention mechanism layer of this module to obtain the preliminary prediction results of different time series components after applying attention and output them to the self-learning weight vector module and the multi-component decomposition module; S5: After the self-learning weight vector module receives the different time series components described in step S1 and the preliminary prediction results described in step S4, the preliminary prediction results are used to calculate and update the parameters of the weights of different time series components and the weights of different time series features in the self-learning weight vector module during back propagation, and the parameters of the weights of different time series components and the weights of different time series features are traversed in a loop, and finally the component dimension weights and feature dimension weights of different time series components are obtained and output to the multi-component decomposition module; S6: The multi-component decomposition module receives the preliminary prediction results described in step S4, and weights the preliminary prediction results according to the component dimension weights and feature dimension weights of the different time series components described in step S5, and reversely adds the weighted calculation results according to the decomposition method of the multi-component decomposition module to obtain the final prediction results.
3. The time series prediction method according to claim 2, characterized in that: The different time series components in step S1 include linear trend components, nonlinear trend components, short-period components, medium- and long-period components, and residual noise components. The steps of extracting different time series components are as follows: S11: Smoothing the acceptable structure of the input step S1 based on the simple moving average (SMA) of the sliding window, realizing the simple moving average by defining the moving average (moving_avg) and obtaining the linear trend component; S12: using double nonlinear mapping to complete the extraction of nonlinear trend components in the acceptable structure described in step S1, specifically, firstly linearly mapping the time dimension or the feature dimension through a fully connected layer Linear layer, then applying the ReLU nonlinear activation function, and finally using a fully connected layer Linear layer to perform linear transformation on the remaining dimensions, and finally obtaining the result of double nonlinear mapping, that is, the nonlinear trend component; S13: first remove the linear trend component and the nonlinear trend component extracted in steps S11 and S12 from the acceptable structure (i.e., the original sequence) to obtain the remaining residual sequence, and then apply a simple moving average (SMA) operation to the remaining residual sequence to extract the short-period component. The value of the sliding window used in this step is smaller than the size of the sliding window used to extract the linear trend component in step S11; S14: removing the short-period component of step S13 from the residual sequence remaining in step S13, thereby obtaining a residual sequence, and extracting medium- and long-period components from the residual sequence using a discrete Fourier transform (DFT) method; S15: Remove the medium and long period components in step S14 from the remaining sequence in step S14 to obtain a residual noise component.
4. The time series prediction method according to claim 2, characterized in that: In step S3, the multi-channel independent modeling weighted prediction module completes the modeling of independent channels through the feature independent linear layer. The formula of the feature independent linear layer is as follows: (Y ( bs,pl ) =X ( bs,sl ) ·W ( sl,pl ) +bias ( bs, 1) )*channels Among them, Y ( bs,pl ) is the output of the independent model for each channel, X ( bs,c,sl ) is the tensor of the input linear layer, W ( sl,pl ) Mapping input time steps to the weight matrix, bias ( bs, 1) is the bias vector of the linear model.
5. The time series prediction method according to claim 2, characterized in that: The processing process of the feature dimension self-attention mechanism layer and the prediction window time dimension self-attention mechanism layer in step S4 is as follows, wherein the input is divided into a time dimension and a feature dimension, and the time dimension is further divided into an input time step (lookback window) time dimension before entering the multi-channel independent modeling weighted prediction module, and a prediction time step time dimension output from the multi-channel independent modeling weighted prediction module: S41: The feature dimension, the input time step time dimension and the prediction time step time dimension are linearly transformed to generate their own query (Q) vector, key (K) vector and value (V) vector respectively; S42: Transpose each key vector in step S42 to obtain K T , then calculate the query vector and K T The dot product (Q·K T ), get the preliminary attention score; S43: Use scaling factor Scaling the output dimension of the dot product in step S42 to obtain a scaled dot product result, i.e., an attention score; S44: Perform a SoftMax operation on the attention score in step S43 to convert it into an attention weight, where the attention weight represents the attention of the query vector of each element on all element key vectors; Among them, for each query position in the sequence, SoftMax transforms the attention score into a probability distribution over all key vector positions, for the input vector as ( i,j ) =[as ( i, 1) ,as ( i, 2) ,...,as ( i,n ) ], the formula of SoftMax is: Where as=[as ( i, 1) ,as ( i, 2) ,...,as ( i,n ) ] is an attention score containing n elements, indicating the relevance of qi to [k1, k2, ..., kn], exp(as ( i,j ) ) is the index of the jth element in the input attention score vector, is the sum of all input attention score indices; S45: Use matrix multiplication (MatMul) to apply the attention weight described in step S44 to the value vector to obtain the preliminary prediction results of different time series components after applying attention. Through aggregation operation, the preliminary prediction results of different time series components after applying attention are output to the learning weight vector module and the multi-component decomposition module.
6. The time series prediction method according to claim 2, characterized in that: The specific process of step S5 is as follows: S51: define the preliminary prediction results of different time series components in step S4 as Si, where i = 1, 2, ..., n, assign a component dimension weight wi to different time series components in the component dimension, and assign a feature dimension weight V[i, j] to each feature in Si in the feature dimension, where j = 1, 2, ..., channels; S52: During the back propagation training process, wi and V[i,j] described in step S51 are updated through the back propagation mechanism, and their gradients are calculated according to the loss function Loss and Therefore, the component dimension weights and feature dimension weights are continuously adjusted during the training process. The component dimension weights and feature dimension weights are shown in the following formula: Among them, η is the learning rate, wi and V[i,j] are the component dimension weights and feature dimension weights of different time series components respectively.
7. The time series prediction method according to claim 6, characterized in that: The specific process of step S6 is as follows: S61: Using the preliminary prediction results of different time series components in step S4 and the component dimension weights wi in step S52, are different time series components weighted by component dimension weights, where By adding up the different time series components after weighting all component dimensions, the comprehensive output Fseries of the component dimension is obtained. S62: Using the preliminary prediction results of different time series components in step S4, weight each feature according to the feature dimension weights in step S52. The specific formula is as follows: in, is the weighted output of the jth feature channel of the i-th different temporal component, and V[i,j] is the corresponding feature dimension weight; Summarize the weighted channels of all feature dimensions to obtain the comprehensive output Fchannels of the feature dimensions. S63: The comprehensive output Fseries of the component dimension and the comprehensive output Fchannels of the feature dimension in steps S61 and S62 are weighted in the component dimension and the feature dimension to obtain the final prediction result Ffinal. The specific formula is as follows: Among them, Si[:,j,:] represents the value of the i-th different timing component on the j-th channel.
8. The time series prediction method according to claim 3, characterized in that: The specific process of step S14 is as follows: S141: The remaining sequence described in step S14 is used as the input for extracting the medium and long period components. The remaining sequence is a time series {x1, x2, ..., xn} of length N, which we split into an even part and an odd part: Even part: x2m, where Odd part:x2m +1 ,in Using the discrete Fourier transform (DFT) formula to expand the even part and the odd part, we can get the following formula: Further simplifying the previous formula, we get: in, That is, the DFT of the even part That is, the DFT of the odd part Therefore, the following formula can be obtained: Through the above splitting, the discrete Fourier transform of the remaining sequence is decomposed into two sequences of length The discrete Fourier transform of the subsequence of , this decomposition process can continue recursively until the length of the sequence is reduced to 1 (the smallest unit), so the discrete Fourier transform can be expressed by the following formula: in, is the rotation factor (or "rotation kernel"), which recursively splits the sequence into two subsequences: even part and odd part. At each recursive level, the discrete Fourier transform combines the results of the even and odd parts by adding and multiplying the rotation factors to form frequency domain data. S142: performing high-frequency filtering on the frequency domain data in step S141 to obtain low-frequency components; S143: Using inverse Fourier transform (IFT) to convert the low-frequency component in step S142 back to the time domain to generate the final medium and long period components, wherein the inverse Fourier transform (IFT) formula is as follows: Among them, xn is the value of the time domain signal at the nth time step, Xk is the complex value of the kth frequency component of the frequency domain data, and N is the total length of the signal (i.e. the number of frequency domain points). is the rotation factor of the inverse Fourier transform, j is the imaginary unit (j 2 =-1).
9. An electronic device, comprising a memory and a processor, characterized in that: Memory: used for storing a computer program for implementing the time series prediction method based on the MPDLinear model; Processor: used to implement the timing prediction method based on the MPDLinear model when executing the computer program.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the time series prediction method based on the MPD Linear model is implemented.
Citation Information
Patent Citations
Lightweight time sequence prediction method based on discrete wavelet transform
CN114219027A
Multivariate time series prediction method based on segmentation strategy and multi-component decomposition algorithm
CN115600656A
Multi-scale entropy gated DWTform meteorological data time sequence prediction method and device
CN117094431A
Method for predicting time sequence in multi-level scene
CN118939956A
Weather prediction method based on interpretable deep learning
CN118981695A
Cited By
Wafer polishing process health state detection method, device, equipment and medium
CN120941284A
Marine engine operation parameter real-time prediction method, system, medium and device based on DLinear algorithm
CN121030660A