Time sequence prediction method based on improved Autoformer model
By introducing RevIN, SGConv and FECM modules into the Autoformer model, the time series prediction model SFRformer is improved, solving the problems of long-distance dependence and distribution offset, and significantly improving the prediction accuracy.
Patent Information
- Application Number
- CN202510225705.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-24
AI Technical Summary
The existing deep learning time series prediction model has limitations in dealing with long-distance dependency and distribution offset problems, resulting in a large deviation from the real data distribution.
An improved Autoformer model, called SFRformer, is proposed. By introducing the RevIN module, SGConv module and FECM module, the remote dependency processing difficulties and distribution offset problems caused by local convolution are solved, so that the model can better capture timing information.
It significantly improves the accuracy of time series prediction, reduces the noise impact caused by Fourier forward and inverse transformation, and effectively deals with long-distance dependency and distribution offset problems.
Smart Images

Figure CN120197747A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of time series prediction, and particularly to a time series prediction method based on an improved Autoformer model. Background Art
[0002] The rapid development of the electric vehicle industry and the connection of a large number of electric vehicles to the power grid will have an inestimable impact on the operation and planning of the power system. At the same time, as a distributed energy storage system, if effectively utilized, it can also provide services such as peak shaving, valley filling, and frequency modulation for the power grid. Therefore, accurate prediction of electric vehicle charging load can promote the effective interaction between electric vehicles and the power grid, realize the orderly progress of charging, optimize the load operation curve of the power grid, ensure the stable operation of the power grid, improve the utilization efficiency of charging facilities, and reduce the negative impact of the huge load brought by electric vehicle charging on the power grid.
[0003] Time series prediction, as a core research direction in the fields of data analysis, statistics, and artificial intelligence, is a key foundation for achieving more complex prediction tasks. Its main task is to accurately predict future trends and patterns from historical data, usually by analyzing the time series characteristics of data points and establishing a prediction model. Time series prediction techniques are not limited to basic trend prediction and pattern recognition, but also support advanced applications such as anomaly detection, periodic analysis, causal relationship inference, and multivariate time series analysis. These techniques have been widely applied in multiple fields, including financial market analysis, weather forecasting, supply chain management, energy demand prediction, and healthcare, and are of great significance for enhancing the prediction ability and decision-making support of intelligent systems.
[0004] The prediction models based on deep learning are the current main prediction models. They use deep neural networks to learn patterns from historical data and show high-precision and high-efficiency performance advantages in fields such as trend prediction and periodic analysis. Currently, the time series prediction models based on deep learning can be divided into two categories: (1) The time series prediction models based on sequence-to-sequence (Seq2Seq). These models process time series data through an encoder-decoder structure, first generating an internal representation of the time series, and then predicting future values. (2) The time series prediction models based on the attention mechanism. These models can more flexibly capture long-term dependencies in the time series by introducing the attention mechanism and directly predict the target value from the input sequence, achieving end-to-end time series prediction. These models have made significant progress in accuracy and speed, enabling time series prediction technology to be widely applied to real-time processing from financial market analysis to weather forecasting, meeting the diverse application scenario requirements.
[0005] Deep learning-based time series prediction models can automatically learn features without manual intervention and demonstrate excellent capabilities when dealing with high-dimensional data. Therefore, this technology has achieved satisfactory experimental results in some object detection problems. Nevertheless, the following deficiencies still exist in this technology:
[0006] 1. Existing mainstream deep learning time series prediction models, such as Transformer and Autoformer, although they have achieved certain results in time series prediction tasks, the prediction results often deviate significantly from the true data distribution.
[0007] 2. Effectively handling long-range dependencies is a key challenge in tasks such as time series prediction, language modeling, and pixel-level image generation. Traditional deep learning models, such as recurrent neural networks (RNNs), Transformers, and convolutional neural networks (CNNs), have some limitations when dealing with long sequence data. For example, RNNs may face the problem of vanishing gradients, the computational complexity of Transformers grows quadratically with the increase in sequence length, and CNNs usually only have local receptive fields in each layer.
[0008] 3. Attention also needs to be paid to the distribution shift problem in time series prediction, that is, the statistical characteristics (such as mean and variance) of time series data change over time, resulting in changes in the time distribution. Summary of the Invention
[0009] Aiming at the deficiencies of the existing technology, the present invention proposes an improved Autoformer model (SFRformer). This model solves the problem that local convolutions, which have been ignored by Transformer-like models, cannot effectively handle long-range dependencies, and at the same time solves the distribution shift problem that often occurs in the analysis of time series data, enabling the model to better capture information in the time series, thereby improving the accuracy of time series prediction.
[0010] To achieve the above object, the present invention provides a time series prediction method based on an improved Autoformer model, including the following steps:
[0011] (1) Obtain a historical time series dataset;
[0012] (2) Construct an SFRformer model, which is based on the Autoformer model and integrates the RevIN module, the SGConv module, and the FECM module;
[0013] (3) Use the historical time series dataset to train the SFRformer model;
[0014] (4) Use the trained SFRformer model for time series prediction.
[0015] Furthermore, the SFRformer model is specifically: at the input end of the Autoformer model, a RevIN module and an SGConv module connected in sequence are introduced to perform regularization processing on the input data and capture context information features; at the output end of the Autoformer model, the same RevIN module is connected for inverse regularization to reintroduce the variables removed by regularization into the sequence; at the same time, an FECM module is introduced between the encoding module and the decoding module of the Autoformer model.
[0016] Furthermore, the FECM module includes a channel attention mechanism and a discrete cosine transform;
[0017] First, the input feature map is split into n subgroups along the channel dimension, denoted as [v0, v1, …, v n-1 , where each v i = R 1×L , i ∈ {0, 1, …, n - 1}, n = N v , N v is the number of channels;
[0018] Then, each subgroup is processed by the corresponding DCT frequency components to obtain the frequency component vector of the subgroup;
[0019] The formula is expressed as:
[0020]
[0021] where: i ∈ {0, 1, …, N v - 1}, j ∈ {0, 1, …, L s - 1}; N v is the number of channels; L s is the sequence length per channel; DCT j () represents the discrete cosine transform operation; is the basis function used in the discrete cosine transform; Freq i represents the frequency component vector obtained by the DCT transformation of the i-th subgroup, and l is the dimension number of the frequency component vector;
[0022] Finally, the frequency component vectors of each subgroup are stacked and synthesized into one path for output.
[0023] Furthermore, it is characterized in that the training process of the SFRformer model is:
[0024] (3.1) Initialize the SFRformer model and set the upper limit of the number of training rounds;
[0025] (3.2) The mean square error and the mean absolute error are used as the loss function of the SFRformer model;
[0026]
[0027] where: w i is the real data at time i; is the data at time i predicted by the SFRformer model; q is the number of data;
[0028]
[0029] where, w k is the real data at time k; is the data at time k predicted by the SFRformer model; q is the number of data;
[0030] (3.3) Use the stochastic gradient descent method or the adaptive optimization method to modify the network weights;
[0031] (3.4) Train the SFRformer model until the upper limit of the training rounds is reached to obtain the trained model parameters.
[0032] The present invention also provides a time series prediction system based on an improved Autoformer model, including:
[0033] A data acquisition module for obtaining a historical time series data set;
[0034] A model construction module for constructing an SFRformer model, which is based on the Autoformer model and integrates the RevIN module, the SGConv module, and the FECM module;
[0035] A model training module for training the SFRformer model using the historical time series data set;
[0036] A model verification module for performing time series prediction using the trained SFRformer model.
[0037] The present invention also provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of the above-mentioned time series prediction method based on an improved Autoformer model are implemented.
[0038] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above-mentioned time series prediction method based on an improved Autoformer model are implemented.
[0039] Advantages of the present invention:
[0040] The present invention proposes a novel hybrid improved Autoformer model - SFRformer. This model optimizes the Encoder-Decoder module in the original Autoformer and adds the FECM module, which can effectively reduce the noise impact caused by the Gibbs effect after the forward and inverse Fourier transforms, and can significantly improve the accuracy of time series prediction. At the same time, this model solves the problem that the local convolution in the Transformer-like models has been ignored, resulting in the inability to effectively handle long-range dependencies, and also solves the distribution shift problem that often occurs in the analysis of time series data, enabling the model to better capture the information in the time series, thereby improving the accuracy of time series prediction. Brief Description of the Drawings
[0041] Figure 1 It is a schematic flowchart of the time series prediction method based on the improved Autoformer model in an embodiment of the present invention.
[0042] Figure 2 It is a schematic diagram of the overall structure of the SFRformer model in an embodiment of the present invention.
[0043] Figure 3 It is a schematic diagram of the structure of the RevIN module in an embodiment of the present invention.
[0044] Figure 4 It is a schematic flowchart of the training process of the SFRformer model in an embodiment of the present invention.
[0045] Figure 5 It is a comparison chart of the short-term time series prediction results of the SFRformer model and other models in an embodiment of the present invention. Detailed Embodiments
[0046] The following further describes the present invention with reference to the drawings and embodiments.
[0047] As Figure 1 shown, the embodiment of the present invention provides a time series prediction method based on an improved Autoformer model, including the following steps:
[0048] S101. Obtain a historical time series data set.
[0049] The present invention takes the time series prediction of charging load data as an example to illustrate the solution. Therefore, this embodiment selects six benchmark public data sets under long-term settings and the power load data of an electric vehicle charging station on the campus of the California Institute of Technology as the data set, and divides them into a training set, a validation set, and a test set according to 7:1:2.
[0050] S102. Construct the SFRformer model, including the RevIN module, the SGConv module, and the optimized Autoformer model.
[0051] In the embodiment of the present invention, based on the Autoformer model, the RevIN module, the SGConv module, and the FECM module are integrated to construct a new prediction model, SFRformer. As Figure 2 shown, at the input end of the Autoformer model (before the encoding module of the Autoformer model), the sequentially connected RevIN module and SGConv module are introduced. The input sequence first undergoes regularization processing through the RevIN module to remove the non-stationary information in the sequence, and then better captures the context information features through the SGConv global convolution module; the FECM module is introduced between the encoding module and the decoding module of the Autoformer model to avoid the noise impact caused by the Gibbs effect resulting from the forward and inverse Fourier transforms in the Auto-correlation module; at the output end of the Autoformer model (after the decoding module of the Autoformer model), the same RevIN module is connected for inverse regularization to reintroduce the unstable variables removed by the previous regularization into the sequence, obtaining the output sequence.
[0052] Among them, the RevIN (Reversible Instance Normalization) module is a normalization and denormalization method for time series prediction, specifically designed to handle the problem of distribution shift. As Figure 3 shown, it can not only normalize the data but also denormalize the normalized data to restore it to the original distribution. This method is particularly suitable for time series data because it can dynamically adjust the normalization parameters of the data to cope with the changes in data distribution. In time series data, distribution shift is a common problem, and the statistical properties (such as mean and variance) of the data change over time, which can affect the model performance, especially when the distribution difference between the training data and the test data is large. This method not only helps to improve the accuracy of time series prediction under distribution shift conditions but also ensures that the data used in the training and prediction stages of the model is consistent in statistical properties.
[0053] The SGConv (Structured Global Convolution) model is used to capture context information features. SGConv is a specially designed convolution model for handling long-range dependence problems in long sequence data. This model uses a global convolution kernel that can cover the entire length of the input sequence, enabling it to capture the long-range correlation information in the sequence.
[0054] FECM (Frequency Enhanced Channel Mechanism) is a frequency-enhanced channel attention mechanism designed for time series prediction tasks. This mechanism uses the Discrete Cosine Transform (DCT) to avoid the Gibbs Phenomenon in the traditional Fourier Transform (FT), thereby more effectively capturing frequency information and avoiding the introduction of high-frequency noise. The core idea of FECM is to simulate the dependence between channels in the frequency domain to enhance the model's ability to capture frequency information in time series data. This mechanism consists of two main steps: the channel attention mechanism and the discrete cosine transform. The channel attention mechanism can learn the interdependence between different channels, while the discrete cosine transform is used to effectively represent frequency information and avoid the Gibbs phenomenon caused by periodic problems.
[0055] To capture more time series information from the feature map, FECM is introduced after the Encoder module of Autoformer to obtain more frequency components, rather than just using global average pooling (GAP) to obtain the lowest frequency component. The specific process of the FECM module is as follows:
[0056] First, FECM divides the input feature map into n subgroups along the channel dimension, denoted as [v0, v1,..., v n-1 , where each v i = R 1×L , i ∈ {0, 1,..., n - 1}, n = N v , N v is the number of channels.
[0057] Subsequently, each subgroup will be processed by the corresponding DCT frequency components, with the frequency range from low to high.
[0058] Expressed by the formula:
[0059]
[0060] where: i ∈ {0, 1,..., N v - 1}, j ∈ {0, 1,..., L s - 1}; N v is the number of channels; L s is the sequence length per channel; DCT j () represents the discrete cosine transform operation; is the basis function used in the discrete cosine transform; Freq i represents the frequency component vector obtained after the i-th subgroup is transformed by DCT, and l is the dimension number of the frequency component vector.
[0061] In this way, the features of each channel interact with all frequency components, comprehensively obtaining important time information from the frequency domain, which will encourage the network to enhance the diversity of feature extraction.
[0062] Finally, the frequency component vectors of each subgroup are stacked and synthesized into one path for output.
[0063] S103. Use the historical time series dataset to train the SFRformer model.
[0064] As Figure 4 shown, use the training set data to train the SFRformer model, and the specific steps are as follows:
[0065] (1) Initialize the SFRformer model and set the upper limit of the number of training rounds.
[0066] Initialize the weights of the SFRformer network model and set the upper limit N of the number of training rounds.
[0067] (2) Use the mean square error and the mean absolute error as the loss function of the SFRformer model.
[0068]
[0069] Among them: w i is the real data at time i; is the data at time i predicted by the SFRformer model; q is the number of data.
[0070]
[0071] Among them, w k is the real data at time k; is the charging load data at time k predicted by the SFRformer model; q is the number of data.
[0072] (3) Use the stochastic gradient descent method or an adaptive optimization method (such as AdamW) to modify the network weights.
[0073] (4) Train the SFRformer model until the upper limit of the number of training rounds is reached to obtain the trained model parameters.
[0074] Use the validation set and the test set to validate and test the trained SFRformer model.
[0075] To verify the effectiveness and accuracy of the SFRformer model, the embodiments of the present invention also selected 2 Transformer - type prediction models (Autoformer and Informer) for comparative experiments with it. The comparison results are as Figure 5As shown, a certain segment of randomly selected data from the test set is used for 96-hour power load forecasting. It can be seen from the figure that although the fluctuation of the charging station's electricity load is relatively large, the SFRformer model still performs well and fits the actual value curve at some peaks.
[0076] To more intuitively compare the prediction accuracy, Table 1 shows the comparison results of the evaluation indicators of different models. The proposed Informer model solves the problem of high computational complexity of the Transformer model in long sequence learning. Although it greatly speeds up the training speed of the Transformer model, due to the use of the sparse attention mechanism, it will limit the information utilization efficiency and affect the prediction effect. Therefore, compared with the Informer model, the MSE indicators of the SFRformer model are reduced by 6.10%, 10.93%, 9.51%, and 6.70% respectively. The Autoformer model proposes a deep decomposition architecture and an autocorrelation mechanism. Compared with the Autoformer model, the MSE indicators of the SFRformer model are reduced by 2.42%, 9.49%, 6.63%, and 5.53% respectively. Through the comparative experiment, it can be seen that the indicators of the SFRformer model are relatively the smallest and the prediction error is relatively the smallest.
[0077] Table 1
[0078]
[0079]
[0080] S104. Use the trained SFRformer model for time series prediction.
[0081] Obtain the current charging load data and input it into the trained SFRformer model to complete the time series prediction.
[0082] The embodiment of the present invention also provides a time series prediction system based on an improved Autoformer model, including:
[0083] A data acquisition module for obtaining a historical time series data set.
[0084] A model building module for building an SFRformer model, which is based on the Autoformer model and integrates a RevIN module, an SGConv module, and an FECM module.
[0085] A model training module for training the SFRformer model using the historical time series data set.
[0086] A model verification module for using the trained SFRformer model for time series prediction.
[0087] An embodiment of the present invention further provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the above-mentioned time series prediction method based on the improved Autoformer model are implemented.
[0088] An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned time series prediction method based on the improved Autoformer model are implemented.
[0089] As mentioned above, it is only a preferred embodiment of the present invention and does not impose any form of limitation on the present invention. Any ordinary technician in the industry can smoothly implement the present invention according to the instructions in the accompanying drawings and the above description. However, any equivalent changes made by those skilled in the art within the scope of the technical solution of the present invention by using the technical content disclosed above, such as slight modifications, decorations, and evolutions, are equivalent embodiments of the present invention. At the same time, any equivalent changes, modifications, and evolutions made to the above embodiments based on the essence of the present invention still fall within the protection scope of the technical solution of the present invention.
Claims
1. A time series prediction method based on an improved Autoformer model, characterized in that: The steps include: (1) Obtain historical time series data sets; (2) constructing an SFRformer model, which is based on the Autoformer model and integrates the RevIN module, the SGConv module, and the FECM module; (3) using the historical time series dataset to train the SFRformer model; (4) Use the trained SFRformer model for time series prediction.
2. The time series prediction method based on the improved Autoformer model according to claim 1, characterized in that: The SFRformer model is specifically as follows: at the input end of the Autoformer model, a RevIN module and an SGConv module connected in sequence are introduced to regularize the input data and capture context information features; at the output end of the Autoformer model, the same RevIN module is connected again to perform inverse regularization, and the variables eliminated by regularization are added back to the sequence; at the same time, a FECM module is introduced between the encoding module and the decoding module of the Autoformer model.
3. The time series prediction method based on the improved Autoformer model according to claim 1, characterized in that: The FECM module includes a channel attention mechanism and a discrete cosine transform; First, the input feature map is split into n subgroups along the channel dimension, denoted as [v0, v1, …, v n-1 ], where each v i =R 1×L ,i∈{0,1,…,n-1},n=N v , N v is the number of channels; then, each subgroup is processed by the corresponding DCT frequency component to obtain the frequency component vector of the subgroup; The formula is: Where: i∈{0, 1, ..., N v -1}, j∈{0, 1, …, L s -1}; N v is the number of channels; L s is the sequence length per channel; DCT j () represents discrete cosine transform operation; is the basis function used in discrete cosine transform; Freq i represents the frequency component vector obtained after DCT transformation of the i-th subgroup, and l is the dimension of the frequency component vector; Finally, the frequency component vectors of each subgroup are stacked and synthesized into one path for output.
4. The time series prediction method based on the improved Autoformer model according to claim 1, characterized in that: The training process of the SFRformer model is: (3.1) Initialize the SFRformer model and set the upper limit of training rounds; (3.2) Using mean square error and mean absolute error as the loss function of the SFRformer model; Where: w i is the real data at time i; is the data at time i predicted by the SFRformer model; q is the number of data; Among them, w k is the real data at time k; is the k-time data predicted by the SFRformer model; q is the number of data; (3.3) Modify network weights using stochastic gradient descent or adaptive optimization methods; (3.4) Train the SFRformer model to reach the upper limit of training rounds and obtain the trained model parameters.
5. The time series prediction system based on the improved Autoformer model is characterized by: include: Data acquisition module, used to obtain historical time series data sets; A model building module is used to build an SFRformer model, which is based on the Autoformer model and integrates the RevIN module, the SGConv module and the FECM module; A model training module, used to train the SFRformer model using the historical time series data set; The model validation module is used to perform time series forecasting using the trained SFRformer model.
6. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the steps of the time series prediction method based on the improved Autoformer model described in any one of claims 1 to 4 are implemented.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the time series prediction method based on the improved Autoformer model described in any one of claims 1 to 4 are implemented.
Citation Information
Cited By
Multivariable time sequence prediction method, device and system and model training method
CN120632637A