A bias error module for mitigating cumulative error and its application method
By introducing a bias error module in multi-step time series data prediction, the residual output is generated using gating units, one-dimensional convolution units and fully connected layers, the problem of error accumulation in iterative prediction is solved, and higher prediction accuracy and model simplification is achieved.
Patent Information
- Application Number
- CN202311257189.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-27
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2043-09-27
AI Technical Summary
In the prior art, multi-step time series data prediction is prone to error accumulation problems during the iteration process, resulting in a decrease in prediction accuracy, and increased model complexity and poor interpretability, making it difficult to apply to different types of iterative prediction models and tasks.
A bias error module is proposed to alleviate cumulative error, including three gating units, one-dimensional convolution unit in the time dimension and a fully connected layer. Through these components, the input data is gated, convolution and fully connected transformed to generate residual outputs, which are used to correct the predicted output of the basic model.
Verified on the industrial dataset of chemical plants, using 60-step time series data in history to predict the iterative multi-step prediction task of the next 20 steps, achieving a reduction of the root mean square error of multi-step prediction by about 40%, while reducing model complexity and improving interpretability.
Smart Images

Figure CN117933305B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of multi-step time series data prediction, and specifically to a bias error module for alleviating cumulative error and its application method. Background Art
[0002] Multi-step time series data prediction refers to using a relevant prediction model to predict multiple consecutive future values based on historical observation data. The iterative strategy uses the predicted value of the previous step as the input for the next step of the model's prediction. Since the predicted value is used as the input, the prediction error will accumulate as the number of iterations increases. Some applications can only use the iterative prediction strategy, such as industrial control system simulation, financial data prediction, and other tasks with exogenous inputs.
[0003] In the prior art, the patent with the publication number CN115936248A proposed an iterative time series prediction model that combines multi-level attention and a bidirectional LSTM neural network. It adaptively adjusts the weights of the input sequence during the iterative prediction process through the attention mechanism to capture the importance of different parts of the sequence, which reduces the accumulation of errors in iterative multi-step prediction. However, the multi-level attention and bidirectional LSTM neural network proposed in this patent significantly increase the complexity of the model, resulting in an increase in computational volume; at the same time, the interpretability of this method is relatively poor, and it is impossible to clearly explain the basis for reducing cumulative error through the attention mechanism.
[0004] In addition, existing methods lack sufficient generality, and it is difficult for these methods to be applied to other different types of iterative prediction models and tasks. Therefore, there is an urgent need to improve them. Summary of the Invention
[0005] The purpose of the present invention is to provide a bias error module for alleviating cumulative error and its application method to solve the problem of error accumulation in multi-step time series prediction tasks based on the iterative strategy mentioned in the above background art.
[0006] To achieve the above-mentioned invention purpose, the present invention adopts the following technical solutions:
[0007] The present invention provides a bias error module for alleviating cumulative error, specifically a bias error module for alleviating cumulative error in iterative multi-step time series prediction, which includes:
[0008] Gated units (GLU), and the number of gated units is set to three. The gated units are used to gate the input data;
[0009] One-dimensional convolutional unit (1Dconv) in the time dimension. The one-dimensional convolutional unit in the time dimension is used to perform convolution operations on the data in the time dimension;
[0010] and a fully connected layer (FC), which is used to perform a fully connected transformation on the data after the convolutional operation.
[0011] The present invention also provides an application method of a bias error module for mitigating cumulative error, which is specifically an application in iterative multi-step time series prediction. Among them, the bias error module used is the above-mentioned bias error module for mitigating the cumulative error of iterative multi-step time series prediction. The application method specifically includes the following steps:
[0012] S1: Construct a basic model. The basic model is any model that accepts a single unit or the original sequence of a multivariate time series as input, performs a single-step prediction task, and outputs a single-step prediction result.
[0013] S2: Construct a bias error module. The bias error module is successively composed of three gated linear units (GLUs), a one-dimensional convolutional unit in the time dimension, and a fully connected layer.
[0014] S3: Initialize the input data and the number of iteration steps. Initialize the input unit or multivariate time series X, where X = {x 1 , x 2 ,..., x n}, where n is the length of the input time series. Initialize the current iteration step t to 0.
[0015] S4: Obtain the single-step prediction output of the basic model. Input the multivariate time series data X into the basic prediction model to obtain the single-step prediction output of the basic model.
[0016] S5: Obtain the iteration step encoding. Encode the iteration step t of the current iterative multi-step prediction in step S1 into the iteration step encoding P t ;
[0017] S6: Obtain the residual output of the bias error module. The input of the bias error module is the multivariate time series input data X (or the hidden layer representation H in the basic model), the single-step prediction output of the basic model at the current step, and the iteration step t of the current iterative multi-step prediction. In the first layer of GLU, the input item X (or H) is first updated element-wise with itself to obtain the output h1. h1 is updated in the second layer of GLU in combination with the prediction output of the basic model to obtain the output h2. h2 is updated in the third layer of GLU in combination with the current iteration step encoding P t to obtain the output h3. h3 is transformed into a residual output R of the same size through a one-dimensional convolutional module and a fully connected layer.
[0018] S7: Obtain the single-step prediction output of the combined model. Add the prediction output of the basic model and the residual output R of the bias error module element-wise to obtain the corrected single-step prediction output of the model.
[0019] S8: Iterative multi-step prediction. Concatenate the corrected model output to the original input X, update the iteration step t, and repeat steps S4 to S7 in sequence until the set multi-step prediction steps are reached or other termination conditions are met.
[0020] Further, in step S1, the base model can be a model for any time series single-step prediction task.
[0021] Further, in step S3, x can be a constant, vector, or multi-dimensional tensor.
[0022] Further, in step S4, the iteration step encoding P of the iteration step t t is given by the following formula:
[0023]
[0024]
[0025] where P t is a vector; t represents the iteration step; i represents the dimension index of the embedding vector; ω k is the frequency parameter; k is used to determine the parity of i: when i is even, when i is odd, d is the dimension of the embedding vector; 10000 is a hyperparameter of the model, which determines the change frequency of the iteration step encoding.
[0026] Further, in step S6, in the first layer of GLU, update the input term X with itself element-wise, which helps to extract and retain key features from the input, as follows:
[0027] h 1 = X ⊙ σ(W 1 *X + b 1 )
[0028] This is inspired by the attention mechanism, which can selectively enhance or weaken different features according to the importance of the input. W 1 and b 1 are the weights and biases of the first layer of GLU, and h 1 is the output of the first layer of GLU.
[0029] Further, in step S6, in the second layer of GLU, update the output h 1 of the first layer of GLU with the single-step prediction output of its base model at the current step element-wise, as follows:
[0030]
[0031] The purpose of this is to enable the bias error module to adjust the output of the residual according to the output of the base model, W 2 and b 2 are the weights and biases of the second-layer GLU, and h 2 is the output of the second-layer GLU.
[0032] Furthermore, in the step S6, in the third-layer GLU, the output h 2 of the second-layer GLU is element-wise updated with the current iteration step encoding P t as follows:
[0033] h 3 = h 2 ⊙ σ(W 3 * P t + b 3 )
[0034] The purpose of this is to enable the bias error module to adjust the output of the residual according to the iteration step. W 3 and b 3 are the weights and biases of the third-layer GLU, and h 3 is the output of the third-layer GLU.
[0035] Furthermore, in the step S6, the output h 3 of the third-layer GLU passes through a one-dimensional convolutional module and a fully connected layer to further capture the local and global relationships of the input sequence and generate a residual output R of the same size as the original input, as follows:
[0036] R = FC(Conv(h 3 ))
[0037] where Conv represents a one-dimensional convolutional operation and FC represents a fully connected layer operation.
[0038] Compared with the prior art, the above one or more technical solutions have the following beneficial effects:
[0039] Through verification on the industrial data set of the rectification section provided by a certain chemical plant, when the bias error module of the present invention performs an iterative multi-step prediction task of predicting the next 20 steps using 60 steps of historical time series data, the method of the present invention reduces the root mean square error of multi-step prediction by approximately 40%. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] The accompanying drawings forming a part of this specification are used to provide a further understanding of the present invention. The schematic embodiments and descriptions thereof of the present invention are used to explain the present invention and do not constitute an improper limitation of the present invention.
[0041] Figure 1 It is a schematic diagram of the application method of the bias error module of the present invention;
[0042] Figure 2 It is a schematic diagram of the structure of the bias error module of the present invention;
[0043] Figure 3 It is a schematic diagram of time series sampling alignment of the present invention. Detailed implementation manners
[0044] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0045] For the cumulative error mitigation method for the multi-step time series prediction task based on the iterative strategy, the present invention has prepared the dataset:
[0046] The industrial data and equipment parameters used in the experiments herein come from the methanol rectification unit of a certain factory. This unit adopts a three-column methanol rectification process, mainly including three parts: a pre-rectification column, a pressurized column, and an atmospheric column. Twenty-three sensor variables are selected for prediction in the selected section. The data is preprocessed:
[0047] (1) Time series sampling alignment: Since different sensors may record data at different frequencies and the data lengths may also be different, in order to perform subsequent analysis and model training, interpolation operations need to be performed on these time series data to ensure that they have the same time interval and sequence length. The interpolation method adopted by the present invention is first-order spline interpolation. By linearly interpolating the known data points, new data points are generated to fill the missing values between the time series, so that the time series of all variables have the same sampling interval and sequence length. After sampling is completed, the time series can be represented as [x 1 , x 2 ,..., x t ,..., x T , where T represents the length of the time series, defined as each discrete time point in the sampled data as a time step, and x t represents the input sample at the t-th time step. The input samples at each time step all contain n variables, and n is the number of sensor variables. Such sampling alignment operations help to maintain the consistency and comparability of the data.
[0048] (2) Data Standardization: The value ranges of different variable sensors may vary significantly, which may pose difficulties to the training of the model. To reduce the impact of different features on model training, data standardization is a common preprocessing method. The goal of data standardization is to transform each feature into a standard distribution with the same scale and range. Among them, the standardization method adopted in the present invention is achieved by calculating the Z-score of each feature, and its standardization formula is as follows:
[0049]
[0050] Where Z is the standardized value, X is the original data, μ is the mean of the feature, and σ is the standard deviation of the feature. By applying this formula, the data can be transformed into a standard normal distribution with a mean of 0 and a standard deviation of 1, thereby eliminating the scale differences between different features and facilitating the stability and accuracy of model training.
[0051] (3) Dataset Partitioning: To verify the generalization ability of the model, the dataset needs to be partitioned into a training set, a validation set, and a test set. A common partitioning ratio is to use 70% of the data for training, 20% for validation, and 10% for testing. Through such partitioning, the rationality of the dataset can be ensured and used for model training, tuning, and evaluation. In this example, the dataset contains a total of 18,146 groups of time series data. After sampling, according to the partitioning ratio, the sizes of the training set, validation set, and test set are 12,702, 3,629, and 1,815 groups respectively. Such a partitioning method can provide sufficient data samples for model training and validation, and ensure the independence of the test set, so as to accurately evaluate the generalization ability of the model.
[0052] Based on the dataset preparation in the above-mentioned Embodiment 1, the present invention proposes a bias error module for mitigating cumulative error and its application method, wherein,
[0053] A bias error module for mitigating cumulative error includes:
[0054] Gated units (GLUs), three of which are provided, and the gated units are used to gate the input data;
[0055] A one-dimensional convolutional unit (1Dconv) in the time dimension, and the one-dimensional convolutional unit in the time dimension is used to perform a convolutional operation on the data in the time dimension;
[0056] And a fully connected layer (FC), and the fully connected layer is used to perform a fully connected transformation on the data after the convolutional operation.
[0057] An application method of a bias error module for mitigating cumulative error specifically includes the following steps:
[0058] Step S1: Construct a basic model;
[0059] In this application example, the basic model is selected as ASTGCN (Attention-based Spatio-Temporal Graph Convolutional Network model), which accepts the historical unit or multivariate time series raw input as its input. Its model parameters and training parameters are as follows in the table:
[0060] Parameter Name Set Value Description learning_rate 0.001 Learning rate during model training time_interval 10s Interval for time series sampling history_len 60 Length of time steps for input historical time series forecast_len 20 Length of output predicted time series Epoch 150 Number of training epochs optimizer Adam Optimizer ASTGCN_STBlock_num 2 Number of ASTGCN spatio-temporal blocks ASTGCN_kenel_size 3 ASTGCN convolution kernel size ASTGCN_num_cheb_filter 64 Number of ASTGCN spatial convolution kernels ASTGCN_num_time_filter 64 Number of ASTGCN temporal convolution kernels ASTGCN_stride 1 Stride of ASTGCN temporal convolution ASTGCN_cheb_K 2 Order of ASTGCN Chebyshev polynomial
[0061] Step S2: Construct a bias error module;
[0062] The bias error module consists of three gated units GLU, a one-dimensional convolutional unit in the time dimension, and a fully connected layer.
[0063] Step S3: Initialize the input data and the number of iteration steps;
[0064] The input data is initialized as the historical multivariate time series X = {x 1 , x 2 ,..., x 60}, where the length of the input time series is 60. At the same time, the current iteration step t is initialized to 0.
[0065] Step S4: Obtain the single-step prediction output of the basic model;
[0066] Use the ASTGCN basic model to predict the historical time series input data X to obtain the single-step prediction output of the basic model
[0067] Step S5: Obtain the iteration step encoding;
[0068] Use the iteration time step t and encode it into the iteration step encoding P through the following formula t .
[0069]
[0070]
[0071] Step S6: Obtain the residual output of the bias error module;
[0072] 1. Input the multivariate time series data X (or the hidden layer representation H in the basic model), the single-step prediction output of the basic model at the current step, and the iteration time step t of the current iteration multi-step prediction into the bias error module.
[0073] 2. In the first layer of GLU, update the input X (or H) with itself element-wise to generate the output h1.
[0074] 3. Within the second - layer GLU, combine h1 with the predicted output of the base model for updating to obtain the output h2.
[0075] 4. Within the third - layer GLU, combine h2 with the iteration - step encoding P t for updating to obtain the output h3.
[0076] 5. Finally, h3 undergoes transformations through a one - dimensional convolutional module and a fully - connected layer to generate a residual output R of the same size as the original input.
[0077] Step S7: Obtain the single - step prediction output of the combined model;
[0078] Sum the predicted output of the base model and the residual output R of the bias - error module element - by - element to obtain the corrected single - step prediction output of the model
[0079] Step S8: Obtain the single - step prediction output of the combined model;
[0080] Concatenate the corrected model output onto the original input X, update the iteration step t, and then repeat steps S4 to S7 in sequence until the multi - step prediction step reaches 20 steps.
[0081] Evaluation metric: Assume that the predicted value of the model for the future H - step time series is and the true value is Then the numerical value of the root - mean - square error RMSE of the multi - step prediction can be calculated by the following formula:
[0082]
[0083] Without the bias block: RMSE: 2.5246;
[0084] With the bias block: RMSE: 1.4860.
[0085] Based on the above conclusions, in a specific application, the present invention uses the attention - based graph convolutional neural network model (ASTGCN) as the base model. After applying the bias - error module of the present invention, through verification on the industrial data set of the rectification section provided by a chemical plant, when performing the iterative multi - step prediction task of predicting the future 20 steps using 60 - step historical time - series data, the method of the present invention reduces the root - mean - square error of the multi - step prediction by approximately 40%.
[0086] The above - mentioned is only the preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and all should be covered within the protection scope of the present invention.
Claims
1. A method for constructing a bias error module to mitigate cumulative error, characterized in that, it includes: Three gating units GLU, which are used to gate the input data; A one-dimensional convolutional unit in the time dimension, which is used to perform convolutional operations on the data in the time dimension; And a fully connected layer, which is used to perform a fully connected transformation on the data after the convolutional operation; The input of the bias error module is the multivariate time series input data X, the single-step prediction output of the basic model at the current step, and the iteration step t of the current iterative multi-step prediction; within the first layer GLU, the input term X first undergoes an element-wise update with itself to obtain the output h1; h1 is updated in combination with the prediction output of the basic model within the second layer GLU to obtain the output h2; h2 is updated in combination with the current iteration step encoding Pt within the third layer GLU to obtain the output h3; h3 passes through the one-dimensional convolutional module and the fully connected layer and is transformed into a residual output R of the same size.
2. A method for applying a bias error module to mitigate cumulative error, characterized by: including the following steps: S1: Construct a basic model; S2: Construct a bias error module; the bias error module is successively composed of three gating units GLU, a one-dimensional convolutional unit in the time dimension, and a fully connected layer; S3: Initialize the input data and the number of iteration steps; initialize the input single or multivariate time series X as where n is the length of the input time series; initialize the current iteration step t as 0; S4: Obtain the single-step prediction output of the base model; input the multivariate time series data X into the base prediction model to obtain the single-step prediction output of the base model ; S5: Obtain the iteration step encoding; encode the iteration step t of the current iterative multi-step prediction in step S1 as the iteration step encoding P t ; S6: Obtain the residual output of the bias error module; the input of the bias error module is the multivariate time series input data X, the single-step prediction output of the basic model at the current step, and the iteration step t of the current iterative multi-step prediction; within the first layer GLU, the input term X first undergoes an element-wise update with itself to obtain the output h1; h1 is updated in combination with the prediction output of the basic model within the second layer GLU to obtain the output h2; h2 is updated in combination with the current iteration step encoding Pt within the third layer GLU to obtain the output h3; h3 passes through the one-dimensional convolutional module and the fully connected layer and is transformed into a residual output R of the same size; S7: Obtain the single-step prediction output of the combined model; add the prediction output of the base model and the residual output R of the bias error module element by element to obtain the corrected single-step prediction output of the model ; S8: Iterative multi-step prediction; The corrected model output is spliced onto the original input X, the iteration step t is updated, and steps S4 to S7 are repeated in sequence until the set multi-step prediction steps are reached or other termination conditions are met.
3. The method for applying a bias error module to mitigate cumulative error according to claim 2, characterized in that: In the step S1, the basic model is a model for any time series single-step prediction task.
4. The method for applying a bias error module to mitigate cumulative error according to claim 2, characterized in that: In the step S3, x is a constant, a vector, or a multi-dimensional tensor.
5. The method for applying a bias error module to mitigate cumulative error according to claim 2, characterized in that: In the step S4, the iteration step encoding P of the iteration step t t is given by the following formula: ; where P t is a vector; t represents the number of iteration steps; i represents the dimension index of the embedding vector; is the frequency parameter; k is used to determine the parity of i: when i is even, and when i is odd, ; d is the dimension of the embedding vector; 10000 is a hyperparameter of the model, which determines the change frequency of the iteration step encoding.
6. The method for applying a bias error module to mitigate cumulative error according to claim 2, characterized in that: In the step S6, within the first layer GLU, the input term X is element-wise updated with itself as follows: ; Among them, W 1 and b 1 are the weights and biases of the first-layer GLU, and h 1 is the output of the first-layer GLU.
7. The method for applying a bias error module to mitigate cumulative error according to claim 6, characterized in that: In the step S6, in the second-layer GLU, the output h of the first-layer GLU 1 and the single-step prediction output of its base model at the current step are updated element-wise as follows: ; Among them, W 2 and b 2 are the weights and biases of the second-layer GLU, and h 2 is the output of the second-layer GLU.
8. The application method of the bias error module for mitigating cumulative error according to claim 7, characterized in that: In the step S6, in the third-layer GLU, the output h of the second-layer GLU 2 is updated element-wise with the encoding P of the current iteration step t as follows: ; Among them, W 3 and b 3 are the weights and biases of the third-layer GLU, and h 3 is the output of the third-layer GLU.
9. The application method of the bias error module for mitigating cumulative error according to claim 8, characterized in that: In the step S6, the output h of the third-layer GLU 3 passes through a one-dimensional convolutional module and a fully-connected layer to further capture the local and global relationships of the input sequence and generate a residual output R of the same size as the original input, as follows: ; wherein, Conv represents a one-dimensional convolution operation, and FC represents a fully connected layer operation.
Citation Information
Patent Citations
Power load prediction method, device and system based on attention network
CN115936248A
Rotary kiln sintering temperature prediction method based on hybrid deep neural network
CN111950191A
Flight operation network delay prediction method based on deep learning combination model
CN116205120A