A time series data prediction method based on deep ensemble learning model and high-low frequency separation
Through deep ensemble learning models and high and low frequency separation technology, decomposing and predicting multiple modal components of time series data, the limitations of the prior art in processing non-stationary and multi-scale data are solved, and more efficient and accurate time series prediction is achieved.
Patent Information
- Application Number
- CN202411673870.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-21
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2044-11-21
AI Technical Summary
Existing time series prediction methods have limitations in dealing with non-stationary, multi-scale features and noisy data, making it difficult to capture all the rules and complexities of the data.
The time series data prediction method based on deep ensemble learning model and high and low frequency separation is adopted, and the time series data is decomposed into multiple modal components through VMD variational modal decomposition. The high-frequency and low-frequency components are predicted using CNN-LSTM and GRU models respectively, and the prediction results of multiple models are fused through adaptive weight adjustment method.
It improves the accuracy and robustness of time series data prediction, can effectively deal with the non-stationarity and multi-scale features of the data, reduces the computational complexity and improves the model training efficiency.
Smart Images

Figure CN119513498B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of time series prediction, and more specifically, relates to a time series data prediction method based on a deep ensemble learning model and high- and low-frequency separation. Background Art
[0002] Time series prediction is an important research topic in many fields, including finance, meteorology, medicine, and industry. Accurate time series prediction can provide strong support for decision-making, optimize resource allocation, and improve system efficiency. However, time series data usually face many challenges. First, time series data are generally non-stationary, with trends and fluctuations changing over time, and usually contain multi-scale features such as long-term trends, seasonal fluctuations, and short-term fluctuations. Second, data may be disturbed by noise and measurement errors, which makes modeling more complicated. In particular, for time series that contain multiple periodicities, sudden events, or external factors, traditional prediction methods often find it difficult to capture all the regularities and complexities of the data. Therefore, how to improve the accuracy of non-stationary time series prediction, especially in the face of measurement noise and long-term and short-term dependencies, has become a core issue facing modern prediction technology.
[0003] Traditional time series prediction mainly uses statistical methods (such as ARIMA, gray prediction model) and machine learning methods (such as artificial neural network, support vector machine). However, these methods have certain limitations when dealing with nonlinear and non-stationary data. Although statistical methods perform well when dealing with stationary data, their prediction accuracy is often low when facing complex and changeable time series data. Although machine learning methods can capture nonlinear relationships in data, their processing capabilities for non-stationary signals are limited, especially when dealing with complex signals containing multiple frequency components, which are prone to overfitting or underfitting. In recent years, some studies have tried to improve the prediction effect through signal processing technology, such as wavelet transform (WT) and empirical mode decomposition (EMD). Although these methods can improve the prediction accuracy to a certain extent, there is still room for improvement in the selection of models for different frequencies. Although WT can effectively decompose the signal, the selection of its basis function and the determination of the decomposition scale are relatively difficult, and the decomposed signal may still not be stable enough. Although EMD can adaptively decompose signals, it has a mode mixing problem, that is, different frequency components may be mixed in the same mode, affecting the accuracy of subsequent analysis.
[0004] Chinese patent document CN113095290A discloses a non-intrusive load decomposition method based on EEMD and GRU. The original power signal is processed by adaptive Gaussian filtering, and the processed power signal is decomposed into multiple modal components using the EEMD algorithm. The features contained in the power signal are amplified. In order to mine the time correlation features between the decomposition time point and the previous multiple time points, a GRU neural network is constructed to process the timing signal. Finally, the test data is input into the trained GRU network to realize load decomposition. In order to further improve the decomposition accuracy and speed, the GRU network parameters are optimized using the Tianniu swarm optimization algorithm.
[0005] Chinese patent document CN116451051A discloses a frequency analysis method and system for time series prediction, the analysis method includes: using discrete Fourier transform DFT to obtain the frequency characteristics of the original input time series; extracting the effective part of the sequence frequency characteristics, filtering the frequency components through the self-attention mechanism to achieve the purpose of eliminating interference items in the frequency components; based on the structure of the convolutional neural network to achieve the highly nonlinear characterization ability of the model, to map the input features to the output features. However, both methods use the same prediction model for components of different frequencies, and there is still room for improvement in prediction accuracy. In addition, the EEMD method may cause modal aliasing problems in the first method.
[0006] In view of this, the present invention designs a time series data prediction method based on a deep ensemble learning model and high- and low-frequency separation to achieve short-term prediction of non-stationary time series data. Summary of the invention
[0007] The present invention aims to overcome at least one defect of the above-mentioned prior art and provides a time series data prediction method based on deep ensemble learning model and high- and low-frequency separation. It combines multiple deep learning models to make the prediction method more flexible and adaptable, and can select appropriate models for prediction for components with different characteristics, thereby improving the overall performance.
[0008] The detailed technical scheme of the present invention is as follows:
[0009] A time series data prediction method based on a deep ensemble learning model and high- and low-frequency separation, the method comprising:
[0010] S1, collect time series data in the system, and then preprocess the obtained data;
[0011] S2, performing VMD variational modal decomposition on the preprocessed original time series data to decompose the time series into K modal components with limited bandwidth;
[0012] S3, combining the maximum information coefficient method and the reconstruction error analysis method to determine the optimal number of decomposition modes K;
[0013] S4, use the zero-crossing rate and center frequency to divide the high and low frequency components for all decomposed modes;
[0014] S5, establishing appropriate prediction models for high-frequency components and low-frequency components respectively;
[0015] S6. All modal prediction results are denormalized and then superimposed to obtain the final time series prediction result.
[0016] Preferably, the preprocessing of the obtained data refers to data cleaning and missing data completion of the obtained original time series signal f(t), wherein the missing data completion adopts a linear interpolation method to fill the missing values, as follows:
[0017] For t 0 The data value f(t 0 ) is missing, respectively take t 1 and t 2 The data value f(t 1 ) and f(t 2 ), where t 1 <t 0 <t 2 , use linear interpolation to estimate t 0 The data value at time f(t 0 ), the calculation formula is as follows:
[0018]
[0019] In formula (1), f(t 0 ) is the time t 0 The data value on f(t 1 ) and f(t 2 ) are respectively at time point t 1 and t 2 The known data value of .
[0020] Preferably, according to the present invention, the VMD variational mode decomposition of the original time series data includes two steps: constructing a constrained variational problem and solving the constrained variational problem;
[0021] Construct the constrained variational problem as follows:
[0022]
[0023]
[0024] In formula (2), u kis the kth modal component after signal decomposition; ω k is the center frequency of the kth modal component; K is the number of modes to be decomposed; represents the partial derivative with respect to t; δ(t) represents the unit pulse function; j is the imaginary unit; πt is the constant factor; * is the convolution operator; f(t) is the original time series signal; t represents the time variable.
[0025] Solving the constrained variational problem is as follows:
[0026] First, the Lagrange multiplier and the second penalty factor are introduced to transform the constrained variational problem into an unconstrained variational problem, and the augmented Lagrangian expression is obtained:
[0027]
[0028]
[0029] In formula (3), α is the quadratic penalty factor, which can ensure the reconstruction accuracy of the signal in the presence of Gaussian noise; λ is the Lagrange multiplier, λ(t) is a time-dependent Lagrange multiplier function, and its main function is to introduce a constraint condition; f(t) is the original time series signal; To reconstruct the time series signal;
[0030] Then, the unconstrained variational problem is solved by the alternating direction multiplier method to achieve effective separation of signal frequencies, where the iterative update formulas for the eigenmode components and the center frequency are:
[0031]
[0032]
[0033] In formula (4) and formula (5), is the frequency domain representation of the kth eigenmode component at the n+1th iteration, f(t),u k Fourier transform of (t),λ(t); is the frequency domain sum of all modal components except the kth mode; ω is the center frequency of the modal component; is the center frequency of the kth modal component at the n+1th iteration;
[0034] For all ω ≥ 0, update
[0035]
[0036] In formula (6), ω is the center frequency of the modal component; is the frequency domain representation of the Lagrange multiplier at the n+1th iteration; is the frequency domain representation of the Lagrange multiplier at the nth iteration; τ is the noise; is the sum of the frequency domain representations of all K modal components;
[0037] Continue to iterate and update until the iteration termination condition is met:
[0038]
[0039] In formula (7), ε is a given threshold, ε = 1e-7; represents the frequency domain representation of the kth eigenmode component at the n+1th iteration; represents the frequency domain representation of the kth eigenmode component at the nth iteration;
[0040] If the conditions are met, the iteration is terminated and the decomposed K modal components are obtained.
[0041] Preferably, according to the present invention, the specific steps of step S3 are as follows:
[0042] S31, calculate the original time series signal f(t) and reconstruct the time series signal The value of the maximum information coefficient MIC between is used to measure the similarity between the reconstructed signal and the original signal; the value of the maximum information coefficient MIC is between 0 and 1. The larger the value, the stronger the dependence between the two variables. The calculation formula of the maximum information coefficient MIC is:
[0043]
[0044] In formula (8), f(t) is the original time series signal; To reconstruct the time series signal; is the value of f(t) and The mutual information between a and b is the distance along f(t) and The number of grids in the direction; B is the preset upper limit of the maximum number of grids; log(min(a,b)) is used to standardize the mutual information value so that the MIC value is between 0 and 1;
[0045] S32. Calculate and reconstruct time series signals The mean square error MSE between the original time series signal f(t) is as follows:
[0046]
[0047] In formula (9), f(t) is the original time series signal; is the reconstructed time series signal; N is the number of data points; i is the i-th data point; is the square of the error at each data point;
[0048] S33. Plot the MIC value and MSE value as a function of the number of modal components K. The MIC value will increase with the increase of K until it reaches a plateau, while the MSE value will decrease with the increase of K within a certain range. Considering the changing trends of the MIC and MSE values, select an optimal K value.
[0049] Preferably, according to the present invention, the specific steps of step S4 are as follows:
[0050] S41, set the zero-crossing rate threshold, use the zero-crossing rate to preliminarily divide the high-frequency and low-frequency components, and the zero-crossing rate calculation formula of each component is as follows:
[0051]
[0052] In formula (10), Z j is the zero-crossing rate; j zero is the number of zero crossing points; J is the signal length;
[0053] S42, setting a center frequency threshold, further dividing the high-frequency component initially divided by the zero-crossing rate using the center frequency, taking the high-frequency component after the center frequency division as the final high-frequency component, and merging the initially divided low-frequency component with the low-frequency component after the center frequency division as the final low-frequency component.
[0054] Preferably, according to the present invention, step S5 is specifically as follows:
[0055] S51, before making a prediction, respectively normalizing the standard deviation of the high-frequency component and the low-frequency component; after dividing the data set, respectively constructing prediction models for the high-frequency component and the low-frequency component;
[0056] S52. For all high-frequency components, the CNN-LSTM model and the GRU model are used for training and prediction respectively. The prediction results of the two models are weighted averaged using the adaptive weight adjustment method to obtain the final prediction results, as follows:
[0057] The prediction results of the CNN-LSTM model and the GRU model are dynamically assigned weights using an adaptive weight adjustment method. The adaptive weight adjustment method first calculates the mean square error (MSE) of each model, and then dynamically adjusts the weights based on these errors, as follows:
[0058] First, calculate the mean square error of the CNN-LSTM model and the GRU model. The calculation formula is as follows:
[0059]
[0060]
[0061] In formula (11) and formula (12), MSE cnn-lstm is the mean square error of the CNN-LSTM model; MSE gru is the mean square error of the GRU model; n is the number of samples in the validation set; is the true value of the i-th sample; is the predicted value of the CNN-LSTM model for the i-th sample; is the predicted value of the GRU model for the i-th sample;
[0062] Then, calculate the overall mean square error:
[0063] MSE total =MSE cnn-lstm +MSE gru (13);
[0064] Finally, according to the ratio of MSE to total error of CNN-LSTM model and GRU model, the corresponding weights are calculated. The weight calculation formula is as follows:
[0065]
[0066]
[0067] In formula (14) and formula (15), α and β are the weights of the CNN-LSTM model and the GRU model, respectively, where the weights α and β will be dynamically adjusted according to the MSE of each model;
[0068] After determining the weights of the CNN-LSTM model and the GRU model, the final prediction value is obtained through weighted combination:
[0069]
[0070] In formula (16), is the final predicted value;
[0071] The above dynamic weight allocation method calculates relative weights so that models with better performance receive higher weights, while models with poorer performance receive lower weights. This weight allocation method allows the model combination to focus more on models that perform well on the validation set. This MSE reverse use mechanism can reduce reliance on a single model, prevent the deviation of a certain model from affecting the overall prediction results, and effectively improve robustness.
[0072] A two-layer LSTM model is used for prediction of low-frequency components because low-frequency components usually involve a longer time span. LSTM can remember information from farther time points, and a single LSTM has a simple structure and is suitable for processing low-frequency data.
[0073] Preferably, according to the present invention, the CNN-LSTM model is composed of two one-dimensional convolutional layers and two LSTM layers. First, the convolutional layer extracts local features from the input time series; then, the CNN-LSTM model uses the LSTM layer to capture the long-term dependencies in the time series and deeply models the time series features; finally, the CNN-LSTM model outputs the predicted value through the fully connected layer;
[0074] The structure of the LSTM layer in the CNN-LSTM model includes three gates and a memory cell. The three gates include a forget gate, an input gate, and an output gate. At time t, the LSTM layer inputs the input value x of the network at the current time. t , the output h at the previous moment t-1 and the cell state of the previous moment; after the input, the first step in LSTM is to pass the forget gate f t The forget gate f decides what information in the cell state to discard. t The calculation method is:
[0075] f t =σ(W f ·[h t-1 ,x t ]+b f )(17)
[0076] In formula (17), f t represents the forget gate, σ() represents the sigmoid function, W f is the weight matrix of the forget gate, h t-1 is the hidden information of the previous moment, x t is the information input at the current moment, b f is the bias of the forget gate;
[0077] The second step in LSTM is the input gate, which determines what new information is added to the cell state. The calculation process is as follows:
[0078] i t =σ(W i ·[h t-1 ,x t ]+b i )(18)
[0079] In formula (18), i t represents the input gate, W i is the weight matrix of the input gate, bi is the input gate bias;
[0080]
[0081] In formula (19), It is obtained by further updating the information in the input gate using the tanh function, where W C is the memory gate weight matrix, b C is the memory gate bias, tanh() is the hyperbolic tangent function, and the symbol · indicates that the corresponding components of the vector are multiplied to obtain a new vector;
[0082]
[0083] In formula (20), C t-1 is the cell state at the previous moment, C t is the cell state at the current moment, is the current candidate cell state;
[0084] The third step in LSTM is the output gate, which determines what information to output. This output is based on the C t , the calculation process is as follows:
[0085] o t =σ(W o [h t-1 ,x t ]+b o )(twenty one)
[0086] h t =o t *tanh(C t )(twenty two)
[0087] In formula (21) and formula (22), o t is the output of the output gate, W o is the weight matrix of the output gate, b o is the output gate bias, h t is the hidden information at this moment in the LSTM model.
[0088] Preferably, the GRU model is composed of two layers of GRU networks. The first layer of GRU captures the dependencies in the time series by returning the output of the entire sequence, and passes the information to the next layer of GRU. The second layer of GRU further compresses and refines the time series features, and finally performs feature mapping through the fully connected layer to output the final prediction value.
[0089] The GRU model is a variant of LSTM. It has fewer parameters and faster convergence speed than LSTM. Compared with LSTM, the main difference of the GRU model is that the forget gate and input gate in the LSTM model are merged into an update gate. Since the GRU model only has an update gate and a reset gate, it can effectively save learning time when training a large amount of data. The GRU model structure includes an update gate and a reset gate:
[0090] Update gate: The update gate is used to control the state information of the previous moment to be passed to the current moment. The formula is as follows:
[0091] z t =σ(W z ·[h t-1 ,x t ]+b z )(twenty three)
[0092] In formula (23), z t is the update gate; σ() represents the sigmoid function; W z is the weight matrix of the update gate; h t-1 is the hidden information of the previous moment; x t The information input at the current moment; b z To update the gate bias;
[0093] Reset gate: The reset gate determines which historical information needs to be forgotten, that is, it determines the hidden information h at the previous moment. t-1 With the current information x t The degree of binding. The formula is as follows:
[0094] r t =σ(W r ·[h t-1 ,x t ]+b r (twenty four)
[0095] In formula (24), r t is the update gate; σ() represents the sigmoid function; W r is the weight matrix of the reset gate; h t-1 is the hidden information of the previous moment; x t The information input at the current moment; b r To reset the gate bias;
[0096] Based on the known update gate and reset gate, the candidate hidden state is further obtained:
[0097]
[0098] In formula (25), represents the candidate hidden state; tanh() represents the hyperbolic tangent function; W h and b h Respectively represent the corresponding weight matrix and bias;
[0099] Finally, the hidden information h at time t in the GRU model t ' can be expressed as:
[0100]
[0101] Compared with the prior art, the present invention has the following beneficial effects:
[0102] (1) The present invention decomposes complex time series data into multiple modal components with a single center frequency through VMD, which can effectively deal with the non-stationarity of the data and improve the accuracy of the prediction.
[0103] (2) The present invention uses the maximum information coefficient MIC and mean square error MSE between the reconstructed time series signal and the original time series signal to determine the modal number K, ensuring that the selected mode has the best prediction performance and avoiding redundancy. After determining K, VMD is used to decompose the original signal into K modal components, and the K components are divided into high and low frequencies using zero crossing rate and center frequency. By using zero crossing rate and center frequency to divide high and low frequency components, the accuracy of analysis and the effectiveness of prediction can be improved.
[0104] (3) The present invention adopts an integrated method of CNN-LSTM and GRU for high-frequency components, which can fully utilize their ability to capture local features and short-term dependencies, and adopts an adaptive module to dynamically adjust the model weights, so that the model combination can better adapt to data changes and enhance the robustness of the model; while LSTM is used for low-frequency components to effectively extract long-term dependencies and specifically improve the prediction effect.
[0105] (4) The present invention combines multiple deep learning models to make the prediction method more flexible and adaptable. It can select appropriate models for prediction for components with different characteristics, thereby improving the overall performance; and all the predicted modal components are superimposed together as the final prediction result. The method of the present invention can handle the big data challenges brought by high-frequency data. Through decomposition and combination strategies, it reduces the computational complexity and improves the model training efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0106] Figure 1 It is a flowchart of the time series data prediction method based on deep integrated learning model and high and low frequency separation described in the present invention.
[0107] Figure 2 1 is a diagram of power load data in an embodiment of the present invention.
[0108] Figure 3 It is an IMF diagram obtained by decomposing the power load data through VMD in the embodiment of the present invention.
[0109] Figure 4 It is a line graph showing changes in MIC and MSE as the number of modes changes in an embodiment of the present invention.
[0110] Figure 5 This is a result diagram of power load data after prediction in an embodiment of the present invention. DETAILED DESCRIPTION
[0111] The present disclosure is further described below in conjunction with the accompanying drawings and embodiments.
[0112] Embodiment 1,
[0113] This embodiment takes power grid load data as an example to provide a power grid load data prediction method based on deep ensemble learning model and high-low frequency separation. Figure 1 As shown, the method includes:
[0114] S1: Collect one year's load data from a power grid, which is measured every hour. Preprocess the data and use linear interpolation to fill in missing values.
[0115] S2: Perform VMD decomposition on the original load data to decompose the load sequence into multiple modal components with limited bandwidth;
[0116] S3: Combine the maximum information coefficient method and reconstruction error analysis method to determine the optimal number of decomposition modes K;
[0117] S4: Use zero-crossing rate and center frequency to separate high- and low-frequency components for all decomposed modes;
[0118] S5: Establish appropriate prediction models for high and low frequency components respectively;
[0119] S6: Superimpose all modal prediction results to obtain the final load prediction result.
[0120] The following is a detailed description of each step:
[0121] In step S1, the original load signal f(t) is preprocessed, including data cleaning and missing data completion. The processed load data signal is as follows: Figure 2 shown.
[0122] For example, 0 The load value at the moment f(t 0 ) is missing, but t is known 1 and t 2 The load values at each moment are f(t 1) and f(t 2 ), where t 1 <t 0 <t 2 , so use linear interpolation to estimate t 0 The load value at the moment f(t 0 ), the calculation formula is as follows:
[0123]
[0124] Where f(t 0 ) is the time t 0 The load value on is calculated by linear interpolation; and f(t 1 ) and f(t 2 ) are respectively at time point t 1 and t 2 The known load value of
[0125] In step S2, the purpose of VMD variational modal decomposition is to find a set of modal component IMFs through an optimization process to minimize the bandwidth between modes, thereby ensuring clear separation of modes. The original load data is f(t), where t is the time variable, and K modal components are obtained after VMD decomposition.
[0126] The implementation of VMD can be divided into two steps: constructing a variational problem and solving the variational problem. First, the variational problem should be constructed according to the goal of VMD. This problem attempts to minimize the bandwidth of each mode while ensuring that the sum of all modes is equal to the original signal. The constrained variational problem is as follows:
[0127]
[0128]
[0129] In the formula, u k is the kth modal component after signal decomposition; ω k is the center frequency of the kth modal component; K is the number of modes to be decomposed; represents the partial derivative with respect to t; δ(t) represents the unit pulse function; j is the imaginary unit; πt is the constant factor; * is the convolution operator; f(t) is the original load signal.
[0130] In order to solve the above constrained optimization problem, the Lagrange multiplier and the second penalty factor are introduced to transform the constrained variation problem into an unconstrained variation problem, and the augmented Lagrangian expression is obtained:
[0131]
[0132] In the formula, α is the quadratic penalty factor, which can ensure the reconstruction accuracy of the signal in the presence of Gaussian noise; λ is the Lagrange multiplier, λ(t) is a time-dependent Lagrange multiplier function, and its main function is to introduce a constraint condition; f(t) is the original load signal; To reconstruct the load signal.
[0133] Then, the unconstrained problem is solved by the alternating direction multiplier method, thereby achieving effective separation of signal frequencies. The iterative update formulas for the eigenmode components and the center frequency are:
[0134]
[0135]
[0136] In the formula, is the frequency domain representation of the kth eigenmode component at the n+1th iteration, f(t),u k Fourier transform of (t),λ(t); is the frequency domain sum of all other modal components except the kth mode; is the center frequency of the kth modal component at the n+1th iteration.
[0137] For all ω ≥ 0, update
[0138]
[0139] Where ω is the center frequency of the modal component; is the frequency domain representation of the Lagrange multiplier at the n+1th iteration; is the frequency domain representation of the Lagrange multiplier at the nth iteration; τ is the noise; is the sum of the frequency domain representations of all K modal components.
[0140] Continue to iterate and update until the iteration termination condition is met:
[0141]
[0142] Wherein, ε is a given threshold value, and in the present invention, ε=1e-7; represents the frequency domain representation of the kth eigenmode component at the n+1th iteration; represents the frequency domain representation of the kth eigenmode component at the nth iteration.
[0143] If the conditions are met, the iteration is terminated and the decomposed K modal components are obtained.
[0144] In step S3, the maximum information coefficient MIC and the mean square error MSE are calculated to jointly determine the number of modal components. The number of modal components directly affects the effect of signal decomposition. The maximum information coefficient MIC is an indicator that measures the correlation between two variables, which can detect nonlinear relationships. Here, the original load signal f(t) and the reconstructed load signal are calculated. The MIC value between is used to measure the similarity between the reconstructed signal and the original signal. The MIC value is between 0 and 1. The larger the value, the stronger the dependency between the two variables. The calculation formula of MIC is:
[0145]
[0146] Where f(t) is the original load signal; To reconstruct the load signal; is the value of f(t) and The mutual information between a and b is the distance along f(t) and The number of grids in the direction; B is a preset maximum number of grids; log(min(a,b)) is used to standardize the mutual information value so that the MIC value is between 0 and 1.
[0147] The MSE is used to evaluate the reconstructed signal. The difference between the original signal f(t) and the MSE is calculated as:
[0148]
[0149] f(t) is the original load signal; is the reconstructed load signal; N is the number of load data points; i is the i-th data point; is the square of the error at each data point.
[0150] By plotting the MIC value and MSE value versus the number of modal components K, as shown in Figure 4 As shown in Figure 1, it can be observed that the MIC value increases with the increase of K until it reaches a plateau; while the MSE value decreases with the increase of K within a certain range. Considering the changing trends of the MIC and MSE values, an optimal K value is selected, and the modal components obtained by VMD decomposition are as follows: Figure 3 shown.
[0151] In step S4, the zero-crossing rate is first used to preliminarily divide the high-frequency and low-frequency components. The zero-crossing rate calculation formula of each component is as follows:
[0152]
[0153] In the formula, Z j is the zero-crossing rate; jzero is the number of zero crossing points; J is the signal length.
[0154] A zero-crossing rate threshold is set here, which is set to 0.1 in this example. After the zero-crossing rate division, the high-frequency component divided initially is further divided using the center frequency, and the high-frequency component divided by the center frequency is used as the final high-frequency component. The threshold of the center frequency is set to the median of the center frequency, and the low-frequency component divided initially is combined with the low-frequency component divided by the center frequency as the final low-frequency component.
[0155] In step S5, high-frequency components usually contain rapid fluctuations and short-term changes, and traditional single models often find it difficult to simultaneously capture local features and long-term dependencies in the data. To this end, an integrated model is proposed in the present invention for predicting high-frequency components, which combines the advantages of the CNN-LSTM model and the GRU model to achieve more accurate predictions. In addition, for low-frequency components, the use of the LSTM model can effectively capture long-term dependencies.
[0156] Before prediction, the standard deviation of high-frequency components and low-frequency components is normalized. After the data set is divided, the prediction models of high-frequency components and low-frequency components are constructed respectively.
[0157] For all high-frequency components, the CNN-LSTM model and the GRU model are used for training and prediction respectively. The prediction results of the two models are weighted averaged using the adaptive weight adjustment method to obtain the final prediction result.
[0158] The CNN-LSTM model in this example consists of two one-dimensional convolutional layers and two LSTM layers. The convolutional layer is used to extract local features from the input time series, and the pooling layer reduces the data dimension by downsampling and retains key features. After the convolutional feature extraction, the model uses the LSTM layer to capture the long-term dependencies in the time series and deeply model the time series features. The final model outputs the predicted value through the fully connected layer, where the structure of the LSTM is as follows:
[0159] The LSTM model consists of three gates: forget gate, input gate, and output gate. In addition to the three gates, the LSTM also has a memory cell to remember the information of the previous moment. At time t, the LSTM has three inputs, including the input value x of the network at the current moment. t , the output h at the previous moment t-1 and the cell state at the previous moment. The first step in LSTM is to use the forget gate f t The forget gate f decides what information in the cell state to discard. t The calculation method is:
[0160] f t =σ(Wf ·[h t-1 ,x t ]+b f )
[0161] where f t represents the forget gate, σ() represents the sigmoid function, W f is the weight matrix of the forget gate, h t-1 is the hidden information of the previous moment, x t is the information input at the current moment, b f is the forget gate bias.
[0162] The second step in LSTM is the input gate, which determines what new information is added to the cell state. The calculation process is as follows:
[0163] i t =σ(W i ·[h t-1 ,x t ]+b i )
[0164]
[0165]
[0166] where i t represents the input gate, W i is the weight matrix of the input gate, b i is the input gate bias. It is obtained by further updating the information in the input gate using the tanh function, where W C is the memory gate weight matrix, b C is the memory gate bias, tanh() is the hyperbolic tangent function, and the symbol · indicates that the corresponding components of the vector are multiplied to obtain a new vector. C t-1 is the cell state at the previous moment, C t is the cell state at the current moment, is the current candidate cell state.
[0167] The third step in LSTM is the output gate, which determines what information to output. This output is based on the C t , the calculation process is as follows:
[0168] o t =σ(W o [h t-1 ,x t ]+b o )
[0169] h t =ot *tanh(C t )
[0170] Among them t is the output of the output gate, W o is the weight matrix of the output gate, b o is the output gate bias, h t For the hidden message of this moment.
[0171] The GRU model in this example consists of a two-layer GRU network. The first layer of GRU captures the dependencies in the time series by returning the output of the entire sequence, and passes the information to the next layer of GRU. The second layer of GRU further compresses and refines the time series features, and finally performs feature mapping through the fully connected layer to output the final prediction value. For the GRU model, it is a variant of LSTM, which has fewer parameters and faster convergence speed than LSTM. Compared with LSTM, the main difference of the GRU model is that the forget gate and the input gate in the LSTM model are merged into an update gate. Since the GRU model only has an update gate and a reset gate, it can effectively save learning time when training a large amount of data. The unit structure of GRU is as follows:
[0172] Update gate: The update gate is used to control the state information of the previous moment to be passed to the current moment. The formula is as follows:
[0173] z t =σ(W z ·[h t-1 ,x t ]+b z )
[0174] In the formula, z t is the update gate; σ() represents the sigmoid function; W z is the weight matrix of the update gate; h t-1 is the hidden information of the previous moment; x t is the information at the current moment; b z To update the gate bias.
[0175] Reset gate: The reset gate determines which historical information needs to be forgotten, that is, it determines the hidden information h at the previous moment. t-1 With the current information x t The degree of binding. The formula is as follows:
[0176] r t =σ(W r ·[h t-1 ,x t ]+b r
[0177] In the formula, rt is the update gate; σ() represents the sigmoid function; W r is the weight matrix of the reset gate; h t-1 is the state information of the previous moment; x t is the information at the current moment; b r To reset the gate bias.
[0178] Based on the known update gate and reset gate, the candidate hidden state is further obtained:
[0179]
[0180] In the formula, Candidate hidden state; tanh() hyperbolic tangent function; W h and b h They represent the corresponding weight matrix and bias respectively.
[0181] Finally, the hidden information h at time t in the GRU model t ' can be expressed as:
[0182]
[0183] After the training set of the modal component is input into the CNN-LSTM and GRU models for training, the validation set is used to dynamically assign weights through the adaptive module to better fuse the prediction results of multiple models. The present invention proposes a method based on adaptive weight adjustment for dynamically assigning weights of prediction results of multiple models. Specifically, the method calculates the mean square error (MSE) of each model and dynamically adjusts the weights according to these errors. First, the mean square error (MSE) of each model is calculated, and the calculation formula is as follows:
[0184]
[0185]
[0186] In the formula, MSE cnn-lstm is the mean square error of the CNN-LSTM model; MSE gru is the mean square error of the GRU model; n is the number of samples in the validation set; is the true value of the i-th sample; is the predicted value of the CNN-LSTM model for the i-th sample; is the predicted value of the GRU model for the i-th sample.
[0187] Then calculate the total mean square error:
[0188] MSE total =MSE cnn-lstm +MSEgru
[0189] Finally, the corresponding weight is calculated based on the ratio of MSE to total error of each model. The weight calculation formula is as follows:
[0190]
[0191]
[0192] Where α and β are the weights of the CNN-LSTM and GRU models, respectively, and the weights α and β are dynamically adjusted according to the MSE of each model.
[0193] After the model weights are determined through the validation set, the final prediction value is obtained through weighted combination:
[0194]
[0195] In the formula, is the final predicted value.
[0196] The above dynamic weight allocation method calculates relative weights so that models with better performance receive higher weights, while models with poorer performance receive lower weights. This weight allocation method allows the model combination to focus more on models that perform well on the validation set. This MSE reverse use mechanism can reduce reliance on a single model, prevent the deviation of a certain model from affecting the overall prediction results, and effectively improve robustness.
[0197] For low-frequency components, only a two-layer LSTM model is used for prediction, because low-frequency components usually involve a longer time span, LSTM can remember information from farther time points, and a single LSTM structure is simple and suitable for processing low-frequency data.
[0198] In step S6, the prediction results of all modal components are denormalized and then superimposed to obtain the final prediction result, such as Figure 5 As shown, in this example, different evaluation indicators are used to comprehensively evaluate the prediction ability of the model, including mean absolute percentage error (MAPE), root mean square error (RMSE), and mean absolute error (MAE). The smaller the values of these indicators, the smaller the error between the model prediction value and the true value, and the higher the prediction accuracy of the model.
Claims
1. A time series load data forecasting method based on deep ensemble learning model and high-low frequency separation, characterized in that: The method comprises: S1, collect the time series load data in the system, and then pre-process the obtained data; S2. Perform VMD variational mode decomposition on the preprocessed original time series load data to decompose the time series into A modal component with limited bandwidth; S3, combined maximum information coefficient method and reconstruction error analysis method to determine the optimal number of decomposition modes , the specific steps are as follows: S31. Calculate the original time series load signal and reconstruct the time series load signal The maximum information coefficient MIC value between , the calculation formula of the maximum information coefficient MIC is: (1) In formula (1), is the original time series load signal; To reconstruct the time series load signal; At a given grid size and Down, and The mutual information between and For along and The number of grids in the direction; The maximum number of grids that can be preset; Used to normalize the mutual information value so that the MIC value is between 0 and 1; S32. Calculate and reconstruct time series load signal and the original time series load signal The mean square error MSE between them is as follows: (2) In formula (2), is the original time series load signal; To reconstruct the time series load signal; is the number of data points; For the data points; is the square of the error at each data point; S33. Plot MIC and MSE values versus the number of modal components Considering the changing trend of MIC and MSE values, an optimal value; S4, use the zero-crossing rate and center frequency to divide the high and low frequency components for all decomposed modes; S5, establishing appropriate prediction models for high-frequency components and low-frequency components respectively; S6. All modal prediction results are denormalized and then superimposed to obtain the final time series prediction result.
2. According to claim 1, a time series load data forecasting method based on deep ensemble learning model and high and low frequency separation is characterized in that: The preprocessing of the obtained data refers to the processing of the original time series load signal obtained. Data cleaning and missing data completion are performed. The missing data completion uses linear interpolation to fill in missing values, as follows: against Data value at time Missing, respectively and Data value at time and ,in , estimated using linear interpolation Data value at time , the calculation formula is as follows: (3) In formula (3), It's time The data value on ; and At the time point and The known data value of .
3. According to claim 1, a time series load data prediction method based on deep ensemble learning model and high and low frequency separation is characterized in that: The VMD variational mode decomposition of the preprocessed original time series load data includes two steps: constructing a constrained variational problem and solving the constrained variational problem; Construct the constrained variational problem as follows: (4) In formula (4), After signal decomposition, modal components; For the The center frequency of the modal components; is the number of modes to be decomposed; Express The partial derivative of represents the unit impulse function; is an imaginary unit; is a constant factor; is the convolution operator; is the original time series load signal; t represents the time variable; Solving the constrained variational problem is as follows: First, the Lagrange multiplier and the second penalty factor are introduced to transform the constrained variational problem into an unconstrained variational problem, and the augmented Lagrangian expression is obtained: (5) In formula (5), is a quadratic penalty factor, which can ensure the reconstruction accuracy of the signal in the presence of Gaussian noise; is the Lagrange multiplier, is a time-dependent Lagrange multiplier function, and its main function is to introduce a constraint condition; is the original time series load signal; To reconstruct the time series load signal; Then, the unconstrained variational problem is solved by the alternating direction multiplier method, where the iterative update formulas for the eigenmode components and the center frequency are: (6) (7) In formula (6) and formula (7), For the The first iteration The frequency domain representation of the eigenmode components is: They are Fourier transform of For the The frequency domain sum of all other modal components except the mode; is the center frequency of the modal component; For the The first iteration The center frequency of the modal components; For all ,renew : (8) In formula (8), is the center frequency of the modal component; For the The frequency domain representation of the Lagrange multipliers at the iteration; For the The frequency domain representation of the Lagrange multipliers at the iteration; for noise; For all The sum of the frequency domain representations of the modal components; Continue to iterate and update until the iteration termination condition is met: (9) In formula (9), For a given threshold, ; Indicates The first iteration Frequency domain representation of the eigenmode components; Indicates The first iteration Frequency domain representation of the eigenmode components; If the conditions are met, the iteration is terminated and the decomposed K modal components are obtained.
4. The method for predicting time series load data based on deep ensemble learning model and high- and low-frequency separation according to claim 1 is characterized in that: The specific steps of step S4 are as follows: S41, set the zero-crossing rate threshold, use the zero-crossing rate to preliminarily divide the high-frequency and low-frequency components, and the zero-crossing rate calculation formula of each component is as follows: (10) In formula (10), is the zero-crossing rate; is the number of zero crossings; is the signal length; S42, setting a center frequency threshold, further dividing the high-frequency component initially divided by the zero-crossing rate using the center frequency, taking the high-frequency component after the center frequency division as the final high-frequency component, and merging the initially divided low-frequency component with the low-frequency component after the center frequency division as the final low-frequency component.
5. The method for predicting time series load data based on deep ensemble learning model and high- and low-frequency separation according to claim 1 is characterized in that: The step S5 is specifically as follows: S51, before making a prediction, respectively normalizing the standard deviation of the high-frequency component and the low-frequency component; after dividing the data set, respectively constructing prediction models for the high-frequency component and the low-frequency component; S52. For all high-frequency components, the CNN-LSTM model and the GRU model are used for training and prediction respectively. The prediction results of the two models are weighted averaged using the adaptive weight adjustment method to obtain the final prediction results, as follows: The prediction results of the CNN-LSTM model and the GRU model are dynamically assigned weights using an adaptive weight adjustment method. The adaptive weight adjustment method first calculates the mean square error (MSE) of each model, and then dynamically adjusts the weights based on these errors, as follows: First, calculate the mean square error of the CNN-LSTM model and the GRU model. The calculation formula is as follows: (11) (12) In formula (11) and formula (12), is the mean square error of the CNN-LSTM model; is the mean square error of the GRU model; is the number of samples in the validation set; For the The true value of samples; For the The predicted value of the CNN-LSTM model for samples; For the The predicted value of the GRU model for samples; Then, calculate the overall mean square error: (13); Finally, according to the ratio of MSE to total error of CNN-LSTM model and GRU model, the corresponding weights are calculated. The weight calculation formula is as follows: (14) (15) In formula (14) and formula (15), and are the weights of the CNN-LSTM model and the GRU model, where the weight and It will be dynamically adjusted based on the MSE of each model; After determining the weights of the CNN-LSTM model and the GRU model, the final prediction value is obtained through weighted combination: (16) In formula (16), is the final predicted value; S53. Use a two-layer LSTM model to predict the low-frequency components.
6. A time series load data forecasting method based on deep ensemble learning model and high and low frequency separation according to claim 5, characterized in that: The CNN-LSTM model consists of two one-dimensional convolutional layers and two LSTM layers. First, the convolutional layer extracts local features from the input time series. Then, the CNN-LSTM model uses the LSTM layer to capture the long-term dependencies in the time series and deeply models the time series features. Finally, the CNN-LSTM model outputs the predicted value through the fully connected layer. The structure of the LSTM layer in the CNN-LSTM model includes three gates and a memory cell, wherein the three gates include a forget gate, an input gate, and an output gate; At this moment, the LSTM layer inputs the input value of the network at the current moment , the output at the previous moment and the cell state at the previous moment; after the input, the first step in LSTM is to pass the forget gate Decide what information in the cell state to discard, the forget gate The calculation method is: (17) In formula (17), represents the forget gate, represents the sigmoid function, is the weight matrix of the forget gate, is the hidden information of the previous moment, The information entered at the current moment, is the bias of the forget gate; The second step in LSTM is the input gate, which determines what new information is added to the cell state. The calculation process is as follows: (18) In formula (18), represents the input gate, is the weight matrix of the input gate, is the input gate bias; (19) In formula (19), It is obtained by further updating the information in the input gate using the tanh function, where is the memory gate weight matrix, is the memory gate bias, Hyperbolic tangent function, symbol Indicates that the corresponding components of the vector are multiplied to obtain a new vector; (20) In formula (20), is the cell state at the previous moment, is the cell state at the current moment, is the current candidate cell state; The third step in LSTM is the output gate, which determines what information to output. This output is based on the previous step. , the calculation process is as follows: (21) (22) In formula (21) and formula (22), is the output of the output gate, is the weight matrix of the output gate, is the output gate bias, is the hidden information at this moment in the LSTM model.
7. The method for predicting time series load data based on deep ensemble learning model and high- and low-frequency separation according to claim 5 is characterized in that: The GRU model consists of two layers of GRU networks. The first layer of GRU captures the dependencies in the time series by returning the output of the entire sequence and passes the information to the next layer of GRU. The second layer of GRU further compresses and refines the time series features, and finally performs feature mapping through the fully connected layer to output the final prediction value. The GRU model structure includes an update gate and a reset gate: Update gate: The update gate is used to control the state information of the previous moment to be passed to the current moment. The formula is as follows: (23) In formula (23), To update the door; Represents the sigmoid function; is the weight matrix of the update gate; is the hidden information of the previous moment; The information entered for the current moment; To update the gate bias; Reset gate: The reset gate determines which historical information needs to be forgotten, that is, it determines the hidden information of the previous moment. Information about the current moment The degree of combination is as follows: (24) In formula (24), To update the door; Represents the sigmoid function; is the weight matrix of the reset gate; is the hidden information of the previous moment; The information entered for the current moment; To reset the gate bias; Based on the known update gate and reset gate, the candidate hidden state is further obtained: (25) In formula (25), represents the candidate hidden state; represents the hyperbolic tangent function; and Respectively represent the corresponding weight matrix and bias; Finally, in the GRU model Hidden information of the moment It can be expressed as: (26)。
Citation Information
Patent Citations
Non-intrusive load decomposition method based on EEMD and GRU
CN113095290A
Frequency analysis method and system for time sequence prediction
CN116451051A
Wind power plant short-term wind speed prediction method integrated with deep learning model
CN110738010A
Reservoir physical property parameter prediction method combined with deep learning
CN110852527A